{
  "id": 370156,
  "title": "[Solved] How to submit, what is wrong ",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/370156",
  "author_name": "Hey24sheep",
  "post_date": "2022-12-03T11:24:17.934000",
  "votes": 12,
  "comment_count": 12,
  "views": 0,
  "content": "<p>I am losing my mind. What is wrong with this competition &amp; my code :\"( </p>\n<p>Am I missing something? </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9062237%2F48c724bc2ea2231d099171a44a2abfa4%2Ferror_Capture.PNG?generation=1670066138086282&amp;alt=media\" alt=\"\"></p>\n<p>Steps I am following are as follows</p>\n<p>1- Convert &amp; Copy files \"dcm\" to <code>SAVE_FOLDER = \"/kaggle/tmp/output/\"</code></p>\n<p>NOTE : I have tried without copying as well. I infer from file without copy directly but that causes \"Notebook Timeout\"</p>\n<pre><code> ():\n    image = read_xray(fpath)\n    fsplits = fpath.split()\n    patient_id = fsplits[-]\n    image_id = fsplits[-].split()[]\n\n    :\n        os.mkdir()\n    :\n        \n\n    \n    cv2.imwrite(, image)\n\ntestfiles = glob.glob()\n\n_ = Parallel(n_jobs=)(\n    delayed(process)(fpath)\n     fpath  tqdm(testfiles)\n)\n</code></pre>\n<p>2 - Infer (trainer + model variables exists)</p>\n<pre><code>test_csv = pd.read_csv()\nog_sub = pd.read_csv().head()\n\ntestfiles = glob.glob()\n\n fpath  tqdm(testfiles):\n    fsplits = fpath.split()\n    patient_id = fsplits[-]\n    image_id = fsplits[-].split()[]\n    prediction_id = (test_csv.loc[test_csv.image_id == (image_id)].prediction_id.values[]).strip()\n    predictions =  \n\n    :\n        datamodule = ImageClassificationData.from_files(predict_files=[fpath,],batch_size=,)\n        predictions = trainer.predict(model, datamodule=datamodule, output=)[][][]\n    :\n        \n        \n\n    newdf = pd.DataFrame([[prediction_id, predictions]], columns=[,])\n    og_sub = pd.concat([newdf, og_sub])\n\nog_sub = og_sub.groupby().mean()  \nog_sub = og_sub.sort_index()\nog_sub.to_csv(, index=)\n\ndisplay(og_sub)\n</code></pre>\n<p><strong>All of these notebook versions you see in the image are succesfull in full commit. But fails during submission.</strong><br>\nThis code above casuses \"Notebook threw exception\". What is wrong here? Am I missing something?</p>\n<p><strong>Solution</strong><br>\nFew things are happening, So I got it solved by doing a few steps. Data is super large on the hidden testset.</p>\n<ul>\n<li>So, convert the data to \"/kaggle/tmp/your_folder\" or \"/kaggle/tmp/output\". Otherwise the notebook will fail (this was happening to me).</li>\n<li>Second issue, converstion is so slow that it takes almost 6+ hours to convert the test set (how do I know? because I tried it in 1 notebook with dummy submission). So read this <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/371033#2060484\" target=\"_blank\">discussion</a>. It's 2+ hours faster.</li>\n<li>Thrid, try to infer on smaller image sizes, or play with img size as per your model size and available ram. It times out if conversion is on 1024+ image size. It depends if you have a larger model then it might give \"ram overflow\" and the notebook will fail.</li>\n</ul>",
  "messages": [
    {
      "id": 2053540,
      "postDate": "2022-12-03T11:24:17.933Z",
      "content": "<p>I am losing my mind. What is wrong with this competition &amp; my code :\"( </p>\n<p>Am I missing something? </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9062237%2F48c724bc2ea2231d099171a44a2abfa4%2Ferror_Capture.PNG?generation=1670066138086282&amp;alt=media\" alt=\"\"></p>\n<p>Steps I am following are as follows</p>\n<p>1- Convert &amp; Copy files \"dcm\" to <code>SAVE_FOLDER = \"/kaggle/tmp/output/\"</code></p>\n<p>NOTE : I have tried without copying as well. I infer from file without copy directly but that causes \"Notebook Timeout\"</p>\n<pre><code> ():\n    image = read_xray(fpath)\n    fsplits = fpath.split()\n    patient_id = fsplits[-]\n    image_id = fsplits[-].split()[]\n\n    :\n        os.mkdir()\n    :\n        \n\n    \n    cv2.imwrite(, image)\n\ntestfiles = glob.glob()\n\n_ = Parallel(n_jobs=)(\n    delayed(process)(fpath)\n     fpath  tqdm(testfiles)\n)\n</code></pre>\n<p>2 - Infer (trainer + model variables exists)</p>\n<pre><code>test_csv = pd.read_csv()\nog_sub = pd.read_csv().head()\n\ntestfiles = glob.glob()\n\n fpath  tqdm(testfiles):\n    fsplits = fpath.split()\n    patient_id = fsplits[-]\n    image_id = fsplits[-].split()[]\n    prediction_id = (test_csv.loc[test_csv.image_id == (image_id)].prediction_id.values[]).strip()\n    predictions =  \n\n    :\n        datamodule = ImageClassificationData.from_files(predict_files=[fpath,],batch_size=,)\n        predictions = trainer.predict(model, datamodule=datamodule, output=)[][][]\n    :\n        \n        \n\n    newdf = pd.DataFrame([[prediction_id, predictions]], columns=[,])\n    og_sub = pd.concat([newdf, og_sub])\n\nog_sub = og_sub.groupby().mean()  \nog_sub = og_sub.sort_index()\nog_sub.to_csv(, index=)\n\ndisplay(og_sub)\n</code></pre>\n<p><strong>All of these notebook versions you see in the image are succesfull in full commit. But fails during submission.</strong><br>\nThis code above casuses \"Notebook threw exception\". What is wrong here? Am I missing something?</p>\n<p><strong>Solution</strong><br>\nFew things are happening, So I got it solved by doing a few steps. Data is super large on the hidden testset.</p>\n<ul>\n<li>So, convert the data to \"/kaggle/tmp/your_folder\" or \"/kaggle/tmp/output\". Otherwise the notebook will fail (this was happening to me).</li>\n<li>Second issue, converstion is so slow that it takes almost 6+ hours to convert the test set (how do I know? because I tried it in 1 notebook with dummy submission). So read this <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/371033#2060484\" target=\"_blank\">discussion</a>. It's 2+ hours faster.</li>\n<li>Thrid, try to infer on smaller image sizes, or play with img size as per your model size and available ram. It times out if conversion is on 1024+ image size. It depends if you have a larger model then it might give \"ram overflow\" and the notebook will fail.</li>\n</ul>",
      "rawMarkdown": "I am losing my mind. What is wrong with this competition & my code :\"( \n\nAm I missing something? \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9062237%2F48c724bc2ea2231d099171a44a2abfa4%2Ferror_Capture.PNG?generation=1670066138086282&alt=media)\n\nSteps I am following are as follows\n\n1- Convert & Copy files \"dcm\" to `SAVE_FOLDER = \"/kaggle/tmp/output/\"`\n\nNOTE : I have tried without copying as well. I infer from file without copy directly but that causes \"Notebook Timeout\"\n\n```python\ndef process(fpath):\n    image = read_xray(fpath)\n    fsplits = fpath.split('/')\n    patient_id = fsplits[-2]\n    image_id = fsplits[-1].split('.')[0]\n\n    try:\n        os.mkdir(f\"{SAVE_FOLDER}/{patient_id}\")\n    except:\n        pass\n\n    # save\n    cv2.imwrite(f\"{SAVE_FOLDER}/{patient_id}/{image_id}.png\", image)\n    \ntestfiles = glob.glob(\"/kaggle/input/rsna-breast-cancer-detection/test_images/*/*.dcm\")\n\n_ = Parallel(n_jobs=4)(\n    delayed(process)(fpath)\n    for fpath in tqdm(testfiles)\n)\n```\n2 - Infer (trainer + model variables exists)\n\n```python\ntest_csv = pd.read_csv(\"/kaggle/input/rsna-breast-cancer-detection/test.csv\")\nog_sub = pd.read_csv(\"/kaggle/input/rsna-breast-cancer-detection/sample_submission.csv\").head(0)\n\ntestfiles = glob.glob(f\"{SAVE_FOLDER}/*/*.png\")\n\nfor fpath in tqdm(testfiles):\n    fsplits = fpath.split('/')\n    patient_id = fsplits[-2]\n    image_id = fsplits[-1].split('.')[0]\n    prediction_id = str(test_csv.loc[test_csv.image_id == int(image_id)].prediction_id.values[0]).strip()\n    predictions = 0.00100000 # dummy, incase of fail\n\n    try:\n        datamodule = ImageClassificationData.from_files(predict_files=[fpath,],batch_size=1,)\n        predictions = trainer.predict(model, datamodule=datamodule, output=\"probabilities\")[0][0][0]\n    except:\n        # predict failed\n        pass\n\n    newdf = pd.DataFrame([[prediction_id, predictions]], columns=['prediction_id','cancer'])\n    og_sub = pd.concat([newdf, og_sub])\n    \nog_sub = og_sub.groupby('prediction_id').mean()  #dummy aggregation method\nog_sub = og_sub.sort_index()\nog_sub.to_csv(\"submission.csv\", index=True)\n\ndisplay(og_sub)\n```\n**All of these notebook versions you see in the image are succesfull in full commit. But fails during submission.**\nThis code above casuses \"Notebook threw exception\". What is wrong here? Am I missing something?\n\n**Solution**\nFew things are happening, So I got it solved by doing a few steps. Data is super large on the hidden testset.\n\n- So, convert the data to \"/kaggle/tmp/your_folder\" or \"/kaggle/tmp/output\". Otherwise the notebook will fail (this was happening to me).\n- Second issue, converstion is so slow that it takes almost 6+ hours to convert the test set (how do I know? because I tried it in 1 notebook with dummy submission). So read this [discussion](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/371033#2060484). It's 2+ hours faster.\n- Thrid, try to infer on smaller image sizes, or play with img size as per your model size and available ram. It times out if conversion is on 1024+ image size. It depends if you have a larger model then it might give \"ram overflow\" and the notebook will fail.",
      "votes": 12
    },
    {
      "id": 2054138,
      "postDate": "2022-12-03T22:04:55.977Z",
      "content": "<p>Maybe set index=False for your submission file. The sample sub file doesn't have an index column.</p>",
      "rawMarkdown": "Maybe set index=False for your submission file. The sample sub file doesn't have an index column.",
      "votes": 1,
      "replies": [
        {
          "id": 2054445,
          "postDate": "2022-12-04T06:41:46.077Z",
          "content": "<p>I did. Both kind of submission files work. If you check my name on leaderboards I am placed that's because I test submitting a sample file with just index=True. Both of the files are accepted. Issue is \"Notebook threw exception\" or \"Timeout\". I am unable to know why either of those happening.</p>",
          "rawMarkdown": "I did. Both kind of submission files work. If you check my name on leaderboards I am placed that's because I test submitting a sample file with just index=True. Both of the files are accepted. Issue is \"Notebook threw exception\" or \"Timeout\". I am unable to know why either of those happening."
        },
        {
          "id": 2055100,
          "postDate": "2022-12-04T18:10:00.777Z",
          "content": "<p>Timeout is self explanatory. Your submission took too long.</p>",
          "rawMarkdown": "Timeout is self explanatory. Your submission took too long."
        },
        {
          "id": 2055495,
          "postDate": "2022-12-05T05:52:46.260Z",
          "content": "<p>Yes, that is self explanatory. </p>",
          "rawMarkdown": "Yes, that is self explanatory. "
        }
      ]
    },
    {
      "id": 2053574,
      "postDate": "2022-12-03T12:14:07.540Z",
      "content": "<p>Try running your code on all the train images.</p>\n<p>My guess is that you are not installing the pylibjpeg &amp; python_gdcm dependencies. (cf 2nd cell of my public notebook)</p>",
      "rawMarkdown": "Try running your code on all the train images.\n\nMy guess is that you are not installing the pylibjpeg & python_gdcm dependencies. (cf 2nd cell of my public notebook)",
      "votes": 1,
      "replies": [
        {
          "id": 2053671,
          "postDate": "2022-12-03T13:33:50.123Z",
          "content": "<p>Sounds good, I'll try it.</p>",
          "rawMarkdown": "Sounds good, I'll try it."
        },
        {
          "id": 2054444,
          "postDate": "2022-12-04T06:40:31.743Z",
          "content": "<p>I have those dependencies installed. </p>",
          "rawMarkdown": "I have those dependencies installed. "
        }
      ]
    },
    {
      "id": 2060450,
      "postDate": "2022-12-10T00:03:08.437Z",
      "content": "<p>Did you solve it? what was the problem?</p>",
      "rawMarkdown": "Did you solve it? what was the problem?",
      "replies": [
        {
          "id": 2060491,
          "postDate": "2022-12-10T03:03:32.030Z",
          "content": "<p>Few things are happening, So I got it solved by doing a few steps. Data is super large on the hidden testset. </p>\n<ul>\n<li>So, convert the data to \"/kaggle/tmp/your_folder\" or \"/kaggle/tmp/output\". Otherwise the notebook will fail (this was happening to me). </li>\n<li>Second, converstion is so slow that it takes almost 6+ hours to convert, So read this <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/371033#2060484\" target=\"_blank\">Discussion</a>. It's 2+ hours faster.</li>\n<li>Thrid, try to infer on smaller image sizes. It times out if conversion is on 1024+ image size.</li>\n</ul>",
          "rawMarkdown": "Few things are happening, So I got it solved by doing a few steps. Data is super large on the hidden testset. \n\n- So, convert the data to \"/kaggle/tmp/your_folder\" or \"/kaggle/tmp/output\". Otherwise the notebook will fail (this was happening to me). \n- Second, converstion is so slow that it takes almost 6+ hours to convert, So read this [Discussion](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/371033#2060484). It's 2+ hours faster.\n- Thrid, try to infer on smaller image sizes. It times out if conversion is on 1024+ image size.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2053652,
      "postDate": "2022-12-03T13:14:06.160Z",
      "content": "<p>I met same problem,too.</p>\n<p>Here is the part of code maybe cause this:</p>\n<pre><code>SAVE_DIR = \"/kaggle/working/\"\n# dcm to png\ndef process(line, size=512, save_folder=\"\", extension=\"png\"):\n    items = line.strip().split(',')\n    patient = items[1]\n    image = items[2]\n    laterality = items[3]\n\n    dcm_path = os.path.join(DATA_DIR, \"test_images\", patient, image+\".dcm\")\n    dicom = pydicom.dcmread(dcm_path)\n    img = dicom.pixel_array\n\n    img = (img - img.min()) / (img.max() - img.min())\n    if dicom.PhotometricInterpretation == \"MONOCHROME1\":\n        img = 1 - img\n    img = cv2.resize(img, (size, size))\n    cv2.imwrite(save_folder + f\"{patient}_{image}_{laterality}.{extension}\", (img * 255).astype(np.uint8))\n\n\ntest_csv = os.path.join(DATA_DIR,\"test.csv\")\nwith open(test_csv, 'r') as f:\n    test_lines = f.readlines()[1:]\nprint(len(test_lines))\n\nSAVE_FOLDER = os.path.join(SAVE_DIR, \"png/\")\nEXTENSION = \"png\"\nos.makedirs(SAVE_FOLDER, exist_ok=True)\n\n_ = Parallel(n_jobs=4)(\n    delayed(process)(line, size=SIZE, save_folder=SAVE_FOLDER, extension=EXTENSION)\n    for line in tqdm(test_lines)\n)\n</code></pre>\n<p>I can run successfully in my notebook, but when I submit always got <strong>Notebook Threw Exception</strong>.<br>\nCan anyone provide some clues? Thx.</p>",
      "rawMarkdown": "I met same problem,too.\n\nHere is the part of code maybe cause this:\n\n\n```\nSAVE_DIR = \"/kaggle/working/\"\n# dcm to png\ndef process(line, size=512, save_folder=\"\", extension=\"png\"):\n    items = line.strip().split(',')\n    patient = items[1]\n    image = items[2]\n    laterality = items[3]\n    \n    dcm_path = os.path.join(DATA_DIR, \"test_images\", patient, image+\".dcm\")\n    dicom = pydicom.dcmread(dcm_path)\n    img = dicom.pixel_array\n\n    img = (img - img.min()) / (img.max() - img.min())\n    if dicom.PhotometricInterpretation == \"MONOCHROME1\":\n        img = 1 - img\n    img = cv2.resize(img, (size, size))\n    cv2.imwrite(save_folder + f\"{patient}_{image}_{laterality}.{extension}\", (img * 255).astype(np.uint8))\n    \n    \ntest_csv = os.path.join(DATA_DIR,\"test.csv\")\nwith open(test_csv, 'r') as f:\n    test_lines = f.readlines()[1:]\nprint(len(test_lines))\n\nSAVE_FOLDER = os.path.join(SAVE_DIR, \"png/\")\nEXTENSION = \"png\"\nos.makedirs(SAVE_FOLDER, exist_ok=True)\n\n_ = Parallel(n_jobs=4)(\n    delayed(process)(line, size=SIZE, save_folder=SAVE_FOLDER, extension=EXTENSION)\n    for line in tqdm(test_lines)\n)\n```\n\n\nI can run successfully in my notebook, but when I submit always got **Notebook Threw Exception**.\nCan anyone provide some clues? Thx.",
      "replies": [
        {
          "id": 2056460,
          "postDate": "2022-12-06T06:20:35.777Z",
          "content": "<p>Finally I submit successfully. Here is the two change for my code:</p>\n<ol>\n<li>Install pylibjpeg as Theo Viel  said. (Although I can run in test env without it.)</li>\n<li>Change png output dir from /kaggle/working/ to /kaggle/tmp/ (Maybe full test png data extend 20G?)</li>\n</ol>",
          "rawMarkdown": "Finally I submit successfully. Here is the two change for my code:\n1. Install pylibjpeg as Theo Viel  said. (Although I can run in test env without it.)\n2. Change png output dir from /kaggle/working/ to /kaggle/tmp/ (Maybe full test png data extend 20G?)",
          "votes": 4
        },
        {
          "id": 2056489,
          "postDate": "2022-12-06T06:58:04.123Z",
          "content": "<p>I have this, now I am facing \"Notebook timeout\" :( trying other params to bypass timeout.</p>",
          "rawMarkdown": "I have this, now I am facing \"Notebook timeout\" :( trying other params to bypass timeout."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2054138,
      "author_name": "David Roberts",
      "author_url": "",
      "post_date": "2022-12-03T22:04:55.977000",
      "content": "<p>Maybe set index=False for your submission file. The sample sub file doesn't have an index column.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2054445,
          "author_name": "Hey24sheep",
          "author_url": "",
          "post_date": "2022-12-04T06:41:46.077000",
          "content": "<p>I did. Both kind of submission files work. If you check my name on leaderboards I am placed that's because I test submitting a sample file with just index=True. Both of the files are accepted. Issue is \"Notebook threw exception\" or \"Timeout\". I am unable to know why either of those happening.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2055100,
          "author_name": "David Roberts",
          "author_url": "",
          "post_date": "2022-12-04T18:10:00.777000",
          "content": "<p>Timeout is self explanatory. Your submission took too long.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2055495,
          "author_name": "Hey24sheep",
          "author_url": "",
          "post_date": "2022-12-05T05:52:46.260000",
          "content": "<p>Yes, that is self explanatory. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2053574,
      "author_name": "Theo Viel",
      "author_url": "",
      "post_date": "2022-12-03T12:14:07.540000",
      "content": "<p>Try running your code on all the train images.</p>\n<p>My guess is that you are not installing the pylibjpeg &amp; python_gdcm dependencies. (cf 2nd cell of my public notebook)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2053671,
          "author_name": "Mr.Fire",
          "author_url": "",
          "post_date": "2022-12-03T13:33:50.123000",
          "content": "<p>Sounds good, I'll try it.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2054444,
          "author_name": "Hey24sheep",
          "author_url": "",
          "post_date": "2022-12-04T06:40:31.743000",
          "content": "<p>I have those dependencies installed. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2060450,
      "author_name": "hyunseoki",
      "author_url": "",
      "post_date": "2022-12-10T00:03:08.437000",
      "content": "<p>Did you solve it? what was the problem?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2060491,
          "author_name": "Hey24sheep",
          "author_url": "",
          "post_date": "2022-12-10T03:03:32.030000",
          "content": "<p>Few things are happening, So I got it solved by doing a few steps. Data is super large on the hidden testset. </p>\n<ul>\n<li>So, convert the data to \"/kaggle/tmp/your_folder\" or \"/kaggle/tmp/output\". Otherwise the notebook will fail (this was happening to me). </li>\n<li>Second, converstion is so slow that it takes almost 6+ hours to convert, So read this <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/371033#2060484\" target=\"_blank\">Discussion</a>. It's 2+ hours faster.</li>\n<li>Thrid, try to infer on smaller image sizes. It times out if conversion is on 1024+ image size.</li>\n</ul>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2053652,
      "author_name": "Mr.Fire",
      "author_url": "",
      "post_date": "2022-12-03T13:14:06.160000",
      "content": "<p>I met same problem,too.</p>\n<p>Here is the part of code maybe cause this:</p>\n<pre><code>SAVE_DIR = \"/kaggle/working/\"\n# dcm to png\ndef process(line, size=512, save_folder=\"\", extension=\"png\"):\n    items = line.strip().split(',')\n    patient = items[1]\n    image = items[2]\n    laterality = items[3]\n\n    dcm_path = os.path.join(DATA_DIR, \"test_images\", patient, image+\".dcm\")\n    dicom = pydicom.dcmread(dcm_path)\n    img = dicom.pixel_array\n\n    img = (img - img.min()) / (img.max() - img.min())\n    if dicom.PhotometricInterpretation == \"MONOCHROME1\":\n        img = 1 - img\n    img = cv2.resize(img, (size, size))\n    cv2.imwrite(save_folder + f\"{patient}_{image}_{laterality}.{extension}\", (img * 255).astype(np.uint8))\n\n\ntest_csv = os.path.join(DATA_DIR,\"test.csv\")\nwith open(test_csv, 'r') as f:\n    test_lines = f.readlines()[1:]\nprint(len(test_lines))\n\nSAVE_FOLDER = os.path.join(SAVE_DIR, \"png/\")\nEXTENSION = \"png\"\nos.makedirs(SAVE_FOLDER, exist_ok=True)\n\n_ = Parallel(n_jobs=4)(\n    delayed(process)(line, size=SIZE, save_folder=SAVE_FOLDER, extension=EXTENSION)\n    for line in tqdm(test_lines)\n)\n</code></pre>\n<p>I can run successfully in my notebook, but when I submit always got <strong>Notebook Threw Exception</strong>.<br>\nCan anyone provide some clues? Thx.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2056460,
          "author_name": "Mr.Fire",
          "author_url": "",
          "post_date": "2022-12-06T06:20:35.777000",
          "content": "<p>Finally I submit successfully. Here is the two change for my code:</p>\n<ol>\n<li>Install pylibjpeg as Theo Viel  said. (Although I can run in test env without it.)</li>\n<li>Change png output dir from /kaggle/working/ to /kaggle/tmp/ (Maybe full test png data extend 20G?)</li>\n</ol>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 2056489,
          "author_name": "Hey24sheep",
          "author_url": "",
          "post_date": "2022-12-06T06:58:04.123000",
          "content": "<p>I have this, now I am facing \"Notebook timeout\" :( trying other params to bypass timeout.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2053540": "I am losing my mind. What is wrong with this competition & my code :\"( \n\nAm I missing something? \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9062237%2F48c724bc2ea2231d099171a44a2abfa4%2Ferror_Capture.PNG?generation=1670066138086282&alt=media)\n\nSteps I am following are as follows\n\n1- Convert & Copy files \"dcm\" to `SAVE_FOLDER = \"/kaggle/tmp/output/\"`\n\nNOTE : I have tried without copying as well. I infer from file without copy directly but that causes \"Notebook Timeout\"\n\n```python\ndef process(fpath):\n    image = read_xray(fpath)\n    fsplits = fpath.split('/')\n    patient_id = fsplits[-2]\n    image_id = fsplits[-1].split('.')[0]\n\n    try:\n        os.mkdir(f\"{SAVE_FOLDER}/{patient_id}\")\n    except:\n        pass\n\n    # save\n    cv2.imwrite(f\"{SAVE_FOLDER}/{patient_id}/{image_id}.png\", image)\n    \ntestfiles = glob.glob(\"/kaggle/input/rsna-breast-cancer-detection/test_images/*/*.dcm\")\n\n_ = Parallel(n_jobs=4)(\n    delayed(process)(fpath)\n    for fpath in tqdm(testfiles)\n)\n```\n2 - Infer (trainer + model variables exists)\n\n```python\ntest_csv = pd.read_csv(\"/kaggle/input/rsna-breast-cancer-detection/test.csv\")\nog_sub = pd.read_csv(\"/kaggle/input/rsna-breast-cancer-detection/sample_submission.csv\").head(0)\n\ntestfiles = glob.glob(f\"{SAVE_FOLDER}/*/*.png\")\n\nfor fpath in tqdm(testfiles):\n    fsplits = fpath.split('/')\n    patient_id = fsplits[-2]\n    image_id = fsplits[-1].split('.')[0]\n    prediction_id = str(test_csv.loc[test_csv.image_id == int(image_id)].prediction_id.values[0]).strip()\n    predictions = 0.00100000 # dummy, incase of fail\n\n    try:\n        datamodule = ImageClassificationData.from_files(predict_files=[fpath,],batch_size=1,)\n        predictions = trainer.predict(model, datamodule=datamodule, output=\"probabilities\")[0][0][0]\n    except:\n        # predict failed\n        pass\n\n    newdf = pd.DataFrame([[prediction_id, predictions]], columns=['prediction_id','cancer'])\n    og_sub = pd.concat([newdf, og_sub])\n    \nog_sub = og_sub.groupby('prediction_id').mean()  #dummy aggregation method\nog_sub = og_sub.sort_index()\nog_sub.to_csv(\"submission.csv\", index=True)\n\ndisplay(og_sub)\n```\n**All of these notebook versions you see in the image are succesfull in full commit. But fails during submission.**\nThis code above casuses \"Notebook threw exception\". What is wrong here? Am I missing something?\n\n**Solution**\nFew things are happening, So I got it solved by doing a few steps. Data is super large on the hidden testset.\n\n- So, convert the data to \"/kaggle/tmp/your_folder\" or \"/kaggle/tmp/output\". Otherwise the notebook will fail (this was happening to me).\n- Second issue, converstion is so slow that it takes almost 6+ hours to convert the test set (how do I know? because I tried it in 1 notebook with dummy submission). So read this [discussion](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/371033#2060484). It's 2+ hours faster.\n- Thrid, try to infer on smaller image sizes, or play with img size as per your model size and available ram. It times out if conversion is on 1024+ image size. It depends if you have a larger model then it might give \"ram overflow\" and the notebook will fail.",
    "2054138": "Maybe set index=False for your submission file. The sample sub file doesn't have an index column.",
    "2053574": "Try running your code on all the train images.\n\nMy guess is that you are not installing the pylibjpeg & python_gdcm dependencies. (cf 2nd cell of my public notebook)",
    "2060450": "Did you solve it? what was the problem?",
    "2053652": "I met same problem,too.\n\nHere is the part of code maybe cause this:\n\n\n```\nSAVE_DIR = \"/kaggle/working/\"\n# dcm to png\ndef process(line, size=512, save_folder=\"\", extension=\"png\"):\n    items = line.strip().split(',')\n    patient = items[1]\n    image = items[2]\n    laterality = items[3]\n    \n    dcm_path = os.path.join(DATA_DIR, \"test_images\", patient, image+\".dcm\")\n    dicom = pydicom.dcmread(dcm_path)\n    img = dicom.pixel_array\n\n    img = (img - img.min()) / (img.max() - img.min())\n    if dicom.PhotometricInterpretation == \"MONOCHROME1\":\n        img = 1 - img\n    img = cv2.resize(img, (size, size))\n    cv2.imwrite(save_folder + f\"{patient}_{image}_{laterality}.{extension}\", (img * 255).astype(np.uint8))\n    \n    \ntest_csv = os.path.join(DATA_DIR,\"test.csv\")\nwith open(test_csv, 'r') as f:\n    test_lines = f.readlines()[1:]\nprint(len(test_lines))\n\nSAVE_FOLDER = os.path.join(SAVE_DIR, \"png/\")\nEXTENSION = \"png\"\nos.makedirs(SAVE_FOLDER, exist_ok=True)\n\n_ = Parallel(n_jobs=4)(\n    delayed(process)(line, size=SIZE, save_folder=SAVE_FOLDER, extension=EXTENSION)\n    for line in tqdm(test_lines)\n)\n```\n\n\nI can run successfully in my notebook, but when I submit always got **Notebook Threw Exception**.\nCan anyone provide some clues? Thx."
  }
}