{
  "id": 373710,
  "title": "Notebook runs ok, but submission fails",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/373710",
  "author_name": "Ben Lebovitz",
  "post_date": "2022-12-22T20:17:06.836000",
  "votes": 2,
  "comment_count": 15,
  "views": 0,
  "content": "<p><a href=\"https://www.kaggle.com/code/beezus666/dicom-to-png-to-predict\" target=\"_blank\">https://www.kaggle.com/code/beezus666/dicom-to-png-to-predict</a><br>\nHi, I just shared the above notebook, based on the work shared by <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> and <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a>. Right now I'm just trying to set up a simple pipeline that does some image manipulation and makes predictions via a pre-trained model imported from another notebook.</p>\n<p><strong>But, I'm running into the same issue that others have faced in that I get a failure on submission…</strong></p>\n<ul>\n<li>When I click submit to the competition, the notebook runs without error, but then in submitting fails</li>\n<li>I pasted my CSV output below for the test set that is made available. I'm reasonably sure this is ok (no duplicates, commas where expected, etc)</li>\n<li>I'm a bit concerned that the notebook runs so quickly. I didn't even turn on the GPU and it completes all its inference in under 5 minutes. I know inference is faster than train, but I was surprised at this speed.</li>\n</ul>\n<p>My CSV output:<br>\nprediction_id,cancer<br>\n10008_L,0.019642680883407593<br>\n10008_R,0.019744178280234337</p>\n<p>Thanks in advance for any help</p>",
  "messages": [
    {
      "id": 2073332,
      "postDate": "2022-12-22T21:57:30.893Z",
      "content": "<p>I know where is the problem :)</p>\n<p>Delete png_file_py and roi_cropped directories from working (!). Scoring system is looking for submission file … and it should be the first one file in working directory.  </p>\n<pre><code>resulting_df.to_csv('submission.csv', index=True)\n!rm -r png_file_py \n!rm -r roi_cropped\n</code></pre>",
      "rawMarkdown": "I know where is the problem :)\n\nDelete png_file_py and roi_cropped directories from working (!). Scoring system is looking for submission file ... and it should be the first one file in working directory.  \n\n\n```\nresulting_df.to_csv('submission.csv', index=True)\n!rm -r png_file_py \n!rm -r roi_cropped\n```",
      "votes": 5,
      "replies": [
        {
          "id": 2073362,
          "postDate": "2022-12-23T00:47:48.067Z",
          "content": "<p>oh… that's it?? heh… I'll try that tomorrow I've used up all 5 of my submissions for today on errors… I was going insane reading about the finer points of pandas index and csv conversion…</p>\n<p>Thanks so much for taking a look at my notebook!</p>",
          "rawMarkdown": "oh... that's it?? heh... I'll try that tomorrow I've used up all 5 of my submissions for today on errors... I was going insane reading about the finer points of pandas index and csv conversion...\n\nThanks so much for taking a look at my notebook!",
          "votes": 1,
          "replies": [
            {
              "id": 2073731,
              "postDate": "2022-12-23T10:49:43.600Z",
              "content": "<p>It could be the reason. If not let me know I will help.</p>",
              "rawMarkdown": "It could be the reason. If not let me know I will help.",
              "votes": 2
            },
            {
              "id": 2073841,
              "postDate": "2022-12-23T13:14:37.167Z",
              "content": "<p>Yeah… no joy… simply deleting the files produces the same results. </p>\n<p>I'll spend the afternoon (US time) rewriting the notebook to use the train set and see where it fails.</p>\n<p>Any help would be greatly appreciated!</p>",
              "rawMarkdown": "Yeah... no joy... simply deleting the files produces the same results. \n\nI'll spend the afternoon (US time) rewriting the notebook to use the train set and see where it fails.\n\nAny help would be greatly appreciated!"
            },
            {
              "id": 2073957,
              "postDate": "2022-12-23T15:31:36.277Z",
              "content": "<p>When you finish rewriting code please share it (set as public - we shoud not break competition rules). I will look into code. </p>",
              "rawMarkdown": "When you finish rewriting code please share it (set as public - we shoud not break competition rules). I will look into code. "
            },
            {
              "id": 2073964,
              "postDate": "2022-12-23T15:41:19.713Z",
              "content": "<pre><code>for result in results:\n        images.append(cv2.rectangle(frame, (int(result['xmin']), int(result['ymin'])), (int(result['xmax']), int(result['ymax'])), (255,0,0), 4))\n\n    ##########\n    # new... create croopped images from test folder save to new folder\n    ##########\n    b_box_values = [v for i, (k, v) in enumerate(results[0].items()) if i &lt; 4]\n    temp_file = cv2.imread(img_file)\n    x1, y1, x2, y2 = map(int, b_box_values)\n    cropped_file = temp_file[y1:y2, x1:x2]\n    cv2.imwrite(f'/kaggle/working/roi_cropped/{os.path.basename(img_file)}', cropped_file)\n</code></pre>\n<p>This could be another source of problem. What when yolo do not provide any results? <br>\nWe have three cases:</p>\n<ol>\n<li>Yolo do not provide any ROI - I know that this is possible (checked on train dataset and … I have such cases)</li>\n<li>Yolo provide one result</li>\n<li>Yolo provide many results</li>\n</ol>\n<p>My workaround (very fast development so forgive me):</p>\n<pre><code>results = model(image)\nbboxes = results.xywh\n\nif len(bboxes[0]) &gt; 0:\n\n            bbox = bboxes[0][0]  ## or take with highest score\n\n            w = bbox[2]\n            h = bbox[3]\n\n            x_min = int(bbox[0] - (w // 2))\n            y_min = int(bbox[1] - (h // 2))\n\n            x_max = int(bbox[0] + (w // 2))\n            y_max = int(bbox[1] + (h // 2))\n\n            tmp_im = image[y_min:y_max, x_min:x_max]\n        else:\n            tmp_im = image\n</code></pre>",
              "rawMarkdown": "```\nfor result in results:\n        images.append(cv2.rectangle(frame, (int(result['xmin']), int(result['ymin'])), (int(result['xmax']), int(result['ymax'])), (255,0,0), 4))\n\n    ##########\n    # new... create croopped images from test folder save to new folder\n    ##########\n    b_box_values = [v for i, (k, v) in enumerate(results[0].items()) if i < 4]\n    temp_file = cv2.imread(img_file)\n    x1, y1, x2, y2 = map(int, b_box_values)\n    cropped_file = temp_file[y1:y2, x1:x2]\n    cv2.imwrite(f'/kaggle/working/roi_cropped/{os.path.basename(img_file)}', cropped_file)\n```\n\nThis could be another source of problem. What when yolo do not provide any results? \nWe have three cases:\n1. Yolo do not provide any ROI - I know that this is possible (checked on train dataset and ... I have such cases)\n2. Yolo provide one result\n3. Yolo provide many results\n\nMy workaround (very fast development so forgive me):\n\n```\nresults = model(image)\nbboxes = results.xywh\n\nif len(bboxes[0]) > 0:\n\n            bbox = bboxes[0][0]  ## or take with highest score\n\n            w = bbox[2]\n            h = bbox[3]\n\n            x_min = int(bbox[0] - (w // 2))\n            y_min = int(bbox[1] - (h // 2))\n\n            x_max = int(bbox[0] + (w // 2))\n            y_max = int(bbox[1] + (h // 2))\n\n            tmp_im = image[y_min:y_max, x_min:x_max]\n        else:\n            tmp_im = image\n```",
              "votes": 1
            },
            {
              "id": 2074050,
              "postDate": "2022-12-23T17:12:59.920Z",
              "content": "<p>I didn't even get down to that section… I changed it over to look at the training images, and turned on the GPU, and now it errors out on the code below because it can't decode DICOM pixel data. </p>\n<p>FYI… from the original version I could never get <code>import dicomsdl</code> to work… I couldn't find where that package was being used and as per the OP it worked ok on the 4 test images so I just went happily on my way</p>\n<p>The code that errors out:</p>\n<pre><code>Parallel(n_jobs=)(\n    delayed(process)(f, size = , save_folder = image_dir_pydicom, dicom_process = )\n    \n     f  dicom_images[:IMAGES_TO_PROCESS]\n)\n</code></pre>\n<p>This calls the process function that does the extract and I get a bunch of errors with joblib and then the final error is this: <br>\nThe following handlers are available to decode the pixel data however they are missing required dependencies: GDCM (req. GDCM), pylibjpeg (req. )</p>\n<p>Anyway, I just published a new version of the notebook with the errors. I'm going to spend the next couple hours hacking on it… help as always is greatly appreciated.</p>",
              "rawMarkdown": "I didn't even get down to that section... I changed it over to look at the training images, and turned on the GPU, and now it errors out on the code below because it can't decode DICOM pixel data. \n\nFYI... from the original version I could never get `import dicomsdl` to work... I couldn't find where that package was being used and as per the OP it worked ok on the 4 test images so I just went happily on my way\n\nThe code that errors out:\n```python\n        \nParallel(n_jobs=4)(\n    delayed(process)(f, size = 512, save_folder = image_dir_pydicom, dicom_process = True)\n    #for f in train_images[:IMAGES_TO_PROCESS]\n    for f in dicom_images[:IMAGES_TO_PROCESS]\n)\n\n\n```This calls the process function that does the extract and I get a bunch of errors with joblib and then the final error is this: \nThe following handlers are available to decode the pixel data however they are missing required dependencies: GDCM (req. GDCM), pylibjpeg (req. )\n\n\nAnyway, I just published a new version of the notebook with the errors. I'm going to spend the next couple hours hacking on it... help as always is greatly appreciated."
            },
            {
              "id": 2074292,
              "postDate": "2022-12-24T00:27:18.587Z",
              "content": "<p>Ok… I'm an idiot… a bunch of the code I copied limited it to 500 images it was processing. Sorry for wasting your time! </p>\n<p>Anyway, the thing is taking awhile to run (about 2 hour so far…) so hopefully it'll actually run and submit correctly and I'll post a new version tomorrow. I think the pipeline I put together is ok and might be useful to others.</p>\n<p>One thing though… I am getting some errors randomly on a ~0.2% of images in processing ROI. I put a little if statement around it to simply use the original image if I can't get the bounding box coordinates see the code below.</p>\n<pre><code> results  (results[], )  (results[]) &gt;= :\n        b_box_values = [v  i, (k, v)  (results[].items())  i &lt; ]    \n    :\n        \n        ()        \n        height, width, channels = frame.shape\n        b_box_values = [, , height, width]   \n</code></pre>",
              "rawMarkdown": "Ok... I'm an idiot... a bunch of the code I copied limited it to 500 images it was processing. Sorry for wasting your time! \n\nAnyway, the thing is taking awhile to run (about 2 hour so far...) so hopefully it'll actually run and submit correctly and I'll post a new version tomorrow. I think the pipeline I put together is ok and might be useful to others.\n\nOne thing though... I am getting some errors randomly on a ~0.2% of images in processing ROI. I put a little if statement around it to simply use the original image if I can't get the bounding box coordinates see the code below.\n\n```python\nif results and isinstance(results[0], dict) and len(results[0]) >= 4:\n        b_box_values = [v for i, (k, v) in enumerate(results[0].items()) if i < 4]    \n    else:\n        # handle the case where the results list is empty or the first element is not a dictionary with at least 4 items\n        print(f'error on {img_file}')        \n        height, width, channels = frame.shape\n        b_box_values = [0, 0, height, width]   \n\n```"
            },
            {
              "id": 2074451,
              "postDate": "2022-12-24T07:33:37.713Z",
              "content": "<p>No problem. Nice to help and talk about solution. </p>",
              "rawMarkdown": "No problem. Nice to help and talk about solution. "
            }
          ]
        }
      ]
    },
    {
      "id": 2073325,
      "postDate": "2022-12-22T21:47:52.570Z",
      "content": "<p>What kind of error do you get? (see in submission summary).</p>",
      "rawMarkdown": "What kind of error do you get? (see in submission summary).",
      "votes": 1,
      "replies": [
        {
          "id": 2073361,
          "postDate": "2022-12-23T00:45:18.480Z",
          "content": "<p>Oh, thanks… I didn't know to look for that. I guess it threw an exception as per below… I'll rewrite it to use the whole training set and likely uncover the error.</p>\n<p>The error: <br>\nNotebook Threw Exception</p>\n<p>Your notebook hit an unhandled error while rerunning your code. Note that the hidden dataset can be larger/smaller/different than the public dataset See more debugging tips</p>",
          "rawMarkdown": "Oh, thanks... I didn't know to look for that. I guess it threw an exception as per below... I'll rewrite it to use the whole training set and likely uncover the error.\n\nThe error: \nNotebook Threw Exception\n\nYour notebook hit an unhandled error while rerunning your code. Note that the hidden dataset can be larger/smaller/different than the public dataset See more debugging tips"
        }
      ]
    },
    {
      "id": 2073285,
      "postDate": "2022-12-22T20:17:06.837Z",
      "content": "<p><a href=\"https://www.kaggle.com/code/beezus666/dicom-to-png-to-predict\" target=\"_blank\">https://www.kaggle.com/code/beezus666/dicom-to-png-to-predict</a><br>\nHi, I just shared the above notebook, based on the work shared by <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> and <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a>. Right now I'm just trying to set up a simple pipeline that does some image manipulation and makes predictions via a pre-trained model imported from another notebook.</p>\n<p><strong>But, I'm running into the same issue that others have faced in that I get a failure on submission…</strong></p>\n<ul>\n<li>When I click submit to the competition, the notebook runs without error, but then in submitting fails</li>\n<li>I pasted my CSV output below for the test set that is made available. I'm reasonably sure this is ok (no duplicates, commas where expected, etc)</li>\n<li>I'm a bit concerned that the notebook runs so quickly. I didn't even turn on the GPU and it completes all its inference in under 5 minutes. I know inference is faster than train, but I was surprised at this speed.</li>\n</ul>\n<p>My CSV output:<br>\nprediction_id,cancer<br>\n10008_L,0.019642680883407593<br>\n10008_R,0.019744178280234337</p>\n<p>Thanks in advance for any help</p>",
      "rawMarkdown": "https://www.kaggle.com/code/beezus666/dicom-to-png-to-predict\nHi, I just shared the above notebook, based on the work shared by @remekkinas and @radek1. Right now I'm just trying to set up a simple pipeline that does some image manipulation and makes predictions via a pre-trained model imported from another notebook.\n\n**But, I'm running into the same issue that others have faced in that I get a failure on submission...**\n- When I click submit to the competition, the notebook runs without error, but then in submitting fails\n- I pasted my CSV output below for the test set that is made available. I'm reasonably sure this is ok (no duplicates, commas where expected, etc)\n- I'm a bit concerned that the notebook runs so quickly. I didn't even turn on the GPU and it completes all its inference in under 5 minutes. I know inference is faster than train, but I was surprised at this speed.\n\nMy CSV output:\nprediction_id,cancer\n10008_L,0.019642680883407593\n10008_R,0.019744178280234337\n\nThanks in advance for any help\n",
      "votes": 2
    },
    {
      "id": 2073312,
      "postDate": "2022-12-22T21:09:40.740Z",
      "content": "<p>One reason could be that your model prediction for some hidden test data is NaN. </p>",
      "rawMarkdown": "One reason could be that your model prediction for some hidden test data is NaN. ",
      "replies": [
        {
          "id": 2073319,
          "postDate": "2022-12-22T21:21:14.787Z",
          "content": "<p>Hmm… but it's just pulling in the images. Also it runs to completion, but fails only on the scoring of it.</p>",
          "rawMarkdown": "Hmm... but it's just pulling in the images. Also it runs to completion, but fails only on the scoring of it."
        },
        {
          "id": 2073321,
          "postDate": "2022-12-22T21:34:49.180Z",
          "content": "<p>Did you run into hidden NaNs?</p>",
          "rawMarkdown": "Did you run into hidden NaNs?",
          "replies": [
            {
              "id": 2073326,
              "postDate": "2022-12-22T21:50:57.350Z",
              "content": "<p>If you use AMP in your training, it is possible to get NaN output at inference. You can implement <code>try except</code> to catch NaN predictions and replace them with, for example 0.5.</p>",
              "rawMarkdown": "If you use AMP in your training, it is possible to get NaN output at inference. You can implement `try except` to catch NaN predictions and replace them with, for example 0.5."
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2073332,
      "author_name": "Remek Kinas",
      "author_url": "",
      "post_date": "2022-12-22T21:57:30.893000",
      "content": "<p>I know where is the problem :)</p>\n<p>Delete png_file_py and roi_cropped directories from working (!). Scoring system is looking for submission file … and it should be the first one file in working directory.  </p>\n<pre><code>resulting_df.to_csv('submission.csv', index=True)\n!rm -r png_file_py \n!rm -r roi_cropped\n</code></pre>",
      "votes": 5,
      "replies": [
        {
          "id": 2073362,
          "author_name": "Ben Lebovitz",
          "author_url": "",
          "post_date": "2022-12-23T00:47:48.067000",
          "content": "<p>oh… that's it?? heh… I'll try that tomorrow I've used up all 5 of my submissions for today on errors… I was going insane reading about the finer points of pandas index and csv conversion…</p>\n<p>Thanks so much for taking a look at my notebook!</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2073731,
              "author_name": "Remek Kinas",
              "author_url": "",
              "post_date": "2022-12-23T10:49:43.600000",
              "content": "<p>It could be the reason. If not let me know I will help.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2073841,
              "author_name": "Ben Lebovitz",
              "author_url": "",
              "post_date": "2022-12-23T13:14:37.167000",
              "content": "<p>Yeah… no joy… simply deleting the files produces the same results. </p>\n<p>I'll spend the afternoon (US time) rewriting the notebook to use the train set and see where it fails.</p>\n<p>Any help would be greatly appreciated!</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2073957,
              "author_name": "Remek Kinas",
              "author_url": "",
              "post_date": "2022-12-23T15:31:36.277000",
              "content": "<p>When you finish rewriting code please share it (set as public - we shoud not break competition rules). I will look into code. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2073964,
              "author_name": "Remek Kinas",
              "author_url": "",
              "post_date": "2022-12-23T15:41:19.713000",
              "content": "<pre><code>for result in results:\n        images.append(cv2.rectangle(frame, (int(result['xmin']), int(result['ymin'])), (int(result['xmax']), int(result['ymax'])), (255,0,0), 4))\n\n    ##########\n    # new... create croopped images from test folder save to new folder\n    ##########\n    b_box_values = [v for i, (k, v) in enumerate(results[0].items()) if i &lt; 4]\n    temp_file = cv2.imread(img_file)\n    x1, y1, x2, y2 = map(int, b_box_values)\n    cropped_file = temp_file[y1:y2, x1:x2]\n    cv2.imwrite(f'/kaggle/working/roi_cropped/{os.path.basename(img_file)}', cropped_file)\n</code></pre>\n<p>This could be another source of problem. What when yolo do not provide any results? <br>\nWe have three cases:</p>\n<ol>\n<li>Yolo do not provide any ROI - I know that this is possible (checked on train dataset and … I have such cases)</li>\n<li>Yolo provide one result</li>\n<li>Yolo provide many results</li>\n</ol>\n<p>My workaround (very fast development so forgive me):</p>\n<pre><code>results = model(image)\nbboxes = results.xywh\n\nif len(bboxes[0]) &gt; 0:\n\n            bbox = bboxes[0][0]  ## or take with highest score\n\n            w = bbox[2]\n            h = bbox[3]\n\n            x_min = int(bbox[0] - (w // 2))\n            y_min = int(bbox[1] - (h // 2))\n\n            x_max = int(bbox[0] + (w // 2))\n            y_max = int(bbox[1] + (h // 2))\n\n            tmp_im = image[y_min:y_max, x_min:x_max]\n        else:\n            tmp_im = image\n</code></pre>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2074050,
              "author_name": "Ben Lebovitz",
              "author_url": "",
              "post_date": "2022-12-23T17:12:59.920000",
              "content": "<p>I didn't even get down to that section… I changed it over to look at the training images, and turned on the GPU, and now it errors out on the code below because it can't decode DICOM pixel data. </p>\n<p>FYI… from the original version I could never get <code>import dicomsdl</code> to work… I couldn't find where that package was being used and as per the OP it worked ok on the 4 test images so I just went happily on my way</p>\n<p>The code that errors out:</p>\n<pre><code>Parallel(n_jobs=)(\n    delayed(process)(f, size = , save_folder = image_dir_pydicom, dicom_process = )\n    \n     f  dicom_images[:IMAGES_TO_PROCESS]\n)\n</code></pre>\n<p>This calls the process function that does the extract and I get a bunch of errors with joblib and then the final error is this: <br>\nThe following handlers are available to decode the pixel data however they are missing required dependencies: GDCM (req. GDCM), pylibjpeg (req. )</p>\n<p>Anyway, I just published a new version of the notebook with the errors. I'm going to spend the next couple hours hacking on it… help as always is greatly appreciated.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2074292,
              "author_name": "Ben Lebovitz",
              "author_url": "",
              "post_date": "2022-12-24T00:27:18.587000",
              "content": "<p>Ok… I'm an idiot… a bunch of the code I copied limited it to 500 images it was processing. Sorry for wasting your time! </p>\n<p>Anyway, the thing is taking awhile to run (about 2 hour so far…) so hopefully it'll actually run and submit correctly and I'll post a new version tomorrow. I think the pipeline I put together is ok and might be useful to others.</p>\n<p>One thing though… I am getting some errors randomly on a ~0.2% of images in processing ROI. I put a little if statement around it to simply use the original image if I can't get the bounding box coordinates see the code below.</p>\n<pre><code> results  (results[], )  (results[]) &gt;= :\n        b_box_values = [v  i, (k, v)  (results[].items())  i &lt; ]    \n    :\n        \n        ()        \n        height, width, channels = frame.shape\n        b_box_values = [, , height, width]   \n</code></pre>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2074451,
              "author_name": "Remek Kinas",
              "author_url": "",
              "post_date": "2022-12-24T07:33:37.713000",
              "content": "<p>No problem. Nice to help and talk about solution. </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2073325,
      "author_name": "Remek Kinas",
      "author_url": "",
      "post_date": "2022-12-22T21:47:52.570000",
      "content": "<p>What kind of error do you get? (see in submission summary).</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2073361,
          "author_name": "Ben Lebovitz",
          "author_url": "",
          "post_date": "2022-12-23T00:45:18.480000",
          "content": "<p>Oh, thanks… I didn't know to look for that. I guess it threw an exception as per below… I'll rewrite it to use the whole training set and likely uncover the error.</p>\n<p>The error: <br>\nNotebook Threw Exception</p>\n<p>Your notebook hit an unhandled error while rerunning your code. Note that the hidden dataset can be larger/smaller/different than the public dataset See more debugging tips</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2073312,
      "author_name": "Rasoul Mojtahedzadeh",
      "author_url": "",
      "post_date": "2022-12-22T21:09:40.740000",
      "content": "<p>One reason could be that your model prediction for some hidden test data is NaN. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 2073319,
          "author_name": "Ben Lebovitz",
          "author_url": "",
          "post_date": "2022-12-22T21:21:14.787000",
          "content": "<p>Hmm… but it's just pulling in the images. Also it runs to completion, but fails only on the scoring of it.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2073321,
          "author_name": "Ben Lebovitz",
          "author_url": "",
          "post_date": "2022-12-22T21:34:49.180000",
          "content": "<p>Did you run into hidden NaNs?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2073326,
              "author_name": "Rasoul Mojtahedzadeh",
              "author_url": "",
              "post_date": "2022-12-22T21:50:57.350000",
              "content": "<p>If you use AMP in your training, it is possible to get NaN output at inference. You can implement <code>try except</code> to catch NaN predictions and replace them with, for example 0.5.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2073332": "I know where is the problem :)\n\nDelete png_file_py and roi_cropped directories from working (!). Scoring system is looking for submission file ... and it should be the first one file in working directory.  \n\n\n```\nresulting_df.to_csv('submission.csv', index=True)\n!rm -r png_file_py \n!rm -r roi_cropped\n```",
    "2073325": "What kind of error do you get? (see in submission summary).",
    "2073285": "https://www.kaggle.com/code/beezus666/dicom-to-png-to-predict\nHi, I just shared the above notebook, based on the work shared by @remekkinas and @radek1. Right now I'm just trying to set up a simple pipeline that does some image manipulation and makes predictions via a pre-trained model imported from another notebook.\n\n**But, I'm running into the same issue that others have faced in that I get a failure on submission...**\n- When I click submit to the competition, the notebook runs without error, but then in submitting fails\n- I pasted my CSV output below for the test set that is made available. I'm reasonably sure this is ok (no duplicates, commas where expected, etc)\n- I'm a bit concerned that the notebook runs so quickly. I didn't even turn on the GPU and it completes all its inference in under 5 minutes. I know inference is faster than train, but I was surprised at this speed.\n\nMy CSV output:\nprediction_id,cancer\n10008_L,0.019642680883407593\n10008_R,0.019744178280234337\n\nThanks in advance for any help\n",
    "2073312": "One reason could be that your model prediction for some hidden test data is NaN. "
  }
}