{
  "id": 388901,
  "title": "Clarification about the notebook submission ",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/388901",
  "author_name": "wallace",
  "post_date": "2023-02-20T04:11:44.713000",
  "votes": 0,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hello,</p>\n<p>I have some questions about this notebook submission. If anyone can help me to clarify them, I will be grateful.</p>\n<p>1) The only thing that I need to do is place my model under the loop below?</p>\n<p>for dirname, _, filenames in os.walk('/kaggle/input/'):<br>\n    for filename in filenames:<br>\n        #model here?</p>\n<p>2) And since the prediction id is given by the patient_id + laterality, I am assuming that we can also read the test.csv or train.csv, right? And the paths are also located in '/kaggle/input/' ? </p>\n<p>3) Is the train.csv also used for the submission? Or do I only need to read from test.csv?</p>\n<p>4) Is it okay to upload our pretrained model weights like a dataset?</p>\n<p>Thank you so much.</p>",
  "messages": [
    {
      "id": 2151594,
      "postDate": "2023-02-20T07:28:48.173Z",
      "content": "<p>Here is what I do:</p>\n<ol>\n<li>Train model on another kaggle notebook/ outside environment and make it into a dataset. Then I upload this dataset into '/kaggle/input' in the inference notebook.</li>\n<li>For inference only test.csv is required. The data in this is just a placeholder. On submission it is replaced with the real test data of ~8000 images.</li>\n<li>Hence prediction_id combination should be done on the basis of test.csv because of point 2.</li>\n<li>I define the model architecture first and then load the weights from the dataset path. (Model.load_wrights('path').</li>\n<li>Use this model then to make predictions. Since there are 2 or more images per Laterality, you may choose to take mean/max of prediction per Laterality.</li>\n<li>Run to check if all OK. then commit and submit the data.</li>\n</ol>",
      "rawMarkdown": "Here is what I do:\n1. Train model on another kaggle notebook/ outside environment and make it into a dataset. Then I upload this dataset into '/kaggle/input' in the inference notebook.\n2. For inference only test.csv is required. The data in this is just a placeholder. On submission it is replaced with the real test data of ~8000 images.\n3. Hence prediction_id combination should be done on the basis of test.csv because of point 2.\n4. I define the model architecture first and then load the weights from the dataset path. (Model.load_wrights('path').\n5. Use this model then to make predictions. Since there are 2 or more images per Laterality, you may choose to take mean/max of prediction per Laterality.\n6. Run to check if all OK. then commit and submit the data.",
      "votes": 1,
      "replies": [
        {
          "id": 2153109,
          "postDate": "2023-02-21T07:52:50.563Z",
          "content": "<p>Thank you very much for your answers.<br>\nSo, my notebook ran successfully but it failed to score. I am wondering if I am using the right location to read the images.  This is the path where I am looking for the images: </p>\n<p>'/kaggle/input/rsna-breast-cancer-detection/test_images'</p>\n<p>Should I use only '/kaggle/input/rsna-breast-cancer-detection/t’ or '/kaggle/input/’ ?</p>\n<p>Thank you.</p>",
          "rawMarkdown": "Thank you very much for your answers.\nSo, my notebook ran successfully but it failed to score. I am wondering if I am using the right location to read the images.  This is the path where I am looking for the images: \n\n'/kaggle/input/rsna-breast-cancer-detection/test_images'\n\nShould I use only '/kaggle/input/rsna-breast-cancer-detection/t’ or '/kaggle/input/’ ?\n\nThank you.\n",
          "replies": [
            {
              "id": 2153138,
              "postDate": "2023-02-21T08:20:30.920Z",
              "content": "<p>I'm don't know your pipeline, so just want to know if u r using the raw .DCM images itself for inference?. My approach:</p>\n<ol>\n<li>Convert the .DCM files to .png format using pydicom since DCM files are not what it actually looks to naked eye. It has to converted based on metadata. The png format images have the same folder tree structure as the DCM folders.</li>\n<li>Add an image path column to test.csv dataframe. I use these paths to feed the model.</li>\n</ol>\n<p>You should try to load and see a single image from the folder. This should give an idea of what the path should be.( In my case it's /kaggle/input/folder/patient_id/img_id)</p>",
              "rawMarkdown": "I'm don't know your pipeline, so just want to know if u r using the raw .DCM images itself for inference?. My approach:\n1. Convert the .DCM files to .png format using pydicom since DCM files are not what it actually looks to naked eye. It has to converted based on metadata. The png format images have the same folder tree structure as the DCM folders.\n2. Add an image path column to test.csv dataframe. I use these paths to feed the model.\n\nYou should try to load and see a single image from the folder. This should give an idea of what the path should be.( In my case it's /kaggle/input/folder/patient_id/img_id)"
            },
            {
              "id": 2154019,
              "postDate": "2023-02-21T19:19:14.960Z",
              "content": "<p>Thank you very much. I found the problem, division by zero.<br>\nIt is working now.</p>",
              "rawMarkdown": "Thank you very much. I found the problem, division by zero.\nIt is working now."
            }
          ]
        }
      ]
    },
    {
      "id": 2151364,
      "postDate": "2023-02-20T04:11:44.713Z",
      "content": "<p>Hello,</p>\n<p>I have some questions about this notebook submission. If anyone can help me to clarify them, I will be grateful.</p>\n<p>1) The only thing that I need to do is place my model under the loop below?</p>\n<p>for dirname, _, filenames in os.walk('/kaggle/input/'):<br>\n    for filename in filenames:<br>\n        #model here?</p>\n<p>2) And since the prediction id is given by the patient_id + laterality, I am assuming that we can also read the test.csv or train.csv, right? And the paths are also located in '/kaggle/input/' ? </p>\n<p>3) Is the train.csv also used for the submission? Or do I only need to read from test.csv?</p>\n<p>4) Is it okay to upload our pretrained model weights like a dataset?</p>\n<p>Thank you so much.</p>",
      "rawMarkdown": "Hello,\n\nI have some questions about this notebook submission. If anyone can help me to clarify them, I will be grateful.\n\n1) The only thing that I need to do is place my model under the loop below?\n\nfor dirname, _, filenames in os.walk('/kaggle/input/'):\n\tfor filename in filenames:\n\t\t#model here?\n\n2) And since the prediction id is given by the patient_id + laterality, I am assuming that we can also read the test.csv or train.csv, right? And the paths are also located in '/kaggle/input/' ? \n\n3) Is the train.csv also used for the submission? Or do I only need to read from test.csv?\n\n4) Is it okay to upload our pretrained model weights like a dataset?\n\nThank you so much.\n\n"
    }
  ],
  "comments": [
    {
      "id": 2151594,
      "author_name": "Sandy",
      "author_url": "",
      "post_date": "2023-02-20T07:28:48.173000",
      "content": "<p>Here is what I do:</p>\n<ol>\n<li>Train model on another kaggle notebook/ outside environment and make it into a dataset. Then I upload this dataset into '/kaggle/input' in the inference notebook.</li>\n<li>For inference only test.csv is required. The data in this is just a placeholder. On submission it is replaced with the real test data of ~8000 images.</li>\n<li>Hence prediction_id combination should be done on the basis of test.csv because of point 2.</li>\n<li>I define the model architecture first and then load the weights from the dataset path. (Model.load_wrights('path').</li>\n<li>Use this model then to make predictions. Since there are 2 or more images per Laterality, you may choose to take mean/max of prediction per Laterality.</li>\n<li>Run to check if all OK. then commit and submit the data.</li>\n</ol>",
      "votes": 1,
      "replies": [
        {
          "id": 2153109,
          "author_name": "wallace",
          "author_url": "",
          "post_date": "2023-02-21T07:52:50.563000",
          "content": "<p>Thank you very much for your answers.<br>\nSo, my notebook ran successfully but it failed to score. I am wondering if I am using the right location to read the images.  This is the path where I am looking for the images: </p>\n<p>'/kaggle/input/rsna-breast-cancer-detection/test_images'</p>\n<p>Should I use only '/kaggle/input/rsna-breast-cancer-detection/t’ or '/kaggle/input/’ ?</p>\n<p>Thank you.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2153138,
              "author_name": "Sandy",
              "author_url": "",
              "post_date": "2023-02-21T08:20:30.920000",
              "content": "<p>I'm don't know your pipeline, so just want to know if u r using the raw .DCM images itself for inference?. My approach:</p>\n<ol>\n<li>Convert the .DCM files to .png format using pydicom since DCM files are not what it actually looks to naked eye. It has to converted based on metadata. The png format images have the same folder tree structure as the DCM folders.</li>\n<li>Add an image path column to test.csv dataframe. I use these paths to feed the model.</li>\n</ol>\n<p>You should try to load and see a single image from the folder. This should give an idea of what the path should be.( In my case it's /kaggle/input/folder/patient_id/img_id)</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2154019,
              "author_name": "wallace",
              "author_url": "",
              "post_date": "2023-02-21T19:19:14.960000",
              "content": "<p>Thank you very much. I found the problem, division by zero.<br>\nIt is working now.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2151594": "Here is what I do:\n1. Train model on another kaggle notebook/ outside environment and make it into a dataset. Then I upload this dataset into '/kaggle/input' in the inference notebook.\n2. For inference only test.csv is required. The data in this is just a placeholder. On submission it is replaced with the real test data of ~8000 images.\n3. Hence prediction_id combination should be done on the basis of test.csv because of point 2.\n4. I define the model architecture first and then load the weights from the dataset path. (Model.load_wrights('path').\n5. Use this model then to make predictions. Since there are 2 or more images per Laterality, you may choose to take mean/max of prediction per Laterality.\n6. Run to check if all OK. then commit and submit the data.",
    "2151364": "Hello,\n\nI have some questions about this notebook submission. If anyone can help me to clarify them, I will be grateful.\n\n1) The only thing that I need to do is place my model under the loop below?\n\nfor dirname, _, filenames in os.walk('/kaggle/input/'):\n\tfor filename in filenames:\n\t\t#model here?\n\n2) And since the prediction id is given by the patient_id + laterality, I am assuming that we can also read the test.csv or train.csv, right? And the paths are also located in '/kaggle/input/' ? \n\n3) Is the train.csv also used for the submission? Or do I only need to read from test.csv?\n\n4) Is it okay to upload our pretrained model weights like a dataset?\n\nThank you so much.\n\n"
  }
}