{
  "id": 427219,
  "title": "New to Kaggle or Machine Learning? Check this out ~",
  "url": "/competitions/rsna-2023-abdominal-trauma-detection/discussion/427219",
  "author_name": "Maggie",
  "post_date": "2023-07-26T23:54:51.706000",
  "votes": 16,
  "comment_count": 11,
  "views": 0,
  "content": "<p>New to machine learning and data science? No question is too basic or too simple. Feel free to start your own thread, or use this thread as a place to post any first-timer clarifying questions for the Kaggle community to help you with!</p>\n<p>If you would consider yourself a beginner but don't know where to get started, let other Kagglers help you take your first steps here!</p>\n<p>New to Kaggle? Take a look at a few videos Dr. Rachael Tatman has put together to learn a bit more about <a href=\"https://www.youtube.com/watch?v=aIus8si_Et0\" target=\"_blank\">site etiquette</a>, or <a href=\"https://www.youtube.com/watch?v=sEJHyuWKd-s\" target=\"_blank\">Kaggle lingo</a>.</p>\n<p>Remember: Kaggle is for everyone. Whether you're teaming up or sharing tips in the competition forum, we expect everyone to follow our <a href=\"https://www.kaggle.com/community-guidelines\" target=\"_blank\">Kaggle community guidelines</a>.</p>\n<p>A tip on sharing content - Kaggle is a collaborative community, whereby sharing techniques, starter notebooks, and ideas in the discussion forums are highly encouraged throughout the competition. However, as the competition draws closer to the final deadline it's customary to keep high-scoring notebooks withheld until after the competition has concluded. This maintains the spirit of the competition, while also allowing individuals to submit their own creative work without jeopardy of a higher-scoring notebook being available for an automatic higher rank (through copy/submit). We disable publishing of public notebooks within the final week of the competition, but encourage you to use your best judgment prior to that deadline.</p>\n<p>Happy Modeling!</p>",
  "messages": [
    {
      "id": 2360705,
      "postDate": "2023-07-26T23:54:51.707Z",
      "content": "<p>New to machine learning and data science? No question is too basic or too simple. Feel free to start your own thread, or use this thread as a place to post any first-timer clarifying questions for the Kaggle community to help you with!</p>\n<p>If you would consider yourself a beginner but don't know where to get started, let other Kagglers help you take your first steps here!</p>\n<p>New to Kaggle? Take a look at a few videos Dr. Rachael Tatman has put together to learn a bit more about <a href=\"https://www.youtube.com/watch?v=aIus8si_Et0\" target=\"_blank\">site etiquette</a>, or <a href=\"https://www.youtube.com/watch?v=sEJHyuWKd-s\" target=\"_blank\">Kaggle lingo</a>.</p>\n<p>Remember: Kaggle is for everyone. Whether you're teaming up or sharing tips in the competition forum, we expect everyone to follow our <a href=\"https://www.kaggle.com/community-guidelines\" target=\"_blank\">Kaggle community guidelines</a>.</p>\n<p>A tip on sharing content - Kaggle is a collaborative community, whereby sharing techniques, starter notebooks, and ideas in the discussion forums are highly encouraged throughout the competition. However, as the competition draws closer to the final deadline it's customary to keep high-scoring notebooks withheld until after the competition has concluded. This maintains the spirit of the competition, while also allowing individuals to submit their own creative work without jeopardy of a higher-scoring notebook being available for an automatic higher rank (through copy/submit). We disable publishing of public notebooks within the final week of the competition, but encourage you to use your best judgment prior to that deadline.</p>\n<p>Happy Modeling!</p>",
      "rawMarkdown": "New to machine learning and data science? No question is too basic or too simple. Feel free to start your own thread, or use this thread as a place to post any first-timer clarifying questions for the Kaggle community to help you with!\n\nIf you would consider yourself a beginner but don't know where to get started, let other Kagglers help you take your first steps here!\n\nNew to Kaggle? Take a look at a few videos Dr. Rachael Tatman has put together to learn a bit more about [site etiquette](https://www.youtube.com/watch?v=aIus8si_Et0), or [Kaggle lingo](https://www.youtube.com/watch?v=sEJHyuWKd-s).\n\nRemember: Kaggle is for everyone. Whether you're teaming up or sharing tips in the competition forum, we expect everyone to follow our [Kaggle community guidelines](https://www.kaggle.com/community-guidelines).\n\nA tip on sharing content - Kaggle is a collaborative community, whereby sharing techniques, starter notebooks, and ideas in the discussion forums are highly encouraged throughout the competition. However, as the competition draws closer to the final deadline it's customary to keep high-scoring notebooks withheld until after the competition has concluded. This maintains the spirit of the competition, while also allowing individuals to submit their own creative work without jeopardy of a higher-scoring notebook being available for an automatic higher rank (through copy/submit). We disable publishing of public notebooks within the final week of the competition, but encourage you to use your best judgment prior to that deadline.\n\n  \n\nHappy Modeling!\n",
      "votes": 16
    },
    {
      "id": 2440969,
      "postDate": "2023-09-15T21:24:37.680Z",
      "content": "<p>Hi, very basic question: I'm trying to run the example training and inference notebooks with no modifications, but the inference notebook can't load the model saved by the training notebook. </p>\n<p>The training notebook saves the model file <code>rsna-atd.keras</code> in my Outputs workspace, but then when I start up the inference notebook I can't see it. I tried clicking \"add data\" but it says there are no data sources found. What am I missing here?</p>",
      "rawMarkdown": "Hi, very basic question: I'm trying to run the example training and inference notebooks with no modifications, but the inference notebook can't load the model saved by the training notebook. \n\nThe training notebook saves the model file `rsna-atd.keras` in my Outputs workspace, but then when I start up the inference notebook I can't see it. I tried clicking \"add data\" but it says there are no data sources found. What am I missing here?",
      "votes": 1
    },
    {
      "id": 2421628,
      "postDate": "2023-09-03T13:10:42.297Z",
      "content": "<p>How can I use the images data without downloading it? </p>",
      "rawMarkdown": "How can I use the images data without downloading it? ",
      "votes": 1,
      "replies": [
        {
          "id": 2422181,
          "postDate": "2023-09-03T18:49:31.073Z",
          "content": "<p>Then you are at kaggle notebook on the right part of screen you can see folder \"Data\". You can click on \"+Add data\" and there choose \"rsna-2023-abdominal-trauma-detection\".</p>\n<p>Then you can use files, for example link to a file can be like \"/kaggle/input/rsna-2023-abdominal-trauma-detection/train_images/122/21217/100.dcm\"</p>\n<p>After that you can preprocess it, for some tip you can check:</p>\n<ol>\n<li><a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/alenic/dataset-size-reduction-400gb-to-7-5gb?scriptVersionId=140444781</a></li>\n<li><a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/jhoward/creating-a-metadata-dataframe-fastai/notebook?scriptVersionId=22141307</a></li>\n<li><a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/awsaf49/rsna-atd-512x512-png-v2-data/notebook</a></li>\n</ol>",
          "rawMarkdown": "Then you are at kaggle notebook on the right part of screen you can see folder \"Data\". You can click on \"+Add data\" and there choose \"rsna-2023-abdominal-trauma-detection\".\n\nThen you can use files, for example link to a file can be like \"/kaggle/input/rsna-2023-abdominal-trauma-detection/train_images/122/21217/100.dcm\"\n\nAfter that you can preprocess it, for some tip you can check:\n1. [https://www.kaggle.com/code/alenic/dataset-size-reduction-400gb-to-7-5gb?scriptVersionId=140444781](url)\n2. [https://www.kaggle.com/code/jhoward/creating-a-metadata-dataframe-fastai/notebook?scriptVersionId=22141307](url)\n3. [https://www.kaggle.com/code/awsaf49/rsna-atd-512x512-png-v2-data/notebook](url)",
          "votes": 2
        }
      ]
    },
    {
      "id": 2419519,
      "postDate": "2023-09-02T04:41:48.013Z",
      "content": "<p>Hi everyone. I have a noobie question.<br>\nI think I read a lot of stuff about my question already, but there is still some mess in my head.</p>\n<p>How to deal with a not consistent amount of input data?<br>\nWe have CT scans per patient, but there can be 512x512x50, 512x512x300, 512x512x800 etc. (by the last position I mean number of CT scans in one folder)</p>\n<p>So, because of that:</p>\n<ol>\n<li>We should just \"resize\" that to some standard size and then give it to NN?<br>\n1a. If it is right - so how exactly we do that? Somewhere I saw using \"scipy.ndimage.zoom\" this function maby do what I need. Or we just simply take a fixed size sample of CT scans (every 2/5/10/.. image depend on input size) and vary it for a more accurate score and appropriate computation time? (if this is right I don't understand how to deal with the situation, then the input size will be less than the fixed size sample - just copy the \"average picture\" to gain a minimum sample size?)</li>\n<li>Also, in the real test set, I assume it can be from 1 to, for example, 1k the amount of CT scans, right?</li>\n</ol>\n<p>If you have some links to read about, it would be great :)</p>",
      "rawMarkdown": "Hi everyone. I have a noobie question.\nI think I read a lot of stuff about my question already, but there is still some mess in my head.\n\nHow to deal with a not consistent amount of input data?\nWe have CT scans per patient, but there can be 512x512x50, 512x512x300, 512x512x800 etc. (by the last position I mean number of CT scans in one folder)\n\nSo, because of that:\n1. We should just \"resize\" that to some standard size and then give it to NN?\n 1a. If it is right - so how exactly we do that? Somewhere I saw using \"scipy.ndimage.zoom\" this function maby do what I need. Or we just simply take a fixed size sample of CT scans (every 2/5/10/.. image depend on input size) and vary it for a more accurate score and appropriate computation time? (if this is right I don't understand how to deal with the situation, then the input size will be less than the fixed size sample - just copy the \"average picture\" to gain a minimum sample size?)\n2. Also, in the real test set, I assume it can be from 1 to, for example, 1k the amount of CT scans, right?\n\nIf you have some links to read about, it would be great :)",
      "votes": 2
    },
    {
      "id": 2463479,
      "postDate": "2023-10-01T11:08:22.113Z",
      "content": "<p>good，it is very helpful</p>",
      "rawMarkdown": "good，it is very helpful"
    },
    {
      "id": 2441965,
      "postDate": "2023-09-16T15:39:53.253Z",
      "content": "<p>Please give some link to higher-scoring public notebook?</p>",
      "rawMarkdown": " Please give some link to higher-scoring public notebook?"
    },
    {
      "id": 2365162,
      "postDate": "2023-07-30T03:26:56.927Z",
      "content": "<p>Thank you, the information would be helpful.</p>",
      "rawMarkdown": "Thank you, the information would be helpful."
    },
    {
      "id": 2364921,
      "postDate": "2023-07-29T18:14:04.063Z",
      "content": "<p><strong>Can not read file in segmentation folder</strong></p>\n<p>According to what i read in the data section, in the segmentation directory, the files are formatted as nifti. But I use nibabel library to read:<br>\nimport nibabel as nib<br>\nimg = nib.load('/kaggle/input/rsna-2023-abdominal-trauma-detection/segmentations/10000')<br>\nAn error occurred: Cannot work out file type of \"/kaggle/input/rsna-2023-abdominal-trauma-detection/segmentations/10000\".<br>\nMaybe it's not in .nii format. So what format are those files and how to read them?<br>\nThen I use this code to read binary file (arcording to i see in information of these file when i downloaded):<br>\nbinary_file_path = '/kaggle/input/rsna-2023-abdominal-trauma-detection/segmentations/21057'<br>\ndt = np.dtype('B')<br>\nLoad the binary data into a numpy array:<br>\nwith open(binary_file_path, \"rb\") as f:<br>\nnumpy_data = np.fromfile(f, dt)<br>\ni received numpy array with shape (1071645024,). But it is not seemly to be suitable with image size (512x512)?</p>",
      "rawMarkdown": "**Can not read file in segmentation folder**\n\nAccording to what i read in the data section, in the segmentation directory, the files are formatted as nifti. But I use nibabel library to read:\nimport nibabel as nib\nimg = nib.load('/kaggle/input/rsna-2023-abdominal-trauma-detection/segmentations/10000')\nAn error occurred: Cannot work out file type of \"/kaggle/input/rsna-2023-abdominal-trauma-detection/segmentations/10000\".\nMaybe it's not in .nii format. So what format are those files and how to read them?\nThen I use this code to read binary file (arcording to i see in information of these file when i downloaded):\nbinary_file_path = '/kaggle/input/rsna-2023-abdominal-trauma-detection/segmentations/21057'\ndt = np.dtype('B')\nLoad the binary data into a numpy array:\nwith open(binary_file_path, \"rb\") as f:\nnumpy_data = np.fromfile(f, dt)\ni received numpy array with shape (1071645024,). But it is not seemly to be suitable with image size (512x512)?"
    },
    {
      "id": 2454200,
      "postDate": "2023-09-24T16:28:46.733Z",
      "content": "<p>Thank you for this information!</p>",
      "rawMarkdown": "Thank you for this information!"
    },
    {
      "id": 2443727,
      "postDate": "2023-09-17T23:29:19.457Z",
      "content": "<p>Thank you for this information!</p>",
      "rawMarkdown": "Thank you for this information!"
    },
    {
      "id": 2426138,
      "postDate": "2023-09-06T12:40:43.140Z",
      "content": "<p>Thanks a lot</p>",
      "rawMarkdown": "Thanks a lot"
    }
  ],
  "comments": [
    {
      "id": 2440969,
      "author_name": "rjwo",
      "author_url": "",
      "post_date": "2023-09-15T21:24:37.680000",
      "content": "<p>Hi, very basic question: I'm trying to run the example training and inference notebooks with no modifications, but the inference notebook can't load the model saved by the training notebook. </p>\n<p>The training notebook saves the model file <code>rsna-atd.keras</code> in my Outputs workspace, but then when I start up the inference notebook I can't see it. I tried clicking \"add data\" but it says there are no data sources found. What am I missing here?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2421628,
      "author_name": "Siddharth Savani",
      "author_url": "",
      "post_date": "2023-09-03T13:10:42.297000",
      "content": "<p>How can I use the images data without downloading it? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2422181,
          "author_name": "pkm",
          "author_url": "",
          "post_date": "2023-09-03T18:49:31.073000",
          "content": "<p>Then you are at kaggle notebook on the right part of screen you can see folder \"Data\". You can click on \"+Add data\" and there choose \"rsna-2023-abdominal-trauma-detection\".</p>\n<p>Then you can use files, for example link to a file can be like \"/kaggle/input/rsna-2023-abdominal-trauma-detection/train_images/122/21217/100.dcm\"</p>\n<p>After that you can preprocess it, for some tip you can check:</p>\n<ol>\n<li><a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/alenic/dataset-size-reduction-400gb-to-7-5gb?scriptVersionId=140444781</a></li>\n<li><a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/jhoward/creating-a-metadata-dataframe-fastai/notebook?scriptVersionId=22141307</a></li>\n<li><a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/awsaf49/rsna-atd-512x512-png-v2-data/notebook</a></li>\n</ol>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2419519,
      "author_name": "pkm",
      "author_url": "",
      "post_date": "2023-09-02T04:41:48.013000",
      "content": "<p>Hi everyone. I have a noobie question.<br>\nI think I read a lot of stuff about my question already, but there is still some mess in my head.</p>\n<p>How to deal with a not consistent amount of input data?<br>\nWe have CT scans per patient, but there can be 512x512x50, 512x512x300, 512x512x800 etc. (by the last position I mean number of CT scans in one folder)</p>\n<p>So, because of that:</p>\n<ol>\n<li>We should just \"resize\" that to some standard size and then give it to NN?<br>\n1a. If it is right - so how exactly we do that? Somewhere I saw using \"scipy.ndimage.zoom\" this function maby do what I need. Or we just simply take a fixed size sample of CT scans (every 2/5/10/.. image depend on input size) and vary it for a more accurate score and appropriate computation time? (if this is right I don't understand how to deal with the situation, then the input size will be less than the fixed size sample - just copy the \"average picture\" to gain a minimum sample size?)</li>\n<li>Also, in the real test set, I assume it can be from 1 to, for example, 1k the amount of CT scans, right?</li>\n</ol>\n<p>If you have some links to read about, it would be great :)</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2463479,
      "author_name": "jakkiabc",
      "author_url": "",
      "post_date": "2023-10-01T11:08:22.113000",
      "content": "<p>good，it is very helpful</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2441965,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-09-16T15:39:53.253000",
      "content": "<p>Please give some link to higher-scoring public notebook?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2365162,
      "author_name": "Yang.Xu21",
      "author_url": "",
      "post_date": "2023-07-30T03:26:56.927000",
      "content": "<p>Thank you, the information would be helpful.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2364921,
      "author_name": "Vietnamese-NHNAM",
      "author_url": "",
      "post_date": "2023-07-29T18:14:04.063000",
      "content": "<p><strong>Can not read file in segmentation folder</strong></p>\n<p>According to what i read in the data section, in the segmentation directory, the files are formatted as nifti. But I use nibabel library to read:<br>\nimport nibabel as nib<br>\nimg = nib.load('/kaggle/input/rsna-2023-abdominal-trauma-detection/segmentations/10000')<br>\nAn error occurred: Cannot work out file type of \"/kaggle/input/rsna-2023-abdominal-trauma-detection/segmentations/10000\".<br>\nMaybe it's not in .nii format. So what format are those files and how to read them?<br>\nThen I use this code to read binary file (arcording to i see in information of these file when i downloaded):<br>\nbinary_file_path = '/kaggle/input/rsna-2023-abdominal-trauma-detection/segmentations/21057'<br>\ndt = np.dtype('B')<br>\nLoad the binary data into a numpy array:<br>\nwith open(binary_file_path, \"rb\") as f:<br>\nnumpy_data = np.fromfile(f, dt)<br>\ni received numpy array with shape (1071645024,). But it is not seemly to be suitable with image size (512x512)?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2454200,
      "author_name": "huanshen1700",
      "author_url": "",
      "post_date": "2023-09-24T16:28:46.733000",
      "content": "<p>Thank you for this information!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2443727,
      "author_name": "Nobo Tana",
      "author_url": "",
      "post_date": "2023-09-17T23:29:19.457000",
      "content": "<p>Thank you for this information!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2426138,
      "author_name": "papapopo",
      "author_url": "",
      "post_date": "2023-09-06T12:40:43.140000",
      "content": "<p>Thanks a lot</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2360705": "New to machine learning and data science? No question is too basic or too simple. Feel free to start your own thread, or use this thread as a place to post any first-timer clarifying questions for the Kaggle community to help you with!\n\nIf you would consider yourself a beginner but don't know where to get started, let other Kagglers help you take your first steps here!\n\nNew to Kaggle? Take a look at a few videos Dr. Rachael Tatman has put together to learn a bit more about [site etiquette](https://www.youtube.com/watch?v=aIus8si_Et0), or [Kaggle lingo](https://www.youtube.com/watch?v=sEJHyuWKd-s).\n\nRemember: Kaggle is for everyone. Whether you're teaming up or sharing tips in the competition forum, we expect everyone to follow our [Kaggle community guidelines](https://www.kaggle.com/community-guidelines).\n\nA tip on sharing content - Kaggle is a collaborative community, whereby sharing techniques, starter notebooks, and ideas in the discussion forums are highly encouraged throughout the competition. However, as the competition draws closer to the final deadline it's customary to keep high-scoring notebooks withheld until after the competition has concluded. This maintains the spirit of the competition, while also allowing individuals to submit their own creative work without jeopardy of a higher-scoring notebook being available for an automatic higher rank (through copy/submit). We disable publishing of public notebooks within the final week of the competition, but encourage you to use your best judgment prior to that deadline.\n\n  \n\nHappy Modeling!\n",
    "2440969": "Hi, very basic question: I'm trying to run the example training and inference notebooks with no modifications, but the inference notebook can't load the model saved by the training notebook. \n\nThe training notebook saves the model file `rsna-atd.keras` in my Outputs workspace, but then when I start up the inference notebook I can't see it. I tried clicking \"add data\" but it says there are no data sources found. What am I missing here?",
    "2421628": "How can I use the images data without downloading it? ",
    "2419519": "Hi everyone. I have a noobie question.\nI think I read a lot of stuff about my question already, but there is still some mess in my head.\n\nHow to deal with a not consistent amount of input data?\nWe have CT scans per patient, but there can be 512x512x50, 512x512x300, 512x512x800 etc. (by the last position I mean number of CT scans in one folder)\n\nSo, because of that:\n1. We should just \"resize\" that to some standard size and then give it to NN?\n 1a. If it is right - so how exactly we do that? Somewhere I saw using \"scipy.ndimage.zoom\" this function maby do what I need. Or we just simply take a fixed size sample of CT scans (every 2/5/10/.. image depend on input size) and vary it for a more accurate score and appropriate computation time? (if this is right I don't understand how to deal with the situation, then the input size will be less than the fixed size sample - just copy the \"average picture\" to gain a minimum sample size?)\n2. Also, in the real test set, I assume it can be from 1 to, for example, 1k the amount of CT scans, right?\n\nIf you have some links to read about, it would be great :)",
    "2463479": "good，it is very helpful",
    "2441965": " Please give some link to higher-scoring public notebook?",
    "2365162": "Thank you, the information would be helpful.",
    "2364921": "**Can not read file in segmentation folder**\n\nAccording to what i read in the data section, in the segmentation directory, the files are formatted as nifti. But I use nibabel library to read:\nimport nibabel as nib\nimg = nib.load('/kaggle/input/rsna-2023-abdominal-trauma-detection/segmentations/10000')\nAn error occurred: Cannot work out file type of \"/kaggle/input/rsna-2023-abdominal-trauma-detection/segmentations/10000\".\nMaybe it's not in .nii format. So what format are those files and how to read them?\nThen I use this code to read binary file (arcording to i see in information of these file when i downloaded):\nbinary_file_path = '/kaggle/input/rsna-2023-abdominal-trauma-detection/segmentations/21057'\ndt = np.dtype('B')\nLoad the binary data into a numpy array:\nwith open(binary_file_path, \"rb\") as f:\nnumpy_data = np.fromfile(f, dt)\ni received numpy array with shape (1071645024,). But it is not seemly to be suitable with image size (512x512)?",
    "2454200": "Thank you for this information!",
    "2443727": "Thank you for this information!",
    "2426138": "Thanks a lot"
  }
}