{
  "id": 450279,
  "title": "79th place - beginner tutorial in applying a previous solution with minimal changes and training time ",
  "url": "/competitions/rsna-2023-abdominal-trauma-detection/discussion/450279",
  "author_name": "Chris Miles",
  "post_date": "2023-10-23T18:58:22.488000",
  "votes": 7,
  "comment_count": 0,
  "views": 0,
  "content": "<p>This post is meant to show other beginners how we can take a previous solution, apply the smallest changes possible, and still achieve a bronze medal with 30 hours of kaggle gpu (silver medal with 30 additional hours of a kaggle gpu). </p>\n<p>Thanks to kaggle and the organizers. <br>\nThanks to <a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">Qishen Ha</a> for their <a href=\"https://www.kaggle.com/competitions/rsna-2022-cervical-spine-fracture-detection/discussion/362787\" target=\"_blank\">1st place solution of the rsna 2022 competition</a>, where the input data was the same, but the targets were fractures in the spinal vertebrae C1-C7. My aim was to learn how this code works and apply it to this competition. Also thanks to <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">Theo Viel</a>, who pointed out Qishen’s solution in his post about <a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/441557\" target=\"_blank\">beating the baseline</a>, and for his notebook about <a href=\"https://www.kaggle.com/code/theoviel/get-started-quicker-dicom-png-conversion\" target=\"_blank\">processing the dicom files into pngs</a>.</p>\n<p>I trained one model to segment all the organs, followed by one model for each organ to classify injury. For extravasation I just predicted an optimized constant value, frequency_of_extravasation x 6. </p>\n<p>I only used kaggle resources (about 30 gpu hours total), scoring .615 private leaderboard. This is with only 15 epochs for the final classification models. After training those models for 45 total epochs (taking about 30 extra kaggle gpu hours), we get .548 private leaderboard, which puts us right on the edge for a silver medal at 56th place. </p>\n<h2>Adapting <a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">Qishen Ha</a>’s approach step by step</h2>\n<h3>Stage 1: Segmentation</h3>\n<p>The first step is to build a model that will segment out the relevant organs. <br>\nFor reference, this is <a href=\"https://www.kaggle.com/code/haqishen/rsna-2022-1st-place-solution-train-stage1\" target=\"_blank\">Qishen’s notebook</a> which builds such a model for spinal vertebrae C1-C7. As an input into that notebook, he has built a dataset with the studies processed into 3D images of size 128x128x128 to save time. </p>\n<p>Here is my <a href=\"https://www.kaggle.com/code/chrisrichardmiles/rsna23-dicom-to-3d-array-128x128x128-fixed\" target=\"_blank\">notebook that processes all training studies into 3D images of size 128x128x128</a>. I have combined Qishen’s original code with Theo Veil’s previously mentioned dicom processing code. The output of this notebook is an input into the next notebook, which trains the segmenter model. </p>\n<p>Here is the <a href=\"https://www.kaggle.com/code/chrisrichardmiles/rsna-2023-train-stage1-seg-mask?scriptVersionId=145931960\" target=\"_blank\">stage 1 segmentation model training notebook</a>. I link to version 7 because linking to a notebook that timed out crashes your browser. In version 9, I train one fold for 15 epochs (which is what I get after the 12 hours of kaggle gpu). From visualizing the output masks, it seems to be good enough.<br>\nHere is the <a href=\"https://www.kaggle.com/datasets/chrisrichardmiles/rsna23-train-stage1\" target=\"_blank\">dataset with the model output</a>,<br>\nInput size: 128x128x128<br>\nEpochs: 15</p>\n<h4>Stage 1.5: Segmentation inference and building 2.5D image input into stage 2 model</h4>\n<p>In order to make the stage 2 training efficient, we precompute the segmentation masks by using the stage 1 segmentation model to infer the segmentation mask for each study_id. After finding the mask for the entire study_id, we build 15 “2.5 dimensional” images for each organ. For each organ we find the min and max value across the z axis (spinal axis), and take 15 images evenly spaced across this z range. For each image we also include 2 images from above and 2 from below for extra information. We also include the segmentation mask so that the model knows where the region of interest is. So for each organ (liver, left_kidney, right_kidney, spleen, bowel), our final result is an array of shape (15, 6, 224, 224). We stack all organ’s outputs and save one file. So for each study_id, we save a file of shape (75, 6, 224, 224). This will be used to create 15 inputs into the stage 2 classifier model for each organ. </p>\n<p>Since the output of kaggle notebooks is limited to 20GB, we use 30 notebooks to get the segmentation masks for all the training data. Here is <a href=\"https://www.kaggle.com/code/chrisrichardmiles/fork-of-rsna23-s1-inf-30parts-2/output\" target=\"_blank\">#2 of 30 as an example</a>. All 30 must be put as an input into the stage 2 model training notebook.</p>\n<h3>Stage 2 models: [classification]</h3>\n<p>For reference, here is Qishen’s <a href=\"https://www.kaggle.com/code/haqishen/rsna-2022-1st-place-solution-train-stage2-type1\" target=\"_blank\">stage 2 training notebook</a>.</p>\n<p>Here is my <a href=\"https://www.kaggle.com/chrisrichardmiles/rsna23-train-stage2-final-5\" target=\"_blank\">stage 2 training notebook</a>. Note that there is code added that is used to continue training, using the best models saved from previous versions of the same notebook. This code should be commented out on the first run. </p>\n<p>Here is the <a href=\"https://www.kaggle.com/code/chrisrichardmiles/fork-of-rsna23-final-inference-4-diff-agg?scriptVersionId=147503119\" target=\"_blank\">final inference notebook</a> which scores .548 private LB. </p>\n<p><strong>Key changes in my stage 2 models, compared to Qishen's model</strong>: </p>\n<ul>\n<li>In Qishen’s notebook, he builds one single model to classify if a vertebra has a fracture. For each vertebrae C1-C7, he makes 105 input samples to train with. This makes sense because each vertebrae looks similar. C2 looks a lot like C5. But for this competition, each organ does not look like the other, so I chose to build  4 different models for liver, kidney, spleen, and bowel. <br>\n<strong>special note about kidney</strong>: Since the segmentation data from the organizers had different labels for left and right kidney, my segmentation masks also had left and right kidney. In order to build a single model for the kidneys, I concatenated the left and right kidney. To be clear I took the left and right kidney arrays (shape (15,6,224,224)) resulting from the input building in stage 1.5, and combined them to get an array of shape (15, 6, 448, 224). </li>\n</ul>\n<p>Here is the dataloader for the stage 2 classifier model: </p>\n<pre><code>class CLSDataset(Dataset):\n    def __init__(self, df, mode, ):\n\n        self.df = df.reset_index()\n        self.mode = mode\n        self. = \n\n    def __len__(self):\n         self.df.shape[]\n\n    def __getitem__(self, index):\n         = self.df.iloc[index]\n\n\n        image_full = .(.cls_inp_path)\n        out = defaultdict(dict)\n         organ, cols, (a, b)  zip(ORGANS, LABELS, ABS): \n            images = []\n               image_full[a: b]: \n                 = .(, , )\n                 = transforms_train(=)['']\n                 = .(, , )\n                images.()\n            images = .stack(images, )\n             organ == 'kidney': \n                images = .concatenate((images[:, :, :, :], images[:, :, :, :]), )\n            out[organ]['images'] = torch.tensor(images).()\n            out[organ][''] = torch.tensor([[cols]] * n_slice_per_c).()\n         out\n</code></pre>\n<p><strong>Note</strong>: </p>\n<ul>\n<li>Even though the batch_size I use for the dataloader is 1, we get 15 training examples for each batch. So the model is treating each 6x224x224 image on its own, but it is processing all 15 images at once, as if the batch size were 15. </li>\n</ul>",
  "messages": [
    {
      "id": 2496147,
      "postDate": "2023-10-23T18:58:22.490Z",
      "content": "<p>This post is meant to show other beginners how we can take a previous solution, apply the smallest changes possible, and still achieve a bronze medal with 30 hours of kaggle gpu (silver medal with 30 additional hours of a kaggle gpu). </p>\n<p>Thanks to kaggle and the organizers. <br>\nThanks to <a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">Qishen Ha</a> for their <a href=\"https://www.kaggle.com/competitions/rsna-2022-cervical-spine-fracture-detection/discussion/362787\" target=\"_blank\">1st place solution of the rsna 2022 competition</a>, where the input data was the same, but the targets were fractures in the spinal vertebrae C1-C7. My aim was to learn how this code works and apply it to this competition. Also thanks to <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">Theo Viel</a>, who pointed out Qishen’s solution in his post about <a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/441557\" target=\"_blank\">beating the baseline</a>, and for his notebook about <a href=\"https://www.kaggle.com/code/theoviel/get-started-quicker-dicom-png-conversion\" target=\"_blank\">processing the dicom files into pngs</a>.</p>\n<p>I trained one model to segment all the organs, followed by one model for each organ to classify injury. For extravasation I just predicted an optimized constant value, frequency_of_extravasation x 6. </p>\n<p>I only used kaggle resources (about 30 gpu hours total), scoring .615 private leaderboard. This is with only 15 epochs for the final classification models. After training those models for 45 total epochs (taking about 30 extra kaggle gpu hours), we get .548 private leaderboard, which puts us right on the edge for a silver medal at 56th place. </p>\n<h2>Adapting <a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">Qishen Ha</a>’s approach step by step</h2>\n<h3>Stage 1: Segmentation</h3>\n<p>The first step is to build a model that will segment out the relevant organs. <br>\nFor reference, this is <a href=\"https://www.kaggle.com/code/haqishen/rsna-2022-1st-place-solution-train-stage1\" target=\"_blank\">Qishen’s notebook</a> which builds such a model for spinal vertebrae C1-C7. As an input into that notebook, he has built a dataset with the studies processed into 3D images of size 128x128x128 to save time. </p>\n<p>Here is my <a href=\"https://www.kaggle.com/code/chrisrichardmiles/rsna23-dicom-to-3d-array-128x128x128-fixed\" target=\"_blank\">notebook that processes all training studies into 3D images of size 128x128x128</a>. I have combined Qishen’s original code with Theo Veil’s previously mentioned dicom processing code. The output of this notebook is an input into the next notebook, which trains the segmenter model. </p>\n<p>Here is the <a href=\"https://www.kaggle.com/code/chrisrichardmiles/rsna-2023-train-stage1-seg-mask?scriptVersionId=145931960\" target=\"_blank\">stage 1 segmentation model training notebook</a>. I link to version 7 because linking to a notebook that timed out crashes your browser. In version 9, I train one fold for 15 epochs (which is what I get after the 12 hours of kaggle gpu). From visualizing the output masks, it seems to be good enough.<br>\nHere is the <a href=\"https://www.kaggle.com/datasets/chrisrichardmiles/rsna23-train-stage1\" target=\"_blank\">dataset with the model output</a>,<br>\nInput size: 128x128x128<br>\nEpochs: 15</p>\n<h4>Stage 1.5: Segmentation inference and building 2.5D image input into stage 2 model</h4>\n<p>In order to make the stage 2 training efficient, we precompute the segmentation masks by using the stage 1 segmentation model to infer the segmentation mask for each study_id. After finding the mask for the entire study_id, we build 15 “2.5 dimensional” images for each organ. For each organ we find the min and max value across the z axis (spinal axis), and take 15 images evenly spaced across this z range. For each image we also include 2 images from above and 2 from below for extra information. We also include the segmentation mask so that the model knows where the region of interest is. So for each organ (liver, left_kidney, right_kidney, spleen, bowel), our final result is an array of shape (15, 6, 224, 224). We stack all organ’s outputs and save one file. So for each study_id, we save a file of shape (75, 6, 224, 224). This will be used to create 15 inputs into the stage 2 classifier model for each organ. </p>\n<p>Since the output of kaggle notebooks is limited to 20GB, we use 30 notebooks to get the segmentation masks for all the training data. Here is <a href=\"https://www.kaggle.com/code/chrisrichardmiles/fork-of-rsna23-s1-inf-30parts-2/output\" target=\"_blank\">#2 of 30 as an example</a>. All 30 must be put as an input into the stage 2 model training notebook.</p>\n<h3>Stage 2 models: [classification]</h3>\n<p>For reference, here is Qishen’s <a href=\"https://www.kaggle.com/code/haqishen/rsna-2022-1st-place-solution-train-stage2-type1\" target=\"_blank\">stage 2 training notebook</a>.</p>\n<p>Here is my <a href=\"https://www.kaggle.com/chrisrichardmiles/rsna23-train-stage2-final-5\" target=\"_blank\">stage 2 training notebook</a>. Note that there is code added that is used to continue training, using the best models saved from previous versions of the same notebook. This code should be commented out on the first run. </p>\n<p>Here is the <a href=\"https://www.kaggle.com/code/chrisrichardmiles/fork-of-rsna23-final-inference-4-diff-agg?scriptVersionId=147503119\" target=\"_blank\">final inference notebook</a> which scores .548 private LB. </p>\n<p><strong>Key changes in my stage 2 models, compared to Qishen's model</strong>: </p>\n<ul>\n<li>In Qishen’s notebook, he builds one single model to classify if a vertebra has a fracture. For each vertebrae C1-C7, he makes 105 input samples to train with. This makes sense because each vertebrae looks similar. C2 looks a lot like C5. But for this competition, each organ does not look like the other, so I chose to build  4 different models for liver, kidney, spleen, and bowel. <br>\n<strong>special note about kidney</strong>: Since the segmentation data from the organizers had different labels for left and right kidney, my segmentation masks also had left and right kidney. In order to build a single model for the kidneys, I concatenated the left and right kidney. To be clear I took the left and right kidney arrays (shape (15,6,224,224)) resulting from the input building in stage 1.5, and combined them to get an array of shape (15, 6, 448, 224). </li>\n</ul>\n<p>Here is the dataloader for the stage 2 classifier model: </p>\n<pre><code>class CLSDataset(Dataset):\n    def __init__(self, df, mode, ):\n\n        self.df = df.reset_index()\n        self.mode = mode\n        self. = \n\n    def __len__(self):\n         self.df.shape[]\n\n    def __getitem__(self, index):\n         = self.df.iloc[index]\n\n\n        image_full = .(.cls_inp_path)\n        out = defaultdict(dict)\n         organ, cols, (a, b)  zip(ORGANS, LABELS, ABS): \n            images = []\n               image_full[a: b]: \n                 = .(, , )\n                 = transforms_train(=)['']\n                 = .(, , )\n                images.()\n            images = .stack(images, )\n             organ == 'kidney': \n                images = .concatenate((images[:, :, :, :], images[:, :, :, :]), )\n            out[organ]['images'] = torch.tensor(images).()\n            out[organ][''] = torch.tensor([[cols]] * n_slice_per_c).()\n         out\n</code></pre>\n<p><strong>Note</strong>: </p>\n<ul>\n<li>Even though the batch_size I use for the dataloader is 1, we get 15 training examples for each batch. So the model is treating each 6x224x224 image on its own, but it is processing all 15 images at once, as if the batch size were 15. </li>\n</ul>",
      "rawMarkdown": "This post is meant to show other beginners how we can take a previous solution, apply the smallest changes possible, and still achieve a bronze medal with 30 hours of kaggle gpu (silver medal with 30 additional hours of a kaggle gpu). \n\nThanks to kaggle and the organizers. \nThanks to [Qishen Ha](https://www.kaggle.com/haqishen) for their [1st place solution of the rsna 2022 competition](https://www.kaggle.com/competitions/rsna-2022-cervical-spine-fracture-detection/discussion/362787), where the input data was the same, but the targets were fractures in the spinal vertebrae C1-C7. My aim was to learn how this code works and apply it to this competition. Also thanks to [Theo Viel](https://www.kaggle.com/theoviel), who pointed out Qishen’s solution in his post about [beating the baseline](https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/441557), and for his notebook about [processing the dicom files into pngs](https://www.kaggle.com/code/theoviel/get-started-quicker-dicom-png-conversion).\n\nI trained one model to segment all the organs, followed by one model for each organ to classify injury. For extravasation I just predicted an optimized constant value, frequency_of_extravasation x 6. \n\nI only used kaggle resources (about 30 gpu hours total), scoring .615 private leaderboard. This is with only 15 epochs for the final classification models. After training those models for 45 total epochs (taking about 30 extra kaggle gpu hours), we get .548 private leaderboard, which puts us right on the edge for a silver medal at 56th place. \n\n## Adapting [Qishen Ha](https://www.kaggle.com/haqishen)’s approach step by step\n\n### Stage 1: Segmentation\nThe first step is to build a model that will segment out the relevant organs. \nFor reference, this is [Qishen’s notebook](https://www.kaggle.com/code/haqishen/rsna-2022-1st-place-solution-train-stage1) which builds such a model for spinal vertebrae C1-C7. As an input into that notebook, he has built a dataset with the studies processed into 3D images of size 128x128x128 to save time. \n\nHere is my [notebook that processes all training studies into 3D images of size 128x128x128](https://www.kaggle.com/code/chrisrichardmiles/rsna23-dicom-to-3d-array-128x128x128-fixed). I have combined Qishen’s original code with Theo Veil’s previously mentioned dicom processing code. The output of this notebook is an input into the next notebook, which trains the segmenter model. \n\nHere is the [stage 1 segmentation model training notebook] (https://www.kaggle.com/code/chrisrichardmiles/rsna-2023-train-stage1-seg-mask?scriptVersionId=145931960). I link to version 7 because linking to a notebook that timed out crashes your browser. In version 9, I train one fold for 15 epochs (which is what I get after the 12 hours of kaggle gpu). From visualizing the output masks, it seems to be good enough.\nHere is the [dataset with the model output](https://www.kaggle.com/datasets/chrisrichardmiles/rsna23-train-stage1),\nInput size: 128x128x128\nEpochs: 15\n\n#### Stage 1.5: Segmentation inference and building 2.5D image input into stage 2 model\nIn order to make the stage 2 training efficient, we precompute the segmentation masks by using the stage 1 segmentation model to infer the segmentation mask for each study_id. After finding the mask for the entire study_id, we build 15 “2.5 dimensional” images for each organ. For each organ we find the min and max value across the z axis (spinal axis), and take 15 images evenly spaced across this z range. For each image we also include 2 images from above and 2 from below for extra information. We also include the segmentation mask so that the model knows where the region of interest is. So for each organ (liver, left_kidney, right_kidney, spleen, bowel), our final result is an array of shape (15, 6, 224, 224). We stack all organ’s outputs and save one file. So for each study_id, we save a file of shape (75, 6, 224, 224). This will be used to create 15 inputs into the stage 2 classifier model for each organ. \n\n\n\nSince the output of kaggle notebooks is limited to 20GB, we use 30 notebooks to get the segmentation masks for all the training data. Here is [#2 of 30 as an example](https://www.kaggle.com/code/chrisrichardmiles/fork-of-rsna23-s1-inf-30parts-2/output). All 30 must be put as an input into the stage 2 model training notebook.\n\n### Stage 2 models: [classification]\nFor reference, here is Qishen’s [stage 2 training notebook](https://www.kaggle.com/code/haqishen/rsna-2022-1st-place-solution-train-stage2-type1).\n\nHere is my [stage 2 training notebook](https://www.kaggle.com/chrisrichardmiles/rsna23-train-stage2-final-5). Note that there is code added that is used to continue training, using the best models saved from previous versions of the same notebook. This code should be commented out on the first run. \n\nHere is the [final inference notebook](https://www.kaggle.com/code/chrisrichardmiles/fork-of-rsna23-final-inference-4-diff-agg?scriptVersionId=147503119) which scores .548 private LB. \n\n**Key changes in my stage 2 models, compared to Qishen's model**: \n* In Qishen’s notebook, he builds one single model to classify if a vertebra has a fracture. For each vertebrae C1-C7, he makes 105 input samples to train with. This makes sense because each vertebrae looks similar. C2 looks a lot like C5. But for this competition, each organ does not look like the other, so I chose to build  4 different models for liver, kidney, spleen, and bowel. \n**special note about kidney**: Since the segmentation data from the organizers had different labels for left and right kidney, my segmentation masks also had left and right kidney. In order to build a single model for the kidneys, I concatenated the left and right kidney. To be clear I took the left and right kidney arrays (shape (15,6,224,224)) resulting from the input building in stage 1.5, and combined them to get an array of shape (15, 6, 448, 224). \n\nHere is the dataloader for the stage 2 classifier model: \n```\nclass CLSDataset(Dataset):\n    def __init__(self, df, mode, transform):\n\n        self.df = df.reset_index()\n        self.mode = mode\n        self.transform = transform\n\n    def __len__(self):\n        return self.df.shape[0]\n\n    def __getitem__(self, index):\n        row = self.df.iloc[index]\n        \n        \n        image_full = np.load(row.cls_inp_path)\n        out = defaultdict(dict)\n        for organ, cols, (a, b) in zip(ORGANS, LABELS, ABS): \n            images = []\n            for image in image_full[a: b]: \n                image = image.transpose(1, 2, 0)\n                image = transforms_train(image=image)['image']\n                image = image.transpose(2, 0, 1)\n                images.append(image)\n            images = np.stack(images, 0)\n            if organ == 'kidney': \n                images = np.concatenate((images[:15, :, :, :], images[15:, :, :, :]), 2)\n            out[organ]['images'] = torch.tensor(images).float()\n            out[organ]['labels'] = torch.tensor([row[cols]] * n_slice_per_c).float()\n        return out\n```\n**Note**: \n* Even though the batch_size I use for the dataloader is 1, we get 15 training examples for each batch. So the model is treating each 6x224x224 image on its own, but it is processing all 15 images at once, as if the batch size were 15. \n\n",
      "votes": 7
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2496147": "This post is meant to show other beginners how we can take a previous solution, apply the smallest changes possible, and still achieve a bronze medal with 30 hours of kaggle gpu (silver medal with 30 additional hours of a kaggle gpu). \n\nThanks to kaggle and the organizers. \nThanks to [Qishen Ha](https://www.kaggle.com/haqishen) for their [1st place solution of the rsna 2022 competition](https://www.kaggle.com/competitions/rsna-2022-cervical-spine-fracture-detection/discussion/362787), where the input data was the same, but the targets were fractures in the spinal vertebrae C1-C7. My aim was to learn how this code works and apply it to this competition. Also thanks to [Theo Viel](https://www.kaggle.com/theoviel), who pointed out Qishen’s solution in his post about [beating the baseline](https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/441557), and for his notebook about [processing the dicom files into pngs](https://www.kaggle.com/code/theoviel/get-started-quicker-dicom-png-conversion).\n\nI trained one model to segment all the organs, followed by one model for each organ to classify injury. For extravasation I just predicted an optimized constant value, frequency_of_extravasation x 6. \n\nI only used kaggle resources (about 30 gpu hours total), scoring .615 private leaderboard. This is with only 15 epochs for the final classification models. After training those models for 45 total epochs (taking about 30 extra kaggle gpu hours), we get .548 private leaderboard, which puts us right on the edge for a silver medal at 56th place. \n\n## Adapting [Qishen Ha](https://www.kaggle.com/haqishen)’s approach step by step\n\n### Stage 1: Segmentation\nThe first step is to build a model that will segment out the relevant organs. \nFor reference, this is [Qishen’s notebook](https://www.kaggle.com/code/haqishen/rsna-2022-1st-place-solution-train-stage1) which builds such a model for spinal vertebrae C1-C7. As an input into that notebook, he has built a dataset with the studies processed into 3D images of size 128x128x128 to save time. \n\nHere is my [notebook that processes all training studies into 3D images of size 128x128x128](https://www.kaggle.com/code/chrisrichardmiles/rsna23-dicom-to-3d-array-128x128x128-fixed). I have combined Qishen’s original code with Theo Veil’s previously mentioned dicom processing code. The output of this notebook is an input into the next notebook, which trains the segmenter model. \n\nHere is the [stage 1 segmentation model training notebook] (https://www.kaggle.com/code/chrisrichardmiles/rsna-2023-train-stage1-seg-mask?scriptVersionId=145931960). I link to version 7 because linking to a notebook that timed out crashes your browser. In version 9, I train one fold for 15 epochs (which is what I get after the 12 hours of kaggle gpu). From visualizing the output masks, it seems to be good enough.\nHere is the [dataset with the model output](https://www.kaggle.com/datasets/chrisrichardmiles/rsna23-train-stage1),\nInput size: 128x128x128\nEpochs: 15\n\n#### Stage 1.5: Segmentation inference and building 2.5D image input into stage 2 model\nIn order to make the stage 2 training efficient, we precompute the segmentation masks by using the stage 1 segmentation model to infer the segmentation mask for each study_id. After finding the mask for the entire study_id, we build 15 “2.5 dimensional” images for each organ. For each organ we find the min and max value across the z axis (spinal axis), and take 15 images evenly spaced across this z range. For each image we also include 2 images from above and 2 from below for extra information. We also include the segmentation mask so that the model knows where the region of interest is. So for each organ (liver, left_kidney, right_kidney, spleen, bowel), our final result is an array of shape (15, 6, 224, 224). We stack all organ’s outputs and save one file. So for each study_id, we save a file of shape (75, 6, 224, 224). This will be used to create 15 inputs into the stage 2 classifier model for each organ. \n\n\n\nSince the output of kaggle notebooks is limited to 20GB, we use 30 notebooks to get the segmentation masks for all the training data. Here is [#2 of 30 as an example](https://www.kaggle.com/code/chrisrichardmiles/fork-of-rsna23-s1-inf-30parts-2/output). All 30 must be put as an input into the stage 2 model training notebook.\n\n### Stage 2 models: [classification]\nFor reference, here is Qishen’s [stage 2 training notebook](https://www.kaggle.com/code/haqishen/rsna-2022-1st-place-solution-train-stage2-type1).\n\nHere is my [stage 2 training notebook](https://www.kaggle.com/chrisrichardmiles/rsna23-train-stage2-final-5). Note that there is code added that is used to continue training, using the best models saved from previous versions of the same notebook. This code should be commented out on the first run. \n\nHere is the [final inference notebook](https://www.kaggle.com/code/chrisrichardmiles/fork-of-rsna23-final-inference-4-diff-agg?scriptVersionId=147503119) which scores .548 private LB. \n\n**Key changes in my stage 2 models, compared to Qishen's model**: \n* In Qishen’s notebook, he builds one single model to classify if a vertebra has a fracture. For each vertebrae C1-C7, he makes 105 input samples to train with. This makes sense because each vertebrae looks similar. C2 looks a lot like C5. But for this competition, each organ does not look like the other, so I chose to build  4 different models for liver, kidney, spleen, and bowel. \n**special note about kidney**: Since the segmentation data from the organizers had different labels for left and right kidney, my segmentation masks also had left and right kidney. In order to build a single model for the kidneys, I concatenated the left and right kidney. To be clear I took the left and right kidney arrays (shape (15,6,224,224)) resulting from the input building in stage 1.5, and combined them to get an array of shape (15, 6, 448, 224). \n\nHere is the dataloader for the stage 2 classifier model: \n```\nclass CLSDataset(Dataset):\n    def __init__(self, df, mode, transform):\n\n        self.df = df.reset_index()\n        self.mode = mode\n        self.transform = transform\n\n    def __len__(self):\n        return self.df.shape[0]\n\n    def __getitem__(self, index):\n        row = self.df.iloc[index]\n        \n        \n        image_full = np.load(row.cls_inp_path)\n        out = defaultdict(dict)\n        for organ, cols, (a, b) in zip(ORGANS, LABELS, ABS): \n            images = []\n            for image in image_full[a: b]: \n                image = image.transpose(1, 2, 0)\n                image = transforms_train(image=image)['image']\n                image = image.transpose(2, 0, 1)\n                images.append(image)\n            images = np.stack(images, 0)\n            if organ == 'kidney': \n                images = np.concatenate((images[:15, :, :, :], images[15:, :, :, :]), 2)\n            out[organ]['images'] = torch.tensor(images).float()\n            out[organ]['labels'] = torch.tensor([row[cols]] * n_slice_per_c).float()\n        return out\n```\n**Note**: \n* Even though the batch_size I use for the dataloader is 1, we get 15 training examples for each batch. So the model is treating each 6x224x224 image on its own, but it is processing all 15 images at once, as if the batch size were 15. \n\n"
  }
}