{
  "id": 362647,
  "title": "[38th Place] Single Stage Single Model Efficientnetv2",
  "url": "/competitions/rsna-2022-cervical-spine-fracture-detection/discussion/362647",
  "author_name": "Yerram Varun",
  "post_date": "2022-10-28T10:09:18.013000",
  "votes": 22,
  "comment_count": 4,
  "views": 0,
  "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4630396%2Fddfb714d23d450f8bf5d50fcac1e8ff5%2Frsna_pipeline.png?generation=1666947796450371&amp;alt=media\" alt=\"\"></p>\n<p>Thanks to RSNA and Kaggle for conducting such an exciting competition, and congratulations to all the winners!</p>\n<p>It is a modified version of the pipeline shared by <a href=\"https://www.kaggle.com/vslaykovsky\" target=\"_blank\">@vslaykovsky</a> in <a href=\"https://www.kaggle.com/competitions/rsna-2022-cervical-spine-fracture-detection/discussion/348462\" target=\"_blank\">this</a> discussion.  </p>\n<h4>CV Strategy</h4>\n<p>I use a 5 fold strategy, grouped by StudyInstanceUID and stratified by the fractures in all vertebrae (A column is created concatenating all the fractures like 0_1_0_0_0_1_0 and then used to stratify)</p>\n<h4>Preprocessing</h4>\n<p>I use windowing to preprocess the Dicom image used in previous RSNA solutions. The three window values were derived from <a href=\"https://arxiv.org/abs/2010.13336\" target=\"_blank\">this</a> paper. </p>\n<pre><code>def window(img, WL=50, WW=350):\n    upper, lower = WL+WW//2, WL-WW//2\n    X = np.clip(img.copy(), lower, upper)\n    X = X - np.min(X)\n    X = X / np.max(X)\n    X = (X*255.0).astype('uint8')\n    return X\n\ndef dicom_load(self, siuid, slic):\n    path = f\"data/train_images/{siuid}/{slic}.dcm\"\n    img = dicom.dcmread(path)\n    img.PhotometricInterpretation = 'YBR_FULL'\n    data = img.pixel_array\n    slope = img[('0028','1053')].value\n    intercept = img[('0028','1052')].value\n    data = data*slope + intercept\n\n    windowed = np.stack((window(data, WL=80, WW=300),\n                        window(data, WL=500, WW=1800),\n                        window(data, WL=400, WW=650)\n                        ), axis=-1)/255.\n    return windowed\n</code></pre>\n<h4>Training Data</h4>\n<p>Each 3-Channel slice is concatenated with +1/-1 Neighbour slices to create a 9-channel Input. This 9 Channel input predicts the fracture probabilities and vertebrae presence of the middle slice.</p>\n<p>During inference, this 3-slice window is shifted by stride 2 for faster Inference.</p>\n<h4>Model</h4>\n<p>I use <code>Timm</code> library to use pretrained models and modify the first layer to take a 9-channel input.</p>\n<pre><code>self.model = timm.create_model('tf_efficientnetv2_l', pretrained=True, num_classes=0)\nself.model.conv_stem = nn.Conv2d(9, 32, kernel_size=(3, 3), stride=(2, 2), padding='same', bias=False)\n</code></pre>\n<h3>What didn't work</h3>\n<p>This is a huge list, but I will narrow it down to 3 approaches to predict patient_overall -&gt;</p>\n<ul>\n<li>In the above approach, the sequential information between the slices is only used in the +1/-1 concatenation. To leverage more context, I generated embeddings for all the slices and then tried to predict the <code>patient_overall</code> using an LSTM/Transformer.</li>\n</ul>\n<p>Possible Reason for not working: This Stage2 model had only 2k data points to learn from, which might not have sufficed.</p>\n<ul>\n<li>Uniformly/Randomly select a fixed number of slices from the patient scans and train a 3D Classifier.</li>\n</ul>\n<p>Possible Reason for not working: Fractures can only be seen in some slices and specific slices need to filtered out.</p>\n<ul>\n<li>Trained an XGBoost to to predict patient overall from the fracture probabilities of C1-C7 vertebrae.</li>\n</ul>\n<p>Possible Reason for not working: Could not beat the prediction performance of <code>1-np.prod(1-c1c7)</code> </p>",
  "messages": [
    {
      "id": 2007561,
      "postDate": "2022-10-28T10:09:18.013Z",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4630396%2Fddfb714d23d450f8bf5d50fcac1e8ff5%2Frsna_pipeline.png?generation=1666947796450371&amp;alt=media\" alt=\"\"></p>\n<p>Thanks to RSNA and Kaggle for conducting such an exciting competition, and congratulations to all the winners!</p>\n<p>It is a modified version of the pipeline shared by <a href=\"https://www.kaggle.com/vslaykovsky\" target=\"_blank\">@vslaykovsky</a> in <a href=\"https://www.kaggle.com/competitions/rsna-2022-cervical-spine-fracture-detection/discussion/348462\" target=\"_blank\">this</a> discussion.  </p>\n<h4>CV Strategy</h4>\n<p>I use a 5 fold strategy, grouped by StudyInstanceUID and stratified by the fractures in all vertebrae (A column is created concatenating all the fractures like 0_1_0_0_0_1_0 and then used to stratify)</p>\n<h4>Preprocessing</h4>\n<p>I use windowing to preprocess the Dicom image used in previous RSNA solutions. The three window values were derived from <a href=\"https://arxiv.org/abs/2010.13336\" target=\"_blank\">this</a> paper. </p>\n<pre><code>def window(img, WL=50, WW=350):\n    upper, lower = WL+WW//2, WL-WW//2\n    X = np.clip(img.copy(), lower, upper)\n    X = X - np.min(X)\n    X = X / np.max(X)\n    X = (X*255.0).astype('uint8')\n    return X\n\ndef dicom_load(self, siuid, slic):\n    path = f\"data/train_images/{siuid}/{slic}.dcm\"\n    img = dicom.dcmread(path)\n    img.PhotometricInterpretation = 'YBR_FULL'\n    data = img.pixel_array\n    slope = img[('0028','1053')].value\n    intercept = img[('0028','1052')].value\n    data = data*slope + intercept\n\n    windowed = np.stack((window(data, WL=80, WW=300),\n                        window(data, WL=500, WW=1800),\n                        window(data, WL=400, WW=650)\n                        ), axis=-1)/255.\n    return windowed\n</code></pre>\n<h4>Training Data</h4>\n<p>Each 3-Channel slice is concatenated with +1/-1 Neighbour slices to create a 9-channel Input. This 9 Channel input predicts the fracture probabilities and vertebrae presence of the middle slice.</p>\n<p>During inference, this 3-slice window is shifted by stride 2 for faster Inference.</p>\n<h4>Model</h4>\n<p>I use <code>Timm</code> library to use pretrained models and modify the first layer to take a 9-channel input.</p>\n<pre><code>self.model = timm.create_model('tf_efficientnetv2_l', pretrained=True, num_classes=0)\nself.model.conv_stem = nn.Conv2d(9, 32, kernel_size=(3, 3), stride=(2, 2), padding='same', bias=False)\n</code></pre>\n<h3>What didn't work</h3>\n<p>This is a huge list, but I will narrow it down to 3 approaches to predict patient_overall -&gt;</p>\n<ul>\n<li>In the above approach, the sequential information between the slices is only used in the +1/-1 concatenation. To leverage more context, I generated embeddings for all the slices and then tried to predict the <code>patient_overall</code> using an LSTM/Transformer.</li>\n</ul>\n<p>Possible Reason for not working: This Stage2 model had only 2k data points to learn from, which might not have sufficed.</p>\n<ul>\n<li>Uniformly/Randomly select a fixed number of slices from the patient scans and train a 3D Classifier.</li>\n</ul>\n<p>Possible Reason for not working: Fractures can only be seen in some slices and specific slices need to filtered out.</p>\n<ul>\n<li>Trained an XGBoost to to predict patient overall from the fracture probabilities of C1-C7 vertebrae.</li>\n</ul>\n<p>Possible Reason for not working: Could not beat the prediction performance of <code>1-np.prod(1-c1c7)</code> </p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4630396%2Fddfb714d23d450f8bf5d50fcac1e8ff5%2Frsna_pipeline.png?generation=1666947796450371&alt=media)\n\nThanks to RSNA and Kaggle for conducting such an exciting competition, and congratulations to all the winners!\n\nIt is a modified version of the pipeline shared by @vslaykovsky in [this](https://www.kaggle.com/competitions/rsna-2022-cervical-spine-fracture-detection/discussion/348462) discussion.  \n\n#### CV Strategy \nI use a 5 fold strategy, grouped by StudyInstanceUID and stratified by the fractures in all vertebrae (A column is created concatenating all the fractures like 0_1_0_0_0_1_0 and then used to stratify)\n\n#### Preprocessing\nI use windowing to preprocess the Dicom image used in previous RSNA solutions. The three window values were derived from [this](https://arxiv.org/abs/2010.13336) paper. \n\n```\ndef window(img, WL=50, WW=350):\n    upper, lower = WL+WW//2, WL-WW//2\n    X = np.clip(img.copy(), lower, upper)\n    X = X - np.min(X)\n    X = X / np.max(X)\n    X = (X*255.0).astype('uint8')\n    return X\n\ndef dicom_load(self, siuid, slic):\n    path = f\"data/train_images/{siuid}/{slic}.dcm\"\n    img = dicom.dcmread(path)\n    img.PhotometricInterpretation = 'YBR_FULL'\n    data = img.pixel_array\n    slope = img[('0028','1053')].value\n    intercept = img[('0028','1052')].value\n    data = data*slope + intercept\n            \n    windowed = np.stack((window(data, WL=80, WW=300),\n                        window(data, WL=500, WW=1800),\n                        window(data, WL=400, WW=650)\n                        ), axis=-1)/255.\n    return windowed\n   ```\n\n#### Training Data\n\nEach 3-Channel slice is concatenated with +1/-1 Neighbour slices to create a 9-channel Input. This 9 Channel input predicts the fracture probabilities and vertebrae presence of the middle slice.\n\nDuring inference, this 3-slice window is shifted by stride 2 for faster Inference.\n\n#### Model\n\nI use `Timm` library to use pretrained models and modify the first layer to take a 9-channel input.\n\n```\nself.model = timm.create_model('tf_efficientnetv2_l', pretrained=True, num_classes=0)\nself.model.conv_stem = nn.Conv2d(9, 32, kernel_size=(3, 3), stride=(2, 2), padding='same', bias=False)\n```\n\n\n### What didn't work\nThis is a huge list, but I will narrow it down to 3 approaches to predict patient_overall ->\n\n- In the above approach, the sequential information between the slices is only used in the +1/-1 concatenation. To leverage more context, I generated embeddings for all the slices and then tried to predict the `patient_overall` using an LSTM/Transformer.\n\nPossible Reason for not working: This Stage2 model had only 2k data points to learn from, which might not have sufficed.\n\n- Uniformly/Randomly select a fixed number of slices from the patient scans and train a 3D Classifier.\n\nPossible Reason for not working: Fractures can only be seen in some slices and specific slices need to filtered out.\n\n- Trained an XGBoost to to predict patient overall from the fracture probabilities of C1-C7 vertebrae.\n\nPossible Reason for not working: Could not beat the prediction performance of `1-np.prod(1-c1c7)` \n\n\n\n\n\n\n\n\n ",
      "votes": 22
    },
    {
      "id": 2007673,
      "postDate": "2022-10-28T12:17:43.477Z",
      "content": "<p>Really similar approach and not working list. I also read the paper. Maybe next competition we can go up with learning from gold medal solutions. Thanks for sharing your solution and congratulations for silver medal.</p>",
      "rawMarkdown": "Really similar approach and not working list. I also read the paper. Maybe next competition we can go up with learning from gold medal solutions. Thanks for sharing your solution and congratulations for silver medal.",
      "votes": 1,
      "replies": [
        {
          "id": 2007888,
          "postDate": "2022-10-28T15:06:00.797Z",
          "content": "<p>Yes. 3D models look important in a lot of solutions.<br>\nCongratulations to you too!</p>",
          "rawMarkdown": "Yes. 3D models look important in a lot of solutions.\nCongratulations to you too!",
          "votes": 1
        }
      ]
    },
    {
      "id": 2011804,
      "postDate": "2022-10-31T19:55:07.120Z",
      "content": "<p>Thanks for sharing! How did you decide on the specific efficientnet model?</p>",
      "rawMarkdown": "Thanks for sharing! How did you decide on the specific efficientnet model?",
      "replies": [
        {
          "id": 2014969,
          "postDate": "2022-11-03T02:13:11.143Z",
          "content": "<p>I tried training resnet models as well, but they didnt perform as well as efficientnet. </p>",
          "rawMarkdown": "I tried training resnet models as well, but they didnt perform as well as efficientnet. "
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2007673,
      "author_name": "olivepicker",
      "author_url": "",
      "post_date": "2022-10-28T12:17:43.477000",
      "content": "<p>Really similar approach and not working list. I also read the paper. Maybe next competition we can go up with learning from gold medal solutions. Thanks for sharing your solution and congratulations for silver medal.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2007888,
          "author_name": "Yerram Varun",
          "author_url": "",
          "post_date": "2022-10-28T15:06:00.797000",
          "content": "<p>Yes. 3D models look important in a lot of solutions.<br>\nCongratulations to you too!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2011804,
      "author_name": "aspiring",
      "author_url": "",
      "post_date": "2022-10-31T19:55:07.120000",
      "content": "<p>Thanks for sharing! How did you decide on the specific efficientnet model?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2014969,
          "author_name": "Yerram Varun",
          "author_url": "",
          "post_date": "2022-11-03T02:13:11.143000",
          "content": "<p>I tried training resnet models as well, but they didnt perform as well as efficientnet. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2007561": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4630396%2Fddfb714d23d450f8bf5d50fcac1e8ff5%2Frsna_pipeline.png?generation=1666947796450371&alt=media)\n\nThanks to RSNA and Kaggle for conducting such an exciting competition, and congratulations to all the winners!\n\nIt is a modified version of the pipeline shared by @vslaykovsky in [this](https://www.kaggle.com/competitions/rsna-2022-cervical-spine-fracture-detection/discussion/348462) discussion.  \n\n#### CV Strategy \nI use a 5 fold strategy, grouped by StudyInstanceUID and stratified by the fractures in all vertebrae (A column is created concatenating all the fractures like 0_1_0_0_0_1_0 and then used to stratify)\n\n#### Preprocessing\nI use windowing to preprocess the Dicom image used in previous RSNA solutions. The three window values were derived from [this](https://arxiv.org/abs/2010.13336) paper. \n\n```\ndef window(img, WL=50, WW=350):\n    upper, lower = WL+WW//2, WL-WW//2\n    X = np.clip(img.copy(), lower, upper)\n    X = X - np.min(X)\n    X = X / np.max(X)\n    X = (X*255.0).astype('uint8')\n    return X\n\ndef dicom_load(self, siuid, slic):\n    path = f\"data/train_images/{siuid}/{slic}.dcm\"\n    img = dicom.dcmread(path)\n    img.PhotometricInterpretation = 'YBR_FULL'\n    data = img.pixel_array\n    slope = img[('0028','1053')].value\n    intercept = img[('0028','1052')].value\n    data = data*slope + intercept\n            \n    windowed = np.stack((window(data, WL=80, WW=300),\n                        window(data, WL=500, WW=1800),\n                        window(data, WL=400, WW=650)\n                        ), axis=-1)/255.\n    return windowed\n   ```\n\n#### Training Data\n\nEach 3-Channel slice is concatenated with +1/-1 Neighbour slices to create a 9-channel Input. This 9 Channel input predicts the fracture probabilities and vertebrae presence of the middle slice.\n\nDuring inference, this 3-slice window is shifted by stride 2 for faster Inference.\n\n#### Model\n\nI use `Timm` library to use pretrained models and modify the first layer to take a 9-channel input.\n\n```\nself.model = timm.create_model('tf_efficientnetv2_l', pretrained=True, num_classes=0)\nself.model.conv_stem = nn.Conv2d(9, 32, kernel_size=(3, 3), stride=(2, 2), padding='same', bias=False)\n```\n\n\n### What didn't work\nThis is a huge list, but I will narrow it down to 3 approaches to predict patient_overall ->\n\n- In the above approach, the sequential information between the slices is only used in the +1/-1 concatenation. To leverage more context, I generated embeddings for all the slices and then tried to predict the `patient_overall` using an LSTM/Transformer.\n\nPossible Reason for not working: This Stage2 model had only 2k data points to learn from, which might not have sufficed.\n\n- Uniformly/Randomly select a fixed number of slices from the patient scans and train a 3D Classifier.\n\nPossible Reason for not working: Fractures can only be seen in some slices and specific slices need to filtered out.\n\n- Trained an XGBoost to to predict patient overall from the fracture probabilities of C1-C7 vertebrae.\n\nPossible Reason for not working: Could not beat the prediction performance of `1-np.prod(1-c1c7)` \n\n\n\n\n\n\n\n\n ",
    "2007673": "Really similar approach and not working list. I also read the paper. Maybe next competition we can go up with learning from gold medal solutions. Thanks for sharing your solution and congratulations for silver medal.",
    "2011804": "Thanks for sharing! How did you decide on the specific efficientnet model?"
  }
}