{
  "id": 117211,
  "title": "#9 Solution with CODE - Team BIG HEAD: Training model bonanza, TTA and L2 stacking.",
  "url": "/competitions/rsna-intracranial-hemorrhage-detection/discussion/117211",
  "author_name": "Andrés Miguel Torrubia Sáez",
  "post_date": "2019-11-14T00:05:35.533000",
  "votes": 44,
  "comment_count": 23,
  "views": 0,
  "content": "<p><strong>CODE UPDATE</strong></p>\n\n<p>Code is here: <a href=\"http://github.com/antorsae/rsna-intracranial-hemorrhage-detection-team-bighead\">http://github.com/antorsae/rsna-intracranial-hemorrhage-detection-team-bighead</a></p>\n\n<p><strong>SOLUTION OVERVIEW</strong></p>\n\n<p>Our solution consists of pretty weak models (CV 0.07x in most of them) using L2 stacking (5 folds) trained with both <em>xgboost</em> and <em>catboost</em> and ensembled via averaging. </p>\n\n<p>We trained ~50 models (10 architectures/losses * 5 folds) in total.</p>\n\n<p>The following table summarizes the architectures, folds, and GPUs to train each model:</p>\n\n<p><img src=\"https://i.imgur.com/pNaJwFP.png\" alt=\"\"></p>\n\n<p><strong>Fastai v1 3-slice networks: standard window and loss</strong></p>\n\n<p>Architectures not highlighted (first five) were implemented using fastai v1 taking 3 consecutive slices (512x512) of a study and feeding them to the vanilla architecture with a fully connected head that outputs 6*3 = 18 logits. Training is done for 15 epochs using 1-cycle-policy. Batch size is allocated dynamically maximizing GPU memory usage. We use random rotations and flips as augmentation.</p>\n\n<p>Loss function is the weighted average of the 3 slices giving more importance to the center slice:\n```\nW_LOSS = 0.1\nGENERAL_WEIGHTS  = FloatTensor([2., 1., 1., 1., 1., 1.])\ngeneral_weights_3slices = torch.cat([GENERAL_WEIGHTS * W_LOSS, GENERAL_WEIGHTS, GENERAL_WEIGHTS * W_LOSS])</p>\n\n<p>def weighted_loss(pred:Tensor,targ:Tensor)-&gt;Tensor:\n    return F.binary_cross_entropy_with_logits(pred, targ.float(), general_weights_3slices.to(device=pred.device))\n```</p>\n\n<p><strong>Fastai v2 3-slice networks: subdural window and subdural focused loss</strong></p>\n\n<p>We decided to use fastai v2 primarily b/c augmentations are done in GPU and a few of the computers we have had CPU bottlenecks doing augmentations, no longer the case with fastai v2.</p>\n\n<p>Architectures highlighted in red we implemented as above with the following differences:\n- Fastai v2 was used: much of a learning process and still has rough edges (some of them we realized after stage 1 finished and we could NOT change code). \n- Window centered at 100 and width of 254 (to take advantage of the range of <code>uint8</code>)\n- Loss weighted on subdural more (10x) than other types:\n```\nSUBDURAL_WEIGHTS = FloatTensor([.8, .4, .4, .4, .4, 4.])\nsubdural_weights_3slices = torch.cat([SUBDURAL_WEIGHTS * W_LOSS, SUBDURAL_WEIGHTS, SUBDURAL_WEIGHTS * W_LOSS])</p>\n\n<p>def subdural_loss(pred:Tensor,targ:Tensor)-&gt;Tensor:\n    return F.binary_cross_entropy_with_logits(pred, targ.float(), subdural_weights_3slices.to(device=pred.device))\n```</p>\n\n<p><strong>Input to L2 Models</strong></p>\n\n<p>Once models are trained we run OOF predictions using TTA with 10 repetitions. And we use the mean and std of those 10 TTA predictions for each architecture as input to both <em>xgboost</em> and <em>catboost</em>, both for the central and surrounding slices.</p>\n\n<p>Two L2 models are trained: <em>xgboost</em> and <em>catboost</em>, and then simply averaged. One submission we did with the fastai v1 models only (they finished sooner) and the other using both.</p>\n\n<p><strong>Things we would have done differently</strong></p>\n\n<ul>\n<li>Class-aware sampling (balance dataset)</li>\n<li>Pseudo-label training</li>\n<li>Fastai v2 head is different than v1 for vision models, the v1 head works better.</li>\n<li>We used a pretty high <em>eps</em> for <em>Adam</em> optimizer in v2, defaults (in v1) work better.</li>\n<li>Learnable window</li>\n<li>TTA with zoom, crops and cut-out</li>\n<li>Add extra channel with distance to center (similar to coord-conv but just radius to center) to make network location-aware.</li>\n<li>L2 model using lightgbm too and averaging 5 folds of L2 (we trained with 4 folds and hence we did not use all training set for L2)</li>\n</ul>\n\n<p>...and the mandatory meme as a tribute to our team name:\n<img src=\"https://i.imgur.com/1bz50NK.png\" alt=\"\"></p>",
  "messages": [
    {
      "id": 672487,
      "postDate": "2019-11-14T00:05:35.533Z",
      "content": "<p><strong>CODE UPDATE</strong></p>\n\n<p>Code is here: <a href=\"http://github.com/antorsae/rsna-intracranial-hemorrhage-detection-team-bighead\">http://github.com/antorsae/rsna-intracranial-hemorrhage-detection-team-bighead</a></p>\n\n<p><strong>SOLUTION OVERVIEW</strong></p>\n\n<p>Our solution consists of pretty weak models (CV 0.07x in most of them) using L2 stacking (5 folds) trained with both <em>xgboost</em> and <em>catboost</em> and ensembled via averaging. </p>\n\n<p>We trained ~50 models (10 architectures/losses * 5 folds) in total.</p>\n\n<p>The following table summarizes the architectures, folds, and GPUs to train each model:</p>\n\n<p><img src=\"https://i.imgur.com/pNaJwFP.png\" alt=\"\"></p>\n\n<p><strong>Fastai v1 3-slice networks: standard window and loss</strong></p>\n\n<p>Architectures not highlighted (first five) were implemented using fastai v1 taking 3 consecutive slices (512x512) of a study and feeding them to the vanilla architecture with a fully connected head that outputs 6*3 = 18 logits. Training is done for 15 epochs using 1-cycle-policy. Batch size is allocated dynamically maximizing GPU memory usage. We use random rotations and flips as augmentation.</p>\n\n<p>Loss function is the weighted average of the 3 slices giving more importance to the center slice:\n```\nW_LOSS = 0.1\nGENERAL_WEIGHTS  = FloatTensor([2., 1., 1., 1., 1., 1.])\ngeneral_weights_3slices = torch.cat([GENERAL_WEIGHTS * W_LOSS, GENERAL_WEIGHTS, GENERAL_WEIGHTS * W_LOSS])</p>\n\n<p>def weighted_loss(pred:Tensor,targ:Tensor)-&gt;Tensor:\n    return F.binary_cross_entropy_with_logits(pred, targ.float(), general_weights_3slices.to(device=pred.device))\n```</p>\n\n<p><strong>Fastai v2 3-slice networks: subdural window and subdural focused loss</strong></p>\n\n<p>We decided to use fastai v2 primarily b/c augmentations are done in GPU and a few of the computers we have had CPU bottlenecks doing augmentations, no longer the case with fastai v2.</p>\n\n<p>Architectures highlighted in red we implemented as above with the following differences:\n- Fastai v2 was used: much of a learning process and still has rough edges (some of them we realized after stage 1 finished and we could NOT change code). \n- Window centered at 100 and width of 254 (to take advantage of the range of <code>uint8</code>)\n- Loss weighted on subdural more (10x) than other types:\n```\nSUBDURAL_WEIGHTS = FloatTensor([.8, .4, .4, .4, .4, 4.])\nsubdural_weights_3slices = torch.cat([SUBDURAL_WEIGHTS * W_LOSS, SUBDURAL_WEIGHTS, SUBDURAL_WEIGHTS * W_LOSS])</p>\n\n<p>def subdural_loss(pred:Tensor,targ:Tensor)-&gt;Tensor:\n    return F.binary_cross_entropy_with_logits(pred, targ.float(), subdural_weights_3slices.to(device=pred.device))\n```</p>\n\n<p><strong>Input to L2 Models</strong></p>\n\n<p>Once models are trained we run OOF predictions using TTA with 10 repetitions. And we use the mean and std of those 10 TTA predictions for each architecture as input to both <em>xgboost</em> and <em>catboost</em>, both for the central and surrounding slices.</p>\n\n<p>Two L2 models are trained: <em>xgboost</em> and <em>catboost</em>, and then simply averaged. One submission we did with the fastai v1 models only (they finished sooner) and the other using both.</p>\n\n<p><strong>Things we would have done differently</strong></p>\n\n<ul>\n<li>Class-aware sampling (balance dataset)</li>\n<li>Pseudo-label training</li>\n<li>Fastai v2 head is different than v1 for vision models, the v1 head works better.</li>\n<li>We used a pretty high <em>eps</em> for <em>Adam</em> optimizer in v2, defaults (in v1) work better.</li>\n<li>Learnable window</li>\n<li>TTA with zoom, crops and cut-out</li>\n<li>Add extra channel with distance to center (similar to coord-conv but just radius to center) to make network location-aware.</li>\n<li>L2 model using lightgbm too and averaging 5 folds of L2 (we trained with 4 folds and hence we did not use all training set for L2)</li>\n</ul>\n\n<p>...and the mandatory meme as a tribute to our team name:\n<img src=\"https://i.imgur.com/1bz50NK.png\" alt=\"\"></p>",
      "rawMarkdown": "**CODE UPDATE**\n\nCode is here: http://github.com/antorsae/rsna-intracranial-hemorrhage-detection-team-bighead\n\n**SOLUTION OVERVIEW**\n\nOur solution consists of pretty weak models (CV 0.07x in most of them) using L2 stacking (5 folds) trained with both _xgboost_ and _catboost_ and ensembled via averaging. \n\nWe trained ~50 models (10 architectures/losses * 5 folds) in total.\n\nThe following table summarizes the architectures, folds, and GPUs to train each model:\n\n![](https://i.imgur.com/pNaJwFP.png)\n\n\n**Fastai v1 3-slice networks: standard window and loss**\n\nArchitectures not highlighted (first five) were implemented using fastai v1 taking 3 consecutive slices (512x512) of a study and feeding them to the vanilla architecture with a fully connected head that outputs 6*3 = 18 logits. Training is done for 15 epochs using 1-cycle-policy. Batch size is allocated dynamically maximizing GPU memory usage. We use random rotations and flips as augmentation.\n\nLoss function is the weighted average of the 3 slices giving more importance to the center slice:\n```\nW_LOSS = 0.1\nGENERAL_WEIGHTS  = FloatTensor([2., 1., 1., 1., 1., 1.])\ngeneral_weights_3slices = torch.cat([GENERAL_WEIGHTS * W_LOSS, GENERAL_WEIGHTS, GENERAL_WEIGHTS * W_LOSS])\n\ndef weighted_loss(pred:Tensor,targ:Tensor)-&gt;Tensor:\n    return F.binary_cross_entropy_with_logits(pred, targ.float(), general_weights_3slices.to(device=pred.device))\n```\n\n**Fastai v2 3-slice networks: subdural window and subdural focused loss**\n\nWe decided to use fastai v2 primarily b/c augmentations are done in GPU and a few of the computers we have had CPU bottlenecks doing augmentations, no longer the case with fastai v2.\n\nArchitectures highlighted in red we implemented as above with the following differences:\n- Fastai v2 was used: much of a learning process and still has rough edges (some of them we realized after stage 1 finished and we could NOT change code). \n- Window centered at 100 and width of 254 (to take advantage of the range of `uint8`)\n- Loss weighted on subdural more (10x) than other types:\n```\nSUBDURAL_WEIGHTS = FloatTensor([.8, .4, .4, .4, .4, 4.])\nsubdural_weights_3slices = torch.cat([SUBDURAL_WEIGHTS * W_LOSS, SUBDURAL_WEIGHTS, SUBDURAL_WEIGHTS * W_LOSS])\n\ndef subdural_loss(pred:Tensor,targ:Tensor)-&gt;Tensor:\n    return F.binary_cross_entropy_with_logits(pred, targ.float(), subdural_weights_3slices.to(device=pred.device))\n```\n\n**Input to L2 Models**\n\nOnce models are trained we run OOF predictions using TTA with 10 repetitions. And we use the mean and std of those 10 TTA predictions for each architecture as input to both _xgboost_ and _catboost_, both for the central and surrounding slices.\n\nTwo L2 models are trained: _xgboost_ and _catboost_, and then simply averaged. One submission we did with the fastai v1 models only (they finished sooner) and the other using both.\n\n**Things we would have done differently**\n\n- Class-aware sampling (balance dataset)\n- Pseudo-label training\n- Fastai v2 head is different than v1 for vision models, the v1 head works better.\n- We used a pretty high _eps_ for _Adam_ optimizer in v2, defaults (in v1) work better.\n- Learnable window\n- TTA with zoom, crops and cut-out\n- Add extra channel with distance to center (similar to coord-conv but just radius to center) to make network location-aware.\n- L2 model using lightgbm too and averaging 5 folds of L2 (we trained with 4 folds and hence we did not use all training set for L2)\n\n...and the mandatory meme as a tribute to our team name:\n![](https://i.imgur.com/1bz50NK.png)",
      "votes": 44
    },
    {
      "id": 674072,
      "postDate": "2019-11-15T21:34:41.560Z",
      "content": "<p>Regarding the stuff you noticed worked better in fastai v1:</p>\n\n<p>The problem with the head was a mistake by me. I removed the initial batchnorm in the head, without testing that change properly. I thought it wouldn’t have any negative impact, but I was wrong.</p>\n\n<p>I’m interested in the eps issue myself, and also noticed it in this comp. I’m not sure still when it should be high and when low. Perhaps the recent Sadam is the better approach <a href=\"https://arxiv.org/abs/1908.00700v2\">https://arxiv.org/abs/1908.00700v2</a></p>",
      "rawMarkdown": "Regarding the stuff you noticed worked better in fastai v1:\n\nThe problem with the head was a mistake by me. I removed the initial batchnorm in the head, without testing that change properly. I thought it wouldn’t have any negative impact, but I was wrong.\n\nI’m interested in the eps issue myself, and also noticed it in this comp. I’m not sure still when it should be high and when low. Perhaps the recent Sadam is the better approach https://arxiv.org/abs/1908.00700v2",
      "votes": 3,
      "replies": [
        {
          "id": 676041,
          "postDate": "2019-11-18T23:29:18.450Z",
          "content": "<p>The batchnorm is changed back to the fastai v1 way in fastai v2 now :) </p>",
          "rawMarkdown": "The batchnorm is changed back to the fastai v1 way in fastai v2 now :) ",
          "votes": 2
        }
      ]
    },
    {
      "id": 674239,
      "postDate": "2019-11-16T04:04:49.267Z",
      "content": "<p>Did I get you right: you take three 1-channel images (slices), concat them and feeding to the network? If yes, then how did you select exactly 3 images from the whole set (if group by SeriesInstanceUID, there will be more than 3 images)</p>",
      "rawMarkdown": "Did I get you right: you take three 1-channel images (slices), concat them and feeding to the network? If yes, then how did you select exactly 3 images from the whole set (if group by SeriesInstanceUID, there will be more than 3 images)",
      "votes": 1,
      "replies": [
        {
          "id": 674297,
          "postDate": "2019-11-16T07:47:31.100Z",
          "content": "<p>The relative height of each slice within the series is stored in the <code>ImagePositionPatient[2]</code> DICOM attribute. We used that to find the slices just above and below, stacked them together in the R, G, B channels, rescaled them according to the DICOM scale/intercept values, and finally windowed it with center=50 width=100.</p>",
          "rawMarkdown": "The relative height of each slice within the series is stored in the `ImagePositionPatient[2]` DICOM attribute. We used that to find the slices just above and below, stacked them together in the R, G, B channels, rescaled them according to the DICOM scale/intercept values, and finally windowed it with center=50 width=100."
        }
      ]
    },
    {
      "id": 673931,
      "postDate": "2019-11-15T17:05:14.480Z",
      "content": "<p><a href=\"/antorsae\">@antorsae</a> \nthanks for posting,many congrats\nCould u explain me on what slices are, i have reading about it but dint get it so far. \nAnd how u are feeding to network and then coming up with final pred for loss calculation</p>",
      "rawMarkdown": "@antorsae \nthanks for posting,many congrats\nCould u explain me on what slices are, i have reading about it but dint get it so far. \nAnd how u are feeding to network and then coming up with final pred for loss calculation",
      "votes": 1,
      "replies": [
        {
          "id": 673960,
          "postDate": "2019-11-15T17:57:56.627Z",
          "content": "<p>Slice is just an \"image\" (DICOM) in the context of a series of slices that comprise a full brain scan.</p>",
          "rawMarkdown": "Slice is just an \"image\" (DICOM) in the context of a series of slices that comprise a full brain scan."
        },
        {
          "id": 674261,
          "postDate": "2019-11-16T05:39:58.587Z",
          "content": "<p>hi <a href=\"/antorsae\">@antorsae</a> \nyes i get that but unable to figure out how we bunch it with training example and feed to the network and how slice selected for  given patientID image ..\n1) selection of slices\n2) feeding to the network <br>\n3) how summarizing the 18 logits that we get at the output.\nIf above three part if u can help understand would be new learning for me. </p>",
          "rawMarkdown": "hi @antorsae \nyes i get that but unable to figure out how we bunch it with training example and feed to the network and how slice selected for  given patientID image ..\n1) selection of slices\n2) feeding to the network  \n3) how summarizing the 18 logits that we get at the output.\nIf above three part if u can help understand would be new learning for me. "
        },
        {
          "id": 674302,
          "postDate": "2019-11-16T08:07:00.363Z",
          "content": "<ol>\n<li><p>Please refer to my answer <a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/117211#674297\">here</a> on how we selected contiguous slices.</p></li>\n<li><p>We stacked the 3 monochrome images together and fed them to the unmodified network as if they were an RGB image, were the central slice is in the \"G\" channel.</p></li>\n<li><p>We used BCE loss on the 18 outputs, but focused on the 6 central predictions by weighting the two adjacent slices 0.1x that of the middle one. In addition, we weighted the <code>any</code> label twice the others.</p></li>\n</ol>",
          "rawMarkdown": "1. Please refer to my answer [here](https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/117211#674297) on how we selected contiguous slices.\n\n2. We stacked the 3 monochrome images together and fed them to the unmodified network as if they were an RGB image, were the central slice is in the \"G\" channel.\n\n3. We used BCE loss on the 18 outputs, but focused on the 6 central predictions by weighting the two adjacent slices 0.1x that of the middle one. In addition, we weighted the `any` label twice the others."
        }
      ]
    },
    {
      "id": 673540,
      "postDate": "2019-11-15T06:05:02.457Z",
      "content": "<p>congrats on the stellar performance and thank you for the write up! 🙂 </p>",
      "rawMarkdown": "congrats on the stellar performance and thank you for the write up! 🙂 ",
      "votes": 1
    },
    {
      "id": 672839,
      "postDate": "2019-11-14T08:15:12.617Z",
      "content": "<p>Good job,  thank you for sharning and congratulations !! <a href=\"/antorsae\">@antorsae</a> </p>",
      "rawMarkdown": "Good job,  thank you for sharning and congratulations !! @antorsae ",
      "votes": 1
    },
    {
      "id": 672820,
      "postDate": "2019-11-14T07:46:25.310Z",
      "content": "<p>Felicidades</p>",
      "rawMarkdown": "Felicidades",
      "votes": 1
    },
    {
      "id": 672721,
      "postDate": "2019-11-14T04:53:45.363Z",
      "content": "<p>Congratulations\nGreat Write-Up\nThanks for Sharing your Valuable Insights &amp; Approach <a href=\"/antorsae\">@antorsae</a> </p>",
      "rawMarkdown": "Congratulations\nGreat Write-Up\nThanks for Sharing your Valuable Insights &amp; Approach @antorsae ",
      "votes": 1
    },
    {
      "id": 672550,
      "postDate": "2019-11-14T01:35:59.133Z",
      "content": "<p>Congrats <a href=\"/antorsae\">@antorsae</a> and team. </p>",
      "rawMarkdown": "Congrats @antorsae and team. ",
      "votes": 1
    },
    {
      "id": 672533,
      "postDate": "2019-11-14T01:03:53.997Z",
      "content": "<p>Congratulations, and thank you for the wonderful write up!</p>",
      "rawMarkdown": "Congratulations, and thank you for the wonderful write up!",
      "votes": 1
    },
    {
      "id": 672510,
      "postDate": "2019-11-14T00:37:49.513Z",
      "content": "<p>Nice work! Congrats!</p>",
      "rawMarkdown": "Nice work! Congrats!",
      "votes": 1
    },
    {
      "id": 672505,
      "postDate": "2019-11-14T00:29:22.183Z",
      "content": "<p>Nice write up! Which do you think contributed <strong>most</strong> boost in score: L2 models, subdural window &amp; focal loss, using three consecutive slices, some combination or something else?</p>",
      "rawMarkdown": "Nice write up! Which do you think contributed **most** boost in score: L2 models, subdural window &amp; focal loss, using three consecutive slices, some combination or something else?",
      "votes": 1,
      "replies": [
        {
          "id": 672975,
          "postDate": "2019-11-14T10:56:20.493Z",
          "content": "<p>In order:\nL2 models\nTTA\n3 slices\nSudural weight\n(We did not use focal loss, just weighted central slice more than surrounding ones)</p>",
          "rawMarkdown": "In order:\nL2 models\nTTA\n3 slices\nSudural weight\n(We did not use focal loss, just weighted central slice more than surrounding ones)"
        }
      ]
    },
    {
      "id": 672576,
      "postDate": "2019-11-14T02:05:22.437Z",
      "content": "<p>So many GPU...... <em>droool</em>. Congrats to you and the team !!</p>",
      "rawMarkdown": "So many GPU...... *droool*. Congrats to you and the team !!",
      "votes": 2
    },
    {
      "id": 672545,
      "postDate": "2019-11-14T01:30:41.903Z",
      "content": "<p><a href=\"/antorsae\">@antorsae</a> </p>\n\n<p>Congrats! very nice results!</p>\n\n<p>\"taking 3 consecutive slices (512x512) of a study and feeding them to the ....\"\nhow do you order the slides? in my previous post, i note that groupby 'patent id' and 'study instance id' seems insufficient. i end up two overlapping volume. thanks!</p>",
      "rawMarkdown": "@antorsae \n\nCongrats! very nice results!\n\n\"taking 3 consecutive slices (512x512) of a study and feeding them to the ....\"\nhow do you order the slides? in my previous post, i note that groupby 'patent id' and 'study instance id' seems insufficient. i end up two overlapping volume. thanks!",
      "votes": 2,
      "replies": [
        {
          "id": 672971,
          "postDate": "2019-11-14T10:53:40.023Z",
          "content": "<p>Order by ImagePositionPatient[2] (z axis), then group by SeriesInstanceUID:</p>\n\n<p><code>\ngs = df.sort_values('ImagePositionPatient_2').groupby('SeriesInstanceUID')\n</code></p>",
          "rawMarkdown": "Order by ImagePositionPatient[2] (z axis), then group by SeriesInstanceUID:\n\n```\ngs = df.sort_values('ImagePositionPatient_2').groupby('SeriesInstanceUID')\n```"
        },
        {
          "id": 675639,
          "postDate": "2019-11-18T10:57:42Z",
          "content": "<p><a href=\"/bacterio\">@bacterio</a> isnt onee PatientID has got only one StudyIUD and seriesInstanceUID. ?</p>",
          "rawMarkdown": "@bacterio isnt onee PatientID has got only one StudyIUD and seriesInstanceUID. ?"
        },
        {
          "id": 675952,
          "postDate": "2019-11-18T20:14:58.227Z",
          "content": "<p><a href=\"/jaideepvalani\">@jaideepvalani</a>: no, each PatientID may have multiple studies. Assuming you have all DICOM metadata in a pandas dataframe you can check it by doing:</p>\n\n<p><code>\nbgs = df.groupby('PatientID')\ns = bgs.SeriesInstanceUID.nunique()\ns[s&gt;1].head()\n</code></p>\n\n<p>(please look at <a href=\"/radek1\">@radek1</a> 's Fastai starter kit to see how to generate such dataframe).</p>\n\n<p><a href=\"/hengck23\">@hengck23</a>: I saw your thread on overlapping volumes. Apparently there is a one-to-one relationship between <code>StudyInstanceUID</code> and <code>SeriesInstanceUID</code> so it should not matter which one you choose. We did use <code>StudyInstanceUID</code> to find neighboring slices, so our model probably faced the same volume interleaving issue you showed. I wonder how other contestants dealt with this if they did.</p>",
          "rawMarkdown": "@jaideepvalani: no, each PatientID may have multiple studies. Assuming you have all DICOM metadata in a pandas dataframe you can check it by doing:\n\n```\nbgs = df.groupby('PatientID')\ns = bgs.SeriesInstanceUID.nunique()\ns[s&gt;1].head()\n```\n\n(please look at @radek1 's Fastai starter kit to see how to generate such dataframe).\n\n@hengck23: I saw your thread on overlapping volumes. Apparently there is a one-to-one relationship between `StudyInstanceUID` and `SeriesInstanceUID` so it should not matter which one you choose. We did use `StudyInstanceUID` to find neighboring slices, so our model probably faced the same volume interleaving issue you showed. I wonder how other contestants dealt with this if they did."
        }
      ]
    },
    {
      "id": 673438,
      "postDate": "2019-11-15T00:49:15.690Z",
      "content": "<p>Congrats and thank you for sharing.</p>",
      "rawMarkdown": "Congrats and thank you for sharing.",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 674072,
      "author_name": "Jeremy Howard",
      "author_url": "",
      "post_date": "2019-11-15T21:34:41.560000",
      "content": "<p>Regarding the stuff you noticed worked better in fastai v1:</p>\n\n<p>The problem with the head was a mistake by me. I removed the initial batchnorm in the head, without testing that change properly. I thought it wouldn’t have any negative impact, but I was wrong.</p>\n\n<p>I’m interested in the eps issue myself, and also noticed it in this comp. I’m not sure still when it should be high and when low. Perhaps the recent Sadam is the better approach <a href=\"https://arxiv.org/abs/1908.00700v2\">https://arxiv.org/abs/1908.00700v2</a></p>",
      "votes": 3,
      "replies": [
        {
          "id": 676041,
          "author_name": "Jeremy Howard",
          "author_url": "",
          "post_date": "2019-11-18T23:29:18.450000",
          "content": "<p>The batchnorm is changed back to the fastai v1 way in fastai v2 now :) </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 674239,
      "author_name": "Mishunyayev Nikita",
      "author_url": "",
      "post_date": "2019-11-16T04:04:49.267000",
      "content": "<p>Did I get you right: you take three 1-channel images (slices), concat them and feeding to the network? If yes, then how did you select exactly 3 images from the whole set (if group by SeriesInstanceUID, there will be more than 3 images)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 674297,
          "author_name": "Javier Martín",
          "author_url": "",
          "post_date": "2019-11-16T07:47:31.100000",
          "content": "<p>The relative height of each slice within the series is stored in the <code>ImagePositionPatient[2]</code> DICOM attribute. We used that to find the slices just above and below, stacked them together in the R, G, B channels, rescaled them according to the DICOM scale/intercept values, and finally windowed it with center=50 width=100.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 673931,
      "author_name": "Jaideep",
      "author_url": "",
      "post_date": "2019-11-15T17:05:14.480000",
      "content": "<p><a href=\"/antorsae\">@antorsae</a> \nthanks for posting,many congrats\nCould u explain me on what slices are, i have reading about it but dint get it so far. \nAnd how u are feeding to network and then coming up with final pred for loss calculation</p>",
      "votes": 1,
      "replies": [
        {
          "id": 673960,
          "author_name": "Andrés Miguel Torrubia Sáez",
          "author_url": "",
          "post_date": "2019-11-15T17:57:56.627000",
          "content": "<p>Slice is just an \"image\" (DICOM) in the context of a series of slices that comprise a full brain scan.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 674261,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2019-11-16T05:39:58.587000",
          "content": "<p>hi <a href=\"/antorsae\">@antorsae</a> \nyes i get that but unable to figure out how we bunch it with training example and feed to the network and how slice selected for  given patientID image ..\n1) selection of slices\n2) feeding to the network <br>\n3) how summarizing the 18 logits that we get at the output.\nIf above three part if u can help understand would be new learning for me. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 674302,
          "author_name": "Javier Martín",
          "author_url": "",
          "post_date": "2019-11-16T08:07:00.363000",
          "content": "<ol>\n<li><p>Please refer to my answer <a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/117211#674297\">here</a> on how we selected contiguous slices.</p></li>\n<li><p>We stacked the 3 monochrome images together and fed them to the unmodified network as if they were an RGB image, were the central slice is in the \"G\" channel.</p></li>\n<li><p>We used BCE loss on the 18 outputs, but focused on the 6 central predictions by weighting the two adjacent slices 0.1x that of the middle one. In addition, we weighted the <code>any</code> label twice the others.</p></li>\n</ol>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 673540,
      "author_name": "Radek Osmulski",
      "author_url": "",
      "post_date": "2019-11-15T06:05:02.457000",
      "content": "<p>congrats on the stellar performance and thank you for the write up! 🙂 </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 672839,
      "author_name": "KhanhVD",
      "author_url": "",
      "post_date": "2019-11-14T08:15:12.617000",
      "content": "<p>Good job,  thank you for sharning and congratulations !! <a href=\"/antorsae\">@antorsae</a> </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 672820,
      "author_name": "Santiago Mota",
      "author_url": "",
      "post_date": "2019-11-14T07:46:25.310000",
      "content": "<p>Felicidades</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 672721,
      "author_name": "Ailurophile",
      "author_url": "",
      "post_date": "2019-11-14T04:53:45.363000",
      "content": "<p>Congratulations\nGreat Write-Up\nThanks for Sharing your Valuable Insights &amp; Approach <a href=\"/antorsae\">@antorsae</a> </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 672550,
      "author_name": "YaGana Sheriff-Hussaini",
      "author_url": "",
      "post_date": "2019-11-14T01:35:59.133000",
      "content": "<p>Congrats <a href=\"/antorsae\">@antorsae</a> and team. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 672533,
      "author_name": "Hilal Shaath",
      "author_url": "",
      "post_date": "2019-11-14T01:03:53.997000",
      "content": "<p>Congratulations, and thank you for the wonderful write up!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 672510,
      "author_name": "Yee Ng",
      "author_url": "",
      "post_date": "2019-11-14T00:37:49.513000",
      "content": "<p>Nice work! Congrats!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 672505,
      "author_name": "Tim Yee",
      "author_url": "",
      "post_date": "2019-11-14T00:29:22.183000",
      "content": "<p>Nice write up! Which do you think contributed <strong>most</strong> boost in score: L2 models, subdural window &amp; focal loss, using three consecutive slices, some combination or something else?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 672975,
          "author_name": "Andrés Miguel Torrubia Sáez",
          "author_url": "",
          "post_date": "2019-11-14T10:56:20.493000",
          "content": "<p>In order:\nL2 models\nTTA\n3 slices\nSudural weight\n(We did not use focal loss, just weighted central slice more than surrounding ones)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 672576,
      "author_name": "yukiya",
      "author_url": "",
      "post_date": "2019-11-14T02:05:22.437000",
      "content": "<p>So many GPU...... <em>droool</em>. Congrats to you and the team !!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 672545,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2019-11-14T01:30:41.903000",
      "content": "<p><a href=\"/antorsae\">@antorsae</a> </p>\n\n<p>Congrats! very nice results!</p>\n\n<p>\"taking 3 consecutive slices (512x512) of a study and feeding them to the ....\"\nhow do you order the slides? in my previous post, i note that groupby 'patent id' and 'study instance id' seems insufficient. i end up two overlapping volume. thanks!</p>",
      "votes": 2,
      "replies": [
        {
          "id": 672971,
          "author_name": "Javier Martín",
          "author_url": "",
          "post_date": "2019-11-14T10:53:40.023000",
          "content": "<p>Order by ImagePositionPatient[2] (z axis), then group by SeriesInstanceUID:</p>\n\n<p><code>\ngs = df.sort_values('ImagePositionPatient_2').groupby('SeriesInstanceUID')\n</code></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 675639,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2019-11-18T10:57:42",
          "content": "<p><a href=\"/bacterio\">@bacterio</a> isnt onee PatientID has got only one StudyIUD and seriesInstanceUID. ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 675952,
          "author_name": "Javier Martín",
          "author_url": "",
          "post_date": "2019-11-18T20:14:58.227000",
          "content": "<p><a href=\"/jaideepvalani\">@jaideepvalani</a>: no, each PatientID may have multiple studies. Assuming you have all DICOM metadata in a pandas dataframe you can check it by doing:</p>\n\n<p><code>\nbgs = df.groupby('PatientID')\ns = bgs.SeriesInstanceUID.nunique()\ns[s&gt;1].head()\n</code></p>\n\n<p>(please look at <a href=\"/radek1\">@radek1</a> 's Fastai starter kit to see how to generate such dataframe).</p>\n\n<p><a href=\"/hengck23\">@hengck23</a>: I saw your thread on overlapping volumes. Apparently there is a one-to-one relationship between <code>StudyInstanceUID</code> and <code>SeriesInstanceUID</code> so it should not matter which one you choose. We did use <code>StudyInstanceUID</code> to find neighboring slices, so our model probably faced the same volume interleaving issue you showed. I wonder how other contestants dealt with this if they did.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 673438,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "2019-11-15T00:49:15.690000",
      "content": "<p>Congrats and thank you for sharing.</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "672487": "**CODE UPDATE**\n\nCode is here: http://github.com/antorsae/rsna-intracranial-hemorrhage-detection-team-bighead\n\n**SOLUTION OVERVIEW**\n\nOur solution consists of pretty weak models (CV 0.07x in most of them) using L2 stacking (5 folds) trained with both _xgboost_ and _catboost_ and ensembled via averaging. \n\nWe trained ~50 models (10 architectures/losses * 5 folds) in total.\n\nThe following table summarizes the architectures, folds, and GPUs to train each model:\n\n![](https://i.imgur.com/pNaJwFP.png)\n\n\n**Fastai v1 3-slice networks: standard window and loss**\n\nArchitectures not highlighted (first five) were implemented using fastai v1 taking 3 consecutive slices (512x512) of a study and feeding them to the vanilla architecture with a fully connected head that outputs 6*3 = 18 logits. Training is done for 15 epochs using 1-cycle-policy. Batch size is allocated dynamically maximizing GPU memory usage. We use random rotations and flips as augmentation.\n\nLoss function is the weighted average of the 3 slices giving more importance to the center slice:\n```\nW_LOSS = 0.1\nGENERAL_WEIGHTS  = FloatTensor([2., 1., 1., 1., 1., 1.])\ngeneral_weights_3slices = torch.cat([GENERAL_WEIGHTS * W_LOSS, GENERAL_WEIGHTS, GENERAL_WEIGHTS * W_LOSS])\n\ndef weighted_loss(pred:Tensor,targ:Tensor)-&gt;Tensor:\n    return F.binary_cross_entropy_with_logits(pred, targ.float(), general_weights_3slices.to(device=pred.device))\n```\n\n**Fastai v2 3-slice networks: subdural window and subdural focused loss**\n\nWe decided to use fastai v2 primarily b/c augmentations are done in GPU and a few of the computers we have had CPU bottlenecks doing augmentations, no longer the case with fastai v2.\n\nArchitectures highlighted in red we implemented as above with the following differences:\n- Fastai v2 was used: much of a learning process and still has rough edges (some of them we realized after stage 1 finished and we could NOT change code). \n- Window centered at 100 and width of 254 (to take advantage of the range of `uint8`)\n- Loss weighted on subdural more (10x) than other types:\n```\nSUBDURAL_WEIGHTS = FloatTensor([.8, .4, .4, .4, .4, 4.])\nsubdural_weights_3slices = torch.cat([SUBDURAL_WEIGHTS * W_LOSS, SUBDURAL_WEIGHTS, SUBDURAL_WEIGHTS * W_LOSS])\n\ndef subdural_loss(pred:Tensor,targ:Tensor)-&gt;Tensor:\n    return F.binary_cross_entropy_with_logits(pred, targ.float(), subdural_weights_3slices.to(device=pred.device))\n```\n\n**Input to L2 Models**\n\nOnce models are trained we run OOF predictions using TTA with 10 repetitions. And we use the mean and std of those 10 TTA predictions for each architecture as input to both _xgboost_ and _catboost_, both for the central and surrounding slices.\n\nTwo L2 models are trained: _xgboost_ and _catboost_, and then simply averaged. One submission we did with the fastai v1 models only (they finished sooner) and the other using both.\n\n**Things we would have done differently**\n\n- Class-aware sampling (balance dataset)\n- Pseudo-label training\n- Fastai v2 head is different than v1 for vision models, the v1 head works better.\n- We used a pretty high _eps_ for _Adam_ optimizer in v2, defaults (in v1) work better.\n- Learnable window\n- TTA with zoom, crops and cut-out\n- Add extra channel with distance to center (similar to coord-conv but just radius to center) to make network location-aware.\n- L2 model using lightgbm too and averaging 5 folds of L2 (we trained with 4 folds and hence we did not use all training set for L2)\n\n...and the mandatory meme as a tribute to our team name:\n![](https://i.imgur.com/1bz50NK.png)",
    "674072": "Regarding the stuff you noticed worked better in fastai v1:\n\nThe problem with the head was a mistake by me. I removed the initial batchnorm in the head, without testing that change properly. I thought it wouldn’t have any negative impact, but I was wrong.\n\nI’m interested in the eps issue myself, and also noticed it in this comp. I’m not sure still when it should be high and when low. Perhaps the recent Sadam is the better approach https://arxiv.org/abs/1908.00700v2",
    "674239": "Did I get you right: you take three 1-channel images (slices), concat them and feeding to the network? If yes, then how did you select exactly 3 images from the whole set (if group by SeriesInstanceUID, there will be more than 3 images)",
    "673931": "@antorsae \nthanks for posting,many congrats\nCould u explain me on what slices are, i have reading about it but dint get it so far. \nAnd how u are feeding to network and then coming up with final pred for loss calculation",
    "673540": "congrats on the stellar performance and thank you for the write up! 🙂 ",
    "672839": "Good job,  thank you for sharning and congratulations !! @antorsae ",
    "672820": "Felicidades",
    "672721": "Congratulations\nGreat Write-Up\nThanks for Sharing your Valuable Insights &amp; Approach @antorsae ",
    "672550": "Congrats @antorsae and team. ",
    "672533": "Congratulations, and thank you for the wonderful write up!",
    "672510": "Nice work! Congrats!",
    "672505": "Nice write up! Which do you think contributed **most** boost in score: L2 models, subdural window &amp; focal loss, using three consecutive slices, some combination or something else?",
    "672576": "So many GPU...... *droool*. Congrats to you and the team !!",
    "672545": "@antorsae \n\nCongrats! very nice results!\n\n\"taking 3 consecutive slices (512x512) of a study and feeding them to the ....\"\nhow do you order the slides? in my previous post, i note that groupby 'patent id' and 'study instance id' seems insufficient. i end up two overlapping volume. thanks!",
    "673438": "Congrats and thank you for sharing."
  }
}