{
  "id": 117480,
  "title": "A surprise Gold, GM, and the real 12th place solution*",
  "url": "/competitions/rsna-intracranial-hemorrhage-detection/discussion/117480",
  "author_name": "yuval reina",
  "post_date": "2019-11-15T20:43:24.431000",
  "votes": 27,
  "comment_count": 6,
  "views": 0,
  "content": "<p>First I want to thank my teammate <a href=\"/zaharch\">@zaharch</a> for a great and successful teamwork (this is our 2nd consecutive gold). Until yesterday we had the top Silver medal - 13th place, and today, for some mysterious reason, we got promoted to 12th and Gold. For me this Gold also means GM (5 golds in the last 5 competitions). \n^ real because <a href=\"/appian\">@appian</a> was promoted to 11th</p>\n\n<p>So lets go for our solution.</p>\n\n<p>As most/all the top solutions we also used a two stage solution:\n1. Base model for feature extraction per image\n2. Shallow model - combining all the output features from a full head scan to predict per image.\nThe 2nd stage also included some post - processing and ensembling.</p>\n\n<h2>Base Model:</h2>\n\n<p>As base model we used a few different models:\n* Densenet 169, 161, 201\n* SE-ResNet101\n* SE - ResNeXt101_32x4d \nFor all models we used 3 folds, for the SE models we also had 5 folds.\nThe SE models with 5 folds gave the best results. \nThe  models where trained for ~4 epochs using the usual augmentations: rotation, flip, zoom, position shift, pixel intensity shift.</p>\n\n<h3>WSO</h3>\n\n<p>As many of the other teams do in their base solutions, we also used 3  windows to handle the large dynamic range of the CT pixels values,  but instead of using fixed windows we let the network find the best windows, as described in <a href=\"https://arxiv.org/pdf/1812.00572.pdf\">Practical Window Setting Optimization for Medical Image Deep Learning</a>. \nThe implementation is quit straight forward: <br>\nadding 3 layers in front of the model:\n<code>\nConv2d(1, 3, kernel_size=(1, 1))\nSigmoid()\nInstanceNorm2d(3)\n</code>\nThe convolution layer was initialize with the soft- tissue, blood, bone values. At the end this layer converged to value close to the usual windows values.</p>\n\n<h3>Feature pooling</h3>\n\n<p>Most of the features at the last layer where zero - we used 8 times pooling to decrease the number of features to ~ 250-300</p>\n\n<h3>TTA</h3>\n\n<p>We created 4 sets of features from augmented images for each train image and 8 sets for each test image.</p>\n\n<h2>Shallow Network</h2>\n\n<p>We used two different shallow networks. (I will describe one here and <a href=\"/zaharch\">@zaharch</a> will describe the 2nd later)</p>\n\n<p>One network was a FCN.</p>\n\n<h3>Input</h3>\n\n<p>Features from all the images of one full head scan, ordered by the Z position.</p>\n\n<h3>Layer</h3>\n\n<ol>\n<li>9 * Num_features 2D convolution - the output is batch_size * num_images * num _channels *  1 * 6</li>\n<li>Squeeze</li>\n<li>1D convolution layer of size 7</li>\n<li>1D convolution layer of size 5</li>\n<li>1D convolution layer of size 3\nWith batch norms, and ReLUs in between.</li>\n</ol>\n\n<p>We trained the shallow network with the TTAed features from the base model and for prediction we used the test features TTA.</p>\n\n<h2>Post Processing and ensembling:</h2>\n\n<p><a href=\"/zaharch\">@zaharch</a> will and as a comment</p>\n\n<h2>Results:</h2>\n\n<p>The base models gave 0.68 - 0.66 on LB (first stage) after fold averaging and with TTA averaging. ** We started using the better models after we already had the shallow network, hence we didn't really submitted a full 5 fold average of base model and the numbers are derived from CV.\nThe best single 5 fold full model (base + shallow), with TTA and fold averaging was the SE-ResNet101 which gave LB 0.6 (first stage). </p>\n\n<p>One drawback we had - we didn't gain much by ensembling many models, maybe we should have used one model and run it more with different seeds.  </p>\n\n<p>And as a last word, I want to thank the organizers and moderators <a href=\"/juliaelliott\">@juliaelliott</a> <a href=\"/philculliton\">@philculliton</a> <a href=\"/lechuck0\">@lechuck0</a> for a great competition and for being flexible and changing the rules to let us use the  metadata which helped all the top teams get really good and interesting solutions. </p>\n\n<h3>Code</h3>\n\n<p><a href=\"https://github.com/nosound2/RSNA-Hemorrhage\">The full code can be found here</a></p>\n\n<h3>More information</h3>\n\n<p>More information about our models can also be found in the following files</p>\n\n<p><a href=\"https://docs.google.com/document/d/1YFwbnmh5QDF77th01eSEqscvb4rWTKyvL0XWSWs-sMg/edit?usp=sharing\">Documentation</a></p>\n\n<p><a href=\"https://drive.google.com/file/d/1Kz_3mkA9volBKNau_u_jzE2YQhkPTIXp/view?usp=sharing\">Presentation</a></p>\n\n<p><a href=\"https://drive.google.com/file/d/1yX6WC9GysdekivPzowU685AeAuE_LHXI/view?usp=sharing\">Video</a></p>",
  "messages": [
    {
      "id": 674039,
      "postDate": "2019-11-15T20:43:24.433Z",
      "content": "<p>First I want to thank my teammate <a href=\"/zaharch\">@zaharch</a> for a great and successful teamwork (this is our 2nd consecutive gold). Until yesterday we had the top Silver medal - 13th place, and today, for some mysterious reason, we got promoted to 12th and Gold. For me this Gold also means GM (5 golds in the last 5 competitions). \n^ real because <a href=\"/appian\">@appian</a> was promoted to 11th</p>\n\n<p>So lets go for our solution.</p>\n\n<p>As most/all the top solutions we also used a two stage solution:\n1. Base model for feature extraction per image\n2. Shallow model - combining all the output features from a full head scan to predict per image.\nThe 2nd stage also included some post - processing and ensembling.</p>\n\n<h2>Base Model:</h2>\n\n<p>As base model we used a few different models:\n* Densenet 169, 161, 201\n* SE-ResNet101\n* SE - ResNeXt101_32x4d \nFor all models we used 3 folds, for the SE models we also had 5 folds.\nThe SE models with 5 folds gave the best results. \nThe  models where trained for ~4 epochs using the usual augmentations: rotation, flip, zoom, position shift, pixel intensity shift.</p>\n\n<h3>WSO</h3>\n\n<p>As many of the other teams do in their base solutions, we also used 3  windows to handle the large dynamic range of the CT pixels values,  but instead of using fixed windows we let the network find the best windows, as described in <a href=\"https://arxiv.org/pdf/1812.00572.pdf\">Practical Window Setting Optimization for Medical Image Deep Learning</a>. \nThe implementation is quit straight forward: <br>\nadding 3 layers in front of the model:\n<code>\nConv2d(1, 3, kernel_size=(1, 1))\nSigmoid()\nInstanceNorm2d(3)\n</code>\nThe convolution layer was initialize with the soft- tissue, blood, bone values. At the end this layer converged to value close to the usual windows values.</p>\n\n<h3>Feature pooling</h3>\n\n<p>Most of the features at the last layer where zero - we used 8 times pooling to decrease the number of features to ~ 250-300</p>\n\n<h3>TTA</h3>\n\n<p>We created 4 sets of features from augmented images for each train image and 8 sets for each test image.</p>\n\n<h2>Shallow Network</h2>\n\n<p>We used two different shallow networks. (I will describe one here and <a href=\"/zaharch\">@zaharch</a> will describe the 2nd later)</p>\n\n<p>One network was a FCN.</p>\n\n<h3>Input</h3>\n\n<p>Features from all the images of one full head scan, ordered by the Z position.</p>\n\n<h3>Layer</h3>\n\n<ol>\n<li>9 * Num_features 2D convolution - the output is batch_size * num_images * num _channels *  1 * 6</li>\n<li>Squeeze</li>\n<li>1D convolution layer of size 7</li>\n<li>1D convolution layer of size 5</li>\n<li>1D convolution layer of size 3\nWith batch norms, and ReLUs in between.</li>\n</ol>\n\n<p>We trained the shallow network with the TTAed features from the base model and for prediction we used the test features TTA.</p>\n\n<h2>Post Processing and ensembling:</h2>\n\n<p><a href=\"/zaharch\">@zaharch</a> will and as a comment</p>\n\n<h2>Results:</h2>\n\n<p>The base models gave 0.68 - 0.66 on LB (first stage) after fold averaging and with TTA averaging. ** We started using the better models after we already had the shallow network, hence we didn't really submitted a full 5 fold average of base model and the numbers are derived from CV.\nThe best single 5 fold full model (base + shallow), with TTA and fold averaging was the SE-ResNet101 which gave LB 0.6 (first stage). </p>\n\n<p>One drawback we had - we didn't gain much by ensembling many models, maybe we should have used one model and run it more with different seeds.  </p>\n\n<p>And as a last word, I want to thank the organizers and moderators <a href=\"/juliaelliott\">@juliaelliott</a> <a href=\"/philculliton\">@philculliton</a> <a href=\"/lechuck0\">@lechuck0</a> for a great competition and for being flexible and changing the rules to let us use the  metadata which helped all the top teams get really good and interesting solutions. </p>\n\n<h3>Code</h3>\n\n<p><a href=\"https://github.com/nosound2/RSNA-Hemorrhage\">The full code can be found here</a></p>\n\n<h3>More information</h3>\n\n<p>More information about our models can also be found in the following files</p>\n\n<p><a href=\"https://docs.google.com/document/d/1YFwbnmh5QDF77th01eSEqscvb4rWTKyvL0XWSWs-sMg/edit?usp=sharing\">Documentation</a></p>\n\n<p><a href=\"https://drive.google.com/file/d/1Kz_3mkA9volBKNau_u_jzE2YQhkPTIXp/view?usp=sharing\">Presentation</a></p>\n\n<p><a href=\"https://drive.google.com/file/d/1yX6WC9GysdekivPzowU685AeAuE_LHXI/view?usp=sharing\">Video</a></p>",
      "rawMarkdown": "First I want to thank my teammate @zaharch for a great and successful teamwork (this is our 2nd consecutive gold). Until yesterday we had the top Silver medal - 13th place, and today, for some mysterious reason, we got promoted to 12th and Gold. For me this Gold also means GM (5 golds in the last 5 competitions). \n^ real because @appian was promoted to 11th\n\nSo lets go for our solution.\n\nAs most/all the top solutions we also used a two stage solution:\n1. Base model for feature extraction per image\n2. Shallow model - combining all the output features from a full head scan to predict per image.\nThe 2nd stage also included some post - processing and ensembling.\n\n## Base Model:\nAs base model we used a few different models:\n* Densenet 169, 161, 201\n* SE-ResNet101\n* SE - ResNeXt101_32x4d \nFor all models we used 3 folds, for the SE models we also had 5 folds.\nThe SE models with 5 folds gave the best results. \nThe  models where trained for ~4 epochs using the usual augmentations: rotation, flip, zoom, position shift, pixel intensity shift.\n\n### WSO\nAs many of the other teams do in their base solutions, we also used 3  windows to handle the large dynamic range of the CT pixels values,  but instead of using fixed windows we let the network find the best windows, as described in [Practical Window Setting Optimization for Medical Image Deep Learning](https://arxiv.org/pdf/1812.00572.pdf). \nThe implementation is quit straight forward:  \nadding 3 layers in front of the model:\n```\nConv2d(1, 3, kernel_size=(1, 1))\nSigmoid()\nInstanceNorm2d(3)\n```\nThe convolution layer was initialize with the soft- tissue, blood, bone values. At the end this layer converged to value close to the usual windows values.\n\n### Feature pooling\nMost of the features at the last layer where zero - we used 8 times pooling to decrease the number of features to ~ 250-300\n\n### TTA\nWe created 4 sets of features from augmented images for each train image and 8 sets for each test image.\n\n## Shallow Network\nWe used two different shallow networks. (I will describe one here and @zaharch will describe the 2nd later)\n\nOne network was a FCN.\n### Input \nFeatures from all the images of one full head scan, ordered by the Z position.\n### Layer\n1. 9 * Num_features 2D convolution - the output is batch_size * num_images * num _channels *  1 * 6\n2. Squeeze\n3. 1D convolution layer of size 7\n4. 1D convolution layer of size 5\n5. 1D convolution layer of size 3\nWith batch norms, and ReLUs in between.\n\nWe trained the shallow network with the TTAed features from the base model and for prediction we used the test features TTA.\n\n## Post Processing and ensembling:\n@zaharch will and as a comment\n\n## Results:\nThe base models gave 0.68 - 0.66 on LB (first stage) after fold averaging and with TTA averaging. ** We started using the better models after we already had the shallow network, hence we didn't really submitted a full 5 fold average of base model and the numbers are derived from CV.\nThe best single 5 fold full model (base + shallow), with TTA and fold averaging was the SE-ResNet101 which gave LB 0.6 (first stage). \n\nOne drawback we had - we didn't gain much by ensembling many models, maybe we should have used one model and run it more with different seeds.  \n\n\nAnd as a last word, I want to thank the organizers and moderators @juliaelliott @philculliton @lechuck0 for a great competition and for being flexible and changing the rules to let us use the  metadata which helped all the top teams get really good and interesting solutions. \n\n### Code\n[The full code can be found here](https://github.com/nosound2/RSNA-Hemorrhage)\n\n### More information\n\nMore information about our models can also be found in the following files\n\n[Documentation](https://docs.google.com/document/d/1YFwbnmh5QDF77th01eSEqscvb4rWTKyvL0XWSWs-sMg/edit?usp=sharing)\n\n[Presentation](https://drive.google.com/file/d/1Kz_3mkA9volBKNau_u_jzE2YQhkPTIXp/view?usp=sharing)\n\n[Video](https://drive.google.com/file/d/1yX6WC9GysdekivPzowU685AeAuE_LHXI/view?usp=sharing)\n\n\n",
      "votes": 27
    },
    {
      "id": 674122,
      "postDate": "2019-11-15T22:59:54.290Z",
      "content": "<p>Thank you, Yuval, for the teamwork, it was an honor working with you. Congratulations with becoming a GM, it was an impressive 5 consecutive golds run!</p>\n\n<p>We published our <a href=\"https://github.com/nosound2/RSNA-Hemorrhage\">full code on github</a>.</p>\n\n<p>In this competition I concentrated on metadata, second step shallow networks and ensembling. </p>\n\n<h2>Metadata</h2>\n\n<p>I parsed all available metadata fields and hand-crafted 22 float, 16 boolean and 10 categorical variables from it. It had a lot of information by itself, but added no value for us above the images information. </p>\n\n<h2>Shallow Network</h2>\n\n<p>The shallow neural network input is <code>(256+num_meta_feats)x60</code>, where 60 is the maximum number of slices per series, and 256 is the backbone output features per image. The first layer was <code>nn.Conv2d(1,64,(256+num_meta_feats,1))</code> followed by 12 <code>nn.Conv1d</code> layers with densenet-inspired skip connections. This shallow network architecture had similar performance to that of Yuval's architecture that he described above.</p>\n\n<h2>Ensembling</h2>\n\n<p>At this stage we had about 15 shallow networks results, from Yuval and mine, running on features from different base models. After trying several more complex ideas which overfitted, we settled on very simple ensembling, which finally added very little above simple averaging of all good models. It was basically a linear combination between Yuval's and mine model means optimized to minimize the log-loss, per target. </p>\n\n<h2>Weights approach</h2>\n\n<p>We noticed that the stage 1 test data is very different from the train data if looked from some metadata fields perspective. For example below is a histogram of mean pixel values for original train vs stage 1 and stage 2 test. There are apparently different groups, which have very different ratios in train vs stage 1. Following this idea we assigned a weight to each series, which assures that weighted histograms of metadata look similar, and also <code>any</code> target is distributed similarly. We used our second final submission on that weighted version, where shallow NNs were trained with weighted sampling, and ensembling was done with weighted loss. It gave a nice boost of about 0.001 on CV and stage 1 LB.</p>\n\n<p>Unfortunately, it was absolutely useless for the second stage, because the stage 2 test data happened to be super similar to the original train distribution, see the picture below. We ended up with two very similar submissions for the second stage.</p>\n\n<p><img src=\"https://imgur.com/SIhpUon.png\" alt=\"Weights\"></p>\n\n<h2>Things that didn't work</h2>\n\n<p>All of the above! But we learn a lot from failures, it was a great competition overall.</p>\n\n<h2>Special thanks</h2>\n\n<p>I want to thank Kaggle and Google for distributing GCP credits for competitions, all my shallow networks I trained on GCP. It could have been very hard for me to participate without it.</p>",
      "rawMarkdown": "Thank you, Yuval, for the teamwork, it was an honor working with you. Congratulations with becoming a GM, it was an impressive 5 consecutive golds run!\n\nWe published our [full code on github](https://github.com/nosound2/RSNA-Hemorrhage).\n\nIn this competition I concentrated on metadata, second step shallow networks and ensembling. \n\n## Metadata\n\nI parsed all available metadata fields and hand-crafted 22 float, 16 boolean and 10 categorical variables from it. It had a lot of information by itself, but added no value for us above the images information. \n\n## Shallow Network\n\nThe shallow neural network input is `(256+num_meta_feats)x60`, where 60 is the maximum number of slices per series, and 256 is the backbone output features per image. The first layer was `nn.Conv2d(1,64,(256+num_meta_feats,1))` followed by 12 `nn.Conv1d` layers with densenet-inspired skip connections. This shallow network architecture had similar performance to that of Yuval's architecture that he described above.\n\n## Ensembling\n\nAt this stage we had about 15 shallow networks results, from Yuval and mine, running on features from different base models. After trying several more complex ideas which overfitted, we settled on very simple ensembling, which finally added very little above simple averaging of all good models. It was basically a linear combination between Yuval's and mine model means optimized to minimize the log-loss, per target. \n\n## Weights approach\n\nWe noticed that the stage 1 test data is very different from the train data if looked from some metadata fields perspective. For example below is a histogram of mean pixel values for original train vs stage 1 and stage 2 test. There are apparently different groups, which have very different ratios in train vs stage 1. Following this idea we assigned a weight to each series, which assures that weighted histograms of metadata look similar, and also `any` target is distributed similarly. We used our second final submission on that weighted version, where shallow NNs were trained with weighted sampling, and ensembling was done with weighted loss. It gave a nice boost of about 0.001 on CV and stage 1 LB.\n\nUnfortunately, it was absolutely useless for the second stage, because the stage 2 test data happened to be super similar to the original train distribution, see the picture below. We ended up with two very similar submissions for the second stage.\n\n![Weights](https://imgur.com/SIhpUon.png)\n\n## Things that didn't work\n\nAll of the above! But we learn a lot from failures, it was a great competition overall.\n\n## Special thanks\n\nI want to thank Kaggle and Google for distributing GCP credits for competitions, all my shallow networks I trained on GCP. It could have been very hard for me to participate without it.",
      "votes": 12,
      "replies": [
        {
          "id": 674231,
          "postDate": "2019-11-16T03:38:05.407Z",
          "content": "<p>Congratulations... Thanks for sharing Code &amp; Insights <a href=\"/zaharch\">@zaharch</a> </p>",
          "rawMarkdown": "Congratulations... Thanks for sharing Code &amp; Insights @zaharch ",
          "votes": 1
        }
      ]
    },
    {
      "id": 674134,
      "postDate": "2019-11-15T23:32:30.750Z",
      "content": "<p>Congratulation <a href=\"/yuval6967\">@yuval6967</a> <a href=\"/zaharch\">@zaharch</a> . Always impressive solution from you guys! <a href=\"/zaharch\">@zaharch</a> You can have a 5 consecutive golds too!!</p>",
      "rawMarkdown": "Congratulation @yuval6967 @zaharch . Always impressive solution from you guys! @zaharch You can have a 5 consecutive golds too!!",
      "votes": 3,
      "replies": [
        {
          "id": 674150,
          "postDate": "2019-11-16T00:00:30.577Z",
          "content": "<p><a href=\"/zaharch\">@zaharch</a> surly can!</p>",
          "rawMarkdown": "@zaharch surly can!",
          "votes": 3
        }
      ]
    },
    {
      "id": 675585,
      "postDate": "2019-11-18T09:08:07.567Z",
      "content": "<p>congratulations team.👍  Great write up. </p>",
      "rawMarkdown": "congratulations team.👍  Great write up. ",
      "votes": 1
    },
    {
      "id": 674230,
      "postDate": "2019-11-16T03:37:09.523Z",
      "content": "<p>Congratulations\nGreat Write-Up\nThanks for Sharing your Valuable Insights &amp; Approach <a href=\"/yuval6967\">@yuval6967</a> </p>",
      "rawMarkdown": "Congratulations\nGreat Write-Up\nThanks for Sharing your Valuable Insights &amp; Approach @yuval6967 ",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 674122,
      "author_name": "nosound",
      "author_url": "",
      "post_date": "2019-11-15T22:59:54.290000",
      "content": "<p>Thank you, Yuval, for the teamwork, it was an honor working with you. Congratulations with becoming a GM, it was an impressive 5 consecutive golds run!</p>\n\n<p>We published our <a href=\"https://github.com/nosound2/RSNA-Hemorrhage\">full code on github</a>.</p>\n\n<p>In this competition I concentrated on metadata, second step shallow networks and ensembling. </p>\n\n<h2>Metadata</h2>\n\n<p>I parsed all available metadata fields and hand-crafted 22 float, 16 boolean and 10 categorical variables from it. It had a lot of information by itself, but added no value for us above the images information. </p>\n\n<h2>Shallow Network</h2>\n\n<p>The shallow neural network input is <code>(256+num_meta_feats)x60</code>, where 60 is the maximum number of slices per series, and 256 is the backbone output features per image. The first layer was <code>nn.Conv2d(1,64,(256+num_meta_feats,1))</code> followed by 12 <code>nn.Conv1d</code> layers with densenet-inspired skip connections. This shallow network architecture had similar performance to that of Yuval's architecture that he described above.</p>\n\n<h2>Ensembling</h2>\n\n<p>At this stage we had about 15 shallow networks results, from Yuval and mine, running on features from different base models. After trying several more complex ideas which overfitted, we settled on very simple ensembling, which finally added very little above simple averaging of all good models. It was basically a linear combination between Yuval's and mine model means optimized to minimize the log-loss, per target. </p>\n\n<h2>Weights approach</h2>\n\n<p>We noticed that the stage 1 test data is very different from the train data if looked from some metadata fields perspective. For example below is a histogram of mean pixel values for original train vs stage 1 and stage 2 test. There are apparently different groups, which have very different ratios in train vs stage 1. Following this idea we assigned a weight to each series, which assures that weighted histograms of metadata look similar, and also <code>any</code> target is distributed similarly. We used our second final submission on that weighted version, where shallow NNs were trained with weighted sampling, and ensembling was done with weighted loss. It gave a nice boost of about 0.001 on CV and stage 1 LB.</p>\n\n<p>Unfortunately, it was absolutely useless for the second stage, because the stage 2 test data happened to be super similar to the original train distribution, see the picture below. We ended up with two very similar submissions for the second stage.</p>\n\n<p><img src=\"https://imgur.com/SIhpUon.png\" alt=\"Weights\"></p>\n\n<h2>Things that didn't work</h2>\n\n<p>All of the above! But we learn a lot from failures, it was a great competition overall.</p>\n\n<h2>Special thanks</h2>\n\n<p>I want to thank Kaggle and Google for distributing GCP credits for competitions, all my shallow networks I trained on GCP. It could have been very hard for me to participate without it.</p>",
      "votes": 12,
      "replies": [
        {
          "id": 674231,
          "author_name": "Ailurophile",
          "author_url": "",
          "post_date": "2019-11-16T03:38:05.407000",
          "content": "<p>Congratulations... Thanks for sharing Code &amp; Insights <a href=\"/zaharch\">@zaharch</a> </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 674134,
      "author_name": "Neuron Engineer",
      "author_url": "",
      "post_date": "2019-11-15T23:32:30.750000",
      "content": "<p>Congratulation <a href=\"/yuval6967\">@yuval6967</a> <a href=\"/zaharch\">@zaharch</a> . Always impressive solution from you guys! <a href=\"/zaharch\">@zaharch</a> You can have a 5 consecutive golds too!!</p>",
      "votes": 3,
      "replies": [
        {
          "id": 674150,
          "author_name": "yuval reina",
          "author_url": "",
          "post_date": "2019-11-16T00:00:30.577000",
          "content": "<p><a href=\"/zaharch\">@zaharch</a> surly can!</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 675585,
      "author_name": "sabari nathan",
      "author_url": "",
      "post_date": "2019-11-18T09:08:07.567000",
      "content": "<p>congratulations team.👍  Great write up. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 674230,
      "author_name": "Ailurophile",
      "author_url": "",
      "post_date": "2019-11-16T03:37:09.523000",
      "content": "<p>Congratulations\nGreat Write-Up\nThanks for Sharing your Valuable Insights &amp; Approach <a href=\"/yuval6967\">@yuval6967</a> </p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "674039": "First I want to thank my teammate @zaharch for a great and successful teamwork (this is our 2nd consecutive gold). Until yesterday we had the top Silver medal - 13th place, and today, for some mysterious reason, we got promoted to 12th and Gold. For me this Gold also means GM (5 golds in the last 5 competitions). \n^ real because @appian was promoted to 11th\n\nSo lets go for our solution.\n\nAs most/all the top solutions we also used a two stage solution:\n1. Base model for feature extraction per image\n2. Shallow model - combining all the output features from a full head scan to predict per image.\nThe 2nd stage also included some post - processing and ensembling.\n\n## Base Model:\nAs base model we used a few different models:\n* Densenet 169, 161, 201\n* SE-ResNet101\n* SE - ResNeXt101_32x4d \nFor all models we used 3 folds, for the SE models we also had 5 folds.\nThe SE models with 5 folds gave the best results. \nThe  models where trained for ~4 epochs using the usual augmentations: rotation, flip, zoom, position shift, pixel intensity shift.\n\n### WSO\nAs many of the other teams do in their base solutions, we also used 3  windows to handle the large dynamic range of the CT pixels values,  but instead of using fixed windows we let the network find the best windows, as described in [Practical Window Setting Optimization for Medical Image Deep Learning](https://arxiv.org/pdf/1812.00572.pdf). \nThe implementation is quit straight forward:  \nadding 3 layers in front of the model:\n```\nConv2d(1, 3, kernel_size=(1, 1))\nSigmoid()\nInstanceNorm2d(3)\n```\nThe convolution layer was initialize with the soft- tissue, blood, bone values. At the end this layer converged to value close to the usual windows values.\n\n### Feature pooling\nMost of the features at the last layer where zero - we used 8 times pooling to decrease the number of features to ~ 250-300\n\n### TTA\nWe created 4 sets of features from augmented images for each train image and 8 sets for each test image.\n\n## Shallow Network\nWe used two different shallow networks. (I will describe one here and @zaharch will describe the 2nd later)\n\nOne network was a FCN.\n### Input \nFeatures from all the images of one full head scan, ordered by the Z position.\n### Layer\n1. 9 * Num_features 2D convolution - the output is batch_size * num_images * num _channels *  1 * 6\n2. Squeeze\n3. 1D convolution layer of size 7\n4. 1D convolution layer of size 5\n5. 1D convolution layer of size 3\nWith batch norms, and ReLUs in between.\n\nWe trained the shallow network with the TTAed features from the base model and for prediction we used the test features TTA.\n\n## Post Processing and ensembling:\n@zaharch will and as a comment\n\n## Results:\nThe base models gave 0.68 - 0.66 on LB (first stage) after fold averaging and with TTA averaging. ** We started using the better models after we already had the shallow network, hence we didn't really submitted a full 5 fold average of base model and the numbers are derived from CV.\nThe best single 5 fold full model (base + shallow), with TTA and fold averaging was the SE-ResNet101 which gave LB 0.6 (first stage). \n\nOne drawback we had - we didn't gain much by ensembling many models, maybe we should have used one model and run it more with different seeds.  \n\n\nAnd as a last word, I want to thank the organizers and moderators @juliaelliott @philculliton @lechuck0 for a great competition and for being flexible and changing the rules to let us use the  metadata which helped all the top teams get really good and interesting solutions. \n\n### Code\n[The full code can be found here](https://github.com/nosound2/RSNA-Hemorrhage)\n\n### More information\n\nMore information about our models can also be found in the following files\n\n[Documentation](https://docs.google.com/document/d/1YFwbnmh5QDF77th01eSEqscvb4rWTKyvL0XWSWs-sMg/edit?usp=sharing)\n\n[Presentation](https://drive.google.com/file/d/1Kz_3mkA9volBKNau_u_jzE2YQhkPTIXp/view?usp=sharing)\n\n[Video](https://drive.google.com/file/d/1yX6WC9GysdekivPzowU685AeAuE_LHXI/view?usp=sharing)\n\n\n",
    "674122": "Thank you, Yuval, for the teamwork, it was an honor working with you. Congratulations with becoming a GM, it was an impressive 5 consecutive golds run!\n\nWe published our [full code on github](https://github.com/nosound2/RSNA-Hemorrhage).\n\nIn this competition I concentrated on metadata, second step shallow networks and ensembling. \n\n## Metadata\n\nI parsed all available metadata fields and hand-crafted 22 float, 16 boolean and 10 categorical variables from it. It had a lot of information by itself, but added no value for us above the images information. \n\n## Shallow Network\n\nThe shallow neural network input is `(256+num_meta_feats)x60`, where 60 is the maximum number of slices per series, and 256 is the backbone output features per image. The first layer was `nn.Conv2d(1,64,(256+num_meta_feats,1))` followed by 12 `nn.Conv1d` layers with densenet-inspired skip connections. This shallow network architecture had similar performance to that of Yuval's architecture that he described above.\n\n## Ensembling\n\nAt this stage we had about 15 shallow networks results, from Yuval and mine, running on features from different base models. After trying several more complex ideas which overfitted, we settled on very simple ensembling, which finally added very little above simple averaging of all good models. It was basically a linear combination between Yuval's and mine model means optimized to minimize the log-loss, per target. \n\n## Weights approach\n\nWe noticed that the stage 1 test data is very different from the train data if looked from some metadata fields perspective. For example below is a histogram of mean pixel values for original train vs stage 1 and stage 2 test. There are apparently different groups, which have very different ratios in train vs stage 1. Following this idea we assigned a weight to each series, which assures that weighted histograms of metadata look similar, and also `any` target is distributed similarly. We used our second final submission on that weighted version, where shallow NNs were trained with weighted sampling, and ensembling was done with weighted loss. It gave a nice boost of about 0.001 on CV and stage 1 LB.\n\nUnfortunately, it was absolutely useless for the second stage, because the stage 2 test data happened to be super similar to the original train distribution, see the picture below. We ended up with two very similar submissions for the second stage.\n\n![Weights](https://imgur.com/SIhpUon.png)\n\n## Things that didn't work\n\nAll of the above! But we learn a lot from failures, it was a great competition overall.\n\n## Special thanks\n\nI want to thank Kaggle and Google for distributing GCP credits for competitions, all my shallow networks I trained on GCP. It could have been very hard for me to participate without it.",
    "674134": "Congratulation @yuval6967 @zaharch . Always impressive solution from you guys! @zaharch You can have a 5 consecutive golds too!!",
    "675585": "congratulations team.👍  Great write up. ",
    "674230": "Congratulations\nGreat Write-Up\nThanks for Sharing your Valuable Insights &amp; Approach @yuval6967 "
  }
}