{
  "id": 193460,
  "title": "7th place - EfficientNet, Transformer and a 2nd opinion [edited]",
  "url": "/competitions/rsna-str-pulmonary-embolism-detection/discussion/193460",
  "author_name": "yuval reina",
  "post_date": "2020-10-27T06:50:57.040000",
  "votes": 54,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I used a two stage model.</p>\n<ol>\n<li>A EfficientNet (B5, B3) used for feature extraction per image</li>\n<li>A transformer used per series to predict the series related classes and the  'PE Present on Image' per image</li>\n</ol>\n<p>This is the 2nd competition I use such network and an extensive description of this model can be found <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/181830\" target=\"_blank\">here</a>  </p>\n<p>The targets for the EfficientNet where: </p>\n<ul>\n<li>The original targets for images where  PE Present on Image = 1</li>\n<li>0 for every other image. Except the Intermediate target which remained the same.</li>\n</ul>\n<p>The loss was weighted BCE - the weights reflecting the competitions matric weights.<br>\nI used flip, rotate, random resize/crop, mean/std shift as augmentation.<br>\nI also use trainable 3 windows to convert the CT image to jpeg (<a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/117480\" target=\"_blank\">WSO</a> )</p>\n<p>The transformer was a 4 layer encoder (using Pytorch's transformer encoder module). Where the relative and absolute places of the images in the series were embedded and added to the features vectors (as is done for positional embedding in NLP transformers such as BERT)<br>\nThe loss function reflected the competition's matric.</p>\n<p>This model gave an <strong>LB of 0.166</strong>  </p>\n<p>Ensembling improved the <strong>LB to 0.162</strong>, but I could only ensemble 2 models in the time frame (I didn't use the public/private LB trick that can add 25%). to improve this I used a <strong>2nd opinion mechanism</strong> - instead of using a 2nd model to inference all the data again, I did what an MD will do, I chose only series where the results where the most uncertain (near 0.5), inference them with another model and ensembled - I did this 3 times for ~ 30-40% of the data each time gaining an equivalence of ensembling 4 models =&gt; <strong>LB 0.157</strong></p>\n<p>As a last step, I checked if any prediction meet the competition's <a href=\"https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/183473\" target=\"_blank\">Label Consistency Requirement </a> and if it didn't, made the minimal changes needed to meet the requirements</p>\n<ul>\n<li><p>The full code can be found in <a href=\"https://github.com/yuval6957/RSNA2020_final.git\" target=\"_blank\">git</a></p></li>\n<li><p>The inference code can be found <a href=\"https://www.kaggle.com/yuval6967/rsna2020-inference-2nd-op-final\" target=\"_blank\">in this notebook</a></p></li>\n<li><p>The models' weights are <a href=\"https://www.kaggle.com/yuval6967/rsna2020-models\" target=\"_blank\">in this</a> public dataset</p></li>\n<li><p>A more detailed description can be found <a href=\"https://github.com/yuval6957/RSNA2020_final/blob/master/Documentation.md\" target=\"_blank\">in this documentation</a> and <a href=\"https://github.com/yuval6957/RSNA2020_final/blob/master/RSNA2020%20presentation.pdf\" target=\"_blank\">this presentation</a></p></li>\n<li><p><a href=\"https://www.youtube.com/watch?v=hVgIawktZgs\" target=\"_blank\">This is a video</a> which present this solution</p></li>\n</ul>",
  "messages": [
    {
      "id": 1061641,
      "postDate": "2020-10-27T06:50:57.040Z",
      "content": "<p>I used a two stage model.</p>\n<ol>\n<li>A EfficientNet (B5, B3) used for feature extraction per image</li>\n<li>A transformer used per series to predict the series related classes and the  'PE Present on Image' per image</li>\n</ol>\n<p>This is the 2nd competition I use such network and an extensive description of this model can be found <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/181830\" target=\"_blank\">here</a>  </p>\n<p>The targets for the EfficientNet where: </p>\n<ul>\n<li>The original targets for images where  PE Present on Image = 1</li>\n<li>0 for every other image. Except the Intermediate target which remained the same.</li>\n</ul>\n<p>The loss was weighted BCE - the weights reflecting the competitions matric weights.<br>\nI used flip, rotate, random resize/crop, mean/std shift as augmentation.<br>\nI also use trainable 3 windows to convert the CT image to jpeg (<a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/117480\" target=\"_blank\">WSO</a> )</p>\n<p>The transformer was a 4 layer encoder (using Pytorch's transformer encoder module). Where the relative and absolute places of the images in the series were embedded and added to the features vectors (as is done for positional embedding in NLP transformers such as BERT)<br>\nThe loss function reflected the competition's matric.</p>\n<p>This model gave an <strong>LB of 0.166</strong>  </p>\n<p>Ensembling improved the <strong>LB to 0.162</strong>, but I could only ensemble 2 models in the time frame (I didn't use the public/private LB trick that can add 25%). to improve this I used a <strong>2nd opinion mechanism</strong> - instead of using a 2nd model to inference all the data again, I did what an MD will do, I chose only series where the results where the most uncertain (near 0.5), inference them with another model and ensembled - I did this 3 times for ~ 30-40% of the data each time gaining an equivalence of ensembling 4 models =&gt; <strong>LB 0.157</strong></p>\n<p>As a last step, I checked if any prediction meet the competition's <a href=\"https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/183473\" target=\"_blank\">Label Consistency Requirement </a> and if it didn't, made the minimal changes needed to meet the requirements</p>\n<ul>\n<li><p>The full code can be found in <a href=\"https://github.com/yuval6957/RSNA2020_final.git\" target=\"_blank\">git</a></p></li>\n<li><p>The inference code can be found <a href=\"https://www.kaggle.com/yuval6967/rsna2020-inference-2nd-op-final\" target=\"_blank\">in this notebook</a></p></li>\n<li><p>The models' weights are <a href=\"https://www.kaggle.com/yuval6967/rsna2020-models\" target=\"_blank\">in this</a> public dataset</p></li>\n<li><p>A more detailed description can be found <a href=\"https://github.com/yuval6957/RSNA2020_final/blob/master/Documentation.md\" target=\"_blank\">in this documentation</a> and <a href=\"https://github.com/yuval6957/RSNA2020_final/blob/master/RSNA2020%20presentation.pdf\" target=\"_blank\">this presentation</a></p></li>\n<li><p><a href=\"https://www.youtube.com/watch?v=hVgIawktZgs\" target=\"_blank\">This is a video</a> which present this solution</p></li>\n</ul>",
      "rawMarkdown": "I used a two stage model.\n1. A EfficientNet (B5, B3) used for feature extraction per image\n2. A transformer used per series to predict the series related classes and the  'PE Present on Image' per image\n\nThis is the 2nd competition I use such network and an extensive description of this model can be found [here](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/181830)  \n\nThe targets for the EfficientNet where: \n* The original targets for images where  PE Present on Image = 1\n* 0 for every other image. Except the Intermediate target which remained the same.\n\nThe loss was weighted BCE - the weights reflecting the competitions matric weights.\nI used flip, rotate, random resize/crop, mean/std shift as augmentation.\nI also use trainable 3 windows to convert the CT image to jpeg ([WSO](https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/117480) )\n\nThe transformer was a 4 layer encoder (using Pytorch's transformer encoder module). Where the relative and absolute places of the images in the series were embedded and added to the features vectors (as is done for positional embedding in NLP transformers such as BERT)\nThe loss function reflected the competition's matric.\n\nThis model gave an **LB of 0.166**  \n\nEnsembling improved the **LB to 0.162**, but I could only ensemble 2 models in the time frame (I didn't use the public/private LB trick that can add 25%). to improve this I used a **2nd opinion mechanism** - instead of using a 2nd model to inference all the data again, I did what an MD will do, I chose only series where the results where the most uncertain (near 0.5), inference them with another model and ensembled - I did this 3 times for ~ 30-40% of the data each time gaining an equivalence of ensembling 4 models => **LB 0.157**\n\nAs a last step, I checked if any prediction meet the competition's [Label Consistency Requirement ](https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/183473) and if it didn't, made the minimal changes needed to meet the requirements\n\n* The full code can be found in [git](https://github.com/yuval6957/RSNA2020_final.git)\n\n* The inference code can be found [in this notebook](https://www.kaggle.com/yuval6967/rsna2020-inference-2nd-op-final)\n\n* The models' weights are [in this](https://www.kaggle.com/yuval6967/rsna2020-models) public dataset\n\n* A more detailed description can be found [in this documentation](https://github.com/yuval6957/RSNA2020_final/blob/master/Documentation.md) and [this presentation](https://github.com/yuval6957/RSNA2020_final/blob/master/RSNA2020%20presentation.pdf)\n\n* [This is a video](https://www.youtube.com/watch?v=hVgIawktZgs) which present this solution\n ",
      "votes": 54
    },
    {
      "id": 1063923,
      "postDate": "2020-10-29T14:03:17.273Z",
      "content": "<p>Thanks for sharing and good explanation!<br>\nRecently, we are seeing a lot of possibilities of transformer into vision tasks.<br>\nThis solution will be a good example.</p>\n<p>Congrats and thanks again. <a href=\"https://www.kaggle.com/yuval6967\" target=\"_blank\">@yuval6967</a> </p>",
      "rawMarkdown": "Thanks for sharing and good explanation!\nRecently, we are seeing a lot of possibilities of transformer into vision tasks.\nThis solution will be a good example.\n\nCongrats and thanks again. @yuval6967 ",
      "votes": 1
    },
    {
      "id": 1061835,
      "postDate": "2020-10-27T11:14:50.790Z",
      "content": "<p>Congratulations on your medal +1</p>",
      "rawMarkdown": "Congratulations on your medal +1",
      "votes": 1
    },
    {
      "id": 1061728,
      "postDate": "2020-10-27T09:13:07.880Z",
      "content": "<p>Congrats on 8th place and solo gold medal <a href=\"https://www.kaggle.com/yuval6967\" target=\"_blank\">@yuval6967</a> and thanks for sharing solution</p>",
      "rawMarkdown": "Congrats on 8th place and solo gold medal @yuval6967 and thanks for sharing solution",
      "votes": 1
    },
    {
      "id": 1061704,
      "postDate": "2020-10-27T08:30:04.373Z",
      "content": "<blockquote>\n  <p>2nd opinion mechanism - instead of using a 2nd model to inference all the data again, I did what an MD will do, I chose only series where the results where the most uncertain (near 0.5), inference them with another model and ensembled - I did this 3 times for ~ 30-40% of the data each time gaining an equivalence of ensembling 4 models =&gt; LB 0.157</p>\n</blockquote>\n<p>Just wow. Thanks for this technique.</p>",
      "rawMarkdown": ">  2nd opinion mechanism - instead of using a 2nd model to inference all the data again, I did what an MD will do, I chose only series where the results where the most uncertain (near 0.5), inference them with another model and ensembled - I did this 3 times for ~ 30-40% of the data each time gaining an equivalence of ensembling 4 models => LB 0.157\n\nJust wow. Thanks for this technique.",
      "votes": 1
    },
    {
      "id": 1061691,
      "postDate": "2020-10-27T08:03:36.050Z",
      "content": "<p>Nice solution, thank you for sharing and congratulation 👍</p>",
      "rawMarkdown": "Nice solution, thank you for sharing and congratulation 👍",
      "votes": 1
    },
    {
      "id": 1061658,
      "postDate": "2020-10-27T07:20:28.500Z",
      "content": "<p><a href=\"https://www.kaggle.com/yuval6967\" target=\"_blank\">@yuval6967</a> Congratulations!</p>",
      "rawMarkdown": "@yuval6967 Congratulations!",
      "votes": 1
    },
    {
      "id": 1061644,
      "postDate": "2020-10-27T06:58:38.027Z",
      "content": "<p>Thank for sharing and congratulations <a href=\"https://www.kaggle.com/yuval6967\" target=\"_blank\">@yuval6967</a> !</p>",
      "rawMarkdown": "Thank for sharing and congratulations @yuval6967 !",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 1063923,
      "author_name": "Heroseo",
      "author_url": "",
      "post_date": "2020-10-29T14:03:17.273000",
      "content": "<p>Thanks for sharing and good explanation!<br>\nRecently, we are seeing a lot of possibilities of transformer into vision tasks.<br>\nThis solution will be a good example.</p>\n<p>Congrats and thanks again. <a href=\"https://www.kaggle.com/yuval6967\" target=\"_blank\">@yuval6967</a> </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1061835,
      "author_name": "Eisa",
      "author_url": "",
      "post_date": "2020-10-27T11:14:50.790000",
      "content": "<p>Congratulations on your medal +1</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1061728,
      "author_name": "KhanhVD",
      "author_url": "",
      "post_date": "2020-10-27T09:13:07.880000",
      "content": "<p>Congrats on 8th place and solo gold medal <a href=\"https://www.kaggle.com/yuval6967\" target=\"_blank\">@yuval6967</a> and thanks for sharing solution</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1061704,
      "author_name": "sin",
      "author_url": "",
      "post_date": "2020-10-27T08:30:04.373000",
      "content": "<blockquote>\n  <p>2nd opinion mechanism - instead of using a 2nd model to inference all the data again, I did what an MD will do, I chose only series where the results where the most uncertain (near 0.5), inference them with another model and ensembled - I did this 3 times for ~ 30-40% of the data each time gaining an equivalence of ensembling 4 models =&gt; LB 0.157</p>\n</blockquote>\n<p>Just wow. Thanks for this technique.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1061691,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-27T08:03:36.050000",
      "content": "<p>Nice solution, thank you for sharing and congratulation 👍</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1061658,
      "author_name": "Alexey Kachalov",
      "author_url": "",
      "post_date": "2020-10-27T07:20:28.500000",
      "content": "<p><a href=\"https://www.kaggle.com/yuval6967\" target=\"_blank\">@yuval6967</a> Congratulations!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1061644,
      "author_name": "DungNB",
      "author_url": "",
      "post_date": "2020-10-27T06:58:38.027000",
      "content": "<p>Thank for sharing and congratulations <a href=\"https://www.kaggle.com/yuval6967\" target=\"_blank\">@yuval6967</a> !</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1061641": "I used a two stage model.\n1. A EfficientNet (B5, B3) used for feature extraction per image\n2. A transformer used per series to predict the series related classes and the  'PE Present on Image' per image\n\nThis is the 2nd competition I use such network and an extensive description of this model can be found [here](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/181830)  \n\nThe targets for the EfficientNet where: \n* The original targets for images where  PE Present on Image = 1\n* 0 for every other image. Except the Intermediate target which remained the same.\n\nThe loss was weighted BCE - the weights reflecting the competitions matric weights.\nI used flip, rotate, random resize/crop, mean/std shift as augmentation.\nI also use trainable 3 windows to convert the CT image to jpeg ([WSO](https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/117480) )\n\nThe transformer was a 4 layer encoder (using Pytorch's transformer encoder module). Where the relative and absolute places of the images in the series were embedded and added to the features vectors (as is done for positional embedding in NLP transformers such as BERT)\nThe loss function reflected the competition's matric.\n\nThis model gave an **LB of 0.166**  \n\nEnsembling improved the **LB to 0.162**, but I could only ensemble 2 models in the time frame (I didn't use the public/private LB trick that can add 25%). to improve this I used a **2nd opinion mechanism** - instead of using a 2nd model to inference all the data again, I did what an MD will do, I chose only series where the results where the most uncertain (near 0.5), inference them with another model and ensembled - I did this 3 times for ~ 30-40% of the data each time gaining an equivalence of ensembling 4 models => **LB 0.157**\n\nAs a last step, I checked if any prediction meet the competition's [Label Consistency Requirement ](https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/183473) and if it didn't, made the minimal changes needed to meet the requirements\n\n* The full code can be found in [git](https://github.com/yuval6957/RSNA2020_final.git)\n\n* The inference code can be found [in this notebook](https://www.kaggle.com/yuval6967/rsna2020-inference-2nd-op-final)\n\n* The models' weights are [in this](https://www.kaggle.com/yuval6967/rsna2020-models) public dataset\n\n* A more detailed description can be found [in this documentation](https://github.com/yuval6957/RSNA2020_final/blob/master/Documentation.md) and [this presentation](https://github.com/yuval6957/RSNA2020_final/blob/master/RSNA2020%20presentation.pdf)\n\n* [This is a video](https://www.youtube.com/watch?v=hVgIawktZgs) which present this solution\n ",
    "1063923": "Thanks for sharing and good explanation!\nRecently, we are seeing a lot of possibilities of transformer into vision tasks.\nThis solution will be a good example.\n\nCongrats and thanks again. @yuval6967 ",
    "1061835": "Congratulations on your medal +1",
    "1061728": "Congrats on 8th place and solo gold medal @yuval6967 and thanks for sharing solution",
    "1061704": ">  2nd opinion mechanism - instead of using a 2nd model to inference all the data again, I did what an MD will do, I chose only series where the results where the most uncertain (near 0.5), inference them with another model and ensembled - I did this 3 times for ~ 30-40% of the data each time gaining an equivalence of ensembling 4 models => LB 0.157\n\nJust wow. Thanks for this technique.",
    "1061691": "Nice solution, thank you for sharing and congratulation 👍",
    "1061658": "@yuval6967 Congratulations!",
    "1061644": "Thank for sharing and congratulations @yuval6967 !"
  }
}