{
  "id": 193404,
  "title": "18th place solution",
  "url": "/competitions/rsna-str-pulmonary-embolism-detection/discussion/193404",
  "author_name": "Theo Viel",
  "post_date": "2020-10-27T00:16:20.892000",
  "votes": 22,
  "comment_count": 12,
  "views": 0,
  "content": "<p><em>I am actually quite surprised I was able to do this well in the competition, I joined way too late and trained very few models. I only had one selected submission that luckily didn't time out. It is a 5-fold efficientnet-b3 followed by a 5-fold sequential model based on its features.</em></p>\n<h3>Introduction</h3>\n<p>My solution consists of a two steps pipeline, similarly to the overpowered baseline :</p>\n<ul>\n<li>Image level efficientnet-b3, trained to classify whether the image has PE. This model is then used to extract features for each slice of the CT scan.</li>\n<li>A sequential is trained on the features extracted by the CNN, it predicts both the image and exam level labels, and directly optimizes the competition metric.</li>\n</ul>\n<p>I joined the competition 9 days before the end, with the overall motivation of doing something similar to the winners of the previous RSNA competition. <br>\nI was able to quickly build the CNN pipeline, and started training a bunch some models. The issue was that those models would be really long to train on my hardware (1x RTX 21080Ti), so I had to improvise.</p>\n<p>Then, it was about quickly engineering the second part of the pipeline and the inference code, which was far from easy. </p>\n<p>Shortly after I joined, the overly powerful baseline was released. I did not end up really using any of the components of it, but it was an additional motivation for me to keep pushing.<br>\nI was able to come up with my first submission one day before the deadline, which <em>somehow</em> scored 23rd on the public leaderboard. </p>\n<h3>Data</h3>\n<p>As I could not fit the 900 Gb dataset on my computer, I solely relied on the <a href=\"https://www.kaggle.com/vaillant/rsna-str-pe-detection-jpeg-256\" target=\"_blank\">256x256 jpgs</a> extracted by <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a>.<br>\nThanks a lot for making it possible for people like me to join the competition.</p>\n<p>I therefore didn't experiment at all with the windowing, and simply loading the data takes a huge part of my runtime.</p>\n<h3>First level : Convolutional Neural Networks</h3>\n<h4>Undersampling</h4>\n<p>The issue with 2D images constructed from CT scans is that images are very similar. Therefore it makes sense not to use every slice per patient. <br>\nAs there is about 400 slices per patient on average, one epoch would take ages and this is not a path I wanted to go. <br>\nTherefore, I only used 30 images per patient at each epoch, this is done using a custom sampler. </p>\n<p>Once this was done, I was able to train the models for 20 epochs in approximately 9 hours, using a 5-fold grouped by patient validation.</p>\n<h4>Models</h4>\n<p>Models were trained as part of a classical binary classification problem, using the binary cross-entropy<br>\nFirst experiments were conducted with a ResNeXt-50 model as it is usually a reliable baseline. <br>\nI then tried to switch to a bigger ResNext-101, but results were not bigger so I quickly gave up with big architectures.<br>\nThe last model I trained is an efficientnet-b3, which was chosen because a batch size of 32 could fit on my GPU. <br>\nIt performed slightly better so I sticked with this model, and I had no time left to train other models.</p>\n<p>The efficientnet was trained for 15 epochs using a linear scheduled learning rate with 0.05 warmup proportion. </p>\n<h4>Augmentations</h4>\n<pre><code>- albu.HorizontalFlip(p=0.5)\n- albu.VerticalFlip(p=0.5)\n- albu.ShiftScaleRotate(shift_limit=0.1, rotate_limit=45, p=0.5)\n- albu.OneOf([albu.RandomGamma(always_apply=True), albu.RandomBrightnessContrast(always_apply=True),], p=0.5)\n- albu.ElasticTransform(alpha=1, sigma=5, alpha_affine=10, border_mode=cv2.BORDER_CONSTANT, p=0.5)\n</code></pre>\n<h3>Second Level</h3>\n<h4>Model</h4>\n<p>The model I used is a MLP + BidiLSTM one that predicts both the image and exam targets using the CNN extracted features as input. <br>\nTwo 2-layer classifiers are plugged on the concatenation of the output of the MLP and of the LSTM.<br>\nI used the concatenation of average and max pooling for the exam level targets.<br>\nIn addition, multi-sample dropout was used for improved convergence. </p>\n<h4>Training</h4>\n<p>The model was trained using the loss function that matches the metric. <br>\nI also used stochastic weighted averaging for the last few epochs, once again to have a bit more robustness.<br>\nA single epoch took approximately a minute. </p>\n<p>The validation scheme is a normal 5-fold, and my CV scores were quite close to the 0.179 score I had on the public LB.</p>\n<h3>Inference</h3>\n<p>My inference code is available here : <a href=\"https://www.kaggle.com/theoviel/pe-inference-2\" target=\"_blank\">https://www.kaggle.com/theoviel/pe-inference-2</a><br>\nI used clipping to make sure the label assignment rules were respected, which dropped my score of approximately <code>0.003</code>.</p>\n<h3>Final words</h3>\n<p>Congratz to the winners, I'm pretty sure my solution is nowhere near what the top 10 has come up with and I'm really glad I was able to finish 18th. I wanted to tackle a medical imaging challenge for a long time but was always hesitating because of the dataset sizes.</p>\n<p>Hopefully next time I don't procrastinate too much and join a bit earlier, I'm pretty sure I'll benefit a lot from teaming up with people and spending more time experimenting.</p>\n<p>Also, the code is available on <strong>GitHub</strong>, although I still have some cleaning to do, and the ReadMe to complete  :  <a href=\"https://github.com/TheoViel/kaggle_pulmonary_embolism_detection\" target=\"_blank\">https://github.com/TheoViel/kaggle_pulmonary_embolism_detection</a></p>\n<p>Thanks for reading ! </p>",
  "messages": [
    {
      "id": 1061332,
      "postDate": "2020-10-27T00:16:20.893Z",
      "content": "<p><em>I am actually quite surprised I was able to do this well in the competition, I joined way too late and trained very few models. I only had one selected submission that luckily didn't time out. It is a 5-fold efficientnet-b3 followed by a 5-fold sequential model based on its features.</em></p>\n<h3>Introduction</h3>\n<p>My solution consists of a two steps pipeline, similarly to the overpowered baseline :</p>\n<ul>\n<li>Image level efficientnet-b3, trained to classify whether the image has PE. This model is then used to extract features for each slice of the CT scan.</li>\n<li>A sequential is trained on the features extracted by the CNN, it predicts both the image and exam level labels, and directly optimizes the competition metric.</li>\n</ul>\n<p>I joined the competition 9 days before the end, with the overall motivation of doing something similar to the winners of the previous RSNA competition. <br>\nI was able to quickly build the CNN pipeline, and started training a bunch some models. The issue was that those models would be really long to train on my hardware (1x RTX 21080Ti), so I had to improvise.</p>\n<p>Then, it was about quickly engineering the second part of the pipeline and the inference code, which was far from easy. </p>\n<p>Shortly after I joined, the overly powerful baseline was released. I did not end up really using any of the components of it, but it was an additional motivation for me to keep pushing.<br>\nI was able to come up with my first submission one day before the deadline, which <em>somehow</em> scored 23rd on the public leaderboard. </p>\n<h3>Data</h3>\n<p>As I could not fit the 900 Gb dataset on my computer, I solely relied on the <a href=\"https://www.kaggle.com/vaillant/rsna-str-pe-detection-jpeg-256\" target=\"_blank\">256x256 jpgs</a> extracted by <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a>.<br>\nThanks a lot for making it possible for people like me to join the competition.</p>\n<p>I therefore didn't experiment at all with the windowing, and simply loading the data takes a huge part of my runtime.</p>\n<h3>First level : Convolutional Neural Networks</h3>\n<h4>Undersampling</h4>\n<p>The issue with 2D images constructed from CT scans is that images are very similar. Therefore it makes sense not to use every slice per patient. <br>\nAs there is about 400 slices per patient on average, one epoch would take ages and this is not a path I wanted to go. <br>\nTherefore, I only used 30 images per patient at each epoch, this is done using a custom sampler. </p>\n<p>Once this was done, I was able to train the models for 20 epochs in approximately 9 hours, using a 5-fold grouped by patient validation.</p>\n<h4>Models</h4>\n<p>Models were trained as part of a classical binary classification problem, using the binary cross-entropy<br>\nFirst experiments were conducted with a ResNeXt-50 model as it is usually a reliable baseline. <br>\nI then tried to switch to a bigger ResNext-101, but results were not bigger so I quickly gave up with big architectures.<br>\nThe last model I trained is an efficientnet-b3, which was chosen because a batch size of 32 could fit on my GPU. <br>\nIt performed slightly better so I sticked with this model, and I had no time left to train other models.</p>\n<p>The efficientnet was trained for 15 epochs using a linear scheduled learning rate with 0.05 warmup proportion. </p>\n<h4>Augmentations</h4>\n<pre><code>- albu.HorizontalFlip(p=0.5)\n- albu.VerticalFlip(p=0.5)\n- albu.ShiftScaleRotate(shift_limit=0.1, rotate_limit=45, p=0.5)\n- albu.OneOf([albu.RandomGamma(always_apply=True), albu.RandomBrightnessContrast(always_apply=True),], p=0.5)\n- albu.ElasticTransform(alpha=1, sigma=5, alpha_affine=10, border_mode=cv2.BORDER_CONSTANT, p=0.5)\n</code></pre>\n<h3>Second Level</h3>\n<h4>Model</h4>\n<p>The model I used is a MLP + BidiLSTM one that predicts both the image and exam targets using the CNN extracted features as input. <br>\nTwo 2-layer classifiers are plugged on the concatenation of the output of the MLP and of the LSTM.<br>\nI used the concatenation of average and max pooling for the exam level targets.<br>\nIn addition, multi-sample dropout was used for improved convergence. </p>\n<h4>Training</h4>\n<p>The model was trained using the loss function that matches the metric. <br>\nI also used stochastic weighted averaging for the last few epochs, once again to have a bit more robustness.<br>\nA single epoch took approximately a minute. </p>\n<p>The validation scheme is a normal 5-fold, and my CV scores were quite close to the 0.179 score I had on the public LB.</p>\n<h3>Inference</h3>\n<p>My inference code is available here : <a href=\"https://www.kaggle.com/theoviel/pe-inference-2\" target=\"_blank\">https://www.kaggle.com/theoviel/pe-inference-2</a><br>\nI used clipping to make sure the label assignment rules were respected, which dropped my score of approximately <code>0.003</code>.</p>\n<h3>Final words</h3>\n<p>Congratz to the winners, I'm pretty sure my solution is nowhere near what the top 10 has come up with and I'm really glad I was able to finish 18th. I wanted to tackle a medical imaging challenge for a long time but was always hesitating because of the dataset sizes.</p>\n<p>Hopefully next time I don't procrastinate too much and join a bit earlier, I'm pretty sure I'll benefit a lot from teaming up with people and spending more time experimenting.</p>\n<p>Also, the code is available on <strong>GitHub</strong>, although I still have some cleaning to do, and the ReadMe to complete  :  <a href=\"https://github.com/TheoViel/kaggle_pulmonary_embolism_detection\" target=\"_blank\">https://github.com/TheoViel/kaggle_pulmonary_embolism_detection</a></p>\n<p>Thanks for reading ! </p>",
      "rawMarkdown": "*I am actually quite surprised I was able to do this well in the competition, I joined way too late and trained very few models. I only had one selected submission that luckily didn't time out. It is a 5-fold efficientnet-b3 followed by a 5-fold sequential model based on its features.*\n\n### Introduction\nMy solution consists of a two steps pipeline, similarly to the overpowered baseline :\n- Image level efficientnet-b3, trained to classify whether the image has PE. This model is then used to extract features for each slice of the CT scan.\n- A sequential is trained on the features extracted by the CNN, it predicts both the image and exam level labels, and directly optimizes the competition metric.\n\nI joined the competition 9 days before the end, with the overall motivation of doing something similar to the winners of the previous RSNA competition. \nI was able to quickly build the CNN pipeline, and started training a bunch some models. The issue was that those models would be really long to train on my hardware (1x RTX 21080Ti), so I had to improvise.\n\nThen, it was about quickly engineering the second part of the pipeline and the inference code, which was far from easy. \n\nShortly after I joined, the overly powerful baseline was released. I did not end up really using any of the components of it, but it was an additional motivation for me to keep pushing.\nI was able to come up with my first submission one day before the deadline, which *somehow* scored 23rd on the public leaderboard. \n\n### Data\n\nAs I could not fit the 900 Gb dataset on my computer, I solely relied on the [256x256 jpgs](https://www.kaggle.com/vaillant/rsna-str-pe-detection-jpeg-256) extracted by @vaillant.\nThanks a lot for making it possible for people like me to join the competition.\n\nI therefore didn't experiment at all with the windowing, and simply loading the data takes a huge part of my runtime.\n\n### First level : Convolutional Neural Networks\n\n#### Undersampling\n\nThe issue with 2D images constructed from CT scans is that images are very similar. Therefore it makes sense not to use every slice per patient. \nAs there is about 400 slices per patient on average, one epoch would take ages and this is not a path I wanted to go. \nTherefore, I only used 30 images per patient at each epoch, this is done using a custom sampler. \n\nOnce this was done, I was able to train the models for 20 epochs in approximately 9 hours, using a 5-fold grouped by patient validation.\n\n#### Models\n\nModels were trained as part of a classical binary classification problem, using the binary cross-entropy\nFirst experiments were conducted with a ResNeXt-50 model as it is usually a reliable baseline. \nI then tried to switch to a bigger ResNext-101, but results were not bigger so I quickly gave up with big architectures.\nThe last model I trained is an efficientnet-b3, which was chosen because a batch size of 32 could fit on my GPU. \nIt performed slightly better so I sticked with this model, and I had no time left to train other models.\n\nThe efficientnet was trained for 15 epochs using a linear scheduled learning rate with 0.05 warmup proportion. \n\n#### Augmentations\n\n```\n- albu.HorizontalFlip(p=0.5)\n- albu.VerticalFlip(p=0.5)\n- albu.ShiftScaleRotate(shift_limit=0.1, rotate_limit=45, p=0.5)\n- albu.OneOf([albu.RandomGamma(always_apply=True), albu.RandomBrightnessContrast(always_apply=True),], p=0.5)\n- albu.ElasticTransform(alpha=1, sigma=5, alpha_affine=10, border_mode=cv2.BORDER_CONSTANT, p=0.5)\n```\n\n### Second Level\n\n#### Model\n\nThe model I used is a MLP + BidiLSTM one that predicts both the image and exam targets using the CNN extracted features as input. \nTwo 2-layer classifiers are plugged on the concatenation of the output of the MLP and of the LSTM.\nI used the concatenation of average and max pooling for the exam level targets.\nIn addition, multi-sample dropout was used for improved convergence. \n\n#### Training\n\nThe model was trained using the loss function that matches the metric. \nI also used stochastic weighted averaging for the last few epochs, once again to have a bit more robustness.\nA single epoch took approximately a minute. \n\nThe validation scheme is a normal 5-fold, and my CV scores were quite close to the 0.179 score I had on the public LB.\n\n\n### Inference\n\nMy inference code is available here : https://www.kaggle.com/theoviel/pe-inference-2\nI used clipping to make sure the label assignment rules were respected, which dropped my score of approximately `0.003`.\n\n\n### Final words\n\nCongratz to the winners, I'm pretty sure my solution is nowhere near what the top 10 has come up with and I'm really glad I was able to finish 18th. I wanted to tackle a medical imaging challenge for a long time but was always hesitating because of the dataset sizes.\n\nHopefully next time I don't procrastinate too much and join a bit earlier, I'm pretty sure I'll benefit a lot from teaming up with people and spending more time experimenting.\n\nAlso, the code is available on **GitHub**, although I still have some cleaning to do, and the ReadMe to complete  :  https://github.com/TheoViel/kaggle_pulmonary_embolism_detection\n\nThanks for reading ! \n",
      "votes": 22
    },
    {
      "id": 1061446,
      "postDate": "2020-10-27T02:36:13.530Z",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> , this is such a smart solution! Can you elaborate more on the custom sampler that you used? What rule did you use to sample the slices for patient images?</p>",
      "rawMarkdown": "Congrats @theoviel , this is such a smart solution! Can you elaborate more on the custom sampler that you used? What rule did you use to sample the slices for patient images?",
      "votes": 1,
      "replies": [
        {
          "id": 1061465,
          "postDate": "2020-10-27T03:00:31.510Z",
          "content": "<p>I took a look at the code. It seems the sampling is done, by first randomly sampling patients and next, sampling slices by idx until an ideal batch size is reached. Very cool and elegant solution to this big data problem! I learnt a lot, thanks.</p>",
          "rawMarkdown": "I took a look at the code. It seems the sampling is done, by first randomly sampling patients and next, sampling slices by idx until an ideal batch size is reached. Very cool and elegant solution to this big data problem! I learnt a lot, thanks.",
          "votes": 2
        },
        {
          "id": 1061468,
          "postDate": "2020-10-27T03:01:17Z",
          "content": "<p>I agree. Very elegant.</p>",
          "rawMarkdown": "I agree. Very elegant.",
          "votes": 2
        },
        {
          "id": 1061661,
          "postDate": "2020-10-27T07:25:53.463Z",
          "content": "<p>Thanks to you two !</p>\n<p>What happens is that before starting the epoch, the sampler does a run on all the shuffled images. It keeps only the 30 first it sees per patient to form the batches.</p>",
          "rawMarkdown": "Thanks to you two !\n\nWhat happens is that before starting the epoch, the sampler does a run on all the shuffled images. It keeps only the 30 first it sees per patient to form the batches.\n",
          "votes": 1
        },
        {
          "id": 1064534,
          "postDate": "2020-10-30T08:07:57.943Z",
          "content": "<p>Hi, great solution. Thanks for sharing! Did you train models from both stages on 30 slices per study?</p>",
          "rawMarkdown": "Hi, great solution. Thanks for sharing! Did you train models from both stages on 30 slices per study?",
          "votes": 1
        },
        {
          "id": 1065482,
          "postDate": "2020-10-31T10:36:21.950Z",
          "content": "<p><a href=\"https://www.kaggle.com/ademyanchuk\" target=\"_blank\">@ademyanchuk</a> Models from the 2nd stage take all the slices as input</p>",
          "rawMarkdown": "@ademyanchuk Models from the 2nd stage take all the slices as input"
        }
      ]
    },
    {
      "id": 1061359,
      "postDate": "2020-10-27T00:42:39.647Z",
      "content": "<p>Thanks for the solution! May I ask what your image level validation score is?</p>",
      "rawMarkdown": "Thanks for the solution! May I ask what your image level validation score is?",
      "votes": 1,
      "replies": [
        {
          "id": 1061664,
          "postDate": "2020-10-27T07:28:44.550Z",
          "content": "<p>You're welcome ! My validation logloss is of about 0.110 on each fold.</p>",
          "rawMarkdown": "You're welcome ! My validation logloss is of about 0.110 on each fold."
        }
      ]
    },
    {
      "id": 1061399,
      "postDate": "2020-10-27T01:37:13.867Z",
      "content": "<p>Great job Theo. I really like how simple your solution is. It is just one stage 1 and one stage 2 model. Very efficient.</p>\n<p>It was smart of you to reduce the training images</p>\n<blockquote>\n  <p>Therefore, I only used 30 images per patient at each epoch, this is done using a custom sampler.</p>\n</blockquote>\n<p>That was important in this competition.</p>",
      "rawMarkdown": "Great job Theo. I really like how simple your solution is. It is just one stage 1 and one stage 2 model. Very efficient.\n\nIt was smart of you to reduce the training images\n> Therefore, I only used 30 images per patient at each epoch, this is done using a custom sampler.\n\nThat was important in this competition.",
      "votes": 2,
      "replies": [
        {
          "id": 1061665,
          "postDate": "2020-10-27T07:29:26.157Z",
          "content": "<p>Thanks Chris ! Turns out having limited computational ressources has some plus sides </p>",
          "rawMarkdown": "Thanks Chris ! Turns out having limited computational ressources has some plus sides ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1061376,
      "postDate": "2020-10-27T01:03:07.770Z",
      "content": "<p>I'm surprised your solution is really similar to the one I was going for but alas, although I joined 13 days before the deadline I simply did not have the experience necessary to finish in time. The only difference being in the sequential model. I was going to terrain an attention model. It was proven to perform better for long sequences which is the case with an avg of 200 scans per patient. A venilla bidirectional gru to encode the features extracted from the cnn then the attention mechanism. That with a slight modification to the attention mechanism that being concatenating the context tensor with the gru's output at that timestamp(i.e. slice index), to help the model predict the label for the correct slice. I'm currently still prototyping the model and haven't tested the idea yet but what do you think?</p>",
      "rawMarkdown": "I'm surprised your solution is really similar to the one I was going for but alas, although I joined 13 days before the deadline I simply did not have the experience necessary to finish in time. The only difference being in the sequential model. I was going to terrain an attention model. It was proven to perform better for long sequences which is the case with an avg of 200 scans per patient. A venilla bidirectional gru to encode the features extracted from the cnn then the attention mechanism. That with a slight modification to the attention mechanism that being concatenating the context tensor with the gru's output at that timestamp(i.e. slice index), to help the model predict the label for the correct slice. I'm currently still prototyping the model and haven't tested the idea yet but what do you think?",
      "replies": [
        {
          "id": 1061813,
          "postDate": "2020-10-27T10:50:50.860Z",
          "content": "<p>If you plan on using attention you might as well want to use transformers as several competitors did. It seems to work well.</p>",
          "rawMarkdown": "If you plan on using attention you might as well want to use transformers as several competitors did. It seems to work well."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1061446,
      "author_name": "Alan Choon",
      "author_url": "",
      "post_date": "2020-10-27T02:36:13.530000",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> , this is such a smart solution! Can you elaborate more on the custom sampler that you used? What rule did you use to sample the slices for patient images?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1061465,
          "author_name": "Alan Choon",
          "author_url": "",
          "post_date": "2020-10-27T03:00:31.510000",
          "content": "<p>I took a look at the code. It seems the sampling is done, by first randomly sampling patients and next, sampling slices by idx until an ideal batch size is reached. Very cool and elegant solution to this big data problem! I learnt a lot, thanks.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1061468,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-10-27T03:01:17",
          "content": "<p>I agree. Very elegant.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1061661,
          "author_name": "Theo Viel",
          "author_url": "",
          "post_date": "2020-10-27T07:25:53.463000",
          "content": "<p>Thanks to you two !</p>\n<p>What happens is that before starting the epoch, the sampler does a run on all the shuffled images. It keeps only the 30 first it sees per patient to form the batches.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1064534,
          "author_name": "A.Demyanchuk",
          "author_url": "",
          "post_date": "2020-10-30T08:07:57.943000",
          "content": "<p>Hi, great solution. Thanks for sharing! Did you train models from both stages on 30 slices per study?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1065482,
          "author_name": "Theo Viel",
          "author_url": "",
          "post_date": "2020-10-31T10:36:21.950000",
          "content": "<p><a href=\"https://www.kaggle.com/ademyanchuk\" target=\"_blank\">@ademyanchuk</a> Models from the 2nd stage take all the slices as input</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1061359,
      "author_name": "sin",
      "author_url": "",
      "post_date": "2020-10-27T00:42:39.647000",
      "content": "<p>Thanks for the solution! May I ask what your image level validation score is?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1061664,
          "author_name": "Theo Viel",
          "author_url": "",
          "post_date": "2020-10-27T07:28:44.550000",
          "content": "<p>You're welcome ! My validation logloss is of about 0.110 on each fold.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1061399,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-10-27T01:37:13.867000",
      "content": "<p>Great job Theo. I really like how simple your solution is. It is just one stage 1 and one stage 2 model. Very efficient.</p>\n<p>It was smart of you to reduce the training images</p>\n<blockquote>\n  <p>Therefore, I only used 30 images per patient at each epoch, this is done using a custom sampler.</p>\n</blockquote>\n<p>That was important in this competition.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1061665,
          "author_name": "Theo Viel",
          "author_url": "",
          "post_date": "2020-10-27T07:29:26.157000",
          "content": "<p>Thanks Chris ! Turns out having limited computational ressources has some plus sides </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1061376,
      "author_name": "DarkCube",
      "author_url": "",
      "post_date": "2020-10-27T01:03:07.770000",
      "content": "<p>I'm surprised your solution is really similar to the one I was going for but alas, although I joined 13 days before the deadline I simply did not have the experience necessary to finish in time. The only difference being in the sequential model. I was going to terrain an attention model. It was proven to perform better for long sequences which is the case with an avg of 200 scans per patient. A venilla bidirectional gru to encode the features extracted from the cnn then the attention mechanism. That with a slight modification to the attention mechanism that being concatenating the context tensor with the gru's output at that timestamp(i.e. slice index), to help the model predict the label for the correct slice. I'm currently still prototyping the model and haven't tested the idea yet but what do you think?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1061813,
          "author_name": "Theo Viel",
          "author_url": "",
          "post_date": "2020-10-27T10:50:50.860000",
          "content": "<p>If you plan on using attention you might as well want to use transformers as several competitors did. It seems to work well.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1061332": "*I am actually quite surprised I was able to do this well in the competition, I joined way too late and trained very few models. I only had one selected submission that luckily didn't time out. It is a 5-fold efficientnet-b3 followed by a 5-fold sequential model based on its features.*\n\n### Introduction\nMy solution consists of a two steps pipeline, similarly to the overpowered baseline :\n- Image level efficientnet-b3, trained to classify whether the image has PE. This model is then used to extract features for each slice of the CT scan.\n- A sequential is trained on the features extracted by the CNN, it predicts both the image and exam level labels, and directly optimizes the competition metric.\n\nI joined the competition 9 days before the end, with the overall motivation of doing something similar to the winners of the previous RSNA competition. \nI was able to quickly build the CNN pipeline, and started training a bunch some models. The issue was that those models would be really long to train on my hardware (1x RTX 21080Ti), so I had to improvise.\n\nThen, it was about quickly engineering the second part of the pipeline and the inference code, which was far from easy. \n\nShortly after I joined, the overly powerful baseline was released. I did not end up really using any of the components of it, but it was an additional motivation for me to keep pushing.\nI was able to come up with my first submission one day before the deadline, which *somehow* scored 23rd on the public leaderboard. \n\n### Data\n\nAs I could not fit the 900 Gb dataset on my computer, I solely relied on the [256x256 jpgs](https://www.kaggle.com/vaillant/rsna-str-pe-detection-jpeg-256) extracted by @vaillant.\nThanks a lot for making it possible for people like me to join the competition.\n\nI therefore didn't experiment at all with the windowing, and simply loading the data takes a huge part of my runtime.\n\n### First level : Convolutional Neural Networks\n\n#### Undersampling\n\nThe issue with 2D images constructed from CT scans is that images are very similar. Therefore it makes sense not to use every slice per patient. \nAs there is about 400 slices per patient on average, one epoch would take ages and this is not a path I wanted to go. \nTherefore, I only used 30 images per patient at each epoch, this is done using a custom sampler. \n\nOnce this was done, I was able to train the models for 20 epochs in approximately 9 hours, using a 5-fold grouped by patient validation.\n\n#### Models\n\nModels were trained as part of a classical binary classification problem, using the binary cross-entropy\nFirst experiments were conducted with a ResNeXt-50 model as it is usually a reliable baseline. \nI then tried to switch to a bigger ResNext-101, but results were not bigger so I quickly gave up with big architectures.\nThe last model I trained is an efficientnet-b3, which was chosen because a batch size of 32 could fit on my GPU. \nIt performed slightly better so I sticked with this model, and I had no time left to train other models.\n\nThe efficientnet was trained for 15 epochs using a linear scheduled learning rate with 0.05 warmup proportion. \n\n#### Augmentations\n\n```\n- albu.HorizontalFlip(p=0.5)\n- albu.VerticalFlip(p=0.5)\n- albu.ShiftScaleRotate(shift_limit=0.1, rotate_limit=45, p=0.5)\n- albu.OneOf([albu.RandomGamma(always_apply=True), albu.RandomBrightnessContrast(always_apply=True),], p=0.5)\n- albu.ElasticTransform(alpha=1, sigma=5, alpha_affine=10, border_mode=cv2.BORDER_CONSTANT, p=0.5)\n```\n\n### Second Level\n\n#### Model\n\nThe model I used is a MLP + BidiLSTM one that predicts both the image and exam targets using the CNN extracted features as input. \nTwo 2-layer classifiers are plugged on the concatenation of the output of the MLP and of the LSTM.\nI used the concatenation of average and max pooling for the exam level targets.\nIn addition, multi-sample dropout was used for improved convergence. \n\n#### Training\n\nThe model was trained using the loss function that matches the metric. \nI also used stochastic weighted averaging for the last few epochs, once again to have a bit more robustness.\nA single epoch took approximately a minute. \n\nThe validation scheme is a normal 5-fold, and my CV scores were quite close to the 0.179 score I had on the public LB.\n\n\n### Inference\n\nMy inference code is available here : https://www.kaggle.com/theoviel/pe-inference-2\nI used clipping to make sure the label assignment rules were respected, which dropped my score of approximately `0.003`.\n\n\n### Final words\n\nCongratz to the winners, I'm pretty sure my solution is nowhere near what the top 10 has come up with and I'm really glad I was able to finish 18th. I wanted to tackle a medical imaging challenge for a long time but was always hesitating because of the dataset sizes.\n\nHopefully next time I don't procrastinate too much and join a bit earlier, I'm pretty sure I'll benefit a lot from teaming up with people and spending more time experimenting.\n\nAlso, the code is available on **GitHub**, although I still have some cleaning to do, and the ReadMe to complete  :  https://github.com/TheoViel/kaggle_pulmonary_embolism_detection\n\nThanks for reading ! \n",
    "1061446": "Congrats @theoviel , this is such a smart solution! Can you elaborate more on the custom sampler that you used? What rule did you use to sample the slices for patient images?",
    "1061359": "Thanks for the solution! May I ask what your image level validation score is?",
    "1061399": "Great job Theo. I really like how simple your solution is. It is just one stage 1 and one stage 2 model. Very efficient.\n\nIt was smart of you to reduce the training images\n> Therefore, I only used 30 images per patient at each epoch, this is done using a custom sampler.\n\nThat was important in this competition.",
    "1061376": "I'm surprised your solution is really similar to the one I was going for but alas, although I joined 13 days before the deadline I simply did not have the experience necessary to finish in time. The only difference being in the sequential model. I was going to terrain an attention model. It was proven to perform better for long sequences which is the case with an avg of 200 scans per patient. A venilla bidirectional gru to encode the features extracted from the cnn then the attention mechanism. That with a slight modification to the attention mechanism that being concatenating the context tensor with the gru's output at that timestamp(i.e. slice index), to help the model predict the label for the correct slice. I'm currently still prototyping the model and haven't tested the idea yet but what do you think?"
  }
}