{
  "id": 193469,
  "title": "Segmentation model with external dataset (+0.11 on CV/LB)",
  "url": "/competitions/rsna-str-pulmonary-embolism-detection/discussion/193469",
  "author_name": "narsil (jobs-in-data.com)",
  "post_date": "2020-10-27T07:26:07.429000",
  "votes": 11,
  "comment_count": 2,
  "views": 0,
  "content": "<p>First of all - big thanks to everyone in this competition - Kaggle is truly an amazing place to learn. In particular, big thanks to my teammate <a href=\"https://www.kaggle.com/allvor\" target=\"_blank\">@allvor</a> who provided great ideas throughout the competition.</p>\n<p>To start - any competition I participate in, I like to do error analysis. Seeing what my model predicts and comparing it to the ground truth can lead to good insights on how the model works and how It could be improved. In this competition, since I don’t any medical background, it was very difficult to even spot areas related to image-level PE, not to mention other labels (maybe except RV/LV label – when after some research I knew more or less where to look). Looking for data to enable error analysis led me to find this:</p>\n<ul>\n<li><a href=\"https://www.nature.com/articles/sdata2018180\" target=\"_blank\">Nature article - A new dataset of computed-tomography angiography images for computer-aided detection of pulmonary embolism</a></li>\n<li><a href=\"https://www.kaggle.com/andrewmvd/pulmonary-embolism-in-ct-images?select=FUMPE\" target=\"_blank\">Kaggle dataset</a></li>\n</ul>\n<p>This is a PE segmentation dataset prepared by two radiologists. Dataset details:</p>\n<ul>\n<li>35 patients, with 32 patients with PE detected and labelled</li>\n<li>Each patient with between 200 to 300 slices each, DICOM format, resolution 512x512, very similar to competition data</li>\n<li>The most important part: each slice with 512x512 binary mask, indicating a presence of PE on a single pixel</li>\n</ul>\n<p>How the masks looked (you can see them in green below):<br>\n<img src=\"https://media.springernature.com/full/springer-static/image/art%3A10.1038%2Fsdata.2018.180/MediaObjects/41597_2018_Article_BFsdata2018180_Fig3_HTML.jpg?as=webp\" alt=\"Visualization of PE masks\"></p>\n<p>Voila - I could finally understand where PE is hiding and how it looks! But - what I could also do is <br>\nto train the segmentation model and use its prediction as a 2nd channel. My idea was that predictions from segmentation model could serve as a kind of \"mask\" for NN - to indicate where it needs to look. As these are predictions from the model, this would not be a binary mask, but more like 'continuous' mask. At the very least they should be able to locate the common places where PE is visible on lung images. Such an approach was applied many times in previous lung competitions and proved to improve results.</p>\n<p>Below I attach a code snippet how I used it:</p>\n<pre><code>masking_model = torch.load('saved_models/unet/best_first_unet_1000healthy.pth')['model'].to(dev).eval()\n\nfor i, sample_batch in enumerate(train_loader):\n    optimizer.zero_grad()            \n    cur_img = sample_batch['img_augmented'].to(dev)\n    pred_mask = masking_model(cur_img)\n    cur_img_masked = torch.cat([pred_mask, cur_img], axis = 1)\n    preds = torch.stack(model(cur_img_masked), axis = 1).view(cur_sample_len, -1)        \n</code></pre>\n<p>I also attach a model trained on 512x512 images – you may try to augment your pipeline with this and check if the score improves.</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/narsil/rsna-pe-competition-segmentation-model-trained\" target=\"_blank\">Model link</a></li>\n</ul>\n<p>After I included this in the model my CV / LB went from 0.36 to 0.25 a week before the deadline. Then the “baseline” kernel came and ruined everything 😊 Anyway, the improvement was so large probably because my base NN was bad – but I make this model available for you to run your own experiments and maybe push the PE detection even further - for the good of everyone.</p>",
  "messages": [
    {
      "id": 1061662,
      "postDate": "2020-10-27T07:26:07.430Z",
      "content": "<p>First of all - big thanks to everyone in this competition - Kaggle is truly an amazing place to learn. In particular, big thanks to my teammate <a href=\"https://www.kaggle.com/allvor\" target=\"_blank\">@allvor</a> who provided great ideas throughout the competition.</p>\n<p>To start - any competition I participate in, I like to do error analysis. Seeing what my model predicts and comparing it to the ground truth can lead to good insights on how the model works and how It could be improved. In this competition, since I don’t any medical background, it was very difficult to even spot areas related to image-level PE, not to mention other labels (maybe except RV/LV label – when after some research I knew more or less where to look). Looking for data to enable error analysis led me to find this:</p>\n<ul>\n<li><a href=\"https://www.nature.com/articles/sdata2018180\" target=\"_blank\">Nature article - A new dataset of computed-tomography angiography images for computer-aided detection of pulmonary embolism</a></li>\n<li><a href=\"https://www.kaggle.com/andrewmvd/pulmonary-embolism-in-ct-images?select=FUMPE\" target=\"_blank\">Kaggle dataset</a></li>\n</ul>\n<p>This is a PE segmentation dataset prepared by two radiologists. Dataset details:</p>\n<ul>\n<li>35 patients, with 32 patients with PE detected and labelled</li>\n<li>Each patient with between 200 to 300 slices each, DICOM format, resolution 512x512, very similar to competition data</li>\n<li>The most important part: each slice with 512x512 binary mask, indicating a presence of PE on a single pixel</li>\n</ul>\n<p>How the masks looked (you can see them in green below):<br>\n<img src=\"https://media.springernature.com/full/springer-static/image/art%3A10.1038%2Fsdata.2018.180/MediaObjects/41597_2018_Article_BFsdata2018180_Fig3_HTML.jpg?as=webp\" alt=\"Visualization of PE masks\"></p>\n<p>Voila - I could finally understand where PE is hiding and how it looks! But - what I could also do is <br>\nto train the segmentation model and use its prediction as a 2nd channel. My idea was that predictions from segmentation model could serve as a kind of \"mask\" for NN - to indicate where it needs to look. As these are predictions from the model, this would not be a binary mask, but more like 'continuous' mask. At the very least they should be able to locate the common places where PE is visible on lung images. Such an approach was applied many times in previous lung competitions and proved to improve results.</p>\n<p>Below I attach a code snippet how I used it:</p>\n<pre><code>masking_model = torch.load('saved_models/unet/best_first_unet_1000healthy.pth')['model'].to(dev).eval()\n\nfor i, sample_batch in enumerate(train_loader):\n    optimizer.zero_grad()            \n    cur_img = sample_batch['img_augmented'].to(dev)\n    pred_mask = masking_model(cur_img)\n    cur_img_masked = torch.cat([pred_mask, cur_img], axis = 1)\n    preds = torch.stack(model(cur_img_masked), axis = 1).view(cur_sample_len, -1)        \n</code></pre>\n<p>I also attach a model trained on 512x512 images – you may try to augment your pipeline with this and check if the score improves.</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/narsil/rsna-pe-competition-segmentation-model-trained\" target=\"_blank\">Model link</a></li>\n</ul>\n<p>After I included this in the model my CV / LB went from 0.36 to 0.25 a week before the deadline. Then the “baseline” kernel came and ruined everything 😊 Anyway, the improvement was so large probably because my base NN was bad – but I make this model available for you to run your own experiments and maybe push the PE detection even further - for the good of everyone.</p>",
      "rawMarkdown": "First of all - big thanks to everyone in this competition - Kaggle is truly an amazing place to learn. In particular, big thanks to my teammate @allvor who provided great ideas throughout the competition.\n\nTo start - any competition I participate in, I like to do error analysis. Seeing what my model predicts and comparing it to the ground truth can lead to good insights on how the model works and how It could be improved. In this competition, since I don’t any medical background, it was very difficult to even spot areas related to image-level PE, not to mention other labels (maybe except RV/LV label – when after some research I knew more or less where to look). Looking for data to enable error analysis led me to find this:\n- [Nature article - A new dataset of computed-tomography angiography images for computer-aided detection of pulmonary embolism](https://www.nature.com/articles/sdata2018180)\n- [Kaggle dataset](https://www.kaggle.com/andrewmvd/pulmonary-embolism-in-ct-images?select=FUMPE)\n\nThis is a PE segmentation dataset prepared by two radiologists. Dataset details:\n- 35 patients, with 32 patients with PE detected and labelled\n- Each patient with between 200 to 300 slices each, DICOM format, resolution 512x512, very similar to competition data\n- The most important part: each slice with 512x512 binary mask, indicating a presence of PE on a single pixel\n\nHow the masks looked (you can see them in green below):\n![Visualization of PE masks](https://media.springernature.com/full/springer-static/image/art%3A10.1038%2Fsdata.2018.180/MediaObjects/41597_2018_Article_BFsdata2018180_Fig3_HTML.jpg?as=webp)\n\nVoila - I could finally understand where PE is hiding and how it looks! But - what I could also do is \nto train the segmentation model and use its prediction as a 2nd channel. My idea was that predictions from segmentation model could serve as a kind of \"mask\" for NN - to indicate where it needs to look. As these are predictions from the model, this would not be a binary mask, but more like 'continuous' mask. At the very least they should be able to locate the common places where PE is visible on lung images. Such an approach was applied many times in previous lung competitions and proved to improve results.\n\n Below I attach a code snippet how I used it:\n```\n\nmasking_model = torch.load('saved_models/unet/best_first_unet_1000healthy.pth')['model'].to(dev).eval()\n\nfor i, sample_batch in enumerate(train_loader):\n    optimizer.zero_grad()            \n    cur_img = sample_batch['img_augmented'].to(dev)\n    pred_mask = masking_model(cur_img)\n    cur_img_masked = torch.cat([pred_mask, cur_img], axis = 1)\n    preds = torch.stack(model(cur_img_masked), axis = 1).view(cur_sample_len, -1)        \n```\n\nI also attach a model trained on 512x512 images – you may try to augment your pipeline with this and check if the score improves.\n- [Model link](https://www.kaggle.com/narsil/rsna-pe-competition-segmentation-model-trained)\n\nAfter I included this in the model my CV / LB went from 0.36 to 0.25 a week before the deadline. Then the “baseline” kernel came and ruined everything 😊 Anyway, the improvement was so large probably because my base NN was bad – but I make this model available for you to run your own experiments and maybe push the PE detection even further - for the good of everyone.\n",
      "votes": 11
    },
    {
      "id": 1062585,
      "postDate": "2020-10-28T02:19:08.550Z",
      "content": "<p>really nice idea :)</p>",
      "rawMarkdown": "really nice idea :)",
      "votes": 1,
      "replies": [
        {
          "id": 1062716,
          "postDate": "2020-10-28T06:06:56.030Z",
          "content": "<p>Thanks! Have you tried using it in your top-notch pipeline? Wondering if it would improve your score, even by a small margin</p>",
          "rawMarkdown": "Thanks! Have you tried using it in your top-notch pipeline? Wondering if it would improve your score, even by a small margin"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1062585,
      "author_name": "DatNT",
      "author_url": "",
      "post_date": "2020-10-28T02:19:08.550000",
      "content": "<p>really nice idea :)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1062716,
          "author_name": "narsil (jobs-in-data.com)",
          "author_url": "",
          "post_date": "2020-10-28T06:06:56.030000",
          "content": "<p>Thanks! Have you tried using it in your top-notch pipeline? Wondering if it would improve your score, even by a small margin</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1061662": "First of all - big thanks to everyone in this competition - Kaggle is truly an amazing place to learn. In particular, big thanks to my teammate @allvor who provided great ideas throughout the competition.\n\nTo start - any competition I participate in, I like to do error analysis. Seeing what my model predicts and comparing it to the ground truth can lead to good insights on how the model works and how It could be improved. In this competition, since I don’t any medical background, it was very difficult to even spot areas related to image-level PE, not to mention other labels (maybe except RV/LV label – when after some research I knew more or less where to look). Looking for data to enable error analysis led me to find this:\n- [Nature article - A new dataset of computed-tomography angiography images for computer-aided detection of pulmonary embolism](https://www.nature.com/articles/sdata2018180)\n- [Kaggle dataset](https://www.kaggle.com/andrewmvd/pulmonary-embolism-in-ct-images?select=FUMPE)\n\nThis is a PE segmentation dataset prepared by two radiologists. Dataset details:\n- 35 patients, with 32 patients with PE detected and labelled\n- Each patient with between 200 to 300 slices each, DICOM format, resolution 512x512, very similar to competition data\n- The most important part: each slice with 512x512 binary mask, indicating a presence of PE on a single pixel\n\nHow the masks looked (you can see them in green below):\n![Visualization of PE masks](https://media.springernature.com/full/springer-static/image/art%3A10.1038%2Fsdata.2018.180/MediaObjects/41597_2018_Article_BFsdata2018180_Fig3_HTML.jpg?as=webp)\n\nVoila - I could finally understand where PE is hiding and how it looks! But - what I could also do is \nto train the segmentation model and use its prediction as a 2nd channel. My idea was that predictions from segmentation model could serve as a kind of \"mask\" for NN - to indicate where it needs to look. As these are predictions from the model, this would not be a binary mask, but more like 'continuous' mask. At the very least they should be able to locate the common places where PE is visible on lung images. Such an approach was applied many times in previous lung competitions and proved to improve results.\n\n Below I attach a code snippet how I used it:\n```\n\nmasking_model = torch.load('saved_models/unet/best_first_unet_1000healthy.pth')['model'].to(dev).eval()\n\nfor i, sample_batch in enumerate(train_loader):\n    optimizer.zero_grad()            \n    cur_img = sample_batch['img_augmented'].to(dev)\n    pred_mask = masking_model(cur_img)\n    cur_img_masked = torch.cat([pred_mask, cur_img], axis = 1)\n    preds = torch.stack(model(cur_img_masked), axis = 1).view(cur_sample_len, -1)        \n```\n\nI also attach a model trained on 512x512 images – you may try to augment your pipeline with this and check if the score improves.\n- [Model link](https://www.kaggle.com/narsil/rsna-pe-competition-segmentation-model-trained)\n\nAfter I included this in the model my CV / LB went from 0.36 to 0.25 a week before the deadline. Then the “baseline” kernel came and ruined everything 😊 Anyway, the improvement was so large probably because my base NN was bad – but I make this model available for you to run your own experiments and maybe push the PE detection even further - for the good of everyone.\n",
    "1062585": "really nice idea :)"
  }
}