{
  "id": 193401,
  "title": "[2nd place] Solution Overview & Code",
  "url": "/competitions/rsna-str-pulmonary-embolism-detection/discussion/193401",
  "author_name": "Ian Pan",
  "post_date": "2020-10-27T00:02:10.974000",
  "votes": 135,
  "comment_count": 39,
  "views": 0,
  "content": "<p>Update: code available at <a href=\"https://github.com/i-pan/kaggle-rsna-pe\" target=\"_blank\">https://github.com/i-pan/kaggle-rsna-pe</a></p>\n<p>Congratulations to all the participants and the winners. Special congrats to prize winners <a href=\"https://www.kaggle.com/osciiart\" target=\"_blank\">@osciiart</a> and <a href=\"https://www.kaggle.com/jpbremer\" target=\"_blank\">@jpbremer</a> who are on track to be physician GMs. This was a tough, compute-heavy challenge given the large amount of data and short timeframe. My setup was 4 24 GB Quadro RTX 6000 GPUs. I always feel guilty during competitions like these since I have the luxury of strong compute. Models were trained using DDP in PyTorch 1.6 with automatic mixed precision. </p>\n<p>Even though the results aren't final yet and my submission may be removed for violating the label consistency requirements (hopefully my heuristic for fixing those worked!), I still wanted to share my solution.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F281652%2Ff6a15edbc81d867c0faecd0894d6aa27%2Fpe.png?generation=1603756913344030&amp;alt=media\" alt=\"\"></p>\n<p>Here is a schematic outlining my solution. It has a lot of moving parts, so I apologize for the lengthy summary. </p>\n<h1>Step 1: Feature Extraction</h1>\n<p>Last year's RSNA Intracranial Hemorrhage Detection challenge shared a lot of similarities with this year's challenge. Most of the top solutions combined 2D CNN feature extraction with sequence modeling. The backbone of my solution also relied on a similar setup. </p>\n<p>I first trained ResNeSt50 on 2D images. Images were windowed using the PE-specific window (WL=100, WW=700) that I mentioned in one of my initial posts on DICOM processing. Each \"image\" was 3 channels, with each channel representing an individual slice. Thus the 2D image was a stack of 3 continuous slices. The targets were the 7 PE-related labels (i.e., excluding the RV/LV ratio labels). Note: most of the PE-related labels were exam-level labels. However, I just assigned  the exam labels to each slice positive for PE (negative slices had all zeros), with the understanding that there would be label noise. Also, I predicted the labels for the middle slice among the 3 slices in the image. Models were trained using vanilla binary cross-entropy loss (<code>BCEWithLogitsLoss</code> in PyTorch), 512x512 with 448x448 random crops (single center crop during inference), RandAugment data augmentation, batch size 128, 5 epochs, 2500 steps per epoch, RAdam optimizer with cosine annealing learning rate scheduler. I used generalized mean pooling and reduced the final feature vector to 512-D. Mean loss was around 0.06 (AUC 0.95-0.96 for slice-wise prediction of PE vs. no PE). Features were then extracted for all slices. </p>\n<h1>Step 2: Sequence Modeling</h1>\n<p>Many of last year's solutions used LSTMs/GRUs as the sequence model of choice. For this competition, I used <code>huggingface</code> transformers (specifically, the <code>Transformer</code> class from <code>transformers.modeling_distilbert</code>). I used one 4-layer transformer to produce slice-wise <code>pe_present_on_image</code> predictions and another transformer to predict exam-wise PE-labels. Sequence length was 512 during training, padded/truncated as necessary. During inference, I used the sequence without modifications. </p>\n<p>Important point: <strong>Images from non-PE exams do not contribute to the loss.</strong> At first, I was training the slice-wise transformer on all exams. Then, I decided to train these models on positive exams only. This lowered my CV by about 0.01-0.02. I used a custom weighted loss where I weighted the loss from each example by the proportion of positive PE slices (as described in the metric), though I'm still not sure I wrote it correctly. </p>\n<p>The exam-level transformer was trained using a weighted BCE loss based on the competition label weights. Exam-level validation losses ranged from 0.15-0.17.  </p>\n<h1>Step 3: Time-Distributed CNN</h1>\n<p>To add some variety into my modeling, I then trained a time-distributed CNN, which is just another way of saying I stacked a transformer on top of a CNN feature extraction backbone and trained end-to-end. </p>\n<p>But before doing that, I performed inference using the slice-wise transformer model to get PE scores for every slice (5-fold OOF predictions). Then, when training the TD-CNN, I only trained on the top 30% of slices from each exam, sorted by PE score. These were trained on 3D volumes of size 32x416x416 cropped to 32x364x364 using the same windowing strategy (WL=100, WW=700) in batches of 16. </p>\n<p>I initialized the CNN backbone and the transformer head with trained models from steps 1 and 2 to help with convergence. I forced all the batch normalization layers in the backbone to <code>eval</code> mode as well- this prevents the running mean and variance in each layer from updating and only trains the coefficients. Exam-level validation losses ranged from 0.15-0.17, similar to step 2. </p>\n<h1>Step 4: Heart Slice Prediction</h1>\n<p>RV/LV ratio is a significant portion of the loss. I hand-labeled slices with heart in 1,000 CT scans and trained a model (EfficientNet-B1 pruned, AUC 0.998, 256x256-&gt;224x224 crops) to classify heart slices in each CT scan. I did this because I felt that by focusing a model on the heart, I could get better, more consistent results across scans. It actually wasn't hard to label 1,000 scans- probably a full day's worth of work. I just needed to find the top and bottom heart slices; everything in between thus must also contain the heart.</p>\n<h1>Step 5: RV/LV 3D CNN</h1>\n<p>I trained a 3D CNN to classify RV/LV ratio. Specifically, I used a 101-layer channel separated network, pretrained on 65 million Instagram videos (<a href=\"https://arxiv.org/abs/1904.02811\" target=\"_blank\">https://arxiv.org/abs/1904.02811</a>, <a href=\"https://github.com/facebookresearch/VMZ)\" target=\"_blank\">https://github.com/facebookresearch/VMZ)</a>. </p>\n<p>This model was trained only on heart slices from each exam. Also, it was only trained on positive exams. This is because RV/LV ratio was not labeled for negative exams- both RV/LV labels are 0. Thus, it didn't make sense to me to try and train my model on the entire dataset's labels directly. Models were trained using a weighted BCE loss using the competition weights. I resized the input to 64x256x256-&gt;64x224x224 crops and used mediastinal window (WL=50, WW=350). AUC for RV/LV ratio &gt; 1 was about 0.85. The validation loss ranged from 0.44-0.48. I then used these models to extract 2048-D features from each exam.</p>\n<p>The challenge here was that the likelihood of an exam being PE-positive greatly influenced the RV/LV labels. From step 2, I calculated 5-fold OOF exam-level predictions for all exams. Then, I trained a linear model that took as input the concatenation of the 2048-D 3D CNN feature and the 7 PE exam labels. This linear model was trained across <strong>all exams</strong> using the labels directly. This way, the model could take into account the imaging features from the scan but also adjust the predictions based on the likelihood of PE. Validation losses ranged from 0.22-0.25. </p>\n<p>I didn't validate my entire pipeline that often due to time constraints, instead choosing to focus on the individual component losses and optimizing each one as well as I could. I did periodically do a sanity-check on 200 single-fold exams to make sure that my entire pipeline was working, and my final validation loss was 0.183 for my 0.150 public LB submission. I'm still not confident that I implemented the metric correctly, but I was seeing good correlation (0.159/0.195-&gt;0.156/0.188-&gt; 0.150/0.183).</p>\n<p>At the end, I applied a function to enforce label consistency requirements for each exam. There was about 0.001 change in CV after applying this function. I have some thoughts about this requirement which maybe I'll save for another post. </p>\n<h1>Final Models</h1>\n<p>Overall, I used these models:<br>\n-2x ResNeSt50 feature extractors<br>\n-6x exam transformers (3 for each extractor)<br>\n-6x slice transformers<br>\n-5x ResNeSt50 TD-CNN<br>\n-1x EfficientNet-B1 pruned heart slice classifier<br>\n-5x ip-CSN-101 3D CNN RV/LV feature extractor<br>\n-5x RV/LV linear model </p>\n<p>Inference took about 7-8 hours for the entire test set. </p>\n<p>Code to follow once I clean it up. </p>\n<p>Things that didn't work as well:<br>\n-3D CNN for PE exam-level prediction<br>\n-Pseudolabeling RV/LV ratio for negative exams<br>\n-Using logits/probabilities instead of features for second-stage model<br>\n-TD-CNN for RV/LV ratio prediction<br>\n-Training feature extractor on positive exams only to improve slice-wise PE prediction </p>",
  "messages": [
    {
      "id": 1061321,
      "postDate": "2020-10-27T00:02:10.973Z",
      "content": "<p>Update: code available at <a href=\"https://github.com/i-pan/kaggle-rsna-pe\" target=\"_blank\">https://github.com/i-pan/kaggle-rsna-pe</a></p>\n<p>Congratulations to all the participants and the winners. Special congrats to prize winners <a href=\"https://www.kaggle.com/osciiart\" target=\"_blank\">@osciiart</a> and <a href=\"https://www.kaggle.com/jpbremer\" target=\"_blank\">@jpbremer</a> who are on track to be physician GMs. This was a tough, compute-heavy challenge given the large amount of data and short timeframe. My setup was 4 24 GB Quadro RTX 6000 GPUs. I always feel guilty during competitions like these since I have the luxury of strong compute. Models were trained using DDP in PyTorch 1.6 with automatic mixed precision. </p>\n<p>Even though the results aren't final yet and my submission may be removed for violating the label consistency requirements (hopefully my heuristic for fixing those worked!), I still wanted to share my solution.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F281652%2Ff6a15edbc81d867c0faecd0894d6aa27%2Fpe.png?generation=1603756913344030&amp;alt=media\" alt=\"\"></p>\n<p>Here is a schematic outlining my solution. It has a lot of moving parts, so I apologize for the lengthy summary. </p>\n<h1>Step 1: Feature Extraction</h1>\n<p>Last year's RSNA Intracranial Hemorrhage Detection challenge shared a lot of similarities with this year's challenge. Most of the top solutions combined 2D CNN feature extraction with sequence modeling. The backbone of my solution also relied on a similar setup. </p>\n<p>I first trained ResNeSt50 on 2D images. Images were windowed using the PE-specific window (WL=100, WW=700) that I mentioned in one of my initial posts on DICOM processing. Each \"image\" was 3 channels, with each channel representing an individual slice. Thus the 2D image was a stack of 3 continuous slices. The targets were the 7 PE-related labels (i.e., excluding the RV/LV ratio labels). Note: most of the PE-related labels were exam-level labels. However, I just assigned  the exam labels to each slice positive for PE (negative slices had all zeros), with the understanding that there would be label noise. Also, I predicted the labels for the middle slice among the 3 slices in the image. Models were trained using vanilla binary cross-entropy loss (<code>BCEWithLogitsLoss</code> in PyTorch), 512x512 with 448x448 random crops (single center crop during inference), RandAugment data augmentation, batch size 128, 5 epochs, 2500 steps per epoch, RAdam optimizer with cosine annealing learning rate scheduler. I used generalized mean pooling and reduced the final feature vector to 512-D. Mean loss was around 0.06 (AUC 0.95-0.96 for slice-wise prediction of PE vs. no PE). Features were then extracted for all slices. </p>\n<h1>Step 2: Sequence Modeling</h1>\n<p>Many of last year's solutions used LSTMs/GRUs as the sequence model of choice. For this competition, I used <code>huggingface</code> transformers (specifically, the <code>Transformer</code> class from <code>transformers.modeling_distilbert</code>). I used one 4-layer transformer to produce slice-wise <code>pe_present_on_image</code> predictions and another transformer to predict exam-wise PE-labels. Sequence length was 512 during training, padded/truncated as necessary. During inference, I used the sequence without modifications. </p>\n<p>Important point: <strong>Images from non-PE exams do not contribute to the loss.</strong> At first, I was training the slice-wise transformer on all exams. Then, I decided to train these models on positive exams only. This lowered my CV by about 0.01-0.02. I used a custom weighted loss where I weighted the loss from each example by the proportion of positive PE slices (as described in the metric), though I'm still not sure I wrote it correctly. </p>\n<p>The exam-level transformer was trained using a weighted BCE loss based on the competition label weights. Exam-level validation losses ranged from 0.15-0.17.  </p>\n<h1>Step 3: Time-Distributed CNN</h1>\n<p>To add some variety into my modeling, I then trained a time-distributed CNN, which is just another way of saying I stacked a transformer on top of a CNN feature extraction backbone and trained end-to-end. </p>\n<p>But before doing that, I performed inference using the slice-wise transformer model to get PE scores for every slice (5-fold OOF predictions). Then, when training the TD-CNN, I only trained on the top 30% of slices from each exam, sorted by PE score. These were trained on 3D volumes of size 32x416x416 cropped to 32x364x364 using the same windowing strategy (WL=100, WW=700) in batches of 16. </p>\n<p>I initialized the CNN backbone and the transformer head with trained models from steps 1 and 2 to help with convergence. I forced all the batch normalization layers in the backbone to <code>eval</code> mode as well- this prevents the running mean and variance in each layer from updating and only trains the coefficients. Exam-level validation losses ranged from 0.15-0.17, similar to step 2. </p>\n<h1>Step 4: Heart Slice Prediction</h1>\n<p>RV/LV ratio is a significant portion of the loss. I hand-labeled slices with heart in 1,000 CT scans and trained a model (EfficientNet-B1 pruned, AUC 0.998, 256x256-&gt;224x224 crops) to classify heart slices in each CT scan. I did this because I felt that by focusing a model on the heart, I could get better, more consistent results across scans. It actually wasn't hard to label 1,000 scans- probably a full day's worth of work. I just needed to find the top and bottom heart slices; everything in between thus must also contain the heart.</p>\n<h1>Step 5: RV/LV 3D CNN</h1>\n<p>I trained a 3D CNN to classify RV/LV ratio. Specifically, I used a 101-layer channel separated network, pretrained on 65 million Instagram videos (<a href=\"https://arxiv.org/abs/1904.02811\" target=\"_blank\">https://arxiv.org/abs/1904.02811</a>, <a href=\"https://github.com/facebookresearch/VMZ)\" target=\"_blank\">https://github.com/facebookresearch/VMZ)</a>. </p>\n<p>This model was trained only on heart slices from each exam. Also, it was only trained on positive exams. This is because RV/LV ratio was not labeled for negative exams- both RV/LV labels are 0. Thus, it didn't make sense to me to try and train my model on the entire dataset's labels directly. Models were trained using a weighted BCE loss using the competition weights. I resized the input to 64x256x256-&gt;64x224x224 crops and used mediastinal window (WL=50, WW=350). AUC for RV/LV ratio &gt; 1 was about 0.85. The validation loss ranged from 0.44-0.48. I then used these models to extract 2048-D features from each exam.</p>\n<p>The challenge here was that the likelihood of an exam being PE-positive greatly influenced the RV/LV labels. From step 2, I calculated 5-fold OOF exam-level predictions for all exams. Then, I trained a linear model that took as input the concatenation of the 2048-D 3D CNN feature and the 7 PE exam labels. This linear model was trained across <strong>all exams</strong> using the labels directly. This way, the model could take into account the imaging features from the scan but also adjust the predictions based on the likelihood of PE. Validation losses ranged from 0.22-0.25. </p>\n<p>I didn't validate my entire pipeline that often due to time constraints, instead choosing to focus on the individual component losses and optimizing each one as well as I could. I did periodically do a sanity-check on 200 single-fold exams to make sure that my entire pipeline was working, and my final validation loss was 0.183 for my 0.150 public LB submission. I'm still not confident that I implemented the metric correctly, but I was seeing good correlation (0.159/0.195-&gt;0.156/0.188-&gt; 0.150/0.183).</p>\n<p>At the end, I applied a function to enforce label consistency requirements for each exam. There was about 0.001 change in CV after applying this function. I have some thoughts about this requirement which maybe I'll save for another post. </p>\n<h1>Final Models</h1>\n<p>Overall, I used these models:<br>\n-2x ResNeSt50 feature extractors<br>\n-6x exam transformers (3 for each extractor)<br>\n-6x slice transformers<br>\n-5x ResNeSt50 TD-CNN<br>\n-1x EfficientNet-B1 pruned heart slice classifier<br>\n-5x ip-CSN-101 3D CNN RV/LV feature extractor<br>\n-5x RV/LV linear model </p>\n<p>Inference took about 7-8 hours for the entire test set. </p>\n<p>Code to follow once I clean it up. </p>\n<p>Things that didn't work as well:<br>\n-3D CNN for PE exam-level prediction<br>\n-Pseudolabeling RV/LV ratio for negative exams<br>\n-Using logits/probabilities instead of features for second-stage model<br>\n-TD-CNN for RV/LV ratio prediction<br>\n-Training feature extractor on positive exams only to improve slice-wise PE prediction </p>",
      "rawMarkdown": "Update: code available at https://github.com/i-pan/kaggle-rsna-pe\n\nCongratulations to all the participants and the winners. Special congrats to prize winners @osciiart and @jpbremer who are on track to be physician GMs. This was a tough, compute-heavy challenge given the large amount of data and short timeframe. My setup was 4 24 GB Quadro RTX 6000 GPUs. I always feel guilty during competitions like these since I have the luxury of strong compute. Models were trained using DDP in PyTorch 1.6 with automatic mixed precision. \n\nEven though the results aren't final yet and my submission may be removed for violating the label consistency requirements (hopefully my heuristic for fixing those worked!), I still wanted to share my solution.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F281652%2Ff6a15edbc81d867c0faecd0894d6aa27%2Fpe.png?generation=1603756913344030&alt=media)\n\nHere is a schematic outlining my solution. It has a lot of moving parts, so I apologize for the lengthy summary. \n\n# Step 1: Feature Extraction\n\nLast year's RSNA Intracranial Hemorrhage Detection challenge shared a lot of similarities with this year's challenge. Most of the top solutions combined 2D CNN feature extraction with sequence modeling. The backbone of my solution also relied on a similar setup. \n\nI first trained ResNeSt50 on 2D images. Images were windowed using the PE-specific window (WL=100, WW=700) that I mentioned in one of my initial posts on DICOM processing. Each \"image\" was 3 channels, with each channel representing an individual slice. Thus the 2D image was a stack of 3 continuous slices. The targets were the 7 PE-related labels (i.e., excluding the RV/LV ratio labels). Note: most of the PE-related labels were exam-level labels. However, I just assigned  the exam labels to each slice positive for PE (negative slices had all zeros), with the understanding that there would be label noise. Also, I predicted the labels for the middle slice among the 3 slices in the image. Models were trained using vanilla binary cross-entropy loss (`BCEWithLogitsLoss` in PyTorch), 512x512 with 448x448 random crops (single center crop during inference), RandAugment data augmentation, batch size 128, 5 epochs, 2500 steps per epoch, RAdam optimizer with cosine annealing learning rate scheduler. I used generalized mean pooling and reduced the final feature vector to 512-D. Mean loss was around 0.06 (AUC 0.95-0.96 for slice-wise prediction of PE vs. no PE). Features were then extracted for all slices. \n\n# Step 2: Sequence Modeling\n\nMany of last year's solutions used LSTMs/GRUs as the sequence model of choice. For this competition, I used `huggingface` transformers (specifically, the `Transformer` class from `transformers.modeling_distilbert`). I used one 4-layer transformer to produce slice-wise `pe_present_on_image` predictions and another transformer to predict exam-wise PE-labels. Sequence length was 512 during training, padded/truncated as necessary. During inference, I used the sequence without modifications. \n\nImportant point: **Images from non-PE exams do not contribute to the loss.** At first, I was training the slice-wise transformer on all exams. Then, I decided to train these models on positive exams only. This lowered my CV by about 0.01-0.02. I used a custom weighted loss where I weighted the loss from each example by the proportion of positive PE slices (as described in the metric), though I'm still not sure I wrote it correctly. \n\nThe exam-level transformer was trained using a weighted BCE loss based on the competition label weights. Exam-level validation losses ranged from 0.15-0.17.  \n\n# Step 3: Time-Distributed CNN\n\nTo add some variety into my modeling, I then trained a time-distributed CNN, which is just another way of saying I stacked a transformer on top of a CNN feature extraction backbone and trained end-to-end. \n\nBut before doing that, I performed inference using the slice-wise transformer model to get PE scores for every slice (5-fold OOF predictions). Then, when training the TD-CNN, I only trained on the top 30% of slices from each exam, sorted by PE score. These were trained on 3D volumes of size 32x416x416 cropped to 32x364x364 using the same windowing strategy (WL=100, WW=700) in batches of 16. \n\nI initialized the CNN backbone and the transformer head with trained models from steps 1 and 2 to help with convergence. I forced all the batch normalization layers in the backbone to `eval` mode as well- this prevents the running mean and variance in each layer from updating and only trains the coefficients. Exam-level validation losses ranged from 0.15-0.17, similar to step 2. \n\n# Step 4: Heart Slice Prediction\n\nRV/LV ratio is a significant portion of the loss. I hand-labeled slices with heart in 1,000 CT scans and trained a model (EfficientNet-B1 pruned, AUC 0.998, 256x256->224x224 crops) to classify heart slices in each CT scan. I did this because I felt that by focusing a model on the heart, I could get better, more consistent results across scans. It actually wasn't hard to label 1,000 scans- probably a full day's worth of work. I just needed to find the top and bottom heart slices; everything in between thus must also contain the heart.\n\n# Step 5: RV/LV 3D CNN\n\nI trained a 3D CNN to classify RV/LV ratio. Specifically, I used a 101-layer channel separated network, pretrained on 65 million Instagram videos (https://arxiv.org/abs/1904.02811, https://github.com/facebookresearch/VMZ). \n\nThis model was trained only on heart slices from each exam. Also, it was only trained on positive exams. This is because RV/LV ratio was not labeled for negative exams- both RV/LV labels are 0. Thus, it didn't make sense to me to try and train my model on the entire dataset's labels directly. Models were trained using a weighted BCE loss using the competition weights. I resized the input to 64x256x256->64x224x224 crops and used mediastinal window (WL=50, WW=350). AUC for RV/LV ratio > 1 was about 0.85. The validation loss ranged from 0.44-0.48. I then used these models to extract 2048-D features from each exam.\n\nThe challenge here was that the likelihood of an exam being PE-positive greatly influenced the RV/LV labels. From step 2, I calculated 5-fold OOF exam-level predictions for all exams. Then, I trained a linear model that took as input the concatenation of the 2048-D 3D CNN feature and the 7 PE exam labels. This linear model was trained across **all exams** using the labels directly. This way, the model could take into account the imaging features from the scan but also adjust the predictions based on the likelihood of PE. Validation losses ranged from 0.22-0.25. \n\nI didn't validate my entire pipeline that often due to time constraints, instead choosing to focus on the individual component losses and optimizing each one as well as I could. I did periodically do a sanity-check on 200 single-fold exams to make sure that my entire pipeline was working, and my final validation loss was 0.183 for my 0.150 public LB submission. I'm still not confident that I implemented the metric correctly, but I was seeing good correlation (0.159/0.195->0.156/0.188-> 0.150/0.183).\n\nAt the end, I applied a function to enforce label consistency requirements for each exam. There was about 0.001 change in CV after applying this function. I have some thoughts about this requirement which maybe I'll save for another post. \n\n# Final Models\n\nOverall, I used these models:\n-2x ResNeSt50 feature extractors\n-6x exam transformers (3 for each extractor)\n-6x slice transformers\n-5x ResNeSt50 TD-CNN\n-1x EfficientNet-B1 pruned heart slice classifier\n-5x ip-CSN-101 3D CNN RV/LV feature extractor\n-5x RV/LV linear model \n\nInference took about 7-8 hours for the entire test set. \n\nCode to follow once I clean it up. \n\nThings that didn't work as well:\n-3D CNN for PE exam-level prediction\n-Pseudolabeling RV/LV ratio for negative exams\n-Using logits/probabilities instead of features for second-stage model\n-TD-CNN for RV/LV ratio prediction\n-Training feature extractor on positive exams only to improve slice-wise PE prediction ",
      "votes": 134
    },
    {
      "id": 1061377,
      "postDate": "2020-10-27T01:03:18.117Z",
      "content": "<p>Congratulations! Great solution write up. Thank you. So many ingenuities squeezed into one diagram.</p>\n<p>I have just one question. How did you find time to do all that as an intern in a hospital as busy as yours?</p>",
      "rawMarkdown": "Congratulations! Great solution write up. Thank you. So many ingenuities squeezed into one diagram.\n\nI have just one question. How did you find time to do all that as an intern in a hospital as busy as yours?",
      "votes": 5,
      "replies": [
        {
          "id": 1061927,
          "postDate": "2020-10-27T12:45:34.570Z",
          "content": "<p>Congrats Ian and Yee ! Both very impressive results. Go Radiology go !</p>\n<p>About your question Yee, I stopped thinking years ago that Ian is a human like us.</p>",
          "rawMarkdown": "Congrats Ian and Yee ! Both very impressive results. Go Radiology go !\n \nAbout your question Yee, I stopped thinking years ago that Ian is a human like us.\n",
          "votes": 2
        },
        {
          "id": 1061940,
          "postDate": "2020-10-27T13:00:57.603Z",
          "content": "<blockquote>\n  <p>I stopped thinking years ago that Ian is a human like us.</p>\n</blockquote>\n<p>haha. Agree.</p>",
          "rawMarkdown": "> I stopped thinking years ago that Ian is a human like us.\n\nhaha. Agree.",
          "votes": 1
        },
        {
          "id": 1061973,
          "postDate": "2020-10-27T13:26:36.360Z",
          "content": "<p>Thanks, congrats on your strong finish as well! After the initial setup, it's just a matter of running experiments, most of which is just waiting for models to finish training. At the beginning of the challenge, I laid out my plan for what I was going to do and just focused on executing it. I didn't deviate from that initial plan very much. Fortunately it worked out. </p>",
          "rawMarkdown": "Thanks, congrats on your strong finish as well! After the initial setup, it's just a matter of running experiments, most of which is just waiting for models to finish training. At the beginning of the challenge, I laid out my plan for what I was going to do and just focused on executing it. I didn't deviate from that initial plan very much. Fortunately it worked out. ",
          "votes": 7
        },
        {
          "id": 1062099,
          "postDate": "2020-10-27T15:13:10.787Z",
          "content": "<p>Ian, should have figured this out years ago in bone age, but you are truly on another plane. Echoing Alexandre's sentiment. Glad you're a rad.</p>",
          "rawMarkdown": "Ian, should have figured this out years ago in bone age, but you are truly on another plane. Echoing Alexandre's sentiment. Glad you're a rad.",
          "votes": 1
        },
        {
          "id": 1064377,
          "postDate": "2020-10-30T03:52:25.100Z",
          "content": "<p>Congrats <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a>  just a thought whoch I couldn't test ,using segmented  lung as one channel instead of windowed one </p>",
          "rawMarkdown": "Congrats @vaillant  just a thought whoch I couldn't test ,using segmented  lung as one channel instead of windowed one "
        }
      ]
    },
    {
      "id": 1117660,
      "postDate": "2020-12-18T10:09:45.573Z",
      "content": "<p>Ian, this is amazing and inspiring. I had a look at the first place solution and that sent me down several rabbit holes, and I'm only getting to look at this for the first time now. I suspect this will send me on a 3 month detour!</p>\n<p>I'm hoping I could get your insights on something. There are a lot of specific choices in this setup, and I noticed your comment below about having planned this at the start, and not deviating much from it. In light of that, it would be super helpful to get your thoughts on which specific choices you are most convinced made a strong difference, and which you are still unsure of.</p>",
      "rawMarkdown": "Ian, this is amazing and inspiring. I had a look at the first place solution and that sent me down several rabbit holes, and I'm only getting to look at this for the first time now. I suspect this will send me on a 3 month detour!\n\nI'm hoping I could get your insights on something. There are a lot of specific choices in this setup, and I noticed your comment below about having planned this at the start, and not deviating much from it. In light of that, it would be super helpful to get your thoughts on which specific choices you are most convinced made a strong difference, and which you are still unsure of.",
      "votes": 1
    },
    {
      "id": 1061385,
      "postDate": "2020-10-27T01:18:25.680Z",
      "content": "<p>Excellent work- thanks for sharing your solution. I'm going to come back and read this a few more times!</p>",
      "rawMarkdown": "Excellent work- thanks for sharing your solution. I'm going to come back and read this a few more times!",
      "votes": 1
    },
    {
      "id": 1062116,
      "postDate": "2020-10-27T15:30:21.920Z",
      "content": "<p>It's a great honor for me to have such words from you. Congratulations!</p>",
      "rawMarkdown": "It's a great honor for me to have such words from you. Congratulations!",
      "votes": 2
    },
    {
      "id": 1064518,
      "postDate": "2020-10-30T07:57:47.043Z",
      "content": "<p>Wow, what a useful notebook. I'm using it for further education.</p>",
      "rawMarkdown": "Wow, what a useful notebook. I'm using it for further education."
    },
    {
      "id": 1061566,
      "postDate": "2020-10-27T04:56:01.413Z",
      "content": "<p>congratulations <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> </p>",
      "rawMarkdown": "congratulations @vaillant ",
      "votes": -1
    },
    {
      "id": 1065349,
      "postDate": "2020-10-31T07:33:36.190Z",
      "content": "<p>it is very meaningful result</p>",
      "rawMarkdown": "it is very meaningful result"
    },
    {
      "id": 1065040,
      "postDate": "2020-10-30T19:25:25.150Z",
      "content": "<p>Wow, that was a nice documentation. Really its useful. And my deep Congratulations to winner..!</p>",
      "rawMarkdown": "Wow, that was a nice documentation. Really its useful. And my deep Congratulations to winner..!"
    },
    {
      "id": 1064731,
      "postDate": "2020-10-30T13:14:39.067Z",
      "content": "<p>Yeah this one is really very impressive </p>",
      "rawMarkdown": "Yeah this one is really very impressive "
    },
    {
      "id": 1064460,
      "postDate": "2020-10-30T06:54:29.423Z",
      "content": "<p>Good job! Interesting pipeline</p>",
      "rawMarkdown": "Good job! Interesting pipeline"
    },
    {
      "id": 1064208,
      "postDate": "2020-10-29T20:50:47.477Z",
      "content": "<p>Congratulations, and Thanks for sharing :)</p>",
      "rawMarkdown": "Congratulations, and Thanks for sharing :)"
    },
    {
      "id": 1064098,
      "postDate": "2020-10-29T17:35:57.350Z",
      "content": "<p>Congratulations </p>",
      "rawMarkdown": "Congratulations "
    },
    {
      "id": 1063855,
      "postDate": "2020-10-29T12:22:37.437Z",
      "content": "<p>congrats!!!</p>",
      "rawMarkdown": "congrats!!!"
    },
    {
      "id": 1063565,
      "postDate": "2020-10-29T04:28:47.043Z",
      "content": "<p>Well done!</p>",
      "rawMarkdown": "Well done!"
    },
    {
      "id": 1063366,
      "postDate": "2020-10-28T19:35:38.857Z",
      "content": "<p>Congrats , great job !</p>",
      "rawMarkdown": "Congrats , great job !"
    },
    {
      "id": 1063307,
      "postDate": "2020-10-28T18:14:33.700Z",
      "content": "<p>I am jealous of your setup. :p<br>\nJoke aside, very well done. You said:</p>\n<blockquote>\n  <p>Inference took about 7-8 hours for the entire test set. </p>\n</blockquote>\n<p>Do you have a rough estimate of the training time? Many thanks and well done again!</p>",
      "rawMarkdown": "I am jealous of your setup. :p\nJoke aside, very well done. You said:\n\n> Inference took about 7-8 hours for the entire test set. \n\nDo you have a rough estimate of the training time? Many thanks and well done again!",
      "replies": [
        {
          "id": 1063415,
          "postDate": "2020-10-28T21:58:09.563Z",
          "content": "<p>The feature extractors took about 5.5 hours to train. TD-CNNs took about 7.5-8 hours to train. The 3D CNN took about 30 minutes, and the transformers were very fast- just a few minutes. </p>",
          "rawMarkdown": "The feature extractors took about 5.5 hours to train. TD-CNNs took about 7.5-8 hours to train. The 3D CNN took about 30 minutes, and the transformers were very fast- just a few minutes. ",
          "votes": 2
        }
      ]
    },
    {
      "id": 1062769,
      "postDate": "2020-10-28T06:58:43.090Z",
      "content": "<p>congrats!!</p>",
      "rawMarkdown": "congrats!!"
    },
    {
      "id": 1062633,
      "postDate": "2020-10-28T03:54:24.493Z",
      "content": "<p>Wow.That's an amazing pipeline. Congratulations!</p>",
      "rawMarkdown": "Wow.That's an amazing pipeline. Congratulations!\n"
    },
    {
      "id": 1062596,
      "postDate": "2020-10-28T02:45:00.693Z",
      "content": "<p>Provisionally 2nd place now! Thanks for a very informational write and congratulations! 👍</p>",
      "rawMarkdown": "Provisionally 2nd place now! Thanks for a very informational write and congratulations! 👍"
    },
    {
      "id": 1062111,
      "postDate": "2020-10-27T15:27:15.050Z",
      "content": "<p>Congratulations! Great solution write up. Thank you.</p>",
      "rawMarkdown": "Congratulations! Great solution write up. Thank you."
    },
    {
      "id": 1061680,
      "postDate": "2020-10-27T07:52:23.510Z",
      "content": "<p>Congrats, that is an impressive solution. And thanks for your JPEG contribution  too.</p>",
      "rawMarkdown": "Congrats, that is an impressive solution. And thanks for your JPEG contribution  too."
    },
    {
      "id": 1061675,
      "postDate": "2020-10-27T07:43:35.543Z",
      "content": "<p>I'm trully impressed by your consistency at getting strong results in medical imaging competitions. </p>\n<p>The fact that you're a physician probably plays a role, but that's even more impressive to me since it means that you already have a really busy schedule. </p>\n<p>Anyways, very elegant solution, I hope you don't get removed for that 3rd place is truly deserved.</p>",
      "rawMarkdown": "I'm trully impressed by your consistency at getting strong results in medical imaging competitions. \n\nThe fact that you're a physician probably plays a role, but that's even more impressive to me since it means that you already have a really busy schedule. \n\nAnyways, very elegant solution, I hope you don't get removed for that 3rd place is truly deserved.",
      "replies": [
        {
          "id": 1063416,
          "postDate": "2020-10-28T22:00:14.010Z",
          "content": "<p>Thank you! When I started out on Kaggle, I felt like my medical imaging domain knowledge would give me a surefire edge in these competitions. Kagglers are too smart though, so at best it's only a slight advantage. </p>",
          "rawMarkdown": "Thank you! When I started out on Kaggle, I felt like my medical imaging domain knowledge would give me a surefire edge in these competitions. Kagglers are too smart though, so at best it's only a slight advantage. ",
          "votes": 2
        }
      ]
    },
    {
      "id": 1061580,
      "postDate": "2020-10-27T05:17:56.357Z",
      "content": "<p>Congratulations and thank you for your solution. </p>\n<p>I have a question about feature extraction. Would you tell me the experiments with other models if you do (efficientnet, ResNext, densenet…)? ResNeSt50 is better than others? or ensemble didn't work?</p>",
      "rawMarkdown": "Congratulations and thank you for your solution. \n\nI have a question about feature extraction. Would you tell me the experiments with other models if you do (efficientnet, ResNext, densenet...)? ResNeSt50 is better than others? or ensemble didn't work?",
      "replies": [
        {
          "id": 1061978,
          "postDate": "2020-10-27T13:29:00.607Z",
          "content": "<p>I tried initially with EfficientNet-B4 but in the past I've found that ResNet-like architectures worked better for tasks relying on feature extraction, so I tried ResNeSt which worked better (0.01 decrease in loss). </p>\n<p>I also tried EfficientNet-B3 pruned at lower resolution to increase inference speed but this was significantly worse (0.02 increase in loss). </p>",
          "rawMarkdown": "I tried initially with EfficientNet-B4 but in the past I've found that ResNet-like architectures worked better for tasks relying on feature extraction, so I tried ResNeSt which worked better (0.01 decrease in loss). \n\nI also tried EfficientNet-B3 pruned at lower resolution to increase inference speed but this was significantly worse (0.02 increase in loss). ",
          "votes": 2
        }
      ]
    },
    {
      "id": 1061488,
      "postDate": "2020-10-27T03:16:04.190Z",
      "content": "<p>Great <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> </p>",
      "rawMarkdown": "Great @vaillant "
    },
    {
      "id": 1061457,
      "postDate": "2020-10-27T02:48:58.723Z",
      "content": "<p>thought about hand labeling the heart earlier, but then i gave up :(</p>",
      "rawMarkdown": "thought about hand labeling the heart earlier, but then i gave up :(",
      "replies": [
        {
          "id": 1061970,
          "postDate": "2020-10-27T13:24:38.727Z",
          "content": "<p>Seems like it didn't make too much of a difference. 😅</p>",
          "rawMarkdown": "Seems like it didn't make too much of a difference. 😅"
        }
      ]
    },
    {
      "id": 1061448,
      "postDate": "2020-10-27T02:37:24.063Z",
      "content": "<p>congratulations <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> !!<br>\nand thanks for sharing the images and help throughout the competition 👍</p>",
      "rawMarkdown": "congratulations @vaillant !!\nand thanks for sharing the images and help throughout the competition 👍"
    },
    {
      "id": 1061441,
      "postDate": "2020-10-27T02:33:24.340Z",
      "content": "<p>Congratulations Ian!</p>",
      "rawMarkdown": "Congratulations Ian!"
    },
    {
      "id": 1061439,
      "postDate": "2020-10-27T02:32:32.847Z",
      "content": "<p>Congratulations Ian! Really amazing approach for formulating this problem in a single end-to-end model. You combined amazing clinical/problem insights with cutting edge architecture. I hope institutions can leverage this to build better tools for PE triage! -- Also a double congrats for doing this during your intern year :)</p>",
      "rawMarkdown": "Congratulations Ian! Really amazing approach for formulating this problem in a single end-to-end model. You combined amazing clinical/problem insights with cutting edge architecture. I hope institutions can leverage this to build better tools for PE triage! -- Also a double congrats for doing this during your intern year :)"
    },
    {
      "id": 1892427,
      "postDate": "2022-08-10T05:00:14.293Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1062145,
      "postDate": "2020-10-27T15:51:28.563Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1061377,
      "author_name": "Yee Ng",
      "author_url": "",
      "post_date": "2020-10-27T01:03:18.117000",
      "content": "<p>Congratulations! Great solution write up. Thank you. So many ingenuities squeezed into one diagram.</p>\n<p>I have just one question. How did you find time to do all that as an intern in a hospital as busy as yours?</p>",
      "votes": 5,
      "replies": [
        {
          "id": 1061927,
          "author_name": "Alexandre Cadrin-Chênevert",
          "author_url": "",
          "post_date": "2020-10-27T12:45:34.570000",
          "content": "<p>Congrats Ian and Yee ! Both very impressive results. Go Radiology go !</p>\n<p>About your question Yee, I stopped thinking years ago that Ian is a human like us.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1061940,
          "author_name": "Yee Ng",
          "author_url": "",
          "post_date": "2020-10-27T13:00:57.603000",
          "content": "<blockquote>\n  <p>I stopped thinking years ago that Ian is a human like us.</p>\n</blockquote>\n<p>haha. Agree.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1061973,
          "author_name": "Ian Pan",
          "author_url": "",
          "post_date": "2020-10-27T13:26:36.360000",
          "content": "<p>Thanks, congrats on your strong finish as well! After the initial setup, it's just a matter of running experiments, most of which is just waiting for models to finish training. At the beginning of the challenge, I laid out my plan for what I was going to do and just focused on executing it. I didn't deviate from that initial plan very much. Fortunately it worked out. </p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 1062099,
          "author_name": "Jesse",
          "author_url": "",
          "post_date": "2020-10-27T15:13:10.787000",
          "content": "<p>Ian, should have figured this out years ago in bone age, but you are truly on another plane. Echoing Alexandre's sentiment. Glad you're a rad.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1064377,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-10-30T03:52:25.100000",
          "content": "<p>Congrats <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a>  just a thought whoch I couldn't test ,using segmented  lung as one channel instead of windowed one </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1117660,
      "author_name": "Alexander Soare",
      "author_url": "",
      "post_date": "2020-12-18T10:09:45.573000",
      "content": "<p>Ian, this is amazing and inspiring. I had a look at the first place solution and that sent me down several rabbit holes, and I'm only getting to look at this for the first time now. I suspect this will send me on a 3 month detour!</p>\n<p>I'm hoping I could get your insights on something. There are a lot of specific choices in this setup, and I noticed your comment below about having planned this at the start, and not deviating much from it. In light of that, it would be super helpful to get your thoughts on which specific choices you are most convinced made a strong difference, and which you are still unsure of.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1061385,
      "author_name": "Rob Mulla",
      "author_url": "",
      "post_date": "2020-10-27T01:18:25.680000",
      "content": "<p>Excellent work- thanks for sharing your solution. I'm going to come back and read this a few more times!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1062116,
      "author_name": "OsciiArt",
      "author_url": "",
      "post_date": "2020-10-27T15:30:21.920000",
      "content": "<p>It's a great honor for me to have such words from you. Congratulations!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1064518,
      "author_name": "Albert Azeri",
      "author_url": "",
      "post_date": "2020-10-30T07:57:47.043000",
      "content": "<p>Wow, what a useful notebook. I'm using it for further education.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1061566,
      "author_name": "Akash Singh",
      "author_url": "",
      "post_date": "2020-10-27T04:56:01.413000",
      "content": "<p>congratulations <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> </p>",
      "votes": -1,
      "replies": []
    },
    {
      "id": 1065349,
      "author_name": "Jiwwwooong",
      "author_url": "",
      "post_date": "2020-10-31T07:33:36.190000",
      "content": "<p>it is very meaningful result</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1065040,
      "author_name": "Shrivas@Amrith",
      "author_url": "",
      "post_date": "2020-10-30T19:25:25.150000",
      "content": "<p>Wow, that was a nice documentation. Really its useful. And my deep Congratulations to winner..!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1064731,
      "author_name": "Aman Jain",
      "author_url": "",
      "post_date": "2020-10-30T13:14:39.067000",
      "content": "<p>Yeah this one is really very impressive </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1064460,
      "author_name": "whoami",
      "author_url": "",
      "post_date": "2020-10-30T06:54:29.423000",
      "content": "<p>Good job! Interesting pipeline</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1064208,
      "author_name": "Dreamer",
      "author_url": "",
      "post_date": "2020-10-29T20:50:47.477000",
      "content": "<p>Congratulations, and Thanks for sharing :)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1064098,
      "author_name": "Mohammed Saket",
      "author_url": "",
      "post_date": "2020-10-29T17:35:57.350000",
      "content": "<p>Congratulations </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1063855,
      "author_name": "Yasser Radouani",
      "author_url": "",
      "post_date": "2020-10-29T12:22:37.437000",
      "content": "<p>congrats!!!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1063565,
      "author_name": "Issac Antony",
      "author_url": "",
      "post_date": "2020-10-29T04:28:47.043000",
      "content": "<p>Well done!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1063366,
      "author_name": "Emirhan KALEM",
      "author_url": "",
      "post_date": "2020-10-28T19:35:38.857000",
      "content": "<p>Congrats , great job !</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1063307,
      "author_name": "Yassine Alouini",
      "author_url": "",
      "post_date": "2020-10-28T18:14:33.700000",
      "content": "<p>I am jealous of your setup. :p<br>\nJoke aside, very well done. You said:</p>\n<blockquote>\n  <p>Inference took about 7-8 hours for the entire test set. </p>\n</blockquote>\n<p>Do you have a rough estimate of the training time? Many thanks and well done again!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1063415,
          "author_name": "Ian Pan",
          "author_url": "",
          "post_date": "2020-10-28T21:58:09.563000",
          "content": "<p>The feature extractors took about 5.5 hours to train. TD-CNNs took about 7.5-8 hours to train. The 3D CNN took about 30 minutes, and the transformers were very fast- just a few minutes. </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1062769,
      "author_name": "yui coiuyu",
      "author_url": "",
      "post_date": "2020-10-28T06:58:43.090000",
      "content": "<p>congrats!!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1062633,
      "author_name": "Ronaldo S.A. Batista",
      "author_url": "",
      "post_date": "2020-10-28T03:54:24.493000",
      "content": "<p>Wow.That's an amazing pipeline. Congratulations!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1062596,
      "author_name": "Amrit Virdee",
      "author_url": "",
      "post_date": "2020-10-28T02:45:00.693000",
      "content": "<p>Provisionally 2nd place now! Thanks for a very informational write and congratulations! 👍</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1062111,
      "author_name": "NANDINI SINGH 05",
      "author_url": "",
      "post_date": "2020-10-27T15:27:15.050000",
      "content": "<p>Congratulations! Great solution write up. Thank you.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1061680,
      "author_name": "Vee",
      "author_url": "",
      "post_date": "2020-10-27T07:52:23.510000",
      "content": "<p>Congrats, that is an impressive solution. And thanks for your JPEG contribution  too.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1061675,
      "author_name": "Theo Viel",
      "author_url": "",
      "post_date": "2020-10-27T07:43:35.543000",
      "content": "<p>I'm trully impressed by your consistency at getting strong results in medical imaging competitions. </p>\n<p>The fact that you're a physician probably plays a role, but that's even more impressive to me since it means that you already have a really busy schedule. </p>\n<p>Anyways, very elegant solution, I hope you don't get removed for that 3rd place is truly deserved.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1063416,
          "author_name": "Ian Pan",
          "author_url": "",
          "post_date": "2020-10-28T22:00:14.010000",
          "content": "<p>Thank you! When I started out on Kaggle, I felt like my medical imaging domain knowledge would give me a surefire edge in these competitions. Kagglers are too smart though, so at best it's only a slight advantage. </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1061580,
      "author_name": "Youngno Yoon",
      "author_url": "",
      "post_date": "2020-10-27T05:17:56.357000",
      "content": "<p>Congratulations and thank you for your solution. </p>\n<p>I have a question about feature extraction. Would you tell me the experiments with other models if you do (efficientnet, ResNext, densenet…)? ResNeSt50 is better than others? or ensemble didn't work?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1061978,
          "author_name": "Ian Pan",
          "author_url": "",
          "post_date": "2020-10-27T13:29:00.607000",
          "content": "<p>I tried initially with EfficientNet-B4 but in the past I've found that ResNet-like architectures worked better for tasks relying on feature extraction, so I tried ResNeSt which worked better (0.01 decrease in loss). </p>\n<p>I also tried EfficientNet-B3 pruned at lower resolution to increase inference speed but this was significantly worse (0.02 increase in loss). </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1061488,
      "author_name": "Jaideep",
      "author_url": "",
      "post_date": "2020-10-27T03:16:04.190000",
      "content": "<p>Great <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1061457,
      "author_name": "DatNT",
      "author_url": "",
      "post_date": "2020-10-27T02:48:58.723000",
      "content": "<p>thought about hand labeling the heart earlier, but then i gave up :(</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1061970,
          "author_name": "Ian Pan",
          "author_url": "",
          "post_date": "2020-10-27T13:24:38.727000",
          "content": "<p>Seems like it didn't make too much of a difference. 😅</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1061448,
      "author_name": "Kamal Das",
      "author_url": "",
      "post_date": "2020-10-27T02:37:24.063000",
      "content": "<p>congratulations <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> !!<br>\nand thanks for sharing the images and help throughout the competition 👍</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1061441,
      "author_name": "DungNB",
      "author_url": "",
      "post_date": "2020-10-27T02:33:24.340000",
      "content": "<p>Congratulations Ian!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1061439,
      "author_name": "aksg87",
      "author_url": "",
      "post_date": "2020-10-27T02:32:32.847000",
      "content": "<p>Congratulations Ian! Really amazing approach for formulating this problem in a single end-to-end model. You combined amazing clinical/problem insights with cutting edge architecture. I hope institutions can leverage this to build better tools for PE triage! -- Also a double congrats for doing this during your intern year :)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1892427,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-10T05:00:14.293000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1062145,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-27T15:51:28.563000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1061321": "Update: code available at https://github.com/i-pan/kaggle-rsna-pe\n\nCongratulations to all the participants and the winners. Special congrats to prize winners @osciiart and @jpbremer who are on track to be physician GMs. This was a tough, compute-heavy challenge given the large amount of data and short timeframe. My setup was 4 24 GB Quadro RTX 6000 GPUs. I always feel guilty during competitions like these since I have the luxury of strong compute. Models were trained using DDP in PyTorch 1.6 with automatic mixed precision. \n\nEven though the results aren't final yet and my submission may be removed for violating the label consistency requirements (hopefully my heuristic for fixing those worked!), I still wanted to share my solution.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F281652%2Ff6a15edbc81d867c0faecd0894d6aa27%2Fpe.png?generation=1603756913344030&alt=media)\n\nHere is a schematic outlining my solution. It has a lot of moving parts, so I apologize for the lengthy summary. \n\n# Step 1: Feature Extraction\n\nLast year's RSNA Intracranial Hemorrhage Detection challenge shared a lot of similarities with this year's challenge. Most of the top solutions combined 2D CNN feature extraction with sequence modeling. The backbone of my solution also relied on a similar setup. \n\nI first trained ResNeSt50 on 2D images. Images were windowed using the PE-specific window (WL=100, WW=700) that I mentioned in one of my initial posts on DICOM processing. Each \"image\" was 3 channels, with each channel representing an individual slice. Thus the 2D image was a stack of 3 continuous slices. The targets were the 7 PE-related labels (i.e., excluding the RV/LV ratio labels). Note: most of the PE-related labels were exam-level labels. However, I just assigned  the exam labels to each slice positive for PE (negative slices had all zeros), with the understanding that there would be label noise. Also, I predicted the labels for the middle slice among the 3 slices in the image. Models were trained using vanilla binary cross-entropy loss (`BCEWithLogitsLoss` in PyTorch), 512x512 with 448x448 random crops (single center crop during inference), RandAugment data augmentation, batch size 128, 5 epochs, 2500 steps per epoch, RAdam optimizer with cosine annealing learning rate scheduler. I used generalized mean pooling and reduced the final feature vector to 512-D. Mean loss was around 0.06 (AUC 0.95-0.96 for slice-wise prediction of PE vs. no PE). Features were then extracted for all slices. \n\n# Step 2: Sequence Modeling\n\nMany of last year's solutions used LSTMs/GRUs as the sequence model of choice. For this competition, I used `huggingface` transformers (specifically, the `Transformer` class from `transformers.modeling_distilbert`). I used one 4-layer transformer to produce slice-wise `pe_present_on_image` predictions and another transformer to predict exam-wise PE-labels. Sequence length was 512 during training, padded/truncated as necessary. During inference, I used the sequence without modifications. \n\nImportant point: **Images from non-PE exams do not contribute to the loss.** At first, I was training the slice-wise transformer on all exams. Then, I decided to train these models on positive exams only. This lowered my CV by about 0.01-0.02. I used a custom weighted loss where I weighted the loss from each example by the proportion of positive PE slices (as described in the metric), though I'm still not sure I wrote it correctly. \n\nThe exam-level transformer was trained using a weighted BCE loss based on the competition label weights. Exam-level validation losses ranged from 0.15-0.17.  \n\n# Step 3: Time-Distributed CNN\n\nTo add some variety into my modeling, I then trained a time-distributed CNN, which is just another way of saying I stacked a transformer on top of a CNN feature extraction backbone and trained end-to-end. \n\nBut before doing that, I performed inference using the slice-wise transformer model to get PE scores for every slice (5-fold OOF predictions). Then, when training the TD-CNN, I only trained on the top 30% of slices from each exam, sorted by PE score. These were trained on 3D volumes of size 32x416x416 cropped to 32x364x364 using the same windowing strategy (WL=100, WW=700) in batches of 16. \n\nI initialized the CNN backbone and the transformer head with trained models from steps 1 and 2 to help with convergence. I forced all the batch normalization layers in the backbone to `eval` mode as well- this prevents the running mean and variance in each layer from updating and only trains the coefficients. Exam-level validation losses ranged from 0.15-0.17, similar to step 2. \n\n# Step 4: Heart Slice Prediction\n\nRV/LV ratio is a significant portion of the loss. I hand-labeled slices with heart in 1,000 CT scans and trained a model (EfficientNet-B1 pruned, AUC 0.998, 256x256->224x224 crops) to classify heart slices in each CT scan. I did this because I felt that by focusing a model on the heart, I could get better, more consistent results across scans. It actually wasn't hard to label 1,000 scans- probably a full day's worth of work. I just needed to find the top and bottom heart slices; everything in between thus must also contain the heart.\n\n# Step 5: RV/LV 3D CNN\n\nI trained a 3D CNN to classify RV/LV ratio. Specifically, I used a 101-layer channel separated network, pretrained on 65 million Instagram videos (https://arxiv.org/abs/1904.02811, https://github.com/facebookresearch/VMZ). \n\nThis model was trained only on heart slices from each exam. Also, it was only trained on positive exams. This is because RV/LV ratio was not labeled for negative exams- both RV/LV labels are 0. Thus, it didn't make sense to me to try and train my model on the entire dataset's labels directly. Models were trained using a weighted BCE loss using the competition weights. I resized the input to 64x256x256->64x224x224 crops and used mediastinal window (WL=50, WW=350). AUC for RV/LV ratio > 1 was about 0.85. The validation loss ranged from 0.44-0.48. I then used these models to extract 2048-D features from each exam.\n\nThe challenge here was that the likelihood of an exam being PE-positive greatly influenced the RV/LV labels. From step 2, I calculated 5-fold OOF exam-level predictions for all exams. Then, I trained a linear model that took as input the concatenation of the 2048-D 3D CNN feature and the 7 PE exam labels. This linear model was trained across **all exams** using the labels directly. This way, the model could take into account the imaging features from the scan but also adjust the predictions based on the likelihood of PE. Validation losses ranged from 0.22-0.25. \n\nI didn't validate my entire pipeline that often due to time constraints, instead choosing to focus on the individual component losses and optimizing each one as well as I could. I did periodically do a sanity-check on 200 single-fold exams to make sure that my entire pipeline was working, and my final validation loss was 0.183 for my 0.150 public LB submission. I'm still not confident that I implemented the metric correctly, but I was seeing good correlation (0.159/0.195->0.156/0.188-> 0.150/0.183).\n\nAt the end, I applied a function to enforce label consistency requirements for each exam. There was about 0.001 change in CV after applying this function. I have some thoughts about this requirement which maybe I'll save for another post. \n\n# Final Models\n\nOverall, I used these models:\n-2x ResNeSt50 feature extractors\n-6x exam transformers (3 for each extractor)\n-6x slice transformers\n-5x ResNeSt50 TD-CNN\n-1x EfficientNet-B1 pruned heart slice classifier\n-5x ip-CSN-101 3D CNN RV/LV feature extractor\n-5x RV/LV linear model \n\nInference took about 7-8 hours for the entire test set. \n\nCode to follow once I clean it up. \n\nThings that didn't work as well:\n-3D CNN for PE exam-level prediction\n-Pseudolabeling RV/LV ratio for negative exams\n-Using logits/probabilities instead of features for second-stage model\n-TD-CNN for RV/LV ratio prediction\n-Training feature extractor on positive exams only to improve slice-wise PE prediction ",
    "1061377": "Congratulations! Great solution write up. Thank you. So many ingenuities squeezed into one diagram.\n\nI have just one question. How did you find time to do all that as an intern in a hospital as busy as yours?",
    "1117660": "Ian, this is amazing and inspiring. I had a look at the first place solution and that sent me down several rabbit holes, and I'm only getting to look at this for the first time now. I suspect this will send me on a 3 month detour!\n\nI'm hoping I could get your insights on something. There are a lot of specific choices in this setup, and I noticed your comment below about having planned this at the start, and not deviating much from it. In light of that, it would be super helpful to get your thoughts on which specific choices you are most convinced made a strong difference, and which you are still unsure of.",
    "1061385": "Excellent work- thanks for sharing your solution. I'm going to come back and read this a few more times!",
    "1062116": "It's a great honor for me to have such words from you. Congratulations!",
    "1064518": "Wow, what a useful notebook. I'm using it for further education.",
    "1061566": "congratulations @vaillant ",
    "1065349": "it is very meaningful result",
    "1065040": "Wow, that was a nice documentation. Really its useful. And my deep Congratulations to winner..!",
    "1064731": "Yeah this one is really very impressive ",
    "1064460": "Good job! Interesting pipeline",
    "1064208": "Congratulations, and Thanks for sharing :)",
    "1064098": "Congratulations ",
    "1063855": "congrats!!!",
    "1063565": "Well done!",
    "1063366": "Congrats , great job !",
    "1063307": "I am jealous of your setup. :p\nJoke aside, very well done. You said:\n\n> Inference took about 7-8 hours for the entire test set. \n\nDo you have a rough estimate of the training time? Many thanks and well done again!",
    "1062769": "congrats!!",
    "1062633": "Wow.That's an amazing pipeline. Congratulations!\n",
    "1062596": "Provisionally 2nd place now! Thanks for a very informational write and congratulations! 👍",
    "1062111": "Congratulations! Great solution write up. Thank you.",
    "1061680": "Congrats, that is an impressive solution. And thanks for your JPEG contribution  too.",
    "1061675": "I'm trully impressed by your consistency at getting strong results in medical imaging competitions. \n\nThe fact that you're a physician probably plays a role, but that's even more impressive to me since it means that you already have a really busy schedule. \n\nAnyways, very elegant solution, I hope you don't get removed for that 3rd place is truly deserved.",
    "1061580": "Congratulations and thank you for your solution. \n\nI have a question about feature extraction. Would you tell me the experiments with other models if you do (efficientnet, ResNext, densenet...)? ResNeSt50 is better than others? or ensemble didn't work?",
    "1061488": "Great @vaillant ",
    "1061457": "thought about hand labeling the heart earlier, but then i gave up :(",
    "1061448": "congratulations @vaillant !!\nand thanks for sharing the images and help throughout the competition 👍",
    "1061441": "Congratulations Ian!",
    "1061439": "Congratulations Ian! Really amazing approach for formulating this problem in a single end-to-end model. You combined amazing clinical/problem insights with cutting edge architecture. I hope institutions can leverage this to build better tools for PE triage! -- Also a double congrats for doing this during your intern year :)",
    "1892427": "",
    "1062145": ""
  }
}