{
  "id": 193402,
  "title": "23rd Place - Fast GPU Experimentation Pipeline!",
  "url": "/competitions/rsna-str-pulmonary-embolism-detection/discussion/193402",
  "author_name": "Chris Deotte",
  "post_date": "2020-10-27T00:11:44.622000",
  "votes": 90,
  "comment_count": 56,
  "views": 0,
  "content": "<p>Thank you Radiological Society of North America (RSNA®), Society of Thoracic Radiology (STR), and Kaggle for hosting this fun competition. Thank you Nvidia for providing compute resources.</p>\n<p>This has been one of my favorite competitions. I enjoyed building an elaborate multi-stage pipeline of stacked models! Working with 3D images was fun and provided an additional challenge compared with 2D images. I particularly enjoyed tackling the challenge of building a fast experimentation pipeline when the training data is <code>1_000_000_000_000 bytes</code> of data! One trillion bytes! This is the largest dataset I have ever worked with.</p>\n<h1>RSNA STR Pulmonary Embolism Detection</h1>\n<p>In the figure below, each row is an exam (i.e. study, i.e. single patient). The row of images are CT scan \"slices\" from the 3D image of a patient's chest. In this competition, we need to predict 9 targets for each patient (each row) (like is pe on left side? on right side? etc) and we need to classify every image (is pe present?). If below were all the data, then we would need to predict <code>3 rows * 9 targets = 27 exam targets</code> and <code>15 images * 1 target = 15 image targets</code>. In total we would need to predict 42 targets.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F7d5ce77d471fbf8ef7e746731eab734c%2Fstudy.png?generation=1603731145650332&amp;alt=media\" alt=\"\"></p>\n<h1>Understanding the Metric</h1>\n<p>On the description page, the metric seems very confusing. However it is just a weighted average of 10 log losses (the 9 types of exam predictions and the 1 type of image prediction). (I explain the metric <a href=\"https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/193598\" target=\"_blank\">here</a>). After computing the 10 weights, we find that 50% of our LB score is from the 9 exam predictions and 50% of our LB score is from the image predictions. Furthermore, the log loss for the image predictions is itself a weighted log loss where an image that is part of an exam without pulmonary embolism has weight zero (very important observation!) </p>\n<ul>\n<li>Improving image level predictions is equally important as improving exam predictions.</li>\n<li>There is no penalty for false positives, so we can train our image prediction models with only the 30% of the data from positive exams!</li>\n</ul>\n<h1>Stage 1 - Model One - Image Level Predictions</h1>\n<h1>(CNN EfficientNet B4)</h1>\n<pre><code>inp = tf.keras.Input(shape=(320, 320, 1)) # INPUT IS UINT8\nx = tf.keras.layers.Concatenate()([inp/255., inp/255., inp/255.])\nbase_model = efn.EfficientNetB4(weights='imagenet', include_top=False) \n\nx = base_model(x)\nx = tf.keras.layers.GlobalAveragePooling2D()(x)    \nx = tf.keras.layers.Dense(1, activation='sigmoid')(x)\n\nmodel = tf.keras.Model(inputs=inp, outputs=x)\nopt = tf.keras.optimizers.Adam(lr=0.000005)\nmodel.compile(loss='binary_crossentropy', optimizer = opt)\n\nmodel.fit(X, y, sample_weight = X.groupby('StudyInstanceUID') \n    .pe_present_on_image.transform('mean') * 5.6222 )\n</code></pre>\n<p>I built two models. Model one predicts image level predictions (i.e. <code>pe_present_on_image</code>) and model two predicts patient level predictions (i.e. exams i.e. studies, like <code>leftsided_pe</code> etc)</p>\n<ul>\n<li>EfficientNet B4 pretrained on <code>imagenet</code></li>\n<li><strong>Only Mediastinal window, (ie. level=40, width=400)</strong> i.e. 1 channel <code>uint8</code></li>\n<li>Random crops of 320x320 from 512x512</li>\n<li>Rotation (+-8 deg) Scale (+-0.16) augmentation</li>\n<li>Coarse Dropout (16 holes sized 50x50)</li>\n<li>Mixup (swap slices of similar Z position with other exams)</li>\n<li>Adam optimizer with constant <code>LR = 5e-6</code></li>\n<li>Training sample weight equal to pe proportion in exam</li>\n<li><strong>Only train on 30% of train data with pe present in exam</strong></li>\n<li>40 minute epochs using 4x V100 GPU</li>\n<li>Train 15 epochs with batch size 128</li>\n</ul>\n<p>Below illustrates my augmentations. For display purposes, we illustrate Mixup with a large yellow, green, or blue square so you can see it better. (During runtime, it was an actual second image).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F5515206e0fd5e5a7379f6272e0e56dc3%2Fmixup.png?generation=1603733206613194&amp;alt=media\" alt=\"\">  </p>\n<h1>Stage 1 - Model Two - Patient Level Predictions</h1>\n<h1>(CNN EfficientNet B4)</h1>\n<p>Each patient has an average of 200 images. Among those 200, if PE is present, it is usually on the middle slices. Therefore I only train my patient level model with slices <code>0.35 &lt; z &lt; 0.65</code>. Then to predict the 9 targets for each patient, I only infer <code>0.35 &lt; z &lt; 0.65</code> and then take the 9 average predictions.</p>\n<ul>\n<li>Most details same as model one</li>\n<li><strong>Only Mediastinal window, (ie. level=40, width=400)</strong> i.e. 1 channel <code>uint8</code></li>\n<li>Output layer of 9 sigmoid units</li>\n<li><strong>Only train on 30% of train data with Z Position between <code>0.35 &lt; z &lt; 0.65</code></strong></li>\n<li>Loss <code>weighted_log_loss</code></li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F5ed4f1ae6291e5edc4f422b479214f80%2Fslices.png?generation=1603734372678374&amp;alt=media\" alt=\"\"></p>\n<h1>Experimentation Pipeline</h1>\n<p>How do we discover the details above? All the settings above were discovered by performing dozens of experiments on <strong>smaller images and smaller backbones</strong>. For example, use 128x128 (with 80x80 crops) EfficientNetB0 and/or 256x256 (with 160x160 crops) EfficientNetB2. Using these smaller models, we can test out ideas on a single GPU in minutes! Also note that we only use 1 channel images of <code>uint8</code>. This is 33% less data than converting images to 3 channels of 3 different CT window schemes.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F60349090637a446427fc8f7c8775a70f%2Fexp2.png?generation=1603744581352698&amp;alt=media\" alt=\"\"></p>\n<p>Remember we are only training with 30% original data. Then using crops makes it 12% of data. Then using 256x256 reduces this to 3% of data. And using 128x128 reduces this to 0.75% of data! Even Kaggle notebooks P100 GPU can train quickly on 80x80 crops from 128x128 and 30% train data.</p>\n<p>Once you find a configuration that works well, then run 512x512 with EfficientNetB4 overnight. Using only 2D predictions for image and 2D predictions for patient, <strong>the above two models obtain LB 0.215 and CV 0.235.</strong> </p>\n<p>We will now increase our CV LB by building stage 2 models that use stage 1 predictions as input</p>\n<h1>Stage 2 - Model One - Image Level Predictions</h1>\n<h1>(Random Forest)</h1>\n<pre><code>FEATURES = ['oof']\nfor k in NEIGHBORS:\n    tmp = train.sort_values('PosZ').groupby('StudyInstanceUID')[['oof']]\n    train['b%i'%k] = tmp.shift(k)\n    train['a%i'%k] = tmp.shift(-k)\n    FEATURES += ['a%i'%k, 'b%i'%k]\ntrain.fillna(-1,inplace=True)\n\nmodel = RandomForestClassifier(max_depth=9, n_estimators=100, \n                           n_jobs=20, min_samples_leaf=50)\nmodel.fit(train.loc[idxT,FEATURES],train.loc[idxT,'pe_present_on_image'],\n                sample_weight = 5.6222 * valid.loc[idxT,'weight'])\n</code></pre>\n<p>All images are slices from 3D images. So adjacent images (within the same exam) contain helpful information. Each plot below displays all 200 or so image level predictions from 1 study. The x axis is z position and the y axis is the prediction value (0 to 1). The blue line is the ground truth, the orange line is the prediction from the model described above. The black line is the random forest Stage 2 model.</p>\n<p>For each image level prediction, a random forest model takes as input the prediction and neighbor predictions [1,2,3,4,5,6,7,8,9,10,15,20,25,30,35,40,45,50,60,70,80,90,100,150,200,250] on either side. Then the random forest model predicts a new image level prediction display in black below. Notice when the original prediction is close to 1, then the random forest pushes it up to 1 and when the original prediction is close to 0, then the random forest pushes it down to 0.</p>\n<p><strong>This stage 2 image level model increased LB to 0.204 from 0.215 and CV to 0.224 from 0.235 (gain = 0.011)</strong></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Fcccdbe82e765c3e7d762031b358dd290%2Fi-smooth.png?generation=1603736540573190&amp;alt=media\" alt=\"\"></p>\n<h1>Stage 2 - Model Two - Patient Level Predictions</h1>\n<h1>(GRU + 1D-CNN)</h1>\n<pre><code>inp = L.Input(shape=(64, 1792))\nx = L.Bidirectional(L.GRU(48, return_sequences=True, \n                    kernel_initializer='orthogonal'))(inp)\nx = L.Bidirectional(L.GRU(48, return_sequences=False, \n                    kernel_initializer='orthogonal'))(x)\nx = L.Dense(9, activation='sigmoid')(x)\n\nmodel = tf.keras.Model(inputs=inp, outputs=x)\nopt = tf.keras.optimizers.Adam(lr=0.00005)\nmodel.compile(loss=weighted_log_loss, optimizer = opt)\n</code></pre>\n<p>Similarly we can use adjacent slice information to improve our patient level predictions. Most patients have between 160 and 310 images per study.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Fa4a52476743ad1e1ada257c586c607fb%2Fhist2.png?generation=1603741964817694&amp;alt=media\" alt=\"\"></p>\n<p>Below are plots of the top view (Z position is vertical axis) of the 3D image (not the ordinary slice view of the 3D image). We notice that most of the crucial information is between 25% and 75% in top view. Therefore we extracted 64 images equally spaces between 25% and 75% Z position. Then we took those 64 images and extracted the GAP embeddings from both our stage 1 model one and stage 1 model two. We trained a stage 2 GRU model and a stage 2 1D-CNN model to predict exam level predictions from these 64 GAP embeddings.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Fe46e41fab360d847c3fadb5d6dff9002%2Ftop.png?generation=1603741873425399&amp;alt=media\" alt=\"\"></p>\n<p><strong>This stage 2 exam level model increased LB to 0.183 from 0.204 and CV to 0.203 from 0.224 (gain = 0.021)</strong></p>\n<h1>Other Misc Ideas</h1>\n<p>When using global average pooling 2D in your CNN, location information is lost. Therefore I tried giving my models location information in various ways to help predict targets related to location (i.e. <code>leftsided_pe</code> etc) and related to size (i.e. <code>rv_lv_ratio_gte_1</code> etc). Unfortunately, none of my ideas increased CV or LB. My favorite is below.</p>\n<h2>Locating PE with Class Activation Maps CAM</h2>\n<p>I extracted class activation maps from my stage 1 models and feed the location information into stage 2 exam models. (CAMs explained <a href=\"https://www.kaggle.com/cdeotte/unsupervised-masks-cv-0-60\" target=\"_blank\">here</a>). In the below figure, the ground truth is in the title. The green circles are the CAM of my EfficientNetB4 model. Note that CT scans are flipped so the left side of the image is \"right\" and the right side of the image is \"left. You can see that the CAM does a good job of locating the PE.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F291dc983514a322d5ce4e56b22a9a681%2Fcam.png?generation=1603755601607192&amp;alt=media\" alt=\"\"></p>\n<h1>Thank you</h1>",
  "messages": [
    {
      "id": 1061327,
      "postDate": "2020-10-27T00:11:44.623Z",
      "content": "<p>Thank you Radiological Society of North America (RSNA®), Society of Thoracic Radiology (STR), and Kaggle for hosting this fun competition. Thank you Nvidia for providing compute resources.</p>\n<p>This has been one of my favorite competitions. I enjoyed building an elaborate multi-stage pipeline of stacked models! Working with 3D images was fun and provided an additional challenge compared with 2D images. I particularly enjoyed tackling the challenge of building a fast experimentation pipeline when the training data is <code>1_000_000_000_000 bytes</code> of data! One trillion bytes! This is the largest dataset I have ever worked with.</p>\n<h1>RSNA STR Pulmonary Embolism Detection</h1>\n<p>In the figure below, each row is an exam (i.e. study, i.e. single patient). The row of images are CT scan \"slices\" from the 3D image of a patient's chest. In this competition, we need to predict 9 targets for each patient (each row) (like is pe on left side? on right side? etc) and we need to classify every image (is pe present?). If below were all the data, then we would need to predict <code>3 rows * 9 targets = 27 exam targets</code> and <code>15 images * 1 target = 15 image targets</code>. In total we would need to predict 42 targets.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F7d5ce77d471fbf8ef7e746731eab734c%2Fstudy.png?generation=1603731145650332&amp;alt=media\" alt=\"\"></p>\n<h1>Understanding the Metric</h1>\n<p>On the description page, the metric seems very confusing. However it is just a weighted average of 10 log losses (the 9 types of exam predictions and the 1 type of image prediction). (I explain the metric <a href=\"https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/193598\" target=\"_blank\">here</a>). After computing the 10 weights, we find that 50% of our LB score is from the 9 exam predictions and 50% of our LB score is from the image predictions. Furthermore, the log loss for the image predictions is itself a weighted log loss where an image that is part of an exam without pulmonary embolism has weight zero (very important observation!) </p>\n<ul>\n<li>Improving image level predictions is equally important as improving exam predictions.</li>\n<li>There is no penalty for false positives, so we can train our image prediction models with only the 30% of the data from positive exams!</li>\n</ul>\n<h1>Stage 1 - Model One - Image Level Predictions</h1>\n<h1>(CNN EfficientNet B4)</h1>\n<pre><code>inp = tf.keras.Input(shape=(320, 320, 1)) # INPUT IS UINT8\nx = tf.keras.layers.Concatenate()([inp/255., inp/255., inp/255.])\nbase_model = efn.EfficientNetB4(weights='imagenet', include_top=False) \n\nx = base_model(x)\nx = tf.keras.layers.GlobalAveragePooling2D()(x)    \nx = tf.keras.layers.Dense(1, activation='sigmoid')(x)\n\nmodel = tf.keras.Model(inputs=inp, outputs=x)\nopt = tf.keras.optimizers.Adam(lr=0.000005)\nmodel.compile(loss='binary_crossentropy', optimizer = opt)\n\nmodel.fit(X, y, sample_weight = X.groupby('StudyInstanceUID') \n    .pe_present_on_image.transform('mean') * 5.6222 )\n</code></pre>\n<p>I built two models. Model one predicts image level predictions (i.e. <code>pe_present_on_image</code>) and model two predicts patient level predictions (i.e. exams i.e. studies, like <code>leftsided_pe</code> etc)</p>\n<ul>\n<li>EfficientNet B4 pretrained on <code>imagenet</code></li>\n<li><strong>Only Mediastinal window, (ie. level=40, width=400)</strong> i.e. 1 channel <code>uint8</code></li>\n<li>Random crops of 320x320 from 512x512</li>\n<li>Rotation (+-8 deg) Scale (+-0.16) augmentation</li>\n<li>Coarse Dropout (16 holes sized 50x50)</li>\n<li>Mixup (swap slices of similar Z position with other exams)</li>\n<li>Adam optimizer with constant <code>LR = 5e-6</code></li>\n<li>Training sample weight equal to pe proportion in exam</li>\n<li><strong>Only train on 30% of train data with pe present in exam</strong></li>\n<li>40 minute epochs using 4x V100 GPU</li>\n<li>Train 15 epochs with batch size 128</li>\n</ul>\n<p>Below illustrates my augmentations. For display purposes, we illustrate Mixup with a large yellow, green, or blue square so you can see it better. (During runtime, it was an actual second image).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F5515206e0fd5e5a7379f6272e0e56dc3%2Fmixup.png?generation=1603733206613194&amp;alt=media\" alt=\"\">  </p>\n<h1>Stage 1 - Model Two - Patient Level Predictions</h1>\n<h1>(CNN EfficientNet B4)</h1>\n<p>Each patient has an average of 200 images. Among those 200, if PE is present, it is usually on the middle slices. Therefore I only train my patient level model with slices <code>0.35 &lt; z &lt; 0.65</code>. Then to predict the 9 targets for each patient, I only infer <code>0.35 &lt; z &lt; 0.65</code> and then take the 9 average predictions.</p>\n<ul>\n<li>Most details same as model one</li>\n<li><strong>Only Mediastinal window, (ie. level=40, width=400)</strong> i.e. 1 channel <code>uint8</code></li>\n<li>Output layer of 9 sigmoid units</li>\n<li><strong>Only train on 30% of train data with Z Position between <code>0.35 &lt; z &lt; 0.65</code></strong></li>\n<li>Loss <code>weighted_log_loss</code></li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F5ed4f1ae6291e5edc4f422b479214f80%2Fslices.png?generation=1603734372678374&amp;alt=media\" alt=\"\"></p>\n<h1>Experimentation Pipeline</h1>\n<p>How do we discover the details above? All the settings above were discovered by performing dozens of experiments on <strong>smaller images and smaller backbones</strong>. For example, use 128x128 (with 80x80 crops) EfficientNetB0 and/or 256x256 (with 160x160 crops) EfficientNetB2. Using these smaller models, we can test out ideas on a single GPU in minutes! Also note that we only use 1 channel images of <code>uint8</code>. This is 33% less data than converting images to 3 channels of 3 different CT window schemes.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F60349090637a446427fc8f7c8775a70f%2Fexp2.png?generation=1603744581352698&amp;alt=media\" alt=\"\"></p>\n<p>Remember we are only training with 30% original data. Then using crops makes it 12% of data. Then using 256x256 reduces this to 3% of data. And using 128x128 reduces this to 0.75% of data! Even Kaggle notebooks P100 GPU can train quickly on 80x80 crops from 128x128 and 30% train data.</p>\n<p>Once you find a configuration that works well, then run 512x512 with EfficientNetB4 overnight. Using only 2D predictions for image and 2D predictions for patient, <strong>the above two models obtain LB 0.215 and CV 0.235.</strong> </p>\n<p>We will now increase our CV LB by building stage 2 models that use stage 1 predictions as input</p>\n<h1>Stage 2 - Model One - Image Level Predictions</h1>\n<h1>(Random Forest)</h1>\n<pre><code>FEATURES = ['oof']\nfor k in NEIGHBORS:\n    tmp = train.sort_values('PosZ').groupby('StudyInstanceUID')[['oof']]\n    train['b%i'%k] = tmp.shift(k)\n    train['a%i'%k] = tmp.shift(-k)\n    FEATURES += ['a%i'%k, 'b%i'%k]\ntrain.fillna(-1,inplace=True)\n\nmodel = RandomForestClassifier(max_depth=9, n_estimators=100, \n                           n_jobs=20, min_samples_leaf=50)\nmodel.fit(train.loc[idxT,FEATURES],train.loc[idxT,'pe_present_on_image'],\n                sample_weight = 5.6222 * valid.loc[idxT,'weight'])\n</code></pre>\n<p>All images are slices from 3D images. So adjacent images (within the same exam) contain helpful information. Each plot below displays all 200 or so image level predictions from 1 study. The x axis is z position and the y axis is the prediction value (0 to 1). The blue line is the ground truth, the orange line is the prediction from the model described above. The black line is the random forest Stage 2 model.</p>\n<p>For each image level prediction, a random forest model takes as input the prediction and neighbor predictions [1,2,3,4,5,6,7,8,9,10,15,20,25,30,35,40,45,50,60,70,80,90,100,150,200,250] on either side. Then the random forest model predicts a new image level prediction display in black below. Notice when the original prediction is close to 1, then the random forest pushes it up to 1 and when the original prediction is close to 0, then the random forest pushes it down to 0.</p>\n<p><strong>This stage 2 image level model increased LB to 0.204 from 0.215 and CV to 0.224 from 0.235 (gain = 0.011)</strong></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Fcccdbe82e765c3e7d762031b358dd290%2Fi-smooth.png?generation=1603736540573190&amp;alt=media\" alt=\"\"></p>\n<h1>Stage 2 - Model Two - Patient Level Predictions</h1>\n<h1>(GRU + 1D-CNN)</h1>\n<pre><code>inp = L.Input(shape=(64, 1792))\nx = L.Bidirectional(L.GRU(48, return_sequences=True, \n                    kernel_initializer='orthogonal'))(inp)\nx = L.Bidirectional(L.GRU(48, return_sequences=False, \n                    kernel_initializer='orthogonal'))(x)\nx = L.Dense(9, activation='sigmoid')(x)\n\nmodel = tf.keras.Model(inputs=inp, outputs=x)\nopt = tf.keras.optimizers.Adam(lr=0.00005)\nmodel.compile(loss=weighted_log_loss, optimizer = opt)\n</code></pre>\n<p>Similarly we can use adjacent slice information to improve our patient level predictions. Most patients have between 160 and 310 images per study.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Fa4a52476743ad1e1ada257c586c607fb%2Fhist2.png?generation=1603741964817694&amp;alt=media\" alt=\"\"></p>\n<p>Below are plots of the top view (Z position is vertical axis) of the 3D image (not the ordinary slice view of the 3D image). We notice that most of the crucial information is between 25% and 75% in top view. Therefore we extracted 64 images equally spaces between 25% and 75% Z position. Then we took those 64 images and extracted the GAP embeddings from both our stage 1 model one and stage 1 model two. We trained a stage 2 GRU model and a stage 2 1D-CNN model to predict exam level predictions from these 64 GAP embeddings.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Fe46e41fab360d847c3fadb5d6dff9002%2Ftop.png?generation=1603741873425399&amp;alt=media\" alt=\"\"></p>\n<p><strong>This stage 2 exam level model increased LB to 0.183 from 0.204 and CV to 0.203 from 0.224 (gain = 0.021)</strong></p>\n<h1>Other Misc Ideas</h1>\n<p>When using global average pooling 2D in your CNN, location information is lost. Therefore I tried giving my models location information in various ways to help predict targets related to location (i.e. <code>leftsided_pe</code> etc) and related to size (i.e. <code>rv_lv_ratio_gte_1</code> etc). Unfortunately, none of my ideas increased CV or LB. My favorite is below.</p>\n<h2>Locating PE with Class Activation Maps CAM</h2>\n<p>I extracted class activation maps from my stage 1 models and feed the location information into stage 2 exam models. (CAMs explained <a href=\"https://www.kaggle.com/cdeotte/unsupervised-masks-cv-0-60\" target=\"_blank\">here</a>). In the below figure, the ground truth is in the title. The green circles are the CAM of my EfficientNetB4 model. Note that CT scans are flipped so the left side of the image is \"right\" and the right side of the image is \"left. You can see that the CAM does a good job of locating the PE.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F291dc983514a322d5ce4e56b22a9a681%2Fcam.png?generation=1603755601607192&amp;alt=media\" alt=\"\"></p>\n<h1>Thank you</h1>",
      "rawMarkdown": "Thank you Radiological Society of North America (RSNA®), Society of Thoracic Radiology (STR), and Kaggle for hosting this fun competition. Thank you Nvidia for providing compute resources.\n\nThis has been one of my favorite competitions. I enjoyed building an elaborate multi-stage pipeline of stacked models! Working with 3D images was fun and provided an additional challenge compared with 2D images. I particularly enjoyed tackling the challenge of building a fast experimentation pipeline when the training data is `1_000_000_000_000 bytes` of data! One trillion bytes! This is the largest dataset I have ever worked with.\n\n# RSNA STR Pulmonary Embolism Detection\nIn the figure below, each row is an exam (i.e. study, i.e. single patient). The row of images are CT scan \"slices\" from the 3D image of a patient's chest. In this competition, we need to predict 9 targets for each patient (each row) (like is pe on left side? on right side? etc) and we need to classify every image (is pe present?). If below were all the data, then we would need to predict `3 rows * 9 targets = 27 exam targets` and `15 images * 1 target = 15 image targets`. In total we would need to predict 42 targets.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F7d5ce77d471fbf8ef7e746731eab734c%2Fstudy.png?generation=1603731145650332&alt=media)\n\n# Understanding the Metric\nOn the description page, the metric seems very confusing. However it is just a weighted average of 10 log losses (the 9 types of exam predictions and the 1 type of image prediction). (I explain the metric [here][2]). After computing the 10 weights, we find that 50% of our LB score is from the 9 exam predictions and 50% of our LB score is from the image predictions. Furthermore, the log loss for the image predictions is itself a weighted log loss where an image that is part of an exam without pulmonary embolism has weight zero (very important observation!) \n\n* Improving image level predictions is equally important as improving exam predictions.\n* There is no penalty for false positives, so we can train our image prediction models with only the 30% of the data from positive exams!\n\n# Stage 1 - Model One - Image Level Predictions\n# (CNN EfficientNet B4)\n\n    inp = tf.keras.Input(shape=(320, 320, 1)) # INPUT IS UINT8\n    x = tf.keras.layers.Concatenate()([inp/255., inp/255., inp/255.])\n    base_model = efn.EfficientNetB4(weights='imagenet', include_top=False) \n\n    x = base_model(x)\n    x = tf.keras.layers.GlobalAveragePooling2D()(x)    \n    x = tf.keras.layers.Dense(1, activation='sigmoid')(x)\n    \n    model = tf.keras.Model(inputs=inp, outputs=x)\n    opt = tf.keras.optimizers.Adam(lr=0.000005)\n    model.compile(loss='binary_crossentropy', optimizer = opt)\n\n    model.fit(X, y, sample_weight = X.groupby('StudyInstanceUID') \n        .pe_present_on_image.transform('mean') * 5.6222 )\n\nI built two models. Model one predicts image level predictions (i.e. `pe_present_on_image`) and model two predicts patient level predictions (i.e. exams i.e. studies, like `leftsided_pe` etc)\n* EfficientNet B4 pretrained on `imagenet`\n* **Only Mediastinal window, (ie. level=40, width=400)** i.e. 1 channel `uint8`\n* Random crops of 320x320 from 512x512\n* Rotation (+-8 deg) Scale (+-0.16) augmentation\n* Coarse Dropout (16 holes sized 50x50)\n* Mixup (swap slices of similar Z position with other exams)\n* Adam optimizer with constant `LR = 5e-6`\n* Training sample weight equal to pe proportion in exam\n* **Only train on 30% of train data with pe present in exam**\n* 40 minute epochs using 4x V100 GPU\n* Train 15 epochs with batch size 128\n\nBelow illustrates my augmentations. For display purposes, we illustrate Mixup with a large yellow, green, or blue square so you can see it better. (During runtime, it was an actual second image).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F5515206e0fd5e5a7379f6272e0e56dc3%2Fmixup.png?generation=1603733206613194&alt=media)  \n\n# Stage 1 - Model Two - Patient Level Predictions\n# (CNN EfficientNet B4)\nEach patient has an average of 200 images. Among those 200, if PE is present, it is usually on the middle slices. Therefore I only train my patient level model with slices `0.35 < z < 0.65`. Then to predict the 9 targets for each patient, I only infer `0.35 < z < 0.65` and then take the 9 average predictions.\n\n* Most details same as model one\n* **Only Mediastinal window, (ie. level=40, width=400)** i.e. 1 channel `uint8`\n* Output layer of 9 sigmoid units\n* **Only train on 30% of train data with Z Position between `0.35 < z < 0.65`**\n* Loss `weighted_log_loss`\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F5ed4f1ae6291e5edc4f422b479214f80%2Fslices.png?generation=1603734372678374&alt=media)\n\n# Experimentation Pipeline\nHow do we discover the details above? All the settings above were discovered by performing dozens of experiments on **smaller images and smaller backbones**. For example, use 128x128 (with 80x80 crops) EfficientNetB0 and/or 256x256 (with 160x160 crops) EfficientNetB2. Using these smaller models, we can test out ideas on a single GPU in minutes! Also note that we only use 1 channel images of `uint8`. This is 33% less data than converting images to 3 channels of 3 different CT window schemes.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F60349090637a446427fc8f7c8775a70f%2Fexp2.png?generation=1603744581352698&alt=media)\n\nRemember we are only training with 30% original data. Then using crops makes it 12% of data. Then using 256x256 reduces this to 3% of data. And using 128x128 reduces this to 0.75% of data! Even Kaggle notebooks P100 GPU can train quickly on 80x80 crops from 128x128 and 30% train data.\n\nOnce you find a configuration that works well, then run 512x512 with EfficientNetB4 overnight. Using only 2D predictions for image and 2D predictions for patient, **the above two models obtain LB 0.215 and CV 0.235.** \n\nWe will now increase our CV LB by building stage 2 models that use stage 1 predictions as input\n  \n# Stage 2 - Model One - Image Level Predictions\n# (Random Forest)\n\n    FEATURES = ['oof']\n    for k in NEIGHBORS:\n        tmp = train.sort_values('PosZ').groupby('StudyInstanceUID')[['oof']]\n        train['b%i'%k] = tmp.shift(k)\n        train['a%i'%k] = tmp.shift(-k)\n        FEATURES += ['a%i'%k, 'b%i'%k]\n    train.fillna(-1,inplace=True)\n\n    model = RandomForestClassifier(max_depth=9, n_estimators=100, \n                               n_jobs=20, min_samples_leaf=50)\n    model.fit(train.loc[idxT,FEATURES],train.loc[idxT,'pe_present_on_image'],\n                    sample_weight = 5.6222 * valid.loc[idxT,'weight'])\n\nAll images are slices from 3D images. So adjacent images (within the same exam) contain helpful information. Each plot below displays all 200 or so image level predictions from 1 study. The x axis is z position and the y axis is the prediction value (0 to 1). The blue line is the ground truth, the orange line is the prediction from the model described above. The black line is the random forest Stage 2 model.\n\nFor each image level prediction, a random forest model takes as input the prediction and neighbor predictions [1,2,3,4,5,6,7,8,9,10,15,20,25,30,35,40,45,50,60,70,80,90,100,150,200,250] on either side. Then the random forest model predicts a new image level prediction display in black below. Notice when the original prediction is close to 1, then the random forest pushes it up to 1 and when the original prediction is close to 0, then the random forest pushes it down to 0.\n\n**This stage 2 image level model increased LB to 0.204 from 0.215 and CV to 0.224 from 0.235 (gain = 0.011)**\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Fcccdbe82e765c3e7d762031b358dd290%2Fi-smooth.png?generation=1603736540573190&alt=media)\n\n# Stage 2 - Model Two - Patient Level Predictions\n# (GRU + 1D-CNN)\n\n    inp = L.Input(shape=(64, 1792))\n    x = L.Bidirectional(L.GRU(48, return_sequences=True, \n                        kernel_initializer='orthogonal'))(inp)\n    x = L.Bidirectional(L.GRU(48, return_sequences=False, \n                        kernel_initializer='orthogonal'))(x)\n    x = L.Dense(9, activation='sigmoid')(x)\n        \n    model = tf.keras.Model(inputs=inp, outputs=x)\n    opt = tf.keras.optimizers.Adam(lr=0.00005)\n    model.compile(loss=weighted_log_loss, optimizer = opt)\n\nSimilarly we can use adjacent slice information to improve our patient level predictions. Most patients have between 160 and 310 images per study.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Fa4a52476743ad1e1ada257c586c607fb%2Fhist2.png?generation=1603741964817694&alt=media)\n\nBelow are plots of the top view (Z position is vertical axis) of the 3D image (not the ordinary slice view of the 3D image). We notice that most of the crucial information is between 25% and 75% in top view. Therefore we extracted 64 images equally spaces between 25% and 75% Z position. Then we took those 64 images and extracted the GAP embeddings from both our stage 1 model one and stage 1 model two. We trained a stage 2 GRU model and a stage 2 1D-CNN model to predict exam level predictions from these 64 GAP embeddings.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Fe46e41fab360d847c3fadb5d6dff9002%2Ftop.png?generation=1603741873425399&alt=media)\n\n**This stage 2 exam level model increased LB to 0.183 from 0.204 and CV to 0.203 from 0.224 (gain = 0.021)**\n\n# Other Misc Ideas\nWhen using global average pooling 2D in your CNN, location information is lost. Therefore I tried giving my models location information in various ways to help predict targets related to location (i.e. `leftsided_pe` etc) and related to size (i.e. `rv_lv_ratio_gte_1` etc). Unfortunately, none of my ideas increased CV or LB. My favorite is below.\n\n## Locating PE with Class Activation Maps CAM\n\nI extracted class activation maps from my stage 1 models and feed the location information into stage 2 exam models. (CAMs explained [here][1]). In the below figure, the ground truth is in the title. The green circles are the CAM of my EfficientNetB4 model. Note that CT scans are flipped so the left side of the image is \"right\" and the right side of the image is \"left. You can see that the CAM does a good job of locating the PE.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F291dc983514a322d5ce4e56b22a9a681%2Fcam.png?generation=1603755601607192&alt=media)\n\n# Thank you\n\n[1]: https://www.kaggle.com/cdeotte/unsupervised-masks-cv-0-60\n[2]: https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/193598",
      "votes": 89
    },
    {
      "id": 1062454,
      "postDate": "2020-10-27T21:00:52.477Z",
      "content": "<p>Thank you so much for sharing <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> !! I love your solution discussions! There is so much to learn from them! 💛</p>\n<p>I was very sad that I was not able to submit a solution because I completely underestimated the performance challenge of this competition. I was so happy to see now that you also used only 30 % of the data (yes! :-)) and windowed images. Why did you use level 40 and width 400? We had the idea to use the \"blood windows\" ranging from 5-60 and 60-160 HU to detect chronic and acute PE and segmented lungs. Why have you used a more broad window instead? How did you use only 1 channel? Aren't 3 channels expected by your model?!</p>\n<p>It's great to see how it helped to discover where PE is usually located. That's so elegant! We thought about masking the exam targets for images showing no PE at all but it already felt like \"showing too much\" for solving the exam-level. Now I understand that my problem was that I tried to solve this all with one model instead of splitting the tasks on several different models and stages. I think I was too focused on finding a single solution because of the performance stuff and my little time available. I was also much too cautious to use smaller image sizes and backbones even though I had in mind \"start simple\". </p>\n<p>Thank you again for sharing your ideas and for making it possible to learn and improve!</p>",
      "rawMarkdown": "Thank you so much for sharing @cdeotte !! I love your solution discussions! There is so much to learn from them! 💛\n\nI was very sad that I was not able to submit a solution because I completely underestimated the performance challenge of this competition. I was so happy to see now that you also used only 30 % of the data (yes! :-)) and windowed images. Why did you use level 40 and width 400? We had the idea to use the \"blood windows\" ranging from 5-60 and 60-160 HU to detect chronic and acute PE and segmented lungs. Why have you used a more broad window instead? How did you use only 1 channel? Aren't 3 channels expected by your model?!\n\nIt's great to see how it helped to discover where PE is usually located. That's so elegant! We thought about masking the exam targets for images showing no PE at all but it already felt like \"showing too much\" for solving the exam-level. Now I understand that my problem was that I tried to solve this all with one model instead of splitting the tasks on several different models and stages. I think I was too focused on finding a single solution because of the performance stuff and my little time available. I was also much too cautious to use smaller image sizes and backbones even though I had in mind \"start simple\". \n\nThank you again for sharing your ideas and for making it possible to learn and improve!",
      "votes": 3,
      "replies": [
        {
          "id": 1062465,
          "postDate": "2020-10-27T21:21:21.497Z",
          "content": "<p>Thanks Laura. Using only one \"window\" (1-channel) reduced disk and RAM storage by 66% compared with saving 3 channel images to disk or RAM. (also note that original dicom is <code>uint16</code>, so applying windowing additionally reduces data by 50%). Furthermore, all images are <code>uint8</code> outside of the model which is 75% less RAM than <code>float32</code> images. (i.e. don't divide by 255 outside the model).</p>\n<p>So, when using 128x128 <code>uint8</code> 1-channel images, all the train data can fit in 8.9GB. (that's 542_769 of the 1_790_594 train images where exam has PE). You don't even need to read from hard drive after putting it all in RAM. The data loader is incredibly fast. I'll say that again, the original 1_000_000_000_000 bytes (1TB) of train data can fit in 8.9GB of memory!</p>\n<p>Then the model receives the 1-channel image and converts to 3-channel <code>float32</code> before inputting into EfficientNet as below</p>\n<pre><code># THIS MODEL TRAINS ON RANDOM 80X80 CROPS FROM 128X128\ninp = tf.keras.Input(shape=(80, 80, 1)) # INPUT IS UINT8\nx = tf.keras.layers.Concatenate()([inp/255., inp/255., inp/255.])\nbase_model = efn.EfficientNetB0(weights='imagenet', include_top=False) \nx = base_model(x)\n</code></pre>\n<p>I experimented with using different windowing schemes, but nothing increased my CV LB. I even did creative stuff like making 7 channel images where the additional 6 channels were neighboring slices. But again it didn't help. I also did strange stuff like embedding location information into images. For example, put left side of image in Red channel and right side of image in Blue channel but again it didn't help.</p>",
          "rawMarkdown": "Thanks Laura. Using only one \"window\" (1-channel) reduced disk and RAM storage by 66% compared with saving 3 channel images to disk or RAM. (also note that original dicom is `uint16`, so applying windowing additionally reduces data by 50%). Furthermore, all images are `uint8` outside of the model which is 75% less RAM than `float32` images. (i.e. don't divide by 255 outside the model).\n\nSo, when using 128x128 `uint8` 1-channel images, all the train data can fit in 8.9GB. (that's 542_769 of the 1_790_594 train images where exam has PE). You don't even need to read from hard drive after putting it all in RAM. The data loader is incredibly fast. I'll say that again, the original 1_000_000_000_000 bytes (1TB) of train data can fit in 8.9GB of memory!\n\nThen the model receives the 1-channel image and converts to 3-channel `float32` before inputting into EfficientNet as below\n\n    # THIS MODEL TRAINS ON RANDOM 80X80 CROPS FROM 128X128\n    inp = tf.keras.Input(shape=(80, 80, 1)) # INPUT IS UINT8\n    x = tf.keras.layers.Concatenate()([inp/255., inp/255., inp/255.])\n    base_model = efn.EfficientNetB0(weights='imagenet', include_top=False) \n    x = base_model(x)\n\nI experimented with using different windowing schemes, but nothing increased my CV LB. I even did creative stuff like making 7 channel images where the additional 6 channels were neighboring slices. But again it didn't help. I also did strange stuff like embedding location information into images. For example, put left side of image in Red channel and right side of image in Blue channel but again it didn't help.\n\n",
          "votes": 5
        },
        {
          "id": 1074060,
          "postDate": "2020-11-10T07:38:09.197Z",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>  - I am actually trying to replicate the solution on my own machine to learn working with Dicoms. </p>\n<p>Currently using the preprocessing script shared by Ian Pan <a href=\"https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/182930\" target=\"_blank\">here</a>.</p>\n<p>Am I right in assuming that you converted dicoms to signle channel PNGs first using (40, 400) mediastinal window and then fed them into the model please? :) </p>",
          "rawMarkdown": "Thank you @cdeotte  - I am actually trying to replicate the solution on my own machine to learn working with Dicoms. \n\nCurrently using the preprocessing script shared by Ian Pan [here](https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/182930).\n\nAm I right in assuming that you converted dicoms to signle channel PNGs first using (40, 400) mediastinal window and then fed them into the model please? :) "
        },
        {
          "id": 1074067,
          "postDate": "2020-11-10T07:44:06.497Z",
          "content": "<p>Also, any chance you'd be willing your preprocessing script please? Thanks in advance!</p>",
          "rawMarkdown": "Also, any chance you'd be willing your preprocessing script please? Thanks in advance!"
        }
      ]
    },
    {
      "id": 1061360,
      "postDate": "2020-10-27T00:43:47.620Z",
      "content": "<p>Great work <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> - Impressive solo finish! Cool to see you implemented random forest in your stage 2 model! I wanted to try a tabular model in stacking but we didn't have time. May I ask why you used random forest and not a GBM?</p>",
      "rawMarkdown": "Great work @cdeotte - Impressive solo finish! Cool to see you implemented random forest in your stage 2 model! I wanted to try a tabular model in stacking but we didn't have time. May I ask why you used random forest and not a GBM?",
      "votes": 3,
      "replies": [
        {
          "id": 1061370,
          "postDate": "2020-10-27T00:54:16.060Z",
          "content": "<p>Thanks Rob. Congrats to your team too.</p>\n<p>It's my intuition and CV proved it correct. (GBM had lower CV) My stage 2 random forest has 53 features. They are <code>oof</code> and then 26 forward shifts and 26 backward shifts of oof. I wanted all these features to be treated somewhat similar. I didn't want the model searching for signal in specific shifts.</p>\n<p>In Ion comp, random forest also did better than GBM for the same reason when using signal and shifted signal. (Random forest overfitted less than GBM).</p>",
          "rawMarkdown": "Thanks Rob. Congrats to your team too.\n\nIt's my intuition and CV proved it correct. (GBM had lower CV) My stage 2 random forest has 53 features. They are `oof` and then 26 forward shifts and 26 backward shifts of oof. I wanted all these features to be treated somewhat similar. I didn't want the model searching for signal in specific shifts.\n\nIn Ion comp, random forest also did better than GBM for the same reason when using signal and shifted signal. (Random forest overfitted less than GBM).",
          "votes": 3
        },
        {
          "id": 1061375,
          "postDate": "2020-10-27T01:02:36.953Z",
          "content": "<p>Makes sense, thanks for explaining!</p>",
          "rawMarkdown": "Makes sense, thanks for explaining!",
          "votes": 2
        },
        {
          "id": 1061379,
          "postDate": "2020-10-27T01:05:24.947Z",
          "content": "<p>Interesting! We have had a very similar approach and lgbm worked much better for us. I will give a write up Tomorrow.</p>",
          "rawMarkdown": "Interesting! We have had a very similar approach and lgbm worked much better for us. I will give a write up Tomorrow.",
          "votes": 4
        }
      ]
    },
    {
      "id": 1082264,
      "postDate": "2020-11-17T18:38:07.450Z",
      "content": "<p>Thank you very much for your insights!! I always join your exact explanations and tutorials!</p>\n<p>Could you explain how you trained your Stage 1 Model with the additional GAP - layer?</p>\n<p>I always get errors, when training it with two outputs (x, GAP) while having only one target. I tried to Compile with:</p>\n<p>model.compile(<br>\n        optimizer='adam',<br>\n        loss = [tf.keras.losses.BinaryCrossentropy(from_logits=True), None],<br>\n        metrics=['accuracy'])   <br>\n but when training (model.fit(get_training_dataset(), …) I get th error: 'Dimensions must be equal, but are 7 and 1791 for '{{node Equal_1}} '</p>",
      "rawMarkdown": "Thank you very much for your insights!! I always join your exact explanations and tutorials!\n\nCould you explain how you trained your Stage 1 Model with the additional GAP - layer?\n\nI always get errors, when training it with two outputs (x, GAP) while having only one target. I tried to Compile with:\n\nmodel.compile(\n        optimizer='adam',\n        loss = [tf.keras.losses.BinaryCrossentropy(from_logits=True), None],\n        metrics=['accuracy'])   \n but when training (model.fit(get_training_dataset(), ...) I get th error: 'Dimensions must be equal, but are 7 and 1791 for '{{node Equal_1}} '",
      "votes": 1,
      "replies": [
        {
          "id": 1105403,
          "postDate": "2020-12-07T20:38:41.513Z",
          "content": "<p><a href=\"https://www.kaggle.com/jensvannahl\" target=\"_blank\">@jensvannahl</a> We build the stage 1 model with only 1 output and train it with 1 output. After it is done training, you build a second model that has 2 outputs and transfer the trained weights from the first model to the second model. Then use the second model for inference only (i.e. <code>_, GAP = model2.predict(train)</code> )to extract the GAP layer. (We do not train any model that has 2 outputs).</p>",
          "rawMarkdown": "@jensvannahl We build the stage 1 model with only 1 output and train it with 1 output. After it is done training, you build a second model that has 2 outputs and transfer the trained weights from the first model to the second model. Then use the second model for inference only (i.e. `_, GAP = model2.predict(train)` )to extract the GAP layer. (We do not train any model that has 2 outputs)."
        }
      ]
    },
    {
      "id": 1082257,
      "postDate": "2020-11-17T18:32:22.393Z",
      "content": "<p>Thank you very much for these insights!!! </p>\n<p>Could you please explain, how you trained stage 1 Model mit additional output of the GAP-layer?<br>\nCompiling the </p>\n<p>model = Model (model input, [x, GAP])            I used<br>\nmodel.compile(<br>\n        optimizer='adam',<br>\n        loss = [tf.keras.losses.BinaryCrossentropy(from_logits=True), None],<br>\n        metrics=['accuracy'])</p>\n<p>with no problems, but with</p>\n<p>model.fit(get_training_dataset(), …)   I always get 'Dimensions must be equal, but are 7 and 1792 for '{{node Equal_1}} = Equal[T=DT_FLOAT, incompatible_shape_error=true](Cast_18, Cast_20)' with input shapes: [?,7], [?,1792].'</p>",
      "rawMarkdown": "Thank you very much for these insights!!! \n\nCould you please explain, how you trained stage 1 Model mit additional output of the GAP-layer?\nCompiling the \n\nmodel = Model (model input, [x, GAP])            I used\nmodel.compile(\n        optimizer='adam',\n        loss = [tf.keras.losses.BinaryCrossentropy(from_logits=True), None],\n        metrics=['accuracy'])\n\nwith no problems, but with\n\nmodel.fit(get_training_dataset(), ...)   I always get 'Dimensions must be equal, but are 7 and 1792 for '{{node Equal_1}} = Equal[T=DT_FLOAT, incompatible_shape_error=true](Cast_18, Cast_20)' with input shapes: [?,7], [?,1792].'",
      "votes": 1
    },
    {
      "id": 1073690,
      "postDate": "2020-11-09T20:15:49.650Z",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> and congrats on the finish! I have been following your work since ISIC Melonama and it is always a pleasure to read :) </p>\n<p>Could I please ask how you trained <code>Stage 1 - Model Two - Patient Level Predictions</code>? </p>\n<p>I am a bit confused by: </p>\n<blockquote>\n  <p>then take the 9 average predictions.</p>\n</blockquote>\n<p>Where are these 9 predictions coming from? And what are the labels used to train the model? Sorry if it's a silly question.</p>",
      "rawMarkdown": "Thank you @cdeotte and congrats on the finish! I have been following your work since ISIC Melonama and it is always a pleasure to read :) \n\nCould I please ask how you trained `Stage 1 - Model Two - Patient Level Predictions`? \n\nI am a bit confused by: \n> then take the 9 average predictions.\n\nWhere are these 9 predictions coming from? And what are the labels used to train the model? Sorry if it's a silly question.",
      "votes": 1,
      "replies": [
        {
          "id": 1073741,
          "postDate": "2020-11-09T21:59:29.063Z",
          "content": "<p>Hi Aman. For each patient (i.e. study), I discard all slices that are outside <code>0.35 &lt; z &lt; 0.65</code>. So for example, if a patient has 100 slices, I sort them in Z order and then discard the first 35 and last 35 slices. I only keep the 30% middle slices from <code>0.35 &lt; z_normalized &lt; 0.65</code>. I put all the slices I'm keeping from all the patients in <code>X_train_exam</code>.</p>\n<p>I then train an EfficientNet using on <code>X_train_exam</code> with an output layer of 9 sigmoid units and weighted log loss (based on the metric weights for each of the 9 exam targets).</p>\n<p>Finally for inference, for each patient, I predict the 9 targets for each of a patient's 30% middle slices (i.e. <code>X_test_exam</code>). So if a test dataset patient originally has 100 slices, i will predict for 30 middle slices. I now have 30 predictions for each of the 9 targets. For a specific target such as <code>rv_lv_ratio_gte_1</code>, i will take the average of those 30 predictions and submit that to Kaggle for that patient.</p>\n<p>==========</p>\n<p>Using the above technique, I could make a submission to Kaggle using only Stage 1 models and achieved LB 0.215. After building a Stage 2 model to predict exam level predictions, I no longer used these Stage 1 predictions in my final submission. </p>\n<p>For a while I ensembled my Stage 1 and Stage 2 exam predictions (and that was better than using either alone). But then my Stage 2 model became so accurate that using Stage 1 predictions only made it worse.</p>",
          "rawMarkdown": "Hi Aman. For each patient (i.e. study), I discard all slices that are outside `0.35 < z < 0.65`. So for example, if a patient has 100 slices, I sort them in Z order and then discard the first 35 and last 35 slices. I only keep the 30% middle slices from `0.35 < z_normalized < 0.65`. I put all the slices I'm keeping from all the patients in `X_train_exam`.\n\nI then train an EfficientNet using on `X_train_exam` with an output layer of 9 sigmoid units and weighted log loss (based on the metric weights for each of the 9 exam targets).\n\nFinally for inference, for each patient, I predict the 9 targets for each of a patient's 30% middle slices (i.e. `X_test_exam`). So if a test dataset patient originally has 100 slices, i will predict for 30 middle slices. I now have 30 predictions for each of the 9 targets. For a specific target such as `rv_lv_ratio_gte_1`, i will take the average of those 30 predictions and submit that to Kaggle for that patient.\n\n==========\n\nUsing the above technique, I could make a submission to Kaggle using only Stage 1 models and achieved LB 0.215. After building a Stage 2 model to predict exam level predictions, I no longer used these Stage 1 predictions in my final submission. \n\nFor a while I ensembled my Stage 1 and Stage 2 exam predictions (and that was better than using either alone). But then my Stage 2 model became so accurate that using Stage 1 predictions only made it worse.",
          "votes": 2
        },
        {
          "id": 1073973,
          "postDate": "2020-11-10T05:04:19.190Z",
          "content": "<p>Thanks so much <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> ! Simple baseline and effective! Thanks again for the great explanation. </p>",
          "rawMarkdown": "Thanks so much @cdeotte ! Simple baseline and effective! Thanks again for the great explanation. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1065308,
      "postDate": "2020-10-31T06:23:26.610Z",
      "content": "<p>Great job!</p>",
      "rawMarkdown": "Great job!",
      "votes": 1
    },
    {
      "id": 1065047,
      "postDate": "2020-10-30T19:37:06.947Z",
      "content": "<p>Wow, learnt a lot of new stuff, thanks</p>",
      "rawMarkdown": "Wow, learnt a lot of new stuff, thanks",
      "votes": 1,
      "replies": [
        {
          "id": 1065230,
          "postDate": "2020-10-31T04:01:29.907Z",
          "content": "<p>Thanks Sayid</p>",
          "rawMarkdown": "Thanks Sayid"
        }
      ]
    },
    {
      "id": 1063967,
      "postDate": "2020-10-29T14:46:09.043Z",
      "content": "<p>Great work <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>.. May be a naive question. the first model the pe_present_on_image has been multiplied by 5.6222. Any reason?. </p>",
      "rawMarkdown": "Great work @cdeotte.. May be a naive question. the first model the pe_present_on_image has been multiplied by 5.6222. Any reason?. ",
      "votes": 1,
      "replies": [
        {
          "id": 1064081,
          "postDate": "2020-10-29T17:05:28.097Z",
          "content": "<p>The factor <code>5.6222</code> isn't really needed. Instead I could increase the learning rate by <code>5.6222</code> or just train longer.</p>\n<p>The point is this. When an NN adjusts the model weights during training, it uses the error gradient multiplied by learning rate multiplied by sample weight. So when you adjust the sample weight you effectively adjust your model's learning rate.</p>\n<p>In order to keep the same learning rate that worked before using sample weight, I adjust by 5.6222. The scaling factor <code>5.6222 = train.shape[0]/train.weight.sum()</code></p>\n<p>The sample weight parameter adds a new training weight to each row of the training data. Notice that if each row had default <code>sample weight = 1</code>, then <code>train.shape[0]/train.weight.sum() = 1</code> and no adjustment is needed.</p>",
          "rawMarkdown": "The factor `5.6222` isn't really needed. Instead I could increase the learning rate by `5.6222` or just train longer.\n\nThe point is this. When an NN adjusts the model weights during training, it uses the error gradient multiplied by learning rate multiplied by sample weight. So when you adjust the sample weight you effectively adjust your model's learning rate.\n\nIn order to keep the same learning rate that worked before using sample weight, I adjust by 5.6222. The scaling factor `5.6222 = train.shape[0]/train.weight.sum()`\n\nThe sample weight parameter adds a new training weight to each row of the training data. Notice that if each row had default `sample weight = 1`, then `train.shape[0]/train.weight.sum() = 1` and no adjustment is needed.",
          "votes": 2
        },
        {
          "id": 1064099,
          "postDate": "2020-10-29T17:38:29.563Z",
          "content": "<p>I wonder why are you using weights instead of oversampling. Theoretically, wouldn't it be better to feed N different augmented inputs instead of one multiplied by N. I understand it costs more, but if cost is the main concern why training with small lr? </p>",
          "rawMarkdown": "I wonder why are you using weights instead of oversampling. Theoretically, wouldn't it be better to feed N different augmented inputs instead of one multiplied by N. I understand it costs more, but if cost is the main concern why training with small lr? ",
          "votes": 2
        },
        {
          "id": 1064130,
          "postDate": "2020-10-29T18:35:59.940Z",
          "content": "<p>The metric is more complicated than oversampling. Imagine we have two patients, <code>patient A</code> and <code>patient B</code>. Now say that <code>patient A</code> has 400 slices and <code>patient B</code> has 300 slices. Imagine that <code>patient A</code> has <strong>only</strong> one slice with <code>pe_present_on_image = 1</code>. Imagine <code>patient B</code> has 150 slices with <code>pe_present_on_image = 1</code>.</p>\n<p>According to the competition metic, the weight assoicated with <code>patient A's</code> one slice should be <code>1/400</code> and the weight associated to each of <code>patient B's</code> 150 slices should be <code>150/300</code> each. Notice that we aren't just making <code>pe_present_on_image = 1</code> oversampled.</p>\n<p>There is a fundamental difference between <code>patient A's</code> one positive slice and <code>patient B's</code> 150 positive slices. The pulmonary embolism in <code>patient A</code> will look different because it is smaller and only on one slice. The embolism in <code>patient B</code> will be larger and cover multiple slices. We need to weight these larger embolisms more. We cannot use oversample and just natively weight all embolisms (i.e. <code>pe_present_on_image = 1</code>) more.</p>",
          "rawMarkdown": "The metric is more complicated than oversampling. Imagine we have two patients, `patient A` and `patient B`. Now say that `patient A` has 400 slices and `patient B` has 300 slices. Imagine that `patient A` has **only** one slice with `pe_present_on_image = 1`. Imagine `patient B` has 150 slices with `pe_present_on_image = 1`.\n\nAccording to the competition metic, the weight assoicated with `patient A's` one slice should be `1/400` and the weight associated to each of `patient B's` 150 slices should be `150/300` each. Notice that we aren't just making `pe_present_on_image = 1` oversampled.\n\nThere is a fundamental difference between `patient A's` one positive slice and `patient B's` 150 positive slices. The pulmonary embolism in `patient A` will look different because it is smaller and only on one slice. The embolism in `patient B` will be larger and cover multiple slices. We need to weight these larger embolisms more. We cannot use oversample and just natively weight all embolisms (i.e. `pe_present_on_image = 1`) more.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1063662,
      "postDate": "2020-10-29T07:29:25.387Z",
      "content": "<p>This was a great read.</p>",
      "rawMarkdown": "This was a great read.",
      "votes": 1
    },
    {
      "id": 1063536,
      "postDate": "2020-10-29T03:31:27.337Z",
      "content": "<p>This is an advanced topic, thanks for sharing your insight</p>",
      "rawMarkdown": "This is an advanced topic, thanks for sharing your insight",
      "votes": 1
    },
    {
      "id": 1062940,
      "postDate": "2020-10-28T10:43:46.453Z",
      "content": "<p>Your explanations are always very detailed and contain great visuals to even better illustrate your work. Great stuff and thanks for sharing!</p>",
      "rawMarkdown": "Your explanations are always very detailed and contain great visuals to even better illustrate your work. Great stuff and thanks for sharing!",
      "votes": 1
    },
    {
      "id": 1062615,
      "postDate": "2020-10-28T03:21:36.563Z",
      "content": "<p>Good strategy.  I tried the efficient B &gt; 3 give me \"CSV not found\" or timeout. My bad bc I can't control it. Learn more from you.</p>",
      "rawMarkdown": "Good strategy.  I tried the efficient B > 3 give me \"CSV not found\" or timeout. My bad bc I can't control it. Learn more from you.",
      "votes": 1
    },
    {
      "id": 1062614,
      "postDate": "2020-10-28T03:14:59.403Z",
      "content": "<p>Your write-up is a master class in clean code and experimentation. Thanks for sharing. </p>",
      "rawMarkdown": "Your write-up is a master class in clean code and experimentation. Thanks for sharing. ",
      "votes": 1
    },
    {
      "id": 1062534,
      "postDate": "2020-10-28T00:12:57.453Z",
      "content": "<p>Thanks for sharing your clear and concise solution, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>! </p>\n<p>In the part of Stage 2 model two, you noted that: </p>\n<blockquote>\n  <p>Then we took those 64 images and extracted the GAP embeddings from both our stage 1 model one and stage 1 model two.</p>\n</blockquote>\n<p>This might be a beginner's question, but what exactly is the GAP embedding? </p>",
      "rawMarkdown": "Thanks for sharing your clear and concise solution, @cdeotte! \n\nIn the part of Stage 2 model two, you noted that: \n\n> Then we took those 64 images and extracted the GAP embeddings from both our stage 1 model one and stage 1 model two.\n\nThis might be a beginner's question, but what exactly is the GAP embedding? ",
      "votes": 1,
      "replies": [
        {
          "id": 1062588,
          "postDate": "2020-10-28T02:32:43.290Z",
          "content": "<p><code>GAP embedding</code> stands for output of the global average pooling 2D layer. The model below outputs both the <code>pe_present_on_image</code> prediction and the GAP embedding. </p>\n<pre><code>    inp = tf.keras.Input(shape=(320, 320, 1))\n    x = tf.keras.layers.Concatenate()([inp/255., inp/255., inp/255.])\n    base_model = efn.EfficientNetB4(weights='imagenet', include_top=False) \n    x = base_model(x)\n    GAP = tf.keras.layers.GlobalAveragePooling2D()(x)    \n    x = tf.keras.layers.Dense(1, activation='sigmoid')(GAP)\n    model2 = tf.keras.Model(inputs=inp, outputs=[x, GAP] )\n</code></pre>\n<p>After training the <code>model</code> described in the main discussion post above, we save the weights <code>model.save_weights('model.h5')</code>, then we load these weights into <code>model2</code> and extract the GAP embeddings as follows</p>\n<pre><code>model2.load_weights('model.h5')\n_, GAP = model2.predict(X_train)\n</code></pre>\n<p>Now <code>GAP</code> is a NumPy array of dimension <code>[ len(X_train), 1792]</code> because EfficientNetB4 has 1792 features in its global average pooling 2D output. If <code>X_train</code> is organized as every 64 rows is one exam, then we reshape GAP with <code>GAP = GAP.reshape((-1,64,1792))</code>. We then train our <code>model_GRU</code> with <code>model_GRU.fit(GAP, targets)</code> where <code>targets</code> has dimension <code>[ len(X_train)/64, 9]</code> and contains the exam targets.</p>",
          "rawMarkdown": "`GAP embedding` stands for output of the global average pooling 2D layer. The model below outputs both the `pe_present_on_image` prediction and the GAP embedding. \n\n        inp = tf.keras.Input(shape=(320, 320, 1))\n        x = tf.keras.layers.Concatenate()([inp/255., inp/255., inp/255.])\n        base_model = efn.EfficientNetB4(weights='imagenet', include_top=False) \n        x = base_model(x)\n        GAP = tf.keras.layers.GlobalAveragePooling2D()(x)    \n        x = tf.keras.layers.Dense(1, activation='sigmoid')(GAP)\n        model2 = tf.keras.Model(inputs=inp, outputs=[x, GAP] )\n\nAfter training the `model` described in the main discussion post above, we save the weights `model.save_weights('model.h5')`, then we load these weights into `model2` and extract the GAP embeddings as follows\n\n    model2.load_weights('model.h5')\n    _, GAP = model2.predict(X_train)\n\nNow `GAP` is a NumPy array of dimension `[ len(X_train), 1792]` because EfficientNetB4 has 1792 features in its global average pooling 2D output. If `X_train` is organized as every 64 rows is one exam, then we reshape GAP with `GAP = GAP.reshape((-1,64,1792))`. We then train our `model_GRU` with `model_GRU.fit(GAP, targets)` where `targets` has dimension `[ len(X_train)/64, 9]` and contains the exam targets.",
          "votes": 3
        },
        {
          "id": 1063467,
          "postDate": "2020-10-29T00:44:03.260Z",
          "content": "<p>Thanks for a detailed follow up! I finally fully understood what you did on your solution. </p>",
          "rawMarkdown": "Thanks for a detailed follow up! I finally fully understood what you did on your solution. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1062448,
      "postDate": "2020-10-27T20:52:14.383Z",
      "content": "<p>Good job on your place. Very well deserved!</p>",
      "rawMarkdown": "Good job on your place. Very well deserved!",
      "votes": 1
    },
    {
      "id": 1062426,
      "postDate": "2020-10-27T20:07:50.180Z",
      "content": "<p>Very good solution and approach ! Best work </p>",
      "rawMarkdown": "Very good solution and approach ! Best work ",
      "votes": 1
    },
    {
      "id": 1061695,
      "postDate": "2020-10-27T08:09:10.417Z",
      "content": "<p>learned a lot from seeing your solution <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>. now i know where i went wrong in my approaches. next time I am going to spend some more time on experiments with lesser data and smaller model.</p>",
      "rawMarkdown": "learned a lot from seeing your solution @cdeotte. now i know where i went wrong in my approaches. next time I am going to spend some more time on experiments with lesser data and smaller model.",
      "votes": 1,
      "replies": [
        {
          "id": 1062375,
          "postDate": "2020-10-27T19:15:00.483Z",
          "content": "<p>Me too. I had so much frustration trying to train on all the data. So many hours drumming my fingers while the training ran only to find it timeout.<br>\nNow thanks to Chris I learn the (in retrospect quite obvious though I missed it entirely :) ) lesson to experiment on smaller data and test ideas quickly. Also if I stepped back and thought as Chris did “do I need all of the data” that would have helped a lot. </p>\n<p>Kaggle = (Pain + inspiration)*learning </p>",
          "rawMarkdown": "Me too. I had so much frustration trying to train on all the data. So many hours drumming my fingers while the training ran only to find it timeout.\nNow thanks to Chris I learn the (in retrospect quite obvious though I missed it entirely :) ) lesson to experiment on smaller data and test ideas quickly. Also if I stepped back and thought as Chris did “do I need all of the data” that would have helped a lot. \n\nKaggle = (Pain + inspiration)*learning ",
          "votes": 1
        },
        {
          "id": 1062401,
          "postDate": "2020-10-27T19:37:31.540Z",
          "content": "<p>Good insight. I elaborate further <a href=\"https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/193598\" target=\"_blank\">here</a></p>",
          "rawMarkdown": "Good insight. I elaborate further [here][1]\n\n[1]: https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/193598"
        }
      ]
    },
    {
      "id": 1061434,
      "postDate": "2020-10-27T02:28:01.167Z",
      "content": "<p>Congrats on results <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> and thanks for sharing the detailed solution !!</p>",
      "rawMarkdown": "Congrats on results @cdeotte and thanks for sharing the detailed solution !!",
      "votes": 1
    },
    {
      "id": 1061361,
      "postDate": "2020-10-27T00:44:49.170Z",
      "content": "<p>Congrats Chris! </p>",
      "rawMarkdown": "Congrats Chris! ",
      "votes": 1
    },
    {
      "id": 1061356,
      "postDate": "2020-10-27T00:40:14.820Z",
      "content": "<p>Congrats on results and thanks for the writeup details solution <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a></p>",
      "rawMarkdown": "Congrats on results and thanks for the writeup details solution @cdeotte",
      "votes": 1
    },
    {
      "id": 1061669,
      "postDate": "2020-10-27T07:34:13.450Z",
      "content": "<p>Congratz on the nice finish Chris !</p>\n<p>I'm always impressed by your write-ups, they really make your solution stand out, and are pleasant to read.</p>",
      "rawMarkdown": "Congratz on the nice finish Chris !\n\nI'm always impressed by your write-ups, they really make your solution stand out, and are pleasant to read.",
      "votes": 2,
      "replies": [
        {
          "id": 1062303,
          "postDate": "2020-10-27T18:10:39.117Z",
          "content": "<p>Thanks Theo</p>",
          "rawMarkdown": "Thanks Theo",
          "votes": 1
        }
      ]
    },
    {
      "id": 1061585,
      "postDate": "2020-10-27T05:36:12.317Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> . Thank you for the amazing write up. I am impressed by your simple and easy flowing explanation in all the competitions :). Keep Shining. </p>",
      "rawMarkdown": "Hi @cdeotte . Thank you for the amazing write up. I am impressed by your simple and easy flowing explanation in all the competitions :). Keep Shining. ",
      "votes": 2
    },
    {
      "id": 1061515,
      "postDate": "2020-10-27T03:53:44.040Z",
      "content": "<p>Wow ! Its amazing how you get the time to create and run so many experiments AND document it too !!<br>\nThank you and congratulations </p>",
      "rawMarkdown": "Wow ! Its amazing how you get the time to create and run so many experiments AND document it too !!\nThank you and congratulations ",
      "votes": 2
    },
    {
      "id": 1061424,
      "postDate": "2020-10-27T02:22:04.337Z",
      "content": "<p>Congratulations Dr. <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>. You are truly an inspiration for newbies.. Thanks a lot for sharing the amazing concepts that fueled your solution… </p>",
      "rawMarkdown": "Congratulations Dr. @cdeotte. You are truly an inspiration for newbies.. Thanks a lot for sharing the amazing concepts that fueled your solution... ",
      "votes": 2,
      "replies": [
        {
          "id": 1061432,
          "postDate": "2020-10-27T02:24:58.567Z",
          "content": "<p>Thank you Redwan <a href=\"https://www.kaggle.com/redwankarimsony\" target=\"_blank\">@redwankarimsony</a> . Your notebooks were very helpful. I read all of them and used some of your ideas. Thank you.</p>",
          "rawMarkdown": "Thank you Redwan @redwankarimsony . Your notebooks were very helpful. I read all of them and used some of your ideas. Thank you.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1061380,
      "postDate": "2020-10-27T01:06:57.670Z",
      "content": "<p>Thank you for sharing. Great work <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> . Very informative and I need more time to read.</p>",
      "rawMarkdown": "Thank you for sharing. Great work @cdeotte . Very informative and I need more time to read.",
      "votes": 2
    },
    {
      "id": 1061365,
      "postDate": "2020-10-27T00:49:17.033Z",
      "content": "<p>You show the power of the NVIDIA device, I decided to buy a Tesla V100 after this competition.</p>",
      "rawMarkdown": "You show the power of the NVIDIA device, I decided to buy a Tesla V100 after this competition.",
      "votes": 2,
      "replies": [
        {
          "id": 1061373,
          "postDate": "2020-10-27T00:57:28.197Z",
          "content": "<p>My wallet hurts from me even reading this comment…<br>\nwith that said, I just bought a few 2080ti's myself 😂</p>",
          "rawMarkdown": "My wallet hurts from me even reading this comment...\nwith that said, I just bought a few 2080ti's myself 😂",
          "votes": 2
        },
        {
          "id": 1061445,
          "postDate": "2020-10-27T02:35:42.563Z",
          "content": "<p>Maybe TITAN RTX is a better choice, with 24GRAM, is better for training with large batch size.<br>\nabout 10% expensive than 2080ti?</p>",
          "rawMarkdown": "Maybe TITAN RTX is a better choice, with 24GRAM, is better for training with large batch size.\nabout 10% expensive than 2080ti?",
          "votes": 1
        },
        {
          "id": 1061451,
          "postDate": "2020-10-27T02:39:51.327Z",
          "content": "<p>As a Grand Grand Master of NVIDIA, could you please give us some advice to buy a GPU for Kaggle within 4000$USD? <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>",
          "rawMarkdown": "As a Grand Grand Master of NVIDIA, could you please give us some advice to buy a GPU for Kaggle within 4000$USD? @cdeotte ",
          "votes": 2
        },
        {
          "id": 1061452,
          "postDate": "2020-10-27T02:40:56.263Z",
          "content": "<p>I haven't compared all the GPUs yet myself. The main difference between the V (professional) series and the RTX (consumer) series is that the V series double checks to prevent computational errors and V series has more VRAM which is very helpful. However for building models (which use random augmentations already), an occasional error may not matter. So perhaps RTX is more bang for your buck. But maybe you will need a few of them so you have enough VRAM.</p>",
          "rawMarkdown": "I haven't compared all the GPUs yet myself. The main difference between the V (professional) series and the RTX (consumer) series is that the V series double checks to prevent computational errors and V series has more VRAM which is very helpful. However for building models (which use random augmentations already), an occasional error may not matter. So perhaps RTX is more bang for your buck. But maybe you will need a few of them so you have enough VRAM.",
          "votes": 3
        }
      ]
    },
    {
      "id": 1061331,
      "postDate": "2020-10-27T00:14:09.297Z",
      "content": "<p>There is alot to learn from <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Thanks for sharning</p>",
      "rawMarkdown": "There is alot to learn from @cdeotte Thanks for sharning",
      "votes": 2
    },
    {
      "id": 1062803,
      "postDate": "2020-10-28T07:49:49.430Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 1065142,
      "postDate": "2020-10-30T22:32:10.853Z",
      "content": "<p>Thanks for sharing !</p>",
      "rawMarkdown": "Thanks for sharing !",
      "votes": 1
    },
    {
      "id": 1064056,
      "postDate": "2020-10-29T16:32:00.303Z",
      "content": "<p>Thanks for sharing.</p>",
      "rawMarkdown": "Thanks for sharing.",
      "votes": 1
    },
    {
      "id": 1063675,
      "postDate": "2020-10-29T07:47:32.210Z",
      "content": "<p>This was of great help . Thank You !</p>",
      "rawMarkdown": "This was of great help . Thank You !",
      "votes": 1
    },
    {
      "id": 1063009,
      "postDate": "2020-10-28T12:30:22.323Z",
      "content": "<p>Thanks for Sharing!!!</p>",
      "rawMarkdown": "Thanks for Sharing!!!",
      "votes": 1
    },
    {
      "id": 1064062,
      "postDate": "2020-10-29T16:43:14.883Z",
      "content": "<p>Thanks for sharing!!</p>",
      "rawMarkdown": "Thanks for sharing!!",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 1062454,
      "author_name": "Laura Fink",
      "author_url": "",
      "post_date": "2020-10-27T21:00:52.477000",
      "content": "<p>Thank you so much for sharing <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> !! I love your solution discussions! There is so much to learn from them! 💛</p>\n<p>I was very sad that I was not able to submit a solution because I completely underestimated the performance challenge of this competition. I was so happy to see now that you also used only 30 % of the data (yes! :-)) and windowed images. Why did you use level 40 and width 400? We had the idea to use the \"blood windows\" ranging from 5-60 and 60-160 HU to detect chronic and acute PE and segmented lungs. Why have you used a more broad window instead? How did you use only 1 channel? Aren't 3 channels expected by your model?!</p>\n<p>It's great to see how it helped to discover where PE is usually located. That's so elegant! We thought about masking the exam targets for images showing no PE at all but it already felt like \"showing too much\" for solving the exam-level. Now I understand that my problem was that I tried to solve this all with one model instead of splitting the tasks on several different models and stages. I think I was too focused on finding a single solution because of the performance stuff and my little time available. I was also much too cautious to use smaller image sizes and backbones even though I had in mind \"start simple\". </p>\n<p>Thank you again for sharing your ideas and for making it possible to learn and improve!</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1062465,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-10-27T21:21:21.497000",
          "content": "<p>Thanks Laura. Using only one \"window\" (1-channel) reduced disk and RAM storage by 66% compared with saving 3 channel images to disk or RAM. (also note that original dicom is <code>uint16</code>, so applying windowing additionally reduces data by 50%). Furthermore, all images are <code>uint8</code> outside of the model which is 75% less RAM than <code>float32</code> images. (i.e. don't divide by 255 outside the model).</p>\n<p>So, when using 128x128 <code>uint8</code> 1-channel images, all the train data can fit in 8.9GB. (that's 542_769 of the 1_790_594 train images where exam has PE). You don't even need to read from hard drive after putting it all in RAM. The data loader is incredibly fast. I'll say that again, the original 1_000_000_000_000 bytes (1TB) of train data can fit in 8.9GB of memory!</p>\n<p>Then the model receives the 1-channel image and converts to 3-channel <code>float32</code> before inputting into EfficientNet as below</p>\n<pre><code># THIS MODEL TRAINS ON RANDOM 80X80 CROPS FROM 128X128\ninp = tf.keras.Input(shape=(80, 80, 1)) # INPUT IS UINT8\nx = tf.keras.layers.Concatenate()([inp/255., inp/255., inp/255.])\nbase_model = efn.EfficientNetB0(weights='imagenet', include_top=False) \nx = base_model(x)\n</code></pre>\n<p>I experimented with using different windowing schemes, but nothing increased my CV LB. I even did creative stuff like making 7 channel images where the additional 6 channels were neighboring slices. But again it didn't help. I also did strange stuff like embedding location information into images. For example, put left side of image in Red channel and right side of image in Blue channel but again it didn't help.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1074060,
          "author_name": "Aman Arora",
          "author_url": "",
          "post_date": "2020-11-10T07:38:09.197000",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>  - I am actually trying to replicate the solution on my own machine to learn working with Dicoms. </p>\n<p>Currently using the preprocessing script shared by Ian Pan <a href=\"https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/182930\" target=\"_blank\">here</a>.</p>\n<p>Am I right in assuming that you converted dicoms to signle channel PNGs first using (40, 400) mediastinal window and then fed them into the model please? :) </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1074067,
          "author_name": "Aman Arora",
          "author_url": "",
          "post_date": "2020-11-10T07:44:06.497000",
          "content": "<p>Also, any chance you'd be willing your preprocessing script please? Thanks in advance!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1061360,
      "author_name": "Rob Mulla",
      "author_url": "",
      "post_date": "2020-10-27T00:43:47.620000",
      "content": "<p>Great work <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> - Impressive solo finish! Cool to see you implemented random forest in your stage 2 model! I wanted to try a tabular model in stacking but we didn't have time. May I ask why you used random forest and not a GBM?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1061370,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-10-27T00:54:16.060000",
          "content": "<p>Thanks Rob. Congrats to your team too.</p>\n<p>It's my intuition and CV proved it correct. (GBM had lower CV) My stage 2 random forest has 53 features. They are <code>oof</code> and then 26 forward shifts and 26 backward shifts of oof. I wanted all these features to be treated somewhat similar. I didn't want the model searching for signal in specific shifts.</p>\n<p>In Ion comp, random forest also did better than GBM for the same reason when using signal and shifted signal. (Random forest overfitted less than GBM).</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1061375,
          "author_name": "Rob Mulla",
          "author_url": "",
          "post_date": "2020-10-27T01:02:36.953000",
          "content": "<p>Makes sense, thanks for explaining!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1061379,
          "author_name": "Jan Bre",
          "author_url": "",
          "post_date": "2020-10-27T01:05:24.947000",
          "content": "<p>Interesting! We have had a very similar approach and lgbm worked much better for us. I will give a write up Tomorrow.</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 1082264,
      "author_name": "Jens van Nahl",
      "author_url": "",
      "post_date": "2020-11-17T18:38:07.450000",
      "content": "<p>Thank you very much for your insights!! I always join your exact explanations and tutorials!</p>\n<p>Could you explain how you trained your Stage 1 Model with the additional GAP - layer?</p>\n<p>I always get errors, when training it with two outputs (x, GAP) while having only one target. I tried to Compile with:</p>\n<p>model.compile(<br>\n        optimizer='adam',<br>\n        loss = [tf.keras.losses.BinaryCrossentropy(from_logits=True), None],<br>\n        metrics=['accuracy'])   <br>\n but when training (model.fit(get_training_dataset(), …) I get th error: 'Dimensions must be equal, but are 7 and 1791 for '{{node Equal_1}} '</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1105403,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-12-07T20:38:41.513000",
          "content": "<p><a href=\"https://www.kaggle.com/jensvannahl\" target=\"_blank\">@jensvannahl</a> We build the stage 1 model with only 1 output and train it with 1 output. After it is done training, you build a second model that has 2 outputs and transfer the trained weights from the first model to the second model. Then use the second model for inference only (i.e. <code>_, GAP = model2.predict(train)</code> )to extract the GAP layer. (We do not train any model that has 2 outputs).</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1082257,
      "author_name": "Jens van Nahl",
      "author_url": "",
      "post_date": "2020-11-17T18:32:22.393000",
      "content": "<p>Thank you very much for these insights!!! </p>\n<p>Could you please explain, how you trained stage 1 Model mit additional output of the GAP-layer?<br>\nCompiling the </p>\n<p>model = Model (model input, [x, GAP])            I used<br>\nmodel.compile(<br>\n        optimizer='adam',<br>\n        loss = [tf.keras.losses.BinaryCrossentropy(from_logits=True), None],<br>\n        metrics=['accuracy'])</p>\n<p>with no problems, but with</p>\n<p>model.fit(get_training_dataset(), …)   I always get 'Dimensions must be equal, but are 7 and 1792 for '{{node Equal_1}} = Equal[T=DT_FLOAT, incompatible_shape_error=true](Cast_18, Cast_20)' with input shapes: [?,7], [?,1792].'</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1073690,
      "author_name": "Aman Arora",
      "author_url": "",
      "post_date": "2020-11-09T20:15:49.650000",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> and congrats on the finish! I have been following your work since ISIC Melonama and it is always a pleasure to read :) </p>\n<p>Could I please ask how you trained <code>Stage 1 - Model Two - Patient Level Predictions</code>? </p>\n<p>I am a bit confused by: </p>\n<blockquote>\n  <p>then take the 9 average predictions.</p>\n</blockquote>\n<p>Where are these 9 predictions coming from? And what are the labels used to train the model? Sorry if it's a silly question.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1073741,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-11-09T21:59:29.063000",
          "content": "<p>Hi Aman. For each patient (i.e. study), I discard all slices that are outside <code>0.35 &lt; z &lt; 0.65</code>. So for example, if a patient has 100 slices, I sort them in Z order and then discard the first 35 and last 35 slices. I only keep the 30% middle slices from <code>0.35 &lt; z_normalized &lt; 0.65</code>. I put all the slices I'm keeping from all the patients in <code>X_train_exam</code>.</p>\n<p>I then train an EfficientNet using on <code>X_train_exam</code> with an output layer of 9 sigmoid units and weighted log loss (based on the metric weights for each of the 9 exam targets).</p>\n<p>Finally for inference, for each patient, I predict the 9 targets for each of a patient's 30% middle slices (i.e. <code>X_test_exam</code>). So if a test dataset patient originally has 100 slices, i will predict for 30 middle slices. I now have 30 predictions for each of the 9 targets. For a specific target such as <code>rv_lv_ratio_gte_1</code>, i will take the average of those 30 predictions and submit that to Kaggle for that patient.</p>\n<p>==========</p>\n<p>Using the above technique, I could make a submission to Kaggle using only Stage 1 models and achieved LB 0.215. After building a Stage 2 model to predict exam level predictions, I no longer used these Stage 1 predictions in my final submission. </p>\n<p>For a while I ensembled my Stage 1 and Stage 2 exam predictions (and that was better than using either alone). But then my Stage 2 model became so accurate that using Stage 1 predictions only made it worse.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1073973,
          "author_name": "Aman Arora",
          "author_url": "",
          "post_date": "2020-11-10T05:04:19.190000",
          "content": "<p>Thanks so much <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> ! Simple baseline and effective! Thanks again for the great explanation. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1065308,
      "author_name": "Dhruv Anurag",
      "author_url": "",
      "post_date": "2020-10-31T06:23:26.610000",
      "content": "<p>Great job!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1065047,
      "author_name": "Sayid Khan",
      "author_url": "",
      "post_date": "2020-10-30T19:37:06.947000",
      "content": "<p>Wow, learnt a lot of new stuff, thanks</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1065230,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-10-31T04:01:29.907000",
          "content": "<p>Thanks Sayid</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1063967,
      "author_name": "Manoj Prabhakar",
      "author_url": "",
      "post_date": "2020-10-29T14:46:09.043000",
      "content": "<p>Great work <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>.. May be a naive question. the first model the pe_present_on_image has been multiplied by 5.6222. Any reason?. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1064081,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-10-29T17:05:28.097000",
          "content": "<p>The factor <code>5.6222</code> isn't really needed. Instead I could increase the learning rate by <code>5.6222</code> or just train longer.</p>\n<p>The point is this. When an NN adjusts the model weights during training, it uses the error gradient multiplied by learning rate multiplied by sample weight. So when you adjust the sample weight you effectively adjust your model's learning rate.</p>\n<p>In order to keep the same learning rate that worked before using sample weight, I adjust by 5.6222. The scaling factor <code>5.6222 = train.shape[0]/train.weight.sum()</code></p>\n<p>The sample weight parameter adds a new training weight to each row of the training data. Notice that if each row had default <code>sample weight = 1</code>, then <code>train.shape[0]/train.weight.sum() = 1</code> and no adjustment is needed.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1064099,
          "author_name": "Miroslav Valan",
          "author_url": "",
          "post_date": "2020-10-29T17:38:29.563000",
          "content": "<p>I wonder why are you using weights instead of oversampling. Theoretically, wouldn't it be better to feed N different augmented inputs instead of one multiplied by N. I understand it costs more, but if cost is the main concern why training with small lr? </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1064130,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-10-29T18:35:59.940000",
          "content": "<p>The metric is more complicated than oversampling. Imagine we have two patients, <code>patient A</code> and <code>patient B</code>. Now say that <code>patient A</code> has 400 slices and <code>patient B</code> has 300 slices. Imagine that <code>patient A</code> has <strong>only</strong> one slice with <code>pe_present_on_image = 1</code>. Imagine <code>patient B</code> has 150 slices with <code>pe_present_on_image = 1</code>.</p>\n<p>According to the competition metic, the weight assoicated with <code>patient A's</code> one slice should be <code>1/400</code> and the weight associated to each of <code>patient B's</code> 150 slices should be <code>150/300</code> each. Notice that we aren't just making <code>pe_present_on_image = 1</code> oversampled.</p>\n<p>There is a fundamental difference between <code>patient A's</code> one positive slice and <code>patient B's</code> 150 positive slices. The pulmonary embolism in <code>patient A</code> will look different because it is smaller and only on one slice. The embolism in <code>patient B</code> will be larger and cover multiple slices. We need to weight these larger embolisms more. We cannot use oversample and just natively weight all embolisms (i.e. <code>pe_present_on_image = 1</code>) more.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1063662,
      "author_name": "Zarif Mustaq",
      "author_url": "",
      "post_date": "2020-10-29T07:29:25.387000",
      "content": "<p>This was a great read.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1063536,
      "author_name": "Prakhar",
      "author_url": "",
      "post_date": "2020-10-29T03:31:27.337000",
      "content": "<p>This is an advanced topic, thanks for sharing your insight</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1062940,
      "author_name": "Manch Hui",
      "author_url": "",
      "post_date": "2020-10-28T10:43:46.453000",
      "content": "<p>Your explanations are always very detailed and contain great visuals to even better illustrate your work. Great stuff and thanks for sharing!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1062615,
      "author_name": "( ͡° ͜ʖ ͡°)",
      "author_url": "",
      "post_date": "2020-10-28T03:21:36.563000",
      "content": "<p>Good strategy.  I tried the efficient B &gt; 3 give me \"CSV not found\" or timeout. My bad bc I can't control it. Learn more from you.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1062614,
      "author_name": "Ronaldo S.A. Batista",
      "author_url": "",
      "post_date": "2020-10-28T03:14:59.403000",
      "content": "<p>Your write-up is a master class in clean code and experimentation. Thanks for sharing. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1062534,
      "author_name": "lr",
      "author_url": "",
      "post_date": "2020-10-28T00:12:57.453000",
      "content": "<p>Thanks for sharing your clear and concise solution, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>! </p>\n<p>In the part of Stage 2 model two, you noted that: </p>\n<blockquote>\n  <p>Then we took those 64 images and extracted the GAP embeddings from both our stage 1 model one and stage 1 model two.</p>\n</blockquote>\n<p>This might be a beginner's question, but what exactly is the GAP embedding? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1062588,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-10-28T02:32:43.290000",
          "content": "<p><code>GAP embedding</code> stands for output of the global average pooling 2D layer. The model below outputs both the <code>pe_present_on_image</code> prediction and the GAP embedding. </p>\n<pre><code>    inp = tf.keras.Input(shape=(320, 320, 1))\n    x = tf.keras.layers.Concatenate()([inp/255., inp/255., inp/255.])\n    base_model = efn.EfficientNetB4(weights='imagenet', include_top=False) \n    x = base_model(x)\n    GAP = tf.keras.layers.GlobalAveragePooling2D()(x)    \n    x = tf.keras.layers.Dense(1, activation='sigmoid')(GAP)\n    model2 = tf.keras.Model(inputs=inp, outputs=[x, GAP] )\n</code></pre>\n<p>After training the <code>model</code> described in the main discussion post above, we save the weights <code>model.save_weights('model.h5')</code>, then we load these weights into <code>model2</code> and extract the GAP embeddings as follows</p>\n<pre><code>model2.load_weights('model.h5')\n_, GAP = model2.predict(X_train)\n</code></pre>\n<p>Now <code>GAP</code> is a NumPy array of dimension <code>[ len(X_train), 1792]</code> because EfficientNetB4 has 1792 features in its global average pooling 2D output. If <code>X_train</code> is organized as every 64 rows is one exam, then we reshape GAP with <code>GAP = GAP.reshape((-1,64,1792))</code>. We then train our <code>model_GRU</code> with <code>model_GRU.fit(GAP, targets)</code> where <code>targets</code> has dimension <code>[ len(X_train)/64, 9]</code> and contains the exam targets.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1063467,
          "author_name": "lr",
          "author_url": "",
          "post_date": "2020-10-29T00:44:03.260000",
          "content": "<p>Thanks for a detailed follow up! I finally fully understood what you did on your solution. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1062448,
      "author_name": "Brenda N",
      "author_url": "",
      "post_date": "2020-10-27T20:52:14.383000",
      "content": "<p>Good job on your place. Very well deserved!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1062426,
      "author_name": "Varun Barath",
      "author_url": "",
      "post_date": "2020-10-27T20:07:50.180000",
      "content": "<p>Very good solution and approach ! Best work </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1061695,
      "author_name": "yuvaramsingh",
      "author_url": "",
      "post_date": "2020-10-27T08:09:10.417000",
      "content": "<p>learned a lot from seeing your solution <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>. now i know where i went wrong in my approaches. next time I am going to spend some more time on experiments with lesser data and smaller model.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1062375,
          "author_name": "Mark P",
          "author_url": "",
          "post_date": "2020-10-27T19:15:00.483000",
          "content": "<p>Me too. I had so much frustration trying to train on all the data. So many hours drumming my fingers while the training ran only to find it timeout.<br>\nNow thanks to Chris I learn the (in retrospect quite obvious though I missed it entirely :) ) lesson to experiment on smaller data and test ideas quickly. Also if I stepped back and thought as Chris did “do I need all of the data” that would have helped a lot. </p>\n<p>Kaggle = (Pain + inspiration)*learning </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1062401,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-10-27T19:37:31.540000",
          "content": "<p>Good insight. I elaborate further <a href=\"https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/193598\" target=\"_blank\">here</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1061434,
      "author_name": "Kamal Das",
      "author_url": "",
      "post_date": "2020-10-27T02:28:01.167000",
      "content": "<p>Congrats on results <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> and thanks for sharing the detailed solution !!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1061361,
      "author_name": "sin",
      "author_url": "",
      "post_date": "2020-10-27T00:44:49.170000",
      "content": "<p>Congrats Chris! </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1061356,
      "author_name": "KhanhVD",
      "author_url": "",
      "post_date": "2020-10-27T00:40:14.820000",
      "content": "<p>Congrats on results and thanks for the writeup details solution <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1061669,
      "author_name": "Theo Viel",
      "author_url": "",
      "post_date": "2020-10-27T07:34:13.450000",
      "content": "<p>Congratz on the nice finish Chris !</p>\n<p>I'm always impressed by your write-ups, they really make your solution stand out, and are pleasant to read.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1062303,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-10-27T18:10:39.117000",
          "content": "<p>Thanks Theo</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1061585,
      "author_name": "JamshaidSohail",
      "author_url": "",
      "post_date": "2020-10-27T05:36:12.317000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> . Thank you for the amazing write up. I am impressed by your simple and easy flowing explanation in all the competitions :). Keep Shining. </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1061515,
      "author_name": "Vee",
      "author_url": "",
      "post_date": "2020-10-27T03:53:44.040000",
      "content": "<p>Wow ! Its amazing how you get the time to create and run so many experiments AND document it too !!<br>\nThank you and congratulations </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1061424,
      "author_name": "Redwan Sony",
      "author_url": "",
      "post_date": "2020-10-27T02:22:04.337000",
      "content": "<p>Congratulations Dr. <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>. You are truly an inspiration for newbies.. Thanks a lot for sharing the amazing concepts that fueled your solution… </p>",
      "votes": 2,
      "replies": [
        {
          "id": 1061432,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-10-27T02:24:58.567000",
          "content": "<p>Thank you Redwan <a href=\"https://www.kaggle.com/redwankarimsony\" target=\"_blank\">@redwankarimsony</a> . Your notebooks were very helpful. I read all of them and used some of your ideas. Thank you.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1061380,
      "author_name": "Johnny Lee",
      "author_url": "",
      "post_date": "2020-10-27T01:06:57.670000",
      "content": "<p>Thank you for sharing. Great work <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> . Very informative and I need more time to read.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1061365,
      "author_name": "Dewei Chen",
      "author_url": "",
      "post_date": "2020-10-27T00:49:17.033000",
      "content": "<p>You show the power of the NVIDIA device, I decided to buy a Tesla V100 after this competition.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1061373,
          "author_name": "Stanley Zheng",
          "author_url": "",
          "post_date": "2020-10-27T00:57:28.197000",
          "content": "<p>My wallet hurts from me even reading this comment…<br>\nwith that said, I just bought a few 2080ti's myself 😂</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1061445,
          "author_name": "Dewei Chen",
          "author_url": "",
          "post_date": "2020-10-27T02:35:42.563000",
          "content": "<p>Maybe TITAN RTX is a better choice, with 24GRAM, is better for training with large batch size.<br>\nabout 10% expensive than 2080ti?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1061451,
          "author_name": "Dewei Chen",
          "author_url": "",
          "post_date": "2020-10-27T02:39:51.327000",
          "content": "<p>As a Grand Grand Master of NVIDIA, could you please give us some advice to buy a GPU for Kaggle within 4000$USD? <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1061452,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-10-27T02:40:56.263000",
          "content": "<p>I haven't compared all the GPUs yet myself. The main difference between the V (professional) series and the RTX (consumer) series is that the V series double checks to prevent computational errors and V series has more VRAM which is very helpful. However for building models (which use random augmentations already), an occasional error may not matter. So perhaps RTX is more bang for your buck. But maybe you will need a few of them so you have enough VRAM.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1061331,
      "author_name": "Saurabh dubey",
      "author_url": "",
      "post_date": "2020-10-27T00:14:09.297000",
      "content": "<p>There is alot to learn from <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Thanks for sharning</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1062803,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-28T07:49:49.430000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1065142,
      "author_name": "Camilo",
      "author_url": "",
      "post_date": "2020-10-30T22:32:10.853000",
      "content": "<p>Thanks for sharing !</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1064056,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-29T16:32:00.303000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1063675,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-29T07:47:32.210000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1063009,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-28T12:30:22.323000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1064062,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-29T16:43:14.883000",
      "content": "",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1061327": "Thank you Radiological Society of North America (RSNA®), Society of Thoracic Radiology (STR), and Kaggle for hosting this fun competition. Thank you Nvidia for providing compute resources.\n\nThis has been one of my favorite competitions. I enjoyed building an elaborate multi-stage pipeline of stacked models! Working with 3D images was fun and provided an additional challenge compared with 2D images. I particularly enjoyed tackling the challenge of building a fast experimentation pipeline when the training data is `1_000_000_000_000 bytes` of data! One trillion bytes! This is the largest dataset I have ever worked with.\n\n# RSNA STR Pulmonary Embolism Detection\nIn the figure below, each row is an exam (i.e. study, i.e. single patient). The row of images are CT scan \"slices\" from the 3D image of a patient's chest. In this competition, we need to predict 9 targets for each patient (each row) (like is pe on left side? on right side? etc) and we need to classify every image (is pe present?). If below were all the data, then we would need to predict `3 rows * 9 targets = 27 exam targets` and `15 images * 1 target = 15 image targets`. In total we would need to predict 42 targets.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F7d5ce77d471fbf8ef7e746731eab734c%2Fstudy.png?generation=1603731145650332&alt=media)\n\n# Understanding the Metric\nOn the description page, the metric seems very confusing. However it is just a weighted average of 10 log losses (the 9 types of exam predictions and the 1 type of image prediction). (I explain the metric [here][2]). After computing the 10 weights, we find that 50% of our LB score is from the 9 exam predictions and 50% of our LB score is from the image predictions. Furthermore, the log loss for the image predictions is itself a weighted log loss where an image that is part of an exam without pulmonary embolism has weight zero (very important observation!) \n\n* Improving image level predictions is equally important as improving exam predictions.\n* There is no penalty for false positives, so we can train our image prediction models with only the 30% of the data from positive exams!\n\n# Stage 1 - Model One - Image Level Predictions\n# (CNN EfficientNet B4)\n\n    inp = tf.keras.Input(shape=(320, 320, 1)) # INPUT IS UINT8\n    x = tf.keras.layers.Concatenate()([inp/255., inp/255., inp/255.])\n    base_model = efn.EfficientNetB4(weights='imagenet', include_top=False) \n\n    x = base_model(x)\n    x = tf.keras.layers.GlobalAveragePooling2D()(x)    \n    x = tf.keras.layers.Dense(1, activation='sigmoid')(x)\n    \n    model = tf.keras.Model(inputs=inp, outputs=x)\n    opt = tf.keras.optimizers.Adam(lr=0.000005)\n    model.compile(loss='binary_crossentropy', optimizer = opt)\n\n    model.fit(X, y, sample_weight = X.groupby('StudyInstanceUID') \n        .pe_present_on_image.transform('mean') * 5.6222 )\n\nI built two models. Model one predicts image level predictions (i.e. `pe_present_on_image`) and model two predicts patient level predictions (i.e. exams i.e. studies, like `leftsided_pe` etc)\n* EfficientNet B4 pretrained on `imagenet`\n* **Only Mediastinal window, (ie. level=40, width=400)** i.e. 1 channel `uint8`\n* Random crops of 320x320 from 512x512\n* Rotation (+-8 deg) Scale (+-0.16) augmentation\n* Coarse Dropout (16 holes sized 50x50)\n* Mixup (swap slices of similar Z position with other exams)\n* Adam optimizer with constant `LR = 5e-6`\n* Training sample weight equal to pe proportion in exam\n* **Only train on 30% of train data with pe present in exam**\n* 40 minute epochs using 4x V100 GPU\n* Train 15 epochs with batch size 128\n\nBelow illustrates my augmentations. For display purposes, we illustrate Mixup with a large yellow, green, or blue square so you can see it better. (During runtime, it was an actual second image).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F5515206e0fd5e5a7379f6272e0e56dc3%2Fmixup.png?generation=1603733206613194&alt=media)  \n\n# Stage 1 - Model Two - Patient Level Predictions\n# (CNN EfficientNet B4)\nEach patient has an average of 200 images. Among those 200, if PE is present, it is usually on the middle slices. Therefore I only train my patient level model with slices `0.35 < z < 0.65`. Then to predict the 9 targets for each patient, I only infer `0.35 < z < 0.65` and then take the 9 average predictions.\n\n* Most details same as model one\n* **Only Mediastinal window, (ie. level=40, width=400)** i.e. 1 channel `uint8`\n* Output layer of 9 sigmoid units\n* **Only train on 30% of train data with Z Position between `0.35 < z < 0.65`**\n* Loss `weighted_log_loss`\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F5ed4f1ae6291e5edc4f422b479214f80%2Fslices.png?generation=1603734372678374&alt=media)\n\n# Experimentation Pipeline\nHow do we discover the details above? All the settings above were discovered by performing dozens of experiments on **smaller images and smaller backbones**. For example, use 128x128 (with 80x80 crops) EfficientNetB0 and/or 256x256 (with 160x160 crops) EfficientNetB2. Using these smaller models, we can test out ideas on a single GPU in minutes! Also note that we only use 1 channel images of `uint8`. This is 33% less data than converting images to 3 channels of 3 different CT window schemes.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F60349090637a446427fc8f7c8775a70f%2Fexp2.png?generation=1603744581352698&alt=media)\n\nRemember we are only training with 30% original data. Then using crops makes it 12% of data. Then using 256x256 reduces this to 3% of data. And using 128x128 reduces this to 0.75% of data! Even Kaggle notebooks P100 GPU can train quickly on 80x80 crops from 128x128 and 30% train data.\n\nOnce you find a configuration that works well, then run 512x512 with EfficientNetB4 overnight. Using only 2D predictions for image and 2D predictions for patient, **the above two models obtain LB 0.215 and CV 0.235.** \n\nWe will now increase our CV LB by building stage 2 models that use stage 1 predictions as input\n  \n# Stage 2 - Model One - Image Level Predictions\n# (Random Forest)\n\n    FEATURES = ['oof']\n    for k in NEIGHBORS:\n        tmp = train.sort_values('PosZ').groupby('StudyInstanceUID')[['oof']]\n        train['b%i'%k] = tmp.shift(k)\n        train['a%i'%k] = tmp.shift(-k)\n        FEATURES += ['a%i'%k, 'b%i'%k]\n    train.fillna(-1,inplace=True)\n\n    model = RandomForestClassifier(max_depth=9, n_estimators=100, \n                               n_jobs=20, min_samples_leaf=50)\n    model.fit(train.loc[idxT,FEATURES],train.loc[idxT,'pe_present_on_image'],\n                    sample_weight = 5.6222 * valid.loc[idxT,'weight'])\n\nAll images are slices from 3D images. So adjacent images (within the same exam) contain helpful information. Each plot below displays all 200 or so image level predictions from 1 study. The x axis is z position and the y axis is the prediction value (0 to 1). The blue line is the ground truth, the orange line is the prediction from the model described above. The black line is the random forest Stage 2 model.\n\nFor each image level prediction, a random forest model takes as input the prediction and neighbor predictions [1,2,3,4,5,6,7,8,9,10,15,20,25,30,35,40,45,50,60,70,80,90,100,150,200,250] on either side. Then the random forest model predicts a new image level prediction display in black below. Notice when the original prediction is close to 1, then the random forest pushes it up to 1 and when the original prediction is close to 0, then the random forest pushes it down to 0.\n\n**This stage 2 image level model increased LB to 0.204 from 0.215 and CV to 0.224 from 0.235 (gain = 0.011)**\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Fcccdbe82e765c3e7d762031b358dd290%2Fi-smooth.png?generation=1603736540573190&alt=media)\n\n# Stage 2 - Model Two - Patient Level Predictions\n# (GRU + 1D-CNN)\n\n    inp = L.Input(shape=(64, 1792))\n    x = L.Bidirectional(L.GRU(48, return_sequences=True, \n                        kernel_initializer='orthogonal'))(inp)\n    x = L.Bidirectional(L.GRU(48, return_sequences=False, \n                        kernel_initializer='orthogonal'))(x)\n    x = L.Dense(9, activation='sigmoid')(x)\n        \n    model = tf.keras.Model(inputs=inp, outputs=x)\n    opt = tf.keras.optimizers.Adam(lr=0.00005)\n    model.compile(loss=weighted_log_loss, optimizer = opt)\n\nSimilarly we can use adjacent slice information to improve our patient level predictions. Most patients have between 160 and 310 images per study.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Fa4a52476743ad1e1ada257c586c607fb%2Fhist2.png?generation=1603741964817694&alt=media)\n\nBelow are plots of the top view (Z position is vertical axis) of the 3D image (not the ordinary slice view of the 3D image). We notice that most of the crucial information is between 25% and 75% in top view. Therefore we extracted 64 images equally spaces between 25% and 75% Z position. Then we took those 64 images and extracted the GAP embeddings from both our stage 1 model one and stage 1 model two. We trained a stage 2 GRU model and a stage 2 1D-CNN model to predict exam level predictions from these 64 GAP embeddings.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Fe46e41fab360d847c3fadb5d6dff9002%2Ftop.png?generation=1603741873425399&alt=media)\n\n**This stage 2 exam level model increased LB to 0.183 from 0.204 and CV to 0.203 from 0.224 (gain = 0.021)**\n\n# Other Misc Ideas\nWhen using global average pooling 2D in your CNN, location information is lost. Therefore I tried giving my models location information in various ways to help predict targets related to location (i.e. `leftsided_pe` etc) and related to size (i.e. `rv_lv_ratio_gte_1` etc). Unfortunately, none of my ideas increased CV or LB. My favorite is below.\n\n## Locating PE with Class Activation Maps CAM\n\nI extracted class activation maps from my stage 1 models and feed the location information into stage 2 exam models. (CAMs explained [here][1]). In the below figure, the ground truth is in the title. The green circles are the CAM of my EfficientNetB4 model. Note that CT scans are flipped so the left side of the image is \"right\" and the right side of the image is \"left. You can see that the CAM does a good job of locating the PE.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F291dc983514a322d5ce4e56b22a9a681%2Fcam.png?generation=1603755601607192&alt=media)\n\n# Thank you\n\n[1]: https://www.kaggle.com/cdeotte/unsupervised-masks-cv-0-60\n[2]: https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/193598",
    "1062454": "Thank you so much for sharing @cdeotte !! I love your solution discussions! There is so much to learn from them! 💛\n\nI was very sad that I was not able to submit a solution because I completely underestimated the performance challenge of this competition. I was so happy to see now that you also used only 30 % of the data (yes! :-)) and windowed images. Why did you use level 40 and width 400? We had the idea to use the \"blood windows\" ranging from 5-60 and 60-160 HU to detect chronic and acute PE and segmented lungs. Why have you used a more broad window instead? How did you use only 1 channel? Aren't 3 channels expected by your model?!\n\nIt's great to see how it helped to discover where PE is usually located. That's so elegant! We thought about masking the exam targets for images showing no PE at all but it already felt like \"showing too much\" for solving the exam-level. Now I understand that my problem was that I tried to solve this all with one model instead of splitting the tasks on several different models and stages. I think I was too focused on finding a single solution because of the performance stuff and my little time available. I was also much too cautious to use smaller image sizes and backbones even though I had in mind \"start simple\". \n\nThank you again for sharing your ideas and for making it possible to learn and improve!",
    "1061360": "Great work @cdeotte - Impressive solo finish! Cool to see you implemented random forest in your stage 2 model! I wanted to try a tabular model in stacking but we didn't have time. May I ask why you used random forest and not a GBM?",
    "1082264": "Thank you very much for your insights!! I always join your exact explanations and tutorials!\n\nCould you explain how you trained your Stage 1 Model with the additional GAP - layer?\n\nI always get errors, when training it with two outputs (x, GAP) while having only one target. I tried to Compile with:\n\nmodel.compile(\n        optimizer='adam',\n        loss = [tf.keras.losses.BinaryCrossentropy(from_logits=True), None],\n        metrics=['accuracy'])   \n but when training (model.fit(get_training_dataset(), ...) I get th error: 'Dimensions must be equal, but are 7 and 1791 for '{{node Equal_1}} '",
    "1082257": "Thank you very much for these insights!!! \n\nCould you please explain, how you trained stage 1 Model mit additional output of the GAP-layer?\nCompiling the \n\nmodel = Model (model input, [x, GAP])            I used\nmodel.compile(\n        optimizer='adam',\n        loss = [tf.keras.losses.BinaryCrossentropy(from_logits=True), None],\n        metrics=['accuracy'])\n\nwith no problems, but with\n\nmodel.fit(get_training_dataset(), ...)   I always get 'Dimensions must be equal, but are 7 and 1792 for '{{node Equal_1}} = Equal[T=DT_FLOAT, incompatible_shape_error=true](Cast_18, Cast_20)' with input shapes: [?,7], [?,1792].'",
    "1073690": "Thank you @cdeotte and congrats on the finish! I have been following your work since ISIC Melonama and it is always a pleasure to read :) \n\nCould I please ask how you trained `Stage 1 - Model Two - Patient Level Predictions`? \n\nI am a bit confused by: \n> then take the 9 average predictions.\n\nWhere are these 9 predictions coming from? And what are the labels used to train the model? Sorry if it's a silly question.",
    "1065308": "Great job!",
    "1065047": "Wow, learnt a lot of new stuff, thanks",
    "1063967": "Great work @cdeotte.. May be a naive question. the first model the pe_present_on_image has been multiplied by 5.6222. Any reason?. ",
    "1063662": "This was a great read.",
    "1063536": "This is an advanced topic, thanks for sharing your insight",
    "1062940": "Your explanations are always very detailed and contain great visuals to even better illustrate your work. Great stuff and thanks for sharing!",
    "1062615": "Good strategy.  I tried the efficient B > 3 give me \"CSV not found\" or timeout. My bad bc I can't control it. Learn more from you.",
    "1062614": "Your write-up is a master class in clean code and experimentation. Thanks for sharing. ",
    "1062534": "Thanks for sharing your clear and concise solution, @cdeotte! \n\nIn the part of Stage 2 model two, you noted that: \n\n> Then we took those 64 images and extracted the GAP embeddings from both our stage 1 model one and stage 1 model two.\n\nThis might be a beginner's question, but what exactly is the GAP embedding? ",
    "1062448": "Good job on your place. Very well deserved!",
    "1062426": "Very good solution and approach ! Best work ",
    "1061695": "learned a lot from seeing your solution @cdeotte. now i know where i went wrong in my approaches. next time I am going to spend some more time on experiments with lesser data and smaller model.",
    "1061434": "Congrats on results @cdeotte and thanks for sharing the detailed solution !!",
    "1061361": "Congrats Chris! ",
    "1061356": "Congrats on results and thanks for the writeup details solution @cdeotte",
    "1061669": "Congratz on the nice finish Chris !\n\nI'm always impressed by your write-ups, they really make your solution stand out, and are pleasant to read.",
    "1061585": "Hi @cdeotte . Thank you for the amazing write up. I am impressed by your simple and easy flowing explanation in all the competitions :). Keep Shining. ",
    "1061515": "Wow ! Its amazing how you get the time to create and run so many experiments AND document it too !!\nThank you and congratulations ",
    "1061424": "Congratulations Dr. @cdeotte. You are truly an inspiration for newbies.. Thanks a lot for sharing the amazing concepts that fueled your solution... ",
    "1061380": "Thank you for sharing. Great work @cdeotte . Very informative and I need more time to read.",
    "1061365": "You show the power of the NVIDIA device, I decided to buy a Tesla V100 after this competition.",
    "1061331": "There is alot to learn from @cdeotte Thanks for sharning",
    "1062803": "",
    "1065142": "Thanks for sharing !",
    "1064056": "Thanks for sharing.",
    "1063675": "This was of great help . Thank You !",
    "1063009": "Thanks for Sharing!!!",
    "1064062": "Thanks for sharing!!"
  }
}