{
  "id": 169637,
  "title": "12th Place Solution - Overview with code files",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/169637",
  "author_name": "Manuel Campos",
  "post_date": "2020-07-24T15:02:46.501000",
  "votes": 25,
  "comment_count": 11,
  "views": 0,
  "content": "<p>First of all, I would like to thank all the people and organizations that have made this Competition possible. In capital letters, THANK YOU to all the TEAMS that with their dedication and effort I hope contribute to improve the diagnosis of prostate cancer and thereby improve people's lives. Indeed, my most sincere congratulations to the WINNERS.</p>\n<p>I am very happy as you can imagine. In a few lines I share with you a quick overview of my time in this Challenge.</p>\n<h3>Kaggle Learning</h3>\n<p>I want to comment here what is usually included in the acknowledgments part but I reserve this special section to highlight the work of those competitors who have made my final solution better, 1) because their ability was not present in my initial knowledge or 2) because their performance improves together with the experience of mine. I mean, in no order of priority,</p>\n<ul>\n<li><strong>(Salman)</strong> <a href=\"https://www.kaggle.com/micheomaano\" target=\"_blank\">@micheomaano</a>:</li>\n</ul>\n<ol>\n<li><a href=\"https://www.kaggle.com/micheomaano/tf-record-256-256-48\" target=\"_blank\">Dataset tf-record-256-56-48</a></li>\n<li><a href=\"https://www.kaggle.com/micheomaano/tpu-training-tensorflow-iafoos-method-42x256x256x3\" target=\"_blank\">TPU Training Tensorflow Iafoos Method 42x256x256x3</a></li>\n<li><a href=\"https://www.kaggle.com/micheomaano/pandas-42x256x256x3-inference\" target=\"_blank\">Pandas 42x256x256x3 Inference</a></li>\n</ol>\n<ul>\n<li><strong>(Qishen Ha)</strong> <a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a>:</li>\n</ul>\n<ol>\n<li><a href=\"https://www.kaggle.com/haqishen/train-efficientnet-b0-w-36-tiles-256-lb0-87\" target=\"_blank\">Train EfficientNet-B0 w/ 36 tiles_256 [LB0.87]</a></li>\n<li><a href=\"https://www.kaggle.com/haqishen/panda-inference-w-36-tiles-256\" target=\"_blank\">PANDA Inference w/ 36 tiles_256</a></li>\n</ol>\n<ul>\n<li><strong>(RAHUL SINGH INDA)</strong> <a href=\"https://www.kaggle.com/rsinda\" target=\"_blank\">@rsinda</a>:</li>\n</ul>\n<ol>\n<li><a href=\"https://www.kaggle.com/rsinda/panda-inference-efficientnet-b1\" target=\"_blank\">Panda Inference EfficientNet-b1</a></li>\n</ol>\n<ul>\n<li><strong>(Iafoss)</strong> <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a>: The Best Accelerator in the Competition, ahead of TPUs.</li>\n</ul>\n<h3>Submission Notebook</h3>\n<p>I have shared an original copy of my inference kernel without additional cleaning as well as a dataset that includes the necessary weights of each of the models that are used in obtaining the final submission,</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/coreacasa/12th-place-solution-quick-save-inference\" target=\"_blank\">Quick Save Inference</a></li>\n<li><a href=\"https://www.kaggle.com/coreacasa/pandaenetb042x256x256x3\" target=\"_blank\">Dataset Model Weights for Inference</a></li>\n</ul>\n<h3>[TPU] Kaggle/Google(Colaboratory)</h3>\n<p>For all my trainings I used the free TPU resources offered by Kaggle / Google (Colaboratory). Thank you very much.</p>\n<h3>Training One: My only approach to validation</h3>\n<p>Very closed to Salman's training kernel I just re-ran its code to complete cross validation. I ran a fold up to 60 epochs to see the evolution of the loss and the rest down to 40 epochs.</p>\n<p>Individually the behavior of the folds was more or less similar in final loss values (mse) and in the number of times in which it stopped improving. The issue is that its merge did not improve the individual performance of some of them over LB and their performance was also uneven when they were introduced into an external ensemble.</p>\n<p>The noise of the labels is a probable cause as already discussed in the discussions or perhaps the sensitivity of the qwk metric to even small variations in mse when its jump to LB.</p>\n<ul>\n<li><strong><a href=\"https://www.kaggle.com/coreacasa/code-base-training-one\" target=\"_blank\">code-base-training-one</a></strong> file, training topics:</li>\n</ul>\n<p><code>Size Image</code> 256<br>\n<code>Size Tiles</code> 256<br>\n<code>Tiles</code> 42<br>\n<code>Augmentation</code> horizonal p=0.5 and vertical p=0.5 flips<br>\n<code>Validation</code> StratifiedKFold 5 on isup grade classes<br>\n<code>Arch</code> EfficientNetB0<br>\n<code>Convolutional Base's Weight</code> Imagenet trainable<br>\n<code>On Top</code> GlobalAveragePooling2D, Dropout(0.5), Dense(1024)<br>\n<code>Output</code> Dense(1) regression objective<br>\n<code>Loss</code> mean_squared_error<br>\n<code>Optimizer</code> Adam<br>\n<code>Leaning Rate</code> 5e-04 init<br>\n<code>Reduce LR</code> decreasing 0.5 with patience 3 epochs<br>\n<code>Save</code> weights only with best validation loss epochs<br>\n<code>Batch Size</code> 64</p>\n<h3>Training Two: Art(Instinct) Validation</h3>\n<p>I never tried detecting noisy labels to remove them from training data. In general I am not in favor of losing any existing information, although in principle it could be harmful by elevating the non-regular part of a data generating process. I would rather transform data than remove it.</p>\n<p>I didn't try either any transformation so I thought about training the models with full dataset in order to prevent the possible existence of more noise in some folds than in others, which probably would be increasing the variability in the inference results.</p>\n<p>Art Validation appears here and it is when the art of the data scientist enters and it is his instinct that determines the goodness of fit and stability of performance in generalization against new observations. Yes, this is Alchemy.</p>\n<ul>\n<li><strong><a href=\"https://www.kaggle.com/coreacasa/code-base-training-two-enets\" target=\"_blank\">code-base-training-two-enets</a></strong> file, from which I trained 3 members of the EfficientNet family. Changes on training one training topics:</li>\n</ul>\n<p><code>Tiles</code> 48<br>\n<code>Validation</code> Art Validation on instinct<br>\n<code>Arch</code> EfficientNetB0, EfficientNetB1 and EfficientNetB2<br>\n<code>Convolutional Base's Weight</code> Noisy Student trainable<br>\n<code>Output</code> Dense(5,activation='sigmoid) ordinal regression objective<br>\n<code>Loss</code> sigmoid_cross_entropy_with_logits<br>\n<code>Leaning Rate</code> custom with 5up, 3sustain, 0.8decay<br>\n<code>Limits LR</code> 1e-05min, 4e-04max<br>\n<code>Save</code> weights only with best loss epochs<br>\n<code>Batch Size</code> 32<br>\n<code>Epochs</code> 60</p>\n<ul>\n<li><strong><a href=\"https://www.kaggle.com/coreacasa/code-base-training-two-densenet\" target=\"_blank\">code-base-training-two-densenet</a></strong> file, from which I trained 1 member of the DenseNet family. Changes on training topics of the previous net family:</li>\n</ul>\n<p><code>Arch</code> Densenet121<br>\n<code>Convolutional Base's Weight</code> Imagenet trainable<br>\n<code>Epochs</code> 40</p>\n<h3>Inference: Diversity of Archs, nTiles and TTAs</h3>\n<p>Of the 2 training processes shown above, the following models were available,</p>\n<ol>\n<li>EfficientNetB0 (5 skf), 42x256x256x3</li>\n<li>EfficientNetB0 (1), 48x256x256x3 </li>\n<li>EfficientNetB1 (1), 48x256x256x3 </li>\n<li>EfficientNetB2 (1), 48x256x256x3 </li>\n<li>DenseNet121 (1), 48x256x256x3 </li>\n</ol>\n<p>Having re-run the Salman kernel, from the public notebooks referenced at the beginning I had,</p>\n<ol>\n<li>EfficientNetB0 (1 skf), 36x256x256x3 (Qishen Ha) </li>\n<li>EfficientNetB1 (1 skf), 36x256x256x3 (RAHUL SINGH INDA)</li>\n</ol>\n<ul>\n<li><p><strong>Test Time Augmentation</strong><br>\n<code>Type A: 5xTTA deterministic</code> <br>\n1xoriginal, 1xTranspose, 1xVerticalFlip, 1xHorizontalFlip, 1xTranspose-&gt;VerticalFlip-&gt;HorizontalFlip<br>\n<code>Type B: 4xTTA pseudo deterministic</code> <br>\n1xoriginal, 1xVerticalFlip, 2xHorizontalFlip(p=0.5)-&gt;VerticalFlip(p=0.5)<br>\n<code>Type C: 2xTTA random </code> <br>\n2xHorizontalFlip(p=0.5)-&gt;VerticalFlip(p=0)</p></li>\n<li><p><strong>White Padding Tile Extraction (Qishen modes)</strong><br>\n1x add zero pad and 1x add 256 pad, that is, 2 different extractions for ALL the images.</p></li>\n</ul>\n<h3>Model Selection and Final Ensemble</h3>\n<pre><code>(3/10)*Public-Quishen [TTA Type A]  \n(3/10)*Public-RAHUL SINGH INDA [TTA Type A] \n\n(1/30)*EfficientNetB0-Fold0-Training One [TTA Type C] \n(1/30)*EfficientNetB0-Fold2-Training One [TTA Type C] \n(1/30)*EfficientNetB0-Fold4-Training One [TTA Type C] \n\n(1/15)*EfficientNetB0-Training Two [TTA Type C] \n(1/15)*EfficientNetB1-Training Two [TTA Type C] \n(1/15)*EfficientNetB2-Training Two [TTA Type C] \n\n(1/10)*DenseNet121-Training Two [TTA Type B] \n</code></pre>\n<p>The random component of the TTAs was not seed (I'll be lucky) and the reproducibility of the results may vary with it. I have just re-run my inference kernel and the results are Private Score 0.92983 (0.92960 original) and Public Score 0.89443 (089352 original).</p>\n<p>With this models structure I was only able to test the last day of the competition. For example, this other ensemble got Private Score 0.93047 and Public Score 0.88889, not including random component in TTA.</p>\n<pre><code>(3.5/10)*Public-Quishen [TTA Type A]  \n(3.5/10)*Public-RAHUL SINGH INDA [TTA Type A] \n\n(1/15)*EfficientNetB0-Training Two [TTA Type A] \n(1/15)*EfficientNetB1-Training Two [TTA Type A] \n(1/15)*EfficientNetB2-Training Two [TTA Type A] \n\n(1/10)*DenseNet121-Training Two [TTA Type A] \n</code></pre>\n<p>One more, my last submission and that finished tight after the deadline got Private Score 0.93052 and Public Score 0.89110,</p>\n<pre><code>(3.5/10)*Public-Quishen [TTA Type A]  \n(3.5/10)*Public-RAHUL SINGH INDA [TTA Type A] \n\n(1/30)*EfficientNetB0-Fold0-Training One [TTA Type C] \n(1/30)*EfficientNetB0-Fold2-Training One [TTA Type C] \n(1/30)*EfficientNetB0-Fold4-Training One [TTA Type C] \n\n(1/15)*EfficientNetB0-Training Two [TTA Type C] \n(1/15)*EfficientNetB1-Training Two [TTA Type C] \n(1/15)*EfficientNetB2-Training Two [TTA Type C] \n</code></pre>\n<h3>That is all, Thanks a lot!</h3>\n<p>By the way, I still tremble with fear<br>\nUpdate: No longer!</p>",
  "messages": [
    {
      "id": 943758,
      "postDate": "2020-07-24T15:02:46.500Z",
      "content": "<p>First of all, I would like to thank all the people and organizations that have made this Competition possible. In capital letters, THANK YOU to all the TEAMS that with their dedication and effort I hope contribute to improve the diagnosis of prostate cancer and thereby improve people's lives. Indeed, my most sincere congratulations to the WINNERS.</p>\n<p>I am very happy as you can imagine. In a few lines I share with you a quick overview of my time in this Challenge.</p>\n<h3>Kaggle Learning</h3>\n<p>I want to comment here what is usually included in the acknowledgments part but I reserve this special section to highlight the work of those competitors who have made my final solution better, 1) because their ability was not present in my initial knowledge or 2) because their performance improves together with the experience of mine. I mean, in no order of priority,</p>\n<ul>\n<li><strong>(Salman)</strong> <a href=\"https://www.kaggle.com/micheomaano\" target=\"_blank\">@micheomaano</a>:</li>\n</ul>\n<ol>\n<li><a href=\"https://www.kaggle.com/micheomaano/tf-record-256-256-48\" target=\"_blank\">Dataset tf-record-256-56-48</a></li>\n<li><a href=\"https://www.kaggle.com/micheomaano/tpu-training-tensorflow-iafoos-method-42x256x256x3\" target=\"_blank\">TPU Training Tensorflow Iafoos Method 42x256x256x3</a></li>\n<li><a href=\"https://www.kaggle.com/micheomaano/pandas-42x256x256x3-inference\" target=\"_blank\">Pandas 42x256x256x3 Inference</a></li>\n</ol>\n<ul>\n<li><strong>(Qishen Ha)</strong> <a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a>:</li>\n</ul>\n<ol>\n<li><a href=\"https://www.kaggle.com/haqishen/train-efficientnet-b0-w-36-tiles-256-lb0-87\" target=\"_blank\">Train EfficientNet-B0 w/ 36 tiles_256 [LB0.87]</a></li>\n<li><a href=\"https://www.kaggle.com/haqishen/panda-inference-w-36-tiles-256\" target=\"_blank\">PANDA Inference w/ 36 tiles_256</a></li>\n</ol>\n<ul>\n<li><strong>(RAHUL SINGH INDA)</strong> <a href=\"https://www.kaggle.com/rsinda\" target=\"_blank\">@rsinda</a>:</li>\n</ul>\n<ol>\n<li><a href=\"https://www.kaggle.com/rsinda/panda-inference-efficientnet-b1\" target=\"_blank\">Panda Inference EfficientNet-b1</a></li>\n</ol>\n<ul>\n<li><strong>(Iafoss)</strong> <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a>: The Best Accelerator in the Competition, ahead of TPUs.</li>\n</ul>\n<h3>Submission Notebook</h3>\n<p>I have shared an original copy of my inference kernel without additional cleaning as well as a dataset that includes the necessary weights of each of the models that are used in obtaining the final submission,</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/coreacasa/12th-place-solution-quick-save-inference\" target=\"_blank\">Quick Save Inference</a></li>\n<li><a href=\"https://www.kaggle.com/coreacasa/pandaenetb042x256x256x3\" target=\"_blank\">Dataset Model Weights for Inference</a></li>\n</ul>\n<h3>[TPU] Kaggle/Google(Colaboratory)</h3>\n<p>For all my trainings I used the free TPU resources offered by Kaggle / Google (Colaboratory). Thank you very much.</p>\n<h3>Training One: My only approach to validation</h3>\n<p>Very closed to Salman's training kernel I just re-ran its code to complete cross validation. I ran a fold up to 60 epochs to see the evolution of the loss and the rest down to 40 epochs.</p>\n<p>Individually the behavior of the folds was more or less similar in final loss values (mse) and in the number of times in which it stopped improving. The issue is that its merge did not improve the individual performance of some of them over LB and their performance was also uneven when they were introduced into an external ensemble.</p>\n<p>The noise of the labels is a probable cause as already discussed in the discussions or perhaps the sensitivity of the qwk metric to even small variations in mse when its jump to LB.</p>\n<ul>\n<li><strong><a href=\"https://www.kaggle.com/coreacasa/code-base-training-one\" target=\"_blank\">code-base-training-one</a></strong> file, training topics:</li>\n</ul>\n<p><code>Size Image</code> 256<br>\n<code>Size Tiles</code> 256<br>\n<code>Tiles</code> 42<br>\n<code>Augmentation</code> horizonal p=0.5 and vertical p=0.5 flips<br>\n<code>Validation</code> StratifiedKFold 5 on isup grade classes<br>\n<code>Arch</code> EfficientNetB0<br>\n<code>Convolutional Base's Weight</code> Imagenet trainable<br>\n<code>On Top</code> GlobalAveragePooling2D, Dropout(0.5), Dense(1024)<br>\n<code>Output</code> Dense(1) regression objective<br>\n<code>Loss</code> mean_squared_error<br>\n<code>Optimizer</code> Adam<br>\n<code>Leaning Rate</code> 5e-04 init<br>\n<code>Reduce LR</code> decreasing 0.5 with patience 3 epochs<br>\n<code>Save</code> weights only with best validation loss epochs<br>\n<code>Batch Size</code> 64</p>\n<h3>Training Two: Art(Instinct) Validation</h3>\n<p>I never tried detecting noisy labels to remove them from training data. In general I am not in favor of losing any existing information, although in principle it could be harmful by elevating the non-regular part of a data generating process. I would rather transform data than remove it.</p>\n<p>I didn't try either any transformation so I thought about training the models with full dataset in order to prevent the possible existence of more noise in some folds than in others, which probably would be increasing the variability in the inference results.</p>\n<p>Art Validation appears here and it is when the art of the data scientist enters and it is his instinct that determines the goodness of fit and stability of performance in generalization against new observations. Yes, this is Alchemy.</p>\n<ul>\n<li><strong><a href=\"https://www.kaggle.com/coreacasa/code-base-training-two-enets\" target=\"_blank\">code-base-training-two-enets</a></strong> file, from which I trained 3 members of the EfficientNet family. Changes on training one training topics:</li>\n</ul>\n<p><code>Tiles</code> 48<br>\n<code>Validation</code> Art Validation on instinct<br>\n<code>Arch</code> EfficientNetB0, EfficientNetB1 and EfficientNetB2<br>\n<code>Convolutional Base's Weight</code> Noisy Student trainable<br>\n<code>Output</code> Dense(5,activation='sigmoid) ordinal regression objective<br>\n<code>Loss</code> sigmoid_cross_entropy_with_logits<br>\n<code>Leaning Rate</code> custom with 5up, 3sustain, 0.8decay<br>\n<code>Limits LR</code> 1e-05min, 4e-04max<br>\n<code>Save</code> weights only with best loss epochs<br>\n<code>Batch Size</code> 32<br>\n<code>Epochs</code> 60</p>\n<ul>\n<li><strong><a href=\"https://www.kaggle.com/coreacasa/code-base-training-two-densenet\" target=\"_blank\">code-base-training-two-densenet</a></strong> file, from which I trained 1 member of the DenseNet family. Changes on training topics of the previous net family:</li>\n</ul>\n<p><code>Arch</code> Densenet121<br>\n<code>Convolutional Base's Weight</code> Imagenet trainable<br>\n<code>Epochs</code> 40</p>\n<h3>Inference: Diversity of Archs, nTiles and TTAs</h3>\n<p>Of the 2 training processes shown above, the following models were available,</p>\n<ol>\n<li>EfficientNetB0 (5 skf), 42x256x256x3</li>\n<li>EfficientNetB0 (1), 48x256x256x3 </li>\n<li>EfficientNetB1 (1), 48x256x256x3 </li>\n<li>EfficientNetB2 (1), 48x256x256x3 </li>\n<li>DenseNet121 (1), 48x256x256x3 </li>\n</ol>\n<p>Having re-run the Salman kernel, from the public notebooks referenced at the beginning I had,</p>\n<ol>\n<li>EfficientNetB0 (1 skf), 36x256x256x3 (Qishen Ha) </li>\n<li>EfficientNetB1 (1 skf), 36x256x256x3 (RAHUL SINGH INDA)</li>\n</ol>\n<ul>\n<li><p><strong>Test Time Augmentation</strong><br>\n<code>Type A: 5xTTA deterministic</code> <br>\n1xoriginal, 1xTranspose, 1xVerticalFlip, 1xHorizontalFlip, 1xTranspose-&gt;VerticalFlip-&gt;HorizontalFlip<br>\n<code>Type B: 4xTTA pseudo deterministic</code> <br>\n1xoriginal, 1xVerticalFlip, 2xHorizontalFlip(p=0.5)-&gt;VerticalFlip(p=0.5)<br>\n<code>Type C: 2xTTA random </code> <br>\n2xHorizontalFlip(p=0.5)-&gt;VerticalFlip(p=0)</p></li>\n<li><p><strong>White Padding Tile Extraction (Qishen modes)</strong><br>\n1x add zero pad and 1x add 256 pad, that is, 2 different extractions for ALL the images.</p></li>\n</ul>\n<h3>Model Selection and Final Ensemble</h3>\n<pre><code>(3/10)*Public-Quishen [TTA Type A]  \n(3/10)*Public-RAHUL SINGH INDA [TTA Type A] \n\n(1/30)*EfficientNetB0-Fold0-Training One [TTA Type C] \n(1/30)*EfficientNetB0-Fold2-Training One [TTA Type C] \n(1/30)*EfficientNetB0-Fold4-Training One [TTA Type C] \n\n(1/15)*EfficientNetB0-Training Two [TTA Type C] \n(1/15)*EfficientNetB1-Training Two [TTA Type C] \n(1/15)*EfficientNetB2-Training Two [TTA Type C] \n\n(1/10)*DenseNet121-Training Two [TTA Type B] \n</code></pre>\n<p>The random component of the TTAs was not seed (I'll be lucky) and the reproducibility of the results may vary with it. I have just re-run my inference kernel and the results are Private Score 0.92983 (0.92960 original) and Public Score 0.89443 (089352 original).</p>\n<p>With this models structure I was only able to test the last day of the competition. For example, this other ensemble got Private Score 0.93047 and Public Score 0.88889, not including random component in TTA.</p>\n<pre><code>(3.5/10)*Public-Quishen [TTA Type A]  \n(3.5/10)*Public-RAHUL SINGH INDA [TTA Type A] \n\n(1/15)*EfficientNetB0-Training Two [TTA Type A] \n(1/15)*EfficientNetB1-Training Two [TTA Type A] \n(1/15)*EfficientNetB2-Training Two [TTA Type A] \n\n(1/10)*DenseNet121-Training Two [TTA Type A] \n</code></pre>\n<p>One more, my last submission and that finished tight after the deadline got Private Score 0.93052 and Public Score 0.89110,</p>\n<pre><code>(3.5/10)*Public-Quishen [TTA Type A]  \n(3.5/10)*Public-RAHUL SINGH INDA [TTA Type A] \n\n(1/30)*EfficientNetB0-Fold0-Training One [TTA Type C] \n(1/30)*EfficientNetB0-Fold2-Training One [TTA Type C] \n(1/30)*EfficientNetB0-Fold4-Training One [TTA Type C] \n\n(1/15)*EfficientNetB0-Training Two [TTA Type C] \n(1/15)*EfficientNetB1-Training Two [TTA Type C] \n(1/15)*EfficientNetB2-Training Two [TTA Type C] \n</code></pre>\n<h3>That is all, Thanks a lot!</h3>\n<p>By the way, I still tremble with fear<br>\nUpdate: No longer!</p>",
      "rawMarkdown": "First of all, I would like to thank all the people and organizations that have made this Competition possible. In capital letters, THANK YOU to all the TEAMS that with their dedication and effort I hope contribute to improve the diagnosis of prostate cancer and thereby improve people's lives. Indeed, my most sincere congratulations to the WINNERS.\n\nI am very happy as you can imagine. In a few lines I share with you a quick overview of my time in this Challenge.\n\n### Kaggle Learning\nI want to comment here what is usually included in the acknowledgments part but I reserve this special section to highlight the work of those competitors who have made my final solution better, 1) because their ability was not present in my initial knowledge or 2) because their performance improves together with the experience of mine. I mean, in no order of priority,\n\n- **(Salman)** @micheomaano:\n1.  [Dataset tf-record-256-56-48](https://www.kaggle.com/micheomaano/tf-record-256-256-48)\n2. [TPU Training Tensorflow Iafoos Method 42x256x256x3](https://www.kaggle.com/micheomaano/tpu-training-tensorflow-iafoos-method-42x256x256x3)\n3. [Pandas 42x256x256x3 Inference](https://www.kaggle.com/micheomaano/pandas-42x256x256x3-inference)\n\n- **(Qishen Ha)** @haqishen:\n1.  [Train EfficientNet-B0 w/ 36 tiles_256 [LB0.87]](https://www.kaggle.com/haqishen/train-efficientnet-b0-w-36-tiles-256-lb0-87)\n2. [PANDA Inference w/ 36 tiles_256](https://www.kaggle.com/haqishen/panda-inference-w-36-tiles-256)\n\n- **(RAHUL SINGH INDA)** @rsinda:\n1.  [Panda Inference EfficientNet-b1](https://www.kaggle.com/rsinda/panda-inference-efficientnet-b1)\n\n- **(Iafoss)** @iafoss: The Best Accelerator in the Competition, ahead of TPUs.\n\n###Submission Notebook    \n   \nI have shared an original copy of my inference kernel without additional cleaning as well as a dataset that includes the necessary weights of each of the models that are used in obtaining the final submission,\n    \n- [Quick Save Inference](https://www.kaggle.com/coreacasa/12th-place-solution-quick-save-inference)\n- [Dataset Model Weights for Inference](https://www.kaggle.com/coreacasa/pandaenetb042x256x256x3)\n   \n### [TPU] Kaggle/Google(Colaboratory)\nFor all my trainings I used the free TPU resources offered by Kaggle / Google (Colaboratory). Thank you very much.\n\n### Training One: My only approach to validation\n\nVery closed to Salman's training kernel I just re-ran its code to complete cross validation. I ran a fold up to 60 epochs to see the evolution of the loss and the rest down to 40 epochs.\n\nIndividually the behavior of the folds was more or less similar in final loss values (mse) and in the number of times in which it stopped improving. The issue is that its merge did not improve the individual performance of some of them over LB and their performance was also uneven when they were introduced into an external ensemble.\n\nThe noise of the labels is a probable cause as already discussed in the discussions or perhaps the sensitivity of the qwk metric to even small variations in mse when its jump to LB.\n\n- **[code-base-training-one](https://www.kaggle.com/coreacasa/code-base-training-one)** file, training topics:\n\n<code>Size Image</code> 256\n<code>Size Tiles</code> 256\n<code>Tiles</code> 42\n<code>Augmentation</code> horizonal p=0.5 and vertical p=0.5 flips\n<code>Validation</code> StratifiedKFold 5 on isup grade classes\n<code>Arch</code> EfficientNetB0\n<code>Convolutional Base's Weight</code> Imagenet trainable\n<code>On Top</code> GlobalAveragePooling2D, Dropout(0.5), Dense(1024)\n<code>Output</code> Dense(1) regression objective\n<code>Loss</code> mean_squared_error\n<code>Optimizer</code> Adam\n<code>Leaning Rate</code> 5e-04 init\n<code>Reduce LR</code> decreasing 0.5 with patience 3 epochs\n<code>Save</code> weights only with best validation loss epochs\n<code>Batch Size</code> 64\n\n### Training Two: Art(Instinct) Validation\n\nI never tried detecting noisy labels to remove them from training data. In general I am not in favor of losing any existing information, although in principle it could be harmful by elevating the non-regular part of a data generating process. I would rather transform data than remove it.\n\nI didn't try either any transformation so I thought about training the models with full dataset in order to prevent the possible existence of more noise in some folds than in others, which probably would be increasing the variability in the inference results.\n\nArt Validation appears here and it is when the art of the data scientist enters and it is his instinct that determines the goodness of fit and stability of performance in generalization against new observations. Yes, this is Alchemy.\n\n- **[code-base-training-two-enets](https://www.kaggle.com/coreacasa/code-base-training-two-enets)** file, from which I trained 3 members of the EfficientNet family. Changes on training one training topics:\n\n<code>Tiles</code> 48\n<code>Validation</code> Art Validation on instinct\n<code>Arch</code> EfficientNetB0, EfficientNetB1 and EfficientNetB2\n<code>Convolutional Base's Weight</code> Noisy Student trainable\n<code>Output</code> Dense(5,activation='sigmoid) ordinal regression objective\n<code>Loss</code> sigmoid_cross_entropy_with_logits\n<code>Leaning Rate</code> custom with 5up, 3sustain, 0.8decay\n<code>Limits LR</code> 1e-05min, 4e-04max\n<code>Save</code> weights only with best loss epochs\n<code>Batch Size</code> 32\n<code>Epochs</code> 60\n\n- **[code-base-training-two-densenet](https://www.kaggle.com/coreacasa/code-base-training-two-densenet)** file, from which I trained 1 member of the DenseNet family. Changes on training topics of the previous net family:\n\n<code>Arch</code> Densenet121\n<code>Convolutional Base's Weight</code> Imagenet trainable\n<code>Epochs</code> 40\n\n### Inference: Diversity of Archs, nTiles and TTAs\n\nOf the 2 training processes shown above, the following models were available,\n1. EfficientNetB0 (5 skf), 42x256x256x3\n2. EfficientNetB0 (1), 48x256x256x3 \n3. EfficientNetB1 (1), 48x256x256x3 \n4. EfficientNetB2 (1), 48x256x256x3 \n5. DenseNet121 (1), 48x256x256x3 \n\nHaving re-run the Salman kernel, from the public notebooks referenced at the beginning I had,\n1. EfficientNetB0 (1 skf), 36x256x256x3 (Qishen Ha) \n2. EfficientNetB1 (1 skf), 36x256x256x3 (RAHUL SINGH INDA)\n    \n- **Test Time Augmentation**\n<code>Type A: 5xTTA deterministic</code> \n1xoriginal, 1xTranspose, 1xVerticalFlip, 1xHorizontalFlip, 1xTranspose-&gt;VerticalFlip-&gt;HorizontalFlip\n<code>Type B: 4xTTA pseudo deterministic</code> \n1xoriginal, 1xVerticalFlip, 2xHorizontalFlip(p=0.5)-&gt;VerticalFlip(p=0.5)\n<code>Type C: 2xTTA random </code> \n2xHorizontalFlip(p=0.5)-&gt;VerticalFlip(p=0)\n\n- **White Padding Tile Extraction (Qishen modes)**\n1x add zero pad and 1x add 256 pad, that is, 2 different extractions for ALL the images.\n\n### Model Selection and Final Ensemble\n\n<pre><code>(3/10)*Public-Quishen [TTA Type A]  \n(3/10)*Public-RAHUL SINGH INDA [TTA Type A] \n\n(1/30)*EfficientNetB0-Fold0-Training One [TTA Type C] \n(1/30)*EfficientNetB0-Fold2-Training One [TTA Type C] \n(1/30)*EfficientNetB0-Fold4-Training One [TTA Type C] \n\n(1/15)*EfficientNetB0-Training Two [TTA Type C] \n(1/15)*EfficientNetB1-Training Two [TTA Type C] \n(1/15)*EfficientNetB2-Training Two [TTA Type C] \n\n(1/10)*DenseNet121-Training Two [TTA Type B] \n</code></pre>\n\nThe random component of the TTAs was not seed (I'll be lucky) and the reproducibility of the results may vary with it. I have just re-run my inference kernel and the results are Private Score 0.92983 (0.92960 original) and Public Score 0.89443 (089352 original).\n\nWith this models structure I was only able to test the last day of the competition. For example, this other ensemble got Private Score 0.93047 and Public Score 0.88889, not including random component in TTA.\n\n<pre><code>(3.5/10)*Public-Quishen [TTA Type A]  \n(3.5/10)*Public-RAHUL SINGH INDA [TTA Type A] \n\n(1/15)*EfficientNetB0-Training Two [TTA Type A] \n(1/15)*EfficientNetB1-Training Two [TTA Type A] \n(1/15)*EfficientNetB2-Training Two [TTA Type A] \n\n(1/10)*DenseNet121-Training Two [TTA Type A] \n</code></pre>\n\nOne more, my last submission and that finished tight after the deadline got Private Score 0.93052 and Public Score 0.89110,\n\n<pre><code>(3.5/10)*Public-Quishen [TTA Type A]  \n(3.5/10)*Public-RAHUL SINGH INDA [TTA Type A] \n\n(1/30)*EfficientNetB0-Fold0-Training One [TTA Type C] \n(1/30)*EfficientNetB0-Fold2-Training One [TTA Type C] \n(1/30)*EfficientNetB0-Fold4-Training One [TTA Type C] \n\n(1/15)*EfficientNetB0-Training Two [TTA Type C] \n(1/15)*EfficientNetB1-Training Two [TTA Type C] \n(1/15)*EfficientNetB2-Training Two [TTA Type C] \n</code></pre>\n\n\n\n\n    \n### That is all, Thanks a lot!\nBy the way, I still tremble with fear\nUpdate: No longer!",
      "votes": 24
    },
    {
      "id": 1006737,
      "postDate": "2020-09-11T14:16:28.400Z",
      "content": "<p><a href=\"https://www.kaggle.com/coreacasa\" target=\"_blank\">@coreacasa</a> Thanks for sharing the overview. But I couldn't understand the below code<br>\n`PREDS = (1/5)<em>PREDS + (1/5)</em>PREDS1 + (1/5)<em>PREDS2 + (1/5)</em>PREDS3 + (1/5)*PREDS4</p>\n<p>FINAL = np.round( (6/10)<em>PREDS +\n                  (2/60)</em>np.array(predictions10) + (2/60)<em>np.array(predictions12) + \n                  (2/60)</em>np.array(predictions20) + (2/60)<em>np.array(predictions22) +\n                  (2/60)</em>np.array(predictions30) + (2/60)<em>np.array(predictions32) +\n                  (0.5/10)</em>np.array(predictions40) + (0.5/10)<em>np.array(predictions42) +\n                  (1/60)</em>np.array(predictions50) + (1/60)<em>np.array(predictions52) +\n                  (1/60)</em>np.array(predictions60) + (1/60)<em>np.array(predictions62) +\n                  (1/60)</em>np.array(predictions70) + (1/60)*np.array(predictions72) )<br>\n`<br>\nWhat are these numbers 6/10, 2/60,1/60,0.5/10? Would be grateful if you can elaborate on this.</p>\n<p>Thanks</p>",
      "rawMarkdown": "@coreacasa Thanks for sharing the overview. But I couldn't understand the below code\n`PREDS = (1/5)*PREDS + (1/5)*PREDS1 + (1/5)*PREDS2 + (1/5)*PREDS3 + (1/5)*PREDS4\n\nFINAL = np.round( (6/10)*PREDS +\n                  (2/60)*np.array(predictions10) + (2/60)*np.array(predictions12) + \n                  (2/60)*np.array(predictions20) + (2/60)*np.array(predictions22) +\n                  (2/60)*np.array(predictions30) + (2/60)*np.array(predictions32) +\n                  (0.5/10)*np.array(predictions40) + (0.5/10)*np.array(predictions42) +\n                  (1/60)*np.array(predictions50) + (1/60)*np.array(predictions52) +\n                  (1/60)*np.array(predictions60) + (1/60)*np.array(predictions62) +\n                  (1/60)*np.array(predictions70) + (1/60)*np.array(predictions72) )\n`\nWhat are these numbers 6/10, 2/60,1/60,0.5/10? Would be grateful if you can elaborate on this.\n\nThanks",
      "votes": 1,
      "replies": [
        {
          "id": 1006910,
          "postDate": "2020-09-11T16:27:32.320Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/gaur128\" target=\"_blank\">@gaur128</a>,<br>\nThey are weights about the participation of each of the TTA types in the models or about each of the models in the final solution. They are based on the experience of performance on the leaderboard during the competition.</p>",
          "rawMarkdown": "Hi @gaur128,\nThey are weights about the participation of each of the TTA types in the models or about each of the models in the final solution. They are based on the experience of performance on the leaderboard during the competition."
        },
        {
          "id": 1006944,
          "postDate": "2020-09-11T16:59:33.457Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/coreacasa\" target=\"_blank\">@coreacasa</a> Thanks for your reply.<br>\nHow did you come up with these weights? Is there some logic behind it or is it somewhat intuition+experimentation.</p>",
          "rawMarkdown": "Hi @coreacasa Thanks for your reply.\nHow did you come up with these weights? Is there some logic behind it or is it somewhat intuition+experimentation.",
          "votes": 1
        },
        {
          "id": 1006954,
          "postDate": "2020-09-11T17:13:05.657Z",
          "content": "<p>There is no math behind it if that is what you mean!</p>",
          "rawMarkdown": "There is no math behind it if that is what you mean!"
        },
        {
          "id": 1007548,
          "postDate": "2020-09-12T09:02:12.967Z",
          "content": "<p>Got it. Thanks</p>",
          "rawMarkdown": "Got it. Thanks"
        }
      ]
    },
    {
      "id": 945513,
      "postDate": "2020-07-25T22:39:33.883Z",
      "content": "<p><a href=\"/coreacasa\">@coreacasa</a> Congrats. </p>",
      "rawMarkdown": "@coreacasa Congrats. ",
      "votes": 1,
      "replies": [
        {
          "id": 946115,
          "postDate": "2020-07-26T11:17:01.280Z",
          "content": "<p>Congratulations to you and thanks a lot <a href=\"/micheomaano\">@micheomaano</a> for sharing your notebooks during the competition, your code is the soul of the code files that are in this solution.</p>",
          "rawMarkdown": "Congratulations to you and thanks a lot @micheomaano for sharing your notebooks during the competition, your code is the soul of the code files that are in this solution."
        }
      ]
    },
    {
      "id": 944085,
      "postDate": "2020-07-24T20:00:31.027Z",
      "content": "<p>Thank you for sharing <a href=\"/coreacasa\">@coreacasa</a> ! this is amazing. \nIt is interesting that I had some parameters almost like your code base 1. \nYou had 42 tiles for EfficientNetB0. I found 48 tiles worked best for me. but my batch size was 2. \nI was also struggling with TTA so i will check on your code. \nThank you again for sharing! this is very educational and will help us all improve! \nCheers! </p>",
      "rawMarkdown": "Thank you for sharing @coreacasa ! this is amazing. \nIt is interesting that I had some parameters almost like your code base 1. \nYou had 42 tiles for EfficientNetB0. I found 48 tiles worked best for me. but my batch size was 2. \nI was also struggling with TTA so i will check on your code. \nThank you again for sharing! this is very educational and will help us all improve! \nCheers! ",
      "votes": 1,
      "replies": [
        {
          "id": 944887,
          "postDate": "2020-07-25T12:31:50.743Z",
          "content": "<p>Thank you <a href=\"/mpsampat\">@mpsampat</a> !!, do not hesitate to ask any questions about it.</p>",
          "rawMarkdown": "Thank you @mpsampat !!, do not hesitate to ask any questions about it.",
          "votes": 1
        }
      ]
    },
    {
      "id": 944181,
      "postDate": "2020-07-24T22:39:46.110Z",
      "content": "<p>Congrats Manuel on a great solution, on solo Gold, and becoming a Competition Master!</p>",
      "rawMarkdown": "Congrats Manuel on a great solution, on solo Gold, and becoming a Competition Master!",
      "votes": 2,
      "replies": [
        {
          "id": 944882,
          "postDate": "2020-07-25T12:27:48.633Z",
          "content": "<p>Thank you very much <a href=\"/cdeotte\">@cdeotte</a>, it is an honor to read your congratulations here!</p>",
          "rawMarkdown": "Thank you very much @cdeotte, it is an honor to read your congratulations here!",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1006737,
      "author_name": "Gaurav Yadav",
      "author_url": "",
      "post_date": "2020-09-11T14:16:28.400000",
      "content": "<p><a href=\"https://www.kaggle.com/coreacasa\" target=\"_blank\">@coreacasa</a> Thanks for sharing the overview. But I couldn't understand the below code<br>\n`PREDS = (1/5)<em>PREDS + (1/5)</em>PREDS1 + (1/5)<em>PREDS2 + (1/5)</em>PREDS3 + (1/5)*PREDS4</p>\n<p>FINAL = np.round( (6/10)<em>PREDS +\n                  (2/60)</em>np.array(predictions10) + (2/60)<em>np.array(predictions12) + \n                  (2/60)</em>np.array(predictions20) + (2/60)<em>np.array(predictions22) +\n                  (2/60)</em>np.array(predictions30) + (2/60)<em>np.array(predictions32) +\n                  (0.5/10)</em>np.array(predictions40) + (0.5/10)<em>np.array(predictions42) +\n                  (1/60)</em>np.array(predictions50) + (1/60)<em>np.array(predictions52) +\n                  (1/60)</em>np.array(predictions60) + (1/60)<em>np.array(predictions62) +\n                  (1/60)</em>np.array(predictions70) + (1/60)*np.array(predictions72) )<br>\n`<br>\nWhat are these numbers 6/10, 2/60,1/60,0.5/10? Would be grateful if you can elaborate on this.</p>\n<p>Thanks</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1006910,
          "author_name": "Manuel Campos",
          "author_url": "",
          "post_date": "2020-09-11T16:27:32.320000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/gaur128\" target=\"_blank\">@gaur128</a>,<br>\nThey are weights about the participation of each of the TTA types in the models or about each of the models in the final solution. They are based on the experience of performance on the leaderboard during the competition.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1006944,
          "author_name": "Gaurav Yadav",
          "author_url": "",
          "post_date": "2020-09-11T16:59:33.457000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/coreacasa\" target=\"_blank\">@coreacasa</a> Thanks for your reply.<br>\nHow did you come up with these weights? Is there some logic behind it or is it somewhat intuition+experimentation.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1006954,
          "author_name": "Manuel Campos",
          "author_url": "",
          "post_date": "2020-09-11T17:13:05.657000",
          "content": "<p>There is no math behind it if that is what you mean!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1007548,
          "author_name": "Gaurav Yadav",
          "author_url": "",
          "post_date": "2020-09-12T09:02:12.967000",
          "content": "<p>Got it. Thanks</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 945513,
      "author_name": "Salman",
      "author_url": "",
      "post_date": "2020-07-25T22:39:33.883000",
      "content": "<p><a href=\"/coreacasa\">@coreacasa</a> Congrats. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 946115,
          "author_name": "Manuel Campos",
          "author_url": "",
          "post_date": "2020-07-26T11:17:01.280000",
          "content": "<p>Congratulations to you and thanks a lot <a href=\"/micheomaano\">@micheomaano</a> for sharing your notebooks during the competition, your code is the soul of the code files that are in this solution.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 944085,
      "author_name": "Mehul Sampat",
      "author_url": "",
      "post_date": "2020-07-24T20:00:31.027000",
      "content": "<p>Thank you for sharing <a href=\"/coreacasa\">@coreacasa</a> ! this is amazing. \nIt is interesting that I had some parameters almost like your code base 1. \nYou had 42 tiles for EfficientNetB0. I found 48 tiles worked best for me. but my batch size was 2. \nI was also struggling with TTA so i will check on your code. \nThank you again for sharing! this is very educational and will help us all improve! \nCheers! </p>",
      "votes": 1,
      "replies": [
        {
          "id": 944887,
          "author_name": "Manuel Campos",
          "author_url": "",
          "post_date": "2020-07-25T12:31:50.743000",
          "content": "<p>Thank you <a href=\"/mpsampat\">@mpsampat</a> !!, do not hesitate to ask any questions about it.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 944181,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-07-24T22:39:46.110000",
      "content": "<p>Congrats Manuel on a great solution, on solo Gold, and becoming a Competition Master!</p>",
      "votes": 2,
      "replies": [
        {
          "id": 944882,
          "author_name": "Manuel Campos",
          "author_url": "",
          "post_date": "2020-07-25T12:27:48.633000",
          "content": "<p>Thank you very much <a href=\"/cdeotte\">@cdeotte</a>, it is an honor to read your congratulations here!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "943758": "First of all, I would like to thank all the people and organizations that have made this Competition possible. In capital letters, THANK YOU to all the TEAMS that with their dedication and effort I hope contribute to improve the diagnosis of prostate cancer and thereby improve people's lives. Indeed, my most sincere congratulations to the WINNERS.\n\nI am very happy as you can imagine. In a few lines I share with you a quick overview of my time in this Challenge.\n\n### Kaggle Learning\nI want to comment here what is usually included in the acknowledgments part but I reserve this special section to highlight the work of those competitors who have made my final solution better, 1) because their ability was not present in my initial knowledge or 2) because their performance improves together with the experience of mine. I mean, in no order of priority,\n\n- **(Salman)** @micheomaano:\n1.  [Dataset tf-record-256-56-48](https://www.kaggle.com/micheomaano/tf-record-256-256-48)\n2. [TPU Training Tensorflow Iafoos Method 42x256x256x3](https://www.kaggle.com/micheomaano/tpu-training-tensorflow-iafoos-method-42x256x256x3)\n3. [Pandas 42x256x256x3 Inference](https://www.kaggle.com/micheomaano/pandas-42x256x256x3-inference)\n\n- **(Qishen Ha)** @haqishen:\n1.  [Train EfficientNet-B0 w/ 36 tiles_256 [LB0.87]](https://www.kaggle.com/haqishen/train-efficientnet-b0-w-36-tiles-256-lb0-87)\n2. [PANDA Inference w/ 36 tiles_256](https://www.kaggle.com/haqishen/panda-inference-w-36-tiles-256)\n\n- **(RAHUL SINGH INDA)** @rsinda:\n1.  [Panda Inference EfficientNet-b1](https://www.kaggle.com/rsinda/panda-inference-efficientnet-b1)\n\n- **(Iafoss)** @iafoss: The Best Accelerator in the Competition, ahead of TPUs.\n\n###Submission Notebook    \n   \nI have shared an original copy of my inference kernel without additional cleaning as well as a dataset that includes the necessary weights of each of the models that are used in obtaining the final submission,\n    \n- [Quick Save Inference](https://www.kaggle.com/coreacasa/12th-place-solution-quick-save-inference)\n- [Dataset Model Weights for Inference](https://www.kaggle.com/coreacasa/pandaenetb042x256x256x3)\n   \n### [TPU] Kaggle/Google(Colaboratory)\nFor all my trainings I used the free TPU resources offered by Kaggle / Google (Colaboratory). Thank you very much.\n\n### Training One: My only approach to validation\n\nVery closed to Salman's training kernel I just re-ran its code to complete cross validation. I ran a fold up to 60 epochs to see the evolution of the loss and the rest down to 40 epochs.\n\nIndividually the behavior of the folds was more or less similar in final loss values (mse) and in the number of times in which it stopped improving. The issue is that its merge did not improve the individual performance of some of them over LB and their performance was also uneven when they were introduced into an external ensemble.\n\nThe noise of the labels is a probable cause as already discussed in the discussions or perhaps the sensitivity of the qwk metric to even small variations in mse when its jump to LB.\n\n- **[code-base-training-one](https://www.kaggle.com/coreacasa/code-base-training-one)** file, training topics:\n\n<code>Size Image</code> 256\n<code>Size Tiles</code> 256\n<code>Tiles</code> 42\n<code>Augmentation</code> horizonal p=0.5 and vertical p=0.5 flips\n<code>Validation</code> StratifiedKFold 5 on isup grade classes\n<code>Arch</code> EfficientNetB0\n<code>Convolutional Base's Weight</code> Imagenet trainable\n<code>On Top</code> GlobalAveragePooling2D, Dropout(0.5), Dense(1024)\n<code>Output</code> Dense(1) regression objective\n<code>Loss</code> mean_squared_error\n<code>Optimizer</code> Adam\n<code>Leaning Rate</code> 5e-04 init\n<code>Reduce LR</code> decreasing 0.5 with patience 3 epochs\n<code>Save</code> weights only with best validation loss epochs\n<code>Batch Size</code> 64\n\n### Training Two: Art(Instinct) Validation\n\nI never tried detecting noisy labels to remove them from training data. In general I am not in favor of losing any existing information, although in principle it could be harmful by elevating the non-regular part of a data generating process. I would rather transform data than remove it.\n\nI didn't try either any transformation so I thought about training the models with full dataset in order to prevent the possible existence of more noise in some folds than in others, which probably would be increasing the variability in the inference results.\n\nArt Validation appears here and it is when the art of the data scientist enters and it is his instinct that determines the goodness of fit and stability of performance in generalization against new observations. Yes, this is Alchemy.\n\n- **[code-base-training-two-enets](https://www.kaggle.com/coreacasa/code-base-training-two-enets)** file, from which I trained 3 members of the EfficientNet family. Changes on training one training topics:\n\n<code>Tiles</code> 48\n<code>Validation</code> Art Validation on instinct\n<code>Arch</code> EfficientNetB0, EfficientNetB1 and EfficientNetB2\n<code>Convolutional Base's Weight</code> Noisy Student trainable\n<code>Output</code> Dense(5,activation='sigmoid) ordinal regression objective\n<code>Loss</code> sigmoid_cross_entropy_with_logits\n<code>Leaning Rate</code> custom with 5up, 3sustain, 0.8decay\n<code>Limits LR</code> 1e-05min, 4e-04max\n<code>Save</code> weights only with best loss epochs\n<code>Batch Size</code> 32\n<code>Epochs</code> 60\n\n- **[code-base-training-two-densenet](https://www.kaggle.com/coreacasa/code-base-training-two-densenet)** file, from which I trained 1 member of the DenseNet family. Changes on training topics of the previous net family:\n\n<code>Arch</code> Densenet121\n<code>Convolutional Base's Weight</code> Imagenet trainable\n<code>Epochs</code> 40\n\n### Inference: Diversity of Archs, nTiles and TTAs\n\nOf the 2 training processes shown above, the following models were available,\n1. EfficientNetB0 (5 skf), 42x256x256x3\n2. EfficientNetB0 (1), 48x256x256x3 \n3. EfficientNetB1 (1), 48x256x256x3 \n4. EfficientNetB2 (1), 48x256x256x3 \n5. DenseNet121 (1), 48x256x256x3 \n\nHaving re-run the Salman kernel, from the public notebooks referenced at the beginning I had,\n1. EfficientNetB0 (1 skf), 36x256x256x3 (Qishen Ha) \n2. EfficientNetB1 (1 skf), 36x256x256x3 (RAHUL SINGH INDA)\n    \n- **Test Time Augmentation**\n<code>Type A: 5xTTA deterministic</code> \n1xoriginal, 1xTranspose, 1xVerticalFlip, 1xHorizontalFlip, 1xTranspose-&gt;VerticalFlip-&gt;HorizontalFlip\n<code>Type B: 4xTTA pseudo deterministic</code> \n1xoriginal, 1xVerticalFlip, 2xHorizontalFlip(p=0.5)-&gt;VerticalFlip(p=0.5)\n<code>Type C: 2xTTA random </code> \n2xHorizontalFlip(p=0.5)-&gt;VerticalFlip(p=0)\n\n- **White Padding Tile Extraction (Qishen modes)**\n1x add zero pad and 1x add 256 pad, that is, 2 different extractions for ALL the images.\n\n### Model Selection and Final Ensemble\n\n<pre><code>(3/10)*Public-Quishen [TTA Type A]  \n(3/10)*Public-RAHUL SINGH INDA [TTA Type A] \n\n(1/30)*EfficientNetB0-Fold0-Training One [TTA Type C] \n(1/30)*EfficientNetB0-Fold2-Training One [TTA Type C] \n(1/30)*EfficientNetB0-Fold4-Training One [TTA Type C] \n\n(1/15)*EfficientNetB0-Training Two [TTA Type C] \n(1/15)*EfficientNetB1-Training Two [TTA Type C] \n(1/15)*EfficientNetB2-Training Two [TTA Type C] \n\n(1/10)*DenseNet121-Training Two [TTA Type B] \n</code></pre>\n\nThe random component of the TTAs was not seed (I'll be lucky) and the reproducibility of the results may vary with it. I have just re-run my inference kernel and the results are Private Score 0.92983 (0.92960 original) and Public Score 0.89443 (089352 original).\n\nWith this models structure I was only able to test the last day of the competition. For example, this other ensemble got Private Score 0.93047 and Public Score 0.88889, not including random component in TTA.\n\n<pre><code>(3.5/10)*Public-Quishen [TTA Type A]  \n(3.5/10)*Public-RAHUL SINGH INDA [TTA Type A] \n\n(1/15)*EfficientNetB0-Training Two [TTA Type A] \n(1/15)*EfficientNetB1-Training Two [TTA Type A] \n(1/15)*EfficientNetB2-Training Two [TTA Type A] \n\n(1/10)*DenseNet121-Training Two [TTA Type A] \n</code></pre>\n\nOne more, my last submission and that finished tight after the deadline got Private Score 0.93052 and Public Score 0.89110,\n\n<pre><code>(3.5/10)*Public-Quishen [TTA Type A]  \n(3.5/10)*Public-RAHUL SINGH INDA [TTA Type A] \n\n(1/30)*EfficientNetB0-Fold0-Training One [TTA Type C] \n(1/30)*EfficientNetB0-Fold2-Training One [TTA Type C] \n(1/30)*EfficientNetB0-Fold4-Training One [TTA Type C] \n\n(1/15)*EfficientNetB0-Training Two [TTA Type C] \n(1/15)*EfficientNetB1-Training Two [TTA Type C] \n(1/15)*EfficientNetB2-Training Two [TTA Type C] \n</code></pre>\n\n\n\n\n    \n### That is all, Thanks a lot!\nBy the way, I still tremble with fear\nUpdate: No longer!",
    "1006737": "@coreacasa Thanks for sharing the overview. But I couldn't understand the below code\n`PREDS = (1/5)*PREDS + (1/5)*PREDS1 + (1/5)*PREDS2 + (1/5)*PREDS3 + (1/5)*PREDS4\n\nFINAL = np.round( (6/10)*PREDS +\n                  (2/60)*np.array(predictions10) + (2/60)*np.array(predictions12) + \n                  (2/60)*np.array(predictions20) + (2/60)*np.array(predictions22) +\n                  (2/60)*np.array(predictions30) + (2/60)*np.array(predictions32) +\n                  (0.5/10)*np.array(predictions40) + (0.5/10)*np.array(predictions42) +\n                  (1/60)*np.array(predictions50) + (1/60)*np.array(predictions52) +\n                  (1/60)*np.array(predictions60) + (1/60)*np.array(predictions62) +\n                  (1/60)*np.array(predictions70) + (1/60)*np.array(predictions72) )\n`\nWhat are these numbers 6/10, 2/60,1/60,0.5/10? Would be grateful if you can elaborate on this.\n\nThanks",
    "945513": "@coreacasa Congrats. ",
    "944085": "Thank you for sharing @coreacasa ! this is amazing. \nIt is interesting that I had some parameters almost like your code base 1. \nYou had 42 tiles for EfficientNetB0. I found 48 tiles worked best for me. but my batch size was 2. \nI was also struggling with TTA so i will check on your code. \nThank you again for sharing! this is very educational and will help us all improve! \nCheers! ",
    "944181": "Congrats Manuel on a great solution, on solo Gold, and becoming a Competition Master!"
  }
}