{
  "id": 169303,
  "title": "Part of 2nd place solution (only kaggle/colab TPU were used) ",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/169303",
  "author_name": "Tsai29",
  "post_date": "2020-07-23T13:48:39.619000",
  "votes": 49,
  "comment_count": 12,
  "views": 0,
  "content": "<p>First of all, I want to thank kaggle and the organizers for hosting such interesting competition.\nAnd I want to thank my teammates, they really did a great job during the competition. \nSince I don't have my own machine, so I only trained my model on kaggle/colab TPU. But in fact, TPU is pretty enough to get 0.92+ private score. I think the most important part of my single model is the loss function, by using separated losses for each data center can handle the label noise of radboud quite well.\n<br></p>\n\n<h2>strategy</h2>\n\n<p>There are two main challenges of this competition, and these challenges already mentioned in the overview of this competition. <br>\n1. Huge image size\n2. noisy data</p>\n\n<p><strong>Challenge 1</strong> <br>\nI followed what @lafoss shared and add some simple preprocessing to get the proper training images. (Thanks a lot @lafoss) <br></p>\n\n<p><strong>Challenge 2</strong> <br>\nFor the challenge 2, I think there are some clues we could found in the documents shared by host and some topics in discussion section. <br></p>\n\n<p>a. The radboud data is more noisy than karolinska <br>\nb. The qwk of radboud training data is around ~0.85, which is judged by human <br>\nc. @lafoss got his 0.9lb on low radboud cv but high karolinska cv <br></p>\n\n<p>So in my point of view, we shouldn't have radboud cv too high, since this might be some overfitting signals of your model. And the reasonable cv of radboud should keep in 0.84~0.86, since the statement b above. At the beginning, I also can't find any pattern between cv and LB. So I used logcoah loss on both center of data, which is the loss function relativly noise robust than MSE . And after I saw the individual cv of @lafoss's, I realized that I should use some methods to limit the model to learn too much about radboud data and learn more about karolinska data. So I used huber loss(which was suggested by my good teammate <a href=\"/rguo97\">@rguo97</a>)  with delta 1 on radboud instances and MSE on karolinska instances. This did keep my radboud cv won't over 0.86 and the cv of karolinska did increased. And this was the time I got my 0.903LB, at the 15 days before the competition end. Well, this also related to what split your folding is, but I think this is the small pattern I thought is right about cv and lb relation. <br>\nBut of course, this isn't enough for getting 2nd place on private board, I think the main reason we got the 2nd place is because all our model have quite nice single model result, and the diversity of our ensemble model bring us to stable LB/PB relation.</p>\n\n<h2>Data prepare/preprocessing:</h2>\n\n<p>I also turned the slides image into tiles, but instead of using MIL method, I glued the tiles into a big image. The tile size is 128x128 and total 144 tiles will be extract from a slide image. Then I glued the 144 tiles into a big image, which means I turned a slide into 1536x1536 image. But before gluing the image, I will get the foreground of each tile and resize back to 128x128. In the end, I wrote every glued images into tfrecord format after I encoded them in jpeg format.</p>\n\n<h2>Data augmentation</h2>\n\n<ol>\n<li>H&amp;E color jitter (0.25 random prob)</li>\n<li>Random contrast/brightness (0.25 random prob)</li>\n<li>Random saturation (0.15 random prob)</li>\n<li>Image transpose (0.5 random prob)</li>\n<li>Random H/V flip</li>\n<li>Shift(50pixels)/Rotate(10degrees)/Scale(0.05) (0.25 random prob)\n<br></li>\n</ol>\n\n<h2>Model</h2>\n\n<p>Efficientnet b3 + GeM + 0.3 dropout + 128 units dense layer + 1 unit dense layer(regression head)\n<br></p>\n\n<h2>Loss function</h2>\n\n<p>MSE loss on karolinska data\nHuber loss with delta 1 on radboud data</p>\n\n<h2>Training</h2>\n\n<p>One cycle cosine annealing with initial learning rate 5e-4 and minimum learning rate 5e-6, and train for 40 epochs. Also I saved the checkpoint for radboud and karolinska separately, so sometimes I got two best checkpoints on radboud and karolinska ceneter. And I normally used best checkpoint of karolinska only in inference.</p>\n\n<h2>Single model score</h2>\n\n<p>0.905 qwk cv on karolinska/0.853 qwk cv on radboud/0.891 qwk overall cv <br>\n0.903 qwk on LB/0.927 qwk on PB</p>\n\n<h2>Final models</h2>\n\n<p>Please refer this topic for our ensemble model\n<a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/169108\">https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/169108</a></p>\n\n<hr>\n\n<p>Here is the training notebook of my model.\n<a href=\"https://www.kaggle.com/xiejialun/my-part-of-2nd-place-solution-tpu/notebook\">https://www.kaggle.com/xiejialun/my-part-of-2nd-place-solution-tpu/notebook</a></p>",
  "messages": [
    {
      "id": 941916,
      "postDate": "2020-07-23T13:48:39.620Z",
      "content": "<p>First of all, I want to thank kaggle and the organizers for hosting such interesting competition.\nAnd I want to thank my teammates, they really did a great job during the competition. \nSince I don't have my own machine, so I only trained my model on kaggle/colab TPU. But in fact, TPU is pretty enough to get 0.92+ private score. I think the most important part of my single model is the loss function, by using separated losses for each data center can handle the label noise of radboud quite well.\n<br></p>\n\n<h2>strategy</h2>\n\n<p>There are two main challenges of this competition, and these challenges already mentioned in the overview of this competition. <br>\n1. Huge image size\n2. noisy data</p>\n\n<p><strong>Challenge 1</strong> <br>\nI followed what @lafoss shared and add some simple preprocessing to get the proper training images. (Thanks a lot @lafoss) <br></p>\n\n<p><strong>Challenge 2</strong> <br>\nFor the challenge 2, I think there are some clues we could found in the documents shared by host and some topics in discussion section. <br></p>\n\n<p>a. The radboud data is more noisy than karolinska <br>\nb. The qwk of radboud training data is around ~0.85, which is judged by human <br>\nc. @lafoss got his 0.9lb on low radboud cv but high karolinska cv <br></p>\n\n<p>So in my point of view, we shouldn't have radboud cv too high, since this might be some overfitting signals of your model. And the reasonable cv of radboud should keep in 0.84~0.86, since the statement b above. At the beginning, I also can't find any pattern between cv and LB. So I used logcoah loss on both center of data, which is the loss function relativly noise robust than MSE . And after I saw the individual cv of @lafoss's, I realized that I should use some methods to limit the model to learn too much about radboud data and learn more about karolinska data. So I used huber loss(which was suggested by my good teammate <a href=\"/rguo97\">@rguo97</a>)  with delta 1 on radboud instances and MSE on karolinska instances. This did keep my radboud cv won't over 0.86 and the cv of karolinska did increased. And this was the time I got my 0.903LB, at the 15 days before the competition end. Well, this also related to what split your folding is, but I think this is the small pattern I thought is right about cv and lb relation. <br>\nBut of course, this isn't enough for getting 2nd place on private board, I think the main reason we got the 2nd place is because all our model have quite nice single model result, and the diversity of our ensemble model bring us to stable LB/PB relation.</p>\n\n<h2>Data prepare/preprocessing:</h2>\n\n<p>I also turned the slides image into tiles, but instead of using MIL method, I glued the tiles into a big image. The tile size is 128x128 and total 144 tiles will be extract from a slide image. Then I glued the 144 tiles into a big image, which means I turned a slide into 1536x1536 image. But before gluing the image, I will get the foreground of each tile and resize back to 128x128. In the end, I wrote every glued images into tfrecord format after I encoded them in jpeg format.</p>\n\n<h2>Data augmentation</h2>\n\n<ol>\n<li>H&amp;E color jitter (0.25 random prob)</li>\n<li>Random contrast/brightness (0.25 random prob)</li>\n<li>Random saturation (0.15 random prob)</li>\n<li>Image transpose (0.5 random prob)</li>\n<li>Random H/V flip</li>\n<li>Shift(50pixels)/Rotate(10degrees)/Scale(0.05) (0.25 random prob)\n<br></li>\n</ol>\n\n<h2>Model</h2>\n\n<p>Efficientnet b3 + GeM + 0.3 dropout + 128 units dense layer + 1 unit dense layer(regression head)\n<br></p>\n\n<h2>Loss function</h2>\n\n<p>MSE loss on karolinska data\nHuber loss with delta 1 on radboud data</p>\n\n<h2>Training</h2>\n\n<p>One cycle cosine annealing with initial learning rate 5e-4 and minimum learning rate 5e-6, and train for 40 epochs. Also I saved the checkpoint for radboud and karolinska separately, so sometimes I got two best checkpoints on radboud and karolinska ceneter. And I normally used best checkpoint of karolinska only in inference.</p>\n\n<h2>Single model score</h2>\n\n<p>0.905 qwk cv on karolinska/0.853 qwk cv on radboud/0.891 qwk overall cv <br>\n0.903 qwk on LB/0.927 qwk on PB</p>\n\n<h2>Final models</h2>\n\n<p>Please refer this topic for our ensemble model\n<a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/169108\">https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/169108</a></p>\n\n<hr>\n\n<p>Here is the training notebook of my model.\n<a href=\"https://www.kaggle.com/xiejialun/my-part-of-2nd-place-solution-tpu/notebook\">https://www.kaggle.com/xiejialun/my-part-of-2nd-place-solution-tpu/notebook</a></p>",
      "rawMarkdown": "First of all, I want to thank kaggle and the organizers for hosting such interesting competition.\nAnd I want to thank my teammates, they really did a great job during the competition. \nSince I don't have my own machine, so I only trained my model on kaggle/colab TPU. But in fact, TPU is pretty enough to get 0.92+ private score. I think the most important part of my single model is the loss function, by using separated losses for each data center can handle the label noise of radboud quite well.\n<br>\n\n## strategy\nThere are two main challenges of this competition, and these challenges already mentioned in the overview of this competition. <br>\n1. Huge image size\n2. noisy data\n\n**Challenge 1** <br>\nI followed what @lafoss shared and add some simple preprocessing to get the proper training images. (Thanks a lot @lafoss) <br>\n\n**Challenge 2** <br>\nFor the challenge 2, I think there are some clues we could found in the documents shared by host and some topics in discussion section. <br>\n\na. The radboud data is more noisy than karolinska <br>\nb. The qwk of radboud training data is around ~0.85, which is judged by human <br>\nc. @lafoss got his 0.9lb on low radboud cv but high karolinska cv <br>\n\nSo in my point of view, we shouldn't have radboud cv too high, since this might be some overfitting signals of your model. And the reasonable cv of radboud should keep in 0.84~0.86, since the statement b above. At the beginning, I also can't find any pattern between cv and LB. So I used logcoah loss on both center of data, which is the loss function relativly noise robust than MSE . And after I saw the individual cv of @lafoss's, I realized that I should use some methods to limit the model to learn too much about radboud data and learn more about karolinska data. So I used huber loss(which was suggested by my good teammate @rguo97)  with delta 1 on radboud instances and MSE on karolinska instances. This did keep my radboud cv won't over 0.86 and the cv of karolinska did increased. And this was the time I got my 0.903LB, at the 15 days before the competition end. Well, this also related to what split your folding is, but I think this is the small pattern I thought is right about cv and lb relation. <br>\nBut of course, this isn't enough for getting 2nd place on private board, I think the main reason we got the 2nd place is because all our model have quite nice single model result, and the diversity of our ensemble model bring us to stable LB/PB relation.\n\n## Data prepare/preprocessing:\nI also turned the slides image into tiles, but instead of using MIL method, I glued the tiles into a big image. The tile size is 128x128 and total 144 tiles will be extract from a slide image. Then I glued the 144 tiles into a big image, which means I turned a slide into 1536x1536 image. But before gluing the image, I will get the foreground of each tile and resize back to 128x128. In the end, I wrote every glued images into tfrecord format after I encoded them in jpeg format.\n\n## Data augmentation\n1. H&amp;E color jitter (0.25 random prob)\n2. Random contrast/brightness (0.25 random prob)\n3. Random saturation (0.15 random prob)\n4. Image transpose (0.5 random prob)\n5. Random H/V flip\n6. Shift(50pixels)/Rotate(10degrees)/Scale(0.05) (0.25 random prob)\n<br>\n\n## Model\nEfficientnet b3 + GeM + 0.3 dropout + 128 units dense layer + 1 unit dense layer(regression head)\n<br>\n\n## Loss function\nMSE loss on karolinska data\nHuber loss with delta 1 on radboud data\n\n## Training\nOne cycle cosine annealing with initial learning rate 5e-4 and minimum learning rate 5e-6, and train for 40 epochs. Also I saved the checkpoint for radboud and karolinska separately, so sometimes I got two best checkpoints on radboud and karolinska ceneter. And I normally used best checkpoint of karolinska only in inference.\n\n## Single model score\n0.905 qwk cv on karolinska/0.853 qwk cv on radboud/0.891 qwk overall cv <br>\n0.903 qwk on LB/0.927 qwk on PB\n\n## Final models\nPlease refer this topic for our ensemble model\nhttps://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/169108\n\n\n-----------------------------\nHere is the training notebook of my model.\nhttps://www.kaggle.com/xiejialun/my-part-of-2nd-place-solution-tpu/notebook",
      "votes": 49
    },
    {
      "id": 948958,
      "postDate": "2020-07-28T10:43:29.953Z",
      "content": "<p>Well deserved, thank you for sharing the information in such neat and detailed manner. Looking forward to many more in future.👍 </p>",
      "rawMarkdown": "Well deserved, thank you for sharing the information in such neat and detailed manner. Looking forward to many more in future.👍 ",
      "votes": 1,
      "replies": [
        {
          "id": 949132,
          "postDate": "2020-07-28T12:51:05.750Z",
          "content": "<p>Thanks! :)</p>",
          "rawMarkdown": "Thanks! :)\n"
        }
      ]
    },
    {
      "id": 945004,
      "postDate": "2020-07-25T13:50:00.183Z",
      "content": "<p>Thank you for sharing.. \nRadboud is calcuated with Huber_loss is quite good.. It is good strategy..\nCould I ask three things?\n1. Why do you use huber_loss not smoothl1loss.(I want to know difference huber_loss and smoothl1loss)\n2. What is roll of cencer_mask in radboud_loss and Karolinska_loss function..(guess.. kind of mask between radboud and Karolinska?)\n3. ColorStain augmentation is really work? how much improvment is for score?</p>\n\n<p>Thank you for sharing :) Congratulation~</p>",
      "rawMarkdown": "Thank you for sharing.. \nRadboud is calcuated with Huber_loss is quite good.. It is good strategy..\nCould I ask three things?\n1. Why do you use huber_loss not smoothl1loss.(I want to know difference huber_loss and smoothl1loss)\n2. What is roll of cencer_mask in radboud_loss and Karolinska_loss function..(guess.. kind of mask between radboud and Karolinska?)\n3. ColorStain augmentation is really work? how much improvment is for score?\n\nThank you for sharing :) Congratulation~",
      "votes": 1,
      "replies": [
        {
          "id": 945547,
          "postDate": "2020-07-26T00:07:43.897Z",
          "content": "<p>Hi <a href=\"/rvslight\">@rvslight</a> : </p>\n\n<p>Thank you :)</p>\n\n<ol>\n<li><p>Since there are still some data in radboud is worth to learn, and if you also limit the model to learn from the good part of radboud data, the model won't perform very well on karolinska data either. And the huber can let instances with absolute difference less than delta has bigger loss, which will make model pay more attention on clean data of radboud more.</p></li>\n<li><p>Yes, you are correct. The center_mask is the mask for radboud/karolinska instances. Zeros in mask are corresponding to radboud instances, and ones corresponding to karolinska instances.</p></li>\n<li><p>To be honest, ColorStain augmentation didn't improve the cv or the LB, maybe just a little bit improvement. But I noticed that model will remember the color pattern of radboud and karolinska if you didn't perform H&amp;E color jitter on them. I was afraid this might impact the generalization ability of my model. So I added ColorStain augmentation in data pipeline just in case.</p></li>\n</ol>",
          "rawMarkdown": "Hi @rvslight : \n\nThank you :)\n\n1. Since there are still some data in radboud is worth to learn, and if you also limit the model to learn from the good part of radboud data, the model won't perform very well on karolinska data either. And the huber can let instances with absolute difference less than delta has bigger loss, which will make model pay more attention on clean data of radboud more.\n\n2. Yes, you are correct. The center_mask is the mask for radboud/karolinska instances. Zeros in mask are corresponding to radboud instances, and ones corresponding to karolinska instances.\n\n3. To be honest, ColorStain augmentation didn't improve the cv or the LB, maybe just a little bit improvement. But I noticed that model will remember the color pattern of radboud and karolinska if you didn't perform H&amp;E color jitter on them. I was afraid this might impact the generalization ability of my model. So I added ColorStain augmentation in data pipeline just in case.",
          "votes": 1
        }
      ]
    },
    {
      "id": 944941,
      "postDate": "2020-07-25T13:17:53.160Z",
      "content": "<p><a href=\"/xiejialun\">@xiejialun</a> \nCongrats. That's a really great achievement. 🎉 \nI think the main reason you got leverage from TPU was to use <code>.tfrec</code> file format. </p>",
      "rawMarkdown": "@xiejialun \nCongrats. That's a really great achievement. 🎉 \nI think the main reason you got leverage from TPU was to use `.tfrec` file format. ",
      "votes": 1,
      "replies": [
        {
          "id": 945548,
          "postDate": "2020-07-26T00:08:49.700Z",
          "content": "<p>Thanks!\nYeah, .tfrec really speed up the training process on TPU!</p>",
          "rawMarkdown": "Thanks!\nYeah, .tfrec really speed up the training process on TPU!",
          "votes": 1
        }
      ]
    },
    {
      "id": 943116,
      "postDate": "2020-07-24T06:35:49.540Z",
      "content": "<p>Thankyou for sharing</p>",
      "rawMarkdown": "Thankyou for sharing",
      "votes": 1,
      "replies": [
        {
          "id": 944238,
          "postDate": "2020-07-25T00:33:19.220Z",
          "content": "<p>No problem :)</p>",
          "rawMarkdown": "No problem :)\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 942928,
      "postDate": "2020-07-24T04:23:47.847Z",
      "content": "<p>Thank you for sharing and congratulations!</p>",
      "rawMarkdown": "Thank you for sharing and congratulations!",
      "votes": 1,
      "replies": [
        {
          "id": 944239,
          "postDate": "2020-07-25T00:33:30.647Z",
          "content": "<p>Thanks! :)</p>",
          "rawMarkdown": "Thanks! :)\n"
        }
      ]
    },
    {
      "id": 948544,
      "postDate": "2020-07-28T03:55:20.177Z",
      "content": "<p>Hi,</p>\n\n<p>Thank you for posting the notebook. I wrote the code for my first competition (Google landmark detection). But to train the model, each epoch was taking about 12 hours (even after reducing the image size to 200X200). It had a lot of images and classes, probably that's why it's slow. So how would you train for such a competition?</p>\n\n<p>Please let me know.</p>",
      "rawMarkdown": "Hi,\n\nThank you for posting the notebook. I wrote the code for my first competition (Google landmark detection). But to train the model, each epoch was taking about 12 hours (even after reducing the image size to 200X200). It had a lot of images and classes, probably that's why it's slow. So how would you train for such a competition?\n\nPlease let me know.",
      "replies": [
        {
          "id": 949158,
          "postDate": "2020-07-28T12:59:10.800Z",
          "content": "<p>Hi,\nFirst of all, are you sure you opened the GPU before you started to train the model?\nSince with resolution this small, the training time shouldn't be this long.</p>\n\n<p>I think there are few things you could check:\n1. Whether you opened the GPU or not ( you can open the GPU on the right side of kaggle notebook)\n2. Make sure you uses the batch size big enough ( I always use the biggest batch size which GPU allowed, then fine tune the batch size to get balance model performance and training time)\n3. Check whether the data pipeline is work normally. You can get the data from your data pipeline and make sure there aren't any bugs in it before training your model.\n4. Try to reduce the parameters of your model to check whether the speed of training is increasing.</p>\n\n<p>Hope you solve your problem well.</p>",
          "rawMarkdown": "Hi,\nFirst of all, are you sure you opened the GPU before you started to train the model?\nSince with resolution this small, the training time shouldn't be this long.\n\nI think there are few things you could check:\n1. Whether you opened the GPU or not ( you can open the GPU on the right side of kaggle notebook)\n2. Make sure you uses the batch size big enough ( I always use the biggest batch size which GPU allowed, then fine tune the batch size to get balance model performance and training time)\n3. Check whether the data pipeline is work normally. You can get the data from your data pipeline and make sure there aren't any bugs in it before training your model.\n4. Try to reduce the parameters of your model to check whether the speed of training is increasing.\n\nHope you solve your problem well."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 948958,
      "author_name": "Manish Kumar",
      "author_url": "",
      "post_date": "2020-07-28T10:43:29.953000",
      "content": "<p>Well deserved, thank you for sharing the information in such neat and detailed manner. Looking forward to many more in future.👍 </p>",
      "votes": 1,
      "replies": [
        {
          "id": 949132,
          "author_name": "Tsai29",
          "author_url": "",
          "post_date": "2020-07-28T12:51:05.750000",
          "content": "<p>Thanks! :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 945004,
      "author_name": "JJShadow",
      "author_url": "",
      "post_date": "2020-07-25T13:50:00.183000",
      "content": "<p>Thank you for sharing.. \nRadboud is calcuated with Huber_loss is quite good.. It is good strategy..\nCould I ask three things?\n1. Why do you use huber_loss not smoothl1loss.(I want to know difference huber_loss and smoothl1loss)\n2. What is roll of cencer_mask in radboud_loss and Karolinska_loss function..(guess.. kind of mask between radboud and Karolinska?)\n3. ColorStain augmentation is really work? how much improvment is for score?</p>\n\n<p>Thank you for sharing :) Congratulation~</p>",
      "votes": 1,
      "replies": [
        {
          "id": 945547,
          "author_name": "Tsai29",
          "author_url": "",
          "post_date": "2020-07-26T00:07:43.897000",
          "content": "<p>Hi <a href=\"/rvslight\">@rvslight</a> : </p>\n\n<p>Thank you :)</p>\n\n<ol>\n<li><p>Since there are still some data in radboud is worth to learn, and if you also limit the model to learn from the good part of radboud data, the model won't perform very well on karolinska data either. And the huber can let instances with absolute difference less than delta has bigger loss, which will make model pay more attention on clean data of radboud more.</p></li>\n<li><p>Yes, you are correct. The center_mask is the mask for radboud/karolinska instances. Zeros in mask are corresponding to radboud instances, and ones corresponding to karolinska instances.</p></li>\n<li><p>To be honest, ColorStain augmentation didn't improve the cv or the LB, maybe just a little bit improvement. But I noticed that model will remember the color pattern of radboud and karolinska if you didn't perform H&amp;E color jitter on them. I was afraid this might impact the generalization ability of my model. So I added ColorStain augmentation in data pipeline just in case.</p></li>\n</ol>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 944941,
      "author_name": "Innat",
      "author_url": "",
      "post_date": "2020-07-25T13:17:53.160000",
      "content": "<p><a href=\"/xiejialun\">@xiejialun</a> \nCongrats. That's a really great achievement. 🎉 \nI think the main reason you got leverage from TPU was to use <code>.tfrec</code> file format. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 945548,
          "author_name": "Tsai29",
          "author_url": "",
          "post_date": "2020-07-26T00:08:49.700000",
          "content": "<p>Thanks!\nYeah, .tfrec really speed up the training process on TPU!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 943116,
      "author_name": "Kapil Kaushik",
      "author_url": "",
      "post_date": "2020-07-24T06:35:49.540000",
      "content": "<p>Thankyou for sharing</p>",
      "votes": 1,
      "replies": [
        {
          "id": 944238,
          "author_name": "Tsai29",
          "author_url": "",
          "post_date": "2020-07-25T00:33:19.220000",
          "content": "<p>No problem :)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 942928,
      "author_name": "Alin Cijov",
      "author_url": "",
      "post_date": "2020-07-24T04:23:47.847000",
      "content": "<p>Thank you for sharing and congratulations!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 944239,
          "author_name": "Tsai29",
          "author_url": "",
          "post_date": "2020-07-25T00:33:30.647000",
          "content": "<p>Thanks! :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 948544,
      "author_name": "SarvagyaGupta",
      "author_url": "",
      "post_date": "2020-07-28T03:55:20.177000",
      "content": "<p>Hi,</p>\n\n<p>Thank you for posting the notebook. I wrote the code for my first competition (Google landmark detection). But to train the model, each epoch was taking about 12 hours (even after reducing the image size to 200X200). It had a lot of images and classes, probably that's why it's slow. So how would you train for such a competition?</p>\n\n<p>Please let me know.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 949158,
          "author_name": "Tsai29",
          "author_url": "",
          "post_date": "2020-07-28T12:59:10.800000",
          "content": "<p>Hi,\nFirst of all, are you sure you opened the GPU before you started to train the model?\nSince with resolution this small, the training time shouldn't be this long.</p>\n\n<p>I think there are few things you could check:\n1. Whether you opened the GPU or not ( you can open the GPU on the right side of kaggle notebook)\n2. Make sure you uses the batch size big enough ( I always use the biggest batch size which GPU allowed, then fine tune the batch size to get balance model performance and training time)\n3. Check whether the data pipeline is work normally. You can get the data from your data pipeline and make sure there aren't any bugs in it before training your model.\n4. Try to reduce the parameters of your model to check whether the speed of training is increasing.</p>\n\n<p>Hope you solve your problem well.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "941916": "First of all, I want to thank kaggle and the organizers for hosting such interesting competition.\nAnd I want to thank my teammates, they really did a great job during the competition. \nSince I don't have my own machine, so I only trained my model on kaggle/colab TPU. But in fact, TPU is pretty enough to get 0.92+ private score. I think the most important part of my single model is the loss function, by using separated losses for each data center can handle the label noise of radboud quite well.\n<br>\n\n## strategy\nThere are two main challenges of this competition, and these challenges already mentioned in the overview of this competition. <br>\n1. Huge image size\n2. noisy data\n\n**Challenge 1** <br>\nI followed what @lafoss shared and add some simple preprocessing to get the proper training images. (Thanks a lot @lafoss) <br>\n\n**Challenge 2** <br>\nFor the challenge 2, I think there are some clues we could found in the documents shared by host and some topics in discussion section. <br>\n\na. The radboud data is more noisy than karolinska <br>\nb. The qwk of radboud training data is around ~0.85, which is judged by human <br>\nc. @lafoss got his 0.9lb on low radboud cv but high karolinska cv <br>\n\nSo in my point of view, we shouldn't have radboud cv too high, since this might be some overfitting signals of your model. And the reasonable cv of radboud should keep in 0.84~0.86, since the statement b above. At the beginning, I also can't find any pattern between cv and LB. So I used logcoah loss on both center of data, which is the loss function relativly noise robust than MSE . And after I saw the individual cv of @lafoss's, I realized that I should use some methods to limit the model to learn too much about radboud data and learn more about karolinska data. So I used huber loss(which was suggested by my good teammate @rguo97)  with delta 1 on radboud instances and MSE on karolinska instances. This did keep my radboud cv won't over 0.86 and the cv of karolinska did increased. And this was the time I got my 0.903LB, at the 15 days before the competition end. Well, this also related to what split your folding is, but I think this is the small pattern I thought is right about cv and lb relation. <br>\nBut of course, this isn't enough for getting 2nd place on private board, I think the main reason we got the 2nd place is because all our model have quite nice single model result, and the diversity of our ensemble model bring us to stable LB/PB relation.\n\n## Data prepare/preprocessing:\nI also turned the slides image into tiles, but instead of using MIL method, I glued the tiles into a big image. The tile size is 128x128 and total 144 tiles will be extract from a slide image. Then I glued the 144 tiles into a big image, which means I turned a slide into 1536x1536 image. But before gluing the image, I will get the foreground of each tile and resize back to 128x128. In the end, I wrote every glued images into tfrecord format after I encoded them in jpeg format.\n\n## Data augmentation\n1. H&amp;E color jitter (0.25 random prob)\n2. Random contrast/brightness (0.25 random prob)\n3. Random saturation (0.15 random prob)\n4. Image transpose (0.5 random prob)\n5. Random H/V flip\n6. Shift(50pixels)/Rotate(10degrees)/Scale(0.05) (0.25 random prob)\n<br>\n\n## Model\nEfficientnet b3 + GeM + 0.3 dropout + 128 units dense layer + 1 unit dense layer(regression head)\n<br>\n\n## Loss function\nMSE loss on karolinska data\nHuber loss with delta 1 on radboud data\n\n## Training\nOne cycle cosine annealing with initial learning rate 5e-4 and minimum learning rate 5e-6, and train for 40 epochs. Also I saved the checkpoint for radboud and karolinska separately, so sometimes I got two best checkpoints on radboud and karolinska ceneter. And I normally used best checkpoint of karolinska only in inference.\n\n## Single model score\n0.905 qwk cv on karolinska/0.853 qwk cv on radboud/0.891 qwk overall cv <br>\n0.903 qwk on LB/0.927 qwk on PB\n\n## Final models\nPlease refer this topic for our ensemble model\nhttps://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/169108\n\n\n-----------------------------\nHere is the training notebook of my model.\nhttps://www.kaggle.com/xiejialun/my-part-of-2nd-place-solution-tpu/notebook",
    "948958": "Well deserved, thank you for sharing the information in such neat and detailed manner. Looking forward to many more in future.👍 ",
    "945004": "Thank you for sharing.. \nRadboud is calcuated with Huber_loss is quite good.. It is good strategy..\nCould I ask three things?\n1. Why do you use huber_loss not smoothl1loss.(I want to know difference huber_loss and smoothl1loss)\n2. What is roll of cencer_mask in radboud_loss and Karolinska_loss function..(guess.. kind of mask between radboud and Karolinska?)\n3. ColorStain augmentation is really work? how much improvment is for score?\n\nThank you for sharing :) Congratulation~",
    "944941": "@xiejialun \nCongrats. That's a really great achievement. 🎉 \nI think the main reason you got leverage from TPU was to use `.tfrec` file format. ",
    "943116": "Thankyou for sharing",
    "942928": "Thank you for sharing and congratulations!",
    "948544": "Hi,\n\nThank you for posting the notebook. I wrote the code for my first competition (Google landmark detection). But to train the model, each epoch was taking about 12 hours (even after reducing the image size to 200X200). It had a lot of images and classes, probably that's why it's slow. So how would you train for such a competition?\n\nPlease let me know."
  }
}