{
  "id": 169342,
  "title": "Silver medal solution -> 24 place",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/169342",
  "author_name": "Vlad Vaduva",
  "post_date": "2020-07-23T15:24:27.036000",
  "votes": 13,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Hello to all participants, organizers and people who are here to learn from the successful solutions or to exchange opinions</p>\n\n<p>This was an interesting competition with a lot of challenge even for the experimented computer vision programmers. The main challenge was the fact that the provided data was limited and noisy\nI am glad with the result, this being the 3rd consecutive Kaggle medal (after gold at Bengali and bronze at M5).\nI will mention some of the things I tried here and a comment about how that implemented technique worked for me.</p>\n\n<p><strong>Pre-processing techniques</strong>\na) Tiles design\n- 256x256x36 tiles  <em>(best results)</em>\n- 256x256x49 tiles  <em>(a lot of white space and image was too big to sustained with the number of provided training data)</em>\n- 128x128x144 tiles <em>(lower results than the 256x256x36 tiles, probably some patterns are interupting by making the slides smaller)</em>\n- 512x512x9 tiles <em>(also lower results than 256x256x36, the tiles being so big they had a lot of white space when select them)</em>\nAdditional comment: <em>256x256x36 seemed to be the sweet point between making the tiles too big where they will have a lot of empty pixels and 128x128x144  where the tiles patterns are interupted</em></p>\n\n<p>b) Data selection\nDue to noisy labels I design a system to eliminate data which had the biggest probability of being labeled wrong. I used the best 5 fold ensable to predict on training data, average the prediction and eliminate data where the difference between prediction and real label was biggest than a specific threhold\nThe thresholds tested were:\n- 2  (137 data eliminated) \n- 3 (39 data eliminated)\n- 4 (7 data eliminated)</p>\n\n<p>Results comment: <em>Best CV results were obtain where I eliminated data where abs(prediction-truth labels)&gt;3 where 39 data were eliminated</em></p>\n\n<p><strong>Model arhitectures</strong>\n- Efficientnet B0\n- Efficientnet B1\n- Efficientnet B2\n- Efficientnet B3\n- Efficientnet B4\n- SE_Resnext50</p>\n\n<p>Result comment: <em>Best result was on B2 arhitecture, B3-B4 lead to overfit and the rest of them simply did not worked for me</em></p>\n\n<p><strong>Data Augmentation</strong>\n- Transpose\n- Vertical Flip\n- Horizontal Flip\n- RandomRotate90\n- Rotate at a random angle\n- Shift starting position pad when making tiles</p>\n\n<p>Result comment:  <em>All the augmentation were made at 2 levels: tile level + after the tile ensamble\nIn the best results I used: Transpose + Vertical Flip + Horizontal Flip</em></p>\n\n<p><strong>Optimizer + scheduler</strong>\n- Adam + GradualWarmupScheduler + CosineAnnealingLR\n- RangerLars + GradualWarmupScheduler + CosineAnnealingLR\n- RangerLars + OneCycleLR\n- Adam + GradualWarmupScheduler + ReduceLROnPlateau</p>\n\n<p>Result comment: Although I tried to use RangerLars  in a variety of combination of schedulers, best result was obtain with the good old Adam (Adam + GradualWarmupScheduler + CosineAnnealingLR)</p>\n\n<p><strong>TTA</strong>\nTTA composed of:\n- Transpose\n- Vertical Flip\n- Horizontal Flip\n- Shift starting position pad when making tiles</p>\n\n<p>Result comment: <em>Usually I recommend to use on TTA similar techniques with what you are using on training augmentation, otherwise the prediction will see images that were not taught by the model. This is why I used:Transpose + Vertical + Horizontal Flip</em></p>\n\n<p>Congratulations to all participants and to the organizers !!!</p>\n\n<p>See you on the next computer vision competition !</p>",
  "messages": [
    {
      "id": 942099,
      "postDate": "2020-07-23T15:24:27.037Z",
      "content": "<p>Hello to all participants, organizers and people who are here to learn from the successful solutions or to exchange opinions</p>\n\n<p>This was an interesting competition with a lot of challenge even for the experimented computer vision programmers. The main challenge was the fact that the provided data was limited and noisy\nI am glad with the result, this being the 3rd consecutive Kaggle medal (after gold at Bengali and bronze at M5).\nI will mention some of the things I tried here and a comment about how that implemented technique worked for me.</p>\n\n<p><strong>Pre-processing techniques</strong>\na) Tiles design\n- 256x256x36 tiles  <em>(best results)</em>\n- 256x256x49 tiles  <em>(a lot of white space and image was too big to sustained with the number of provided training data)</em>\n- 128x128x144 tiles <em>(lower results than the 256x256x36 tiles, probably some patterns are interupting by making the slides smaller)</em>\n- 512x512x9 tiles <em>(also lower results than 256x256x36, the tiles being so big they had a lot of white space when select them)</em>\nAdditional comment: <em>256x256x36 seemed to be the sweet point between making the tiles too big where they will have a lot of empty pixels and 128x128x144  where the tiles patterns are interupted</em></p>\n\n<p>b) Data selection\nDue to noisy labels I design a system to eliminate data which had the biggest probability of being labeled wrong. I used the best 5 fold ensable to predict on training data, average the prediction and eliminate data where the difference between prediction and real label was biggest than a specific threhold\nThe thresholds tested were:\n- 2  (137 data eliminated) \n- 3 (39 data eliminated)\n- 4 (7 data eliminated)</p>\n\n<p>Results comment: <em>Best CV results were obtain where I eliminated data where abs(prediction-truth labels)&gt;3 where 39 data were eliminated</em></p>\n\n<p><strong>Model arhitectures</strong>\n- Efficientnet B0\n- Efficientnet B1\n- Efficientnet B2\n- Efficientnet B3\n- Efficientnet B4\n- SE_Resnext50</p>\n\n<p>Result comment: <em>Best result was on B2 arhitecture, B3-B4 lead to overfit and the rest of them simply did not worked for me</em></p>\n\n<p><strong>Data Augmentation</strong>\n- Transpose\n- Vertical Flip\n- Horizontal Flip\n- RandomRotate90\n- Rotate at a random angle\n- Shift starting position pad when making tiles</p>\n\n<p>Result comment:  <em>All the augmentation were made at 2 levels: tile level + after the tile ensamble\nIn the best results I used: Transpose + Vertical Flip + Horizontal Flip</em></p>\n\n<p><strong>Optimizer + scheduler</strong>\n- Adam + GradualWarmupScheduler + CosineAnnealingLR\n- RangerLars + GradualWarmupScheduler + CosineAnnealingLR\n- RangerLars + OneCycleLR\n- Adam + GradualWarmupScheduler + ReduceLROnPlateau</p>\n\n<p>Result comment: Although I tried to use RangerLars  in a variety of combination of schedulers, best result was obtain with the good old Adam (Adam + GradualWarmupScheduler + CosineAnnealingLR)</p>\n\n<p><strong>TTA</strong>\nTTA composed of:\n- Transpose\n- Vertical Flip\n- Horizontal Flip\n- Shift starting position pad when making tiles</p>\n\n<p>Result comment: <em>Usually I recommend to use on TTA similar techniques with what you are using on training augmentation, otherwise the prediction will see images that were not taught by the model. This is why I used:Transpose + Vertical + Horizontal Flip</em></p>\n\n<p>Congratulations to all participants and to the organizers !!!</p>\n\n<p>See you on the next computer vision competition !</p>",
      "rawMarkdown": "Hello to all participants, organizers and people who are here to learn from the successful solutions or to exchange opinions\n\nThis was an interesting competition with a lot of challenge even for the experimented computer vision programmers. The main challenge was the fact that the provided data was limited and noisy\nI am glad with the result, this being the 3rd consecutive Kaggle medal (after gold at Bengali and bronze at M5).\nI will mention some of the things I tried here and a comment about how that implemented technique worked for me.\n\n**Pre-processing techniques**\na) Tiles design\n- 256x256x36 tiles  *(best results)*\n- 256x256x49 tiles  *(a lot of white space and image was too big to sustained with the number of provided training data)*\n- 128x128x144 tiles *(lower results than the 256x256x36 tiles, probably some patterns are interupting by making the slides smaller)*\n- 512x512x9 tiles *(also lower results than 256x256x36, the tiles being so big they had a lot of white space when select them)*\nAdditional comment: *256x256x36 seemed to be the sweet point between making the tiles too big where they will have a lot of empty pixels and 128x128x144  where the tiles patterns are interupted*\n\nb) Data selection\nDue to noisy labels I design a system to eliminate data which had the biggest probability of being labeled wrong. I used the best 5 fold ensable to predict on training data, average the prediction and eliminate data where the difference between prediction and real label was biggest than a specific threhold\nThe thresholds tested were:\n- 2  (137 data eliminated) \n- 3 (39 data eliminated)\n- 4 (7 data eliminated)\n\nResults comment: *Best CV results were obtain where I eliminated data where abs(prediction-truth labels)&gt;3 where 39 data were eliminated*\n\n**Model arhitectures**\n- Efficientnet B0\n- Efficientnet B1\n- Efficientnet B2\n- Efficientnet B3\n- Efficientnet B4\n- SE_Resnext50\n\nResult comment: *Best result was on B2 arhitecture, B3-B4 lead to overfit and the rest of them simply did not worked for me*\n\n**Data Augmentation**\n- Transpose\n- Vertical Flip\n- Horizontal Flip\n- RandomRotate90\n- Rotate at a random angle\n- Shift starting position pad when making tiles\n\nResult comment:  *All the augmentation were made at 2 levels: tile level + after the tile ensamble\nIn the best results I used: Transpose + Vertical Flip + Horizontal Flip*\n\n**Optimizer + scheduler**\n- Adam + GradualWarmupScheduler + CosineAnnealingLR\n- RangerLars + GradualWarmupScheduler + CosineAnnealingLR\n- RangerLars + OneCycleLR\n- Adam + GradualWarmupScheduler + ReduceLROnPlateau\n\nResult comment: Although I tried to use RangerLars  in a variety of combination of schedulers, best result was obtain with the good old Adam (Adam + GradualWarmupScheduler + CosineAnnealingLR)\n\n\n**TTA**\nTTA composed of:\n- Transpose\n- Vertical Flip\n- Horizontal Flip\n- Shift starting position pad when making tiles\n\nResult comment: *Usually I recommend to use on TTA similar techniques with what you are using on training augmentation, otherwise the prediction will see images that were not taught by the model. This is why I used:Transpose + Vertical + Horizontal Flip*\n\n\nCongratulations to all participants and to the organizers !!!\n\nSee you on the next computer vision competition !",
      "votes": 13
    },
    {
      "id": 946783,
      "postDate": "2020-07-26T20:26:27.433Z",
      "content": "<p>Congratulations, your solutions always seem to exhibit solid computer vision fundamentals.</p>",
      "rawMarkdown": "Congratulations, your solutions always seem to exhibit solid computer vision fundamentals.",
      "votes": 1,
      "replies": [
        {
          "id": 947996,
          "postDate": "2020-07-27T15:45:12.437Z",
          "content": "<p>Thank you <a href=\"/yousof9\">@yousof9</a> . It is my profession, this is what I do where I work</p>",
          "rawMarkdown": "Thank you @yousof9 . It is my profession, this is what I do where I work",
          "votes": 1
        }
      ]
    },
    {
      "id": 944230,
      "postDate": "2020-07-25T00:10:11.480Z",
      "content": "<p><a href=\"/vladvdv\">@vladvdv</a> Thanks for solution sharing and the kind words under my solution topic :D! I want to consult you a question.😄\nIf I use some lr-scheduler, such as  onecycle and cosineannealing for training a pre-set value of epoch(assume 30). After ending training, the scheduler already reached the min-lr. But I find the model isn't converge well. Is there any proper ways to reuse these model weights for further training? I try set a small lr(magnitude about 1e-06), but the result isn't so good. Thanks in advance :D</p>",
      "rawMarkdown": "@vladvdv Thanks for solution sharing and the kind words under my solution topic :D! I want to consult you a question.😄\nIf I use some lr-scheduler, such as  onecycle and cosineannealing for training a pre-set value of epoch(assume 30). After ending training, the scheduler already reached the min-lr. But I find the model isn't converge well. Is there any proper ways to reuse these model weights for further training? I try set a small lr(magnitude about 1e-06), but the result isn't so good. Thanks in advance :D",
      "votes": 1,
      "replies": [
        {
          "id": 944638,
          "postDate": "2020-07-25T08:14:48.737Z",
          "content": "<p><a href=\"/cnzengshiyuan\">@cnzengshiyuan</a>  . After you trained a model with onecycle or cosineannealing the final lr at the end is very small. If you want to retrain this model after you finish the first training you have 2 possibilities:\n1) Use the same lr like the end of the first training session or even smaller and it should work but it will take some time to see improvements (small lr -&gt; very small steps)\n2) Use a bigger lr but this is very risky. It can drag you out of the local minimum you have found instead of further exploring it and the loss will increase. Another possibility is to increase a little bit the lr and on the retrain come with a cosine anneling, to somehow simulate the CosineAnnealingWarmRestarts </p>",
          "rawMarkdown": "@cnzengshiyuan  . After you trained a model with onecycle or cosineannealing the final lr at the end is very small. If you want to retrain this model after you finish the first training you have 2 possibilities:\n1) Use the same lr like the end of the first training session or even smaller and it should work but it will take some time to see improvements (small lr -&gt; very small steps)\n2) Use a bigger lr but this is very risky. It can drag you out of the local minimum you have found instead of further exploring it and the loss will increase. Another possibility is to increase a little bit the lr and on the retrain come with a cosine anneling, to somehow simulate the CosineAnnealingWarmRestarts \n",
          "votes": 1
        },
        {
          "id": 944857,
          "postDate": "2020-07-25T11:49:34.363Z",
          "content": "<p><a href=\"/vladvdv\">@vladvdv</a> , OK, I will try and do some experiments further. Thanks for your advice :D</p>",
          "rawMarkdown": "@vladvdv , OK, I will try and do some experiments further. Thanks for your advice :D",
          "votes": 1
        },
        {
          "id": 945214,
          "postDate": "2020-07-25T16:55:16.713Z",
          "content": "<p>My pleasure, good luck with the experiments !</p>",
          "rawMarkdown": "My pleasure, good luck with the experiments !",
          "votes": 1
        },
        {
          "id": 945639,
          "postDate": "2020-07-26T03:39:13.760Z",
          "content": "<p>Thanks ! 🤝 🤝 🤝 </p>",
          "rawMarkdown": "Thanks ! 🤝 🤝 🤝 ",
          "votes": 1
        }
      ]
    },
    {
      "id": 942159,
      "postDate": "2020-07-23T16:03:09.400Z",
      "content": "<p>Interesting solution - yours is a Silver medal solution, not a bronze one.</p>\n\n<p>Are you focussing only on Computer vision competitions?</p>",
      "rawMarkdown": "Interesting solution - yours is a Silver medal solution, not a bronze one.\n\nAre you focussing only on Computer vision competitions?",
      "votes": 2,
      "replies": [
        {
          "id": 942171,
          "postDate": "2020-07-23T16:08:34.927Z",
          "content": "<p><a href=\"/kurianbenoy\">@kurianbenoy</a> My bad, I corrected the title now. Thanks for the observation\nComputer vision is my best area of expertise but I sometimes participate on others also if they seem interesting. Any new thing that I can learn is a useful thing</p>",
          "rawMarkdown": "@kurianbenoy My bad, I corrected the title now. Thanks for the observation\nComputer vision is my best area of expertise but I sometimes participate on others also if they seem interesting. Any new thing that I can learn is a useful thing",
          "votes": 2
        },
        {
          "id": 942356,
          "postDate": "2020-07-23T17:53:49.410Z",
          "content": "<p>Congrats👍 </p>",
          "rawMarkdown": "Congrats👍 "
        },
        {
          "id": 942505,
          "postDate": "2020-07-23T19:30:36.523Z",
          "content": "<p>Thank you <a href=\"/spidyweb\">@spidyweb</a> </p>",
          "rawMarkdown": "Thank you @spidyweb "
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 946783,
      "author_name": "Yousef Rabi",
      "author_url": "",
      "post_date": "2020-07-26T20:26:27.433000",
      "content": "<p>Congratulations, your solutions always seem to exhibit solid computer vision fundamentals.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 947996,
          "author_name": "Vlad Vaduva",
          "author_url": "",
          "post_date": "2020-07-27T15:45:12.437000",
          "content": "<p>Thank you <a href=\"/yousof9\">@yousof9</a> . It is my profession, this is what I do where I work</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 944230,
      "author_name": "Shiyuan Zeng",
      "author_url": "",
      "post_date": "2020-07-25T00:10:11.480000",
      "content": "<p><a href=\"/vladvdv\">@vladvdv</a> Thanks for solution sharing and the kind words under my solution topic :D! I want to consult you a question.😄\nIf I use some lr-scheduler, such as  onecycle and cosineannealing for training a pre-set value of epoch(assume 30). After ending training, the scheduler already reached the min-lr. But I find the model isn't converge well. Is there any proper ways to reuse these model weights for further training? I try set a small lr(magnitude about 1e-06), but the result isn't so good. Thanks in advance :D</p>",
      "votes": 1,
      "replies": [
        {
          "id": 944638,
          "author_name": "Vlad Vaduva",
          "author_url": "",
          "post_date": "2020-07-25T08:14:48.737000",
          "content": "<p><a href=\"/cnzengshiyuan\">@cnzengshiyuan</a>  . After you trained a model with onecycle or cosineannealing the final lr at the end is very small. If you want to retrain this model after you finish the first training you have 2 possibilities:\n1) Use the same lr like the end of the first training session or even smaller and it should work but it will take some time to see improvements (small lr -&gt; very small steps)\n2) Use a bigger lr but this is very risky. It can drag you out of the local minimum you have found instead of further exploring it and the loss will increase. Another possibility is to increase a little bit the lr and on the retrain come with a cosine anneling, to somehow simulate the CosineAnnealingWarmRestarts </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 944857,
          "author_name": "Shiyuan Zeng",
          "author_url": "",
          "post_date": "2020-07-25T11:49:34.363000",
          "content": "<p><a href=\"/vladvdv\">@vladvdv</a> , OK, I will try and do some experiments further. Thanks for your advice :D</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 945214,
          "author_name": "Vlad Vaduva",
          "author_url": "",
          "post_date": "2020-07-25T16:55:16.713000",
          "content": "<p>My pleasure, good luck with the experiments !</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 945639,
          "author_name": "Shiyuan Zeng",
          "author_url": "",
          "post_date": "2020-07-26T03:39:13.760000",
          "content": "<p>Thanks ! 🤝 🤝 🤝 </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 942159,
      "author_name": "Kurian Benoy",
      "author_url": "",
      "post_date": "2020-07-23T16:03:09.400000",
      "content": "<p>Interesting solution - yours is a Silver medal solution, not a bronze one.</p>\n\n<p>Are you focussing only on Computer vision competitions?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 942171,
          "author_name": "Vlad Vaduva",
          "author_url": "",
          "post_date": "2020-07-23T16:08:34.927000",
          "content": "<p><a href=\"/kurianbenoy\">@kurianbenoy</a> My bad, I corrected the title now. Thanks for the observation\nComputer vision is my best area of expertise but I sometimes participate on others also if they seem interesting. Any new thing that I can learn is a useful thing</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 942356,
          "author_name": "Rajeev",
          "author_url": "",
          "post_date": "2020-07-23T17:53:49.410000",
          "content": "<p>Congrats👍 </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 942505,
          "author_name": "Vlad Vaduva",
          "author_url": "",
          "post_date": "2020-07-23T19:30:36.523000",
          "content": "<p>Thank you <a href=\"/spidyweb\">@spidyweb</a> </p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "942099": "Hello to all participants, organizers and people who are here to learn from the successful solutions or to exchange opinions\n\nThis was an interesting competition with a lot of challenge even for the experimented computer vision programmers. The main challenge was the fact that the provided data was limited and noisy\nI am glad with the result, this being the 3rd consecutive Kaggle medal (after gold at Bengali and bronze at M5).\nI will mention some of the things I tried here and a comment about how that implemented technique worked for me.\n\n**Pre-processing techniques**\na) Tiles design\n- 256x256x36 tiles  *(best results)*\n- 256x256x49 tiles  *(a lot of white space and image was too big to sustained with the number of provided training data)*\n- 128x128x144 tiles *(lower results than the 256x256x36 tiles, probably some patterns are interupting by making the slides smaller)*\n- 512x512x9 tiles *(also lower results than 256x256x36, the tiles being so big they had a lot of white space when select them)*\nAdditional comment: *256x256x36 seemed to be the sweet point between making the tiles too big where they will have a lot of empty pixels and 128x128x144  where the tiles patterns are interupted*\n\nb) Data selection\nDue to noisy labels I design a system to eliminate data which had the biggest probability of being labeled wrong. I used the best 5 fold ensable to predict on training data, average the prediction and eliminate data where the difference between prediction and real label was biggest than a specific threhold\nThe thresholds tested were:\n- 2  (137 data eliminated) \n- 3 (39 data eliminated)\n- 4 (7 data eliminated)\n\nResults comment: *Best CV results were obtain where I eliminated data where abs(prediction-truth labels)&gt;3 where 39 data were eliminated*\n\n**Model arhitectures**\n- Efficientnet B0\n- Efficientnet B1\n- Efficientnet B2\n- Efficientnet B3\n- Efficientnet B4\n- SE_Resnext50\n\nResult comment: *Best result was on B2 arhitecture, B3-B4 lead to overfit and the rest of them simply did not worked for me*\n\n**Data Augmentation**\n- Transpose\n- Vertical Flip\n- Horizontal Flip\n- RandomRotate90\n- Rotate at a random angle\n- Shift starting position pad when making tiles\n\nResult comment:  *All the augmentation were made at 2 levels: tile level + after the tile ensamble\nIn the best results I used: Transpose + Vertical Flip + Horizontal Flip*\n\n**Optimizer + scheduler**\n- Adam + GradualWarmupScheduler + CosineAnnealingLR\n- RangerLars + GradualWarmupScheduler + CosineAnnealingLR\n- RangerLars + OneCycleLR\n- Adam + GradualWarmupScheduler + ReduceLROnPlateau\n\nResult comment: Although I tried to use RangerLars  in a variety of combination of schedulers, best result was obtain with the good old Adam (Adam + GradualWarmupScheduler + CosineAnnealingLR)\n\n\n**TTA**\nTTA composed of:\n- Transpose\n- Vertical Flip\n- Horizontal Flip\n- Shift starting position pad when making tiles\n\nResult comment: *Usually I recommend to use on TTA similar techniques with what you are using on training augmentation, otherwise the prediction will see images that were not taught by the model. This is why I used:Transpose + Vertical + Horizontal Flip*\n\n\nCongratulations to all participants and to the organizers !!!\n\nSee you on the next computer vision competition !",
    "946783": "Congratulations, your solutions always seem to exhibit solid computer vision fundamentals.",
    "944230": "@vladvdv Thanks for solution sharing and the kind words under my solution topic :D! I want to consult you a question.😄\nIf I use some lr-scheduler, such as  onecycle and cosineannealing for training a pre-set value of epoch(assume 30). After ending training, the scheduler already reached the min-lr. But I find the model isn't converge well. Is there any proper ways to reuse these model weights for further training? I try set a small lr(magnitude about 1e-06), but the result isn't so good. Thanks in advance :D",
    "942159": "Interesting solution - yours is a Silver medal solution, not a bronze one.\n\nAre you focussing only on Computer vision competitions?"
  }
}