{
  "id": 169938,
  "title": "Private Leader Board QWK Score 93.0 ",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/169938",
  "author_name": "Jaideep",
  "post_date": "2020-07-25T18:50:37.544000",
  "votes": -2,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Congratulations to those who were able to retain their top lb positions based on the best selections for private data set . All those like me who couldnt make it to private lb even though with good private lb score unselected,its always a better next time. </p>\n\n<p>I suggest kaggle team &amp; host team to enable automatic selection of  best submissions based on Highest Private lb score to avoid such shake up in future which should also ensure best solution kernel  are not left out because of them not getting selected by participant. In real world i suppose selections for productions are not made like this. </p>\n\n<h1>Tile Extraction</h1>\n\n<p>This is base method published generously by <a href=\"/iafoss\">@iafoss</a>  kudos \nfurther refined that using below kernel\n<a href=\"https://www.kaggle.com/akensert/panda-optimized-tiling-tf-data-dataset\">https://www.kaggle.com/akensert/panda-optimized-tiling-tf-data-dataset</a> \noptimized compute parameter  as below to cover up more n more tissues regions</p>\n\n<p><code>def compute_coords(image,\n                   patch_size=256,\n                   precompute=False,\n                   min_patch_info=0.15,\n                   min_axis_info=0.15,\n                   min_consec_axis_info=0.35,\n                   min_decimal_keep=0.7):\n</code></p>\n\n<h1>Loss</h1>\n\n<p>Multilabel Multiclass as explaned in <a href=\"/haqishen\">@haqishen</a>  kernel </p>\n\n<h1>Highest Image Size during Training</h1>\n\n<p>6 X6 X256</p>\n\n<h1>Sampling</h1>\n\n<p>Performed balanced sampling  of Ra and Ka sources. This helped in stablizing the loss</p>\n\n<h1>Training Methodology</h1>\n\n<ul>\n<li>FastAI framework</li>\n<li>Progressively increased the  number of tiles  like 9,16,25,36 as regularization method to prevent overfitting</li>\n<li>Albumentation augmentations  comprising  of Shift scale rotation, optical distortion,One of (Median Blur,Motion Blur,Blur),  RandomGamma, Clahe  and dihedral rotations.</li>\n<li><p>Since no of tiles for many WSI were more than 36 so i used grouping methodology to select first N slides and last N slides from Tile Pool of each WSI during training iteration at random. Eg if number of tiles of WSI is 48,grouping would select either 0 to N  or  Total Slide cnt -N to Total slides  at random. to cover up entire slide. This systematic grouping method gave some boost.</p></li>\n<li><p>Pseudo Labelling  : </p>\n\n<ol><li>Found out 359 WSI (139 ka,220 Ra) which were  having poor predictions . I concluded it by keeping them in validation   and found very poor qwk ,by excluding all these 359 slides cv QWK was always  above 91</li>\n<li>Replaced for subset of 359 WSI their actual predictions with model predictions such that   difference between actual and predicted   gt than 2</li>\n<li>Kept increasing the pseudo labelling progressively till found QWK reaching 93-94</li></ol>\n\n<p>This approach gave only 88  in public leader board but 93 in private leader board perhaps this \n was reason i dint select it :( . I wish i had invested some more time on  it ,probably would have landed up  further higher rankings. I suppose one of my post suggesting this became basis for top private lb positions during last days. I m glad it helped some.</p></li>\n</ul>\n\n<h1>Inference</h1>\n\n<ul>\n<li>Used image of size 7 X 7 X 280  but tile extracted of size 256 . This gave suddenly a big boost in score when i used it for first time</li>\n<li>During inference i ran two iteration of inference one for Group 0  and one for Group 1. Gave more weightage to Group 1 prediction because i took stats of cancerous tissue mean location using masks provided,mean location happened to lie at centre of image ,assumed same for test also. This also gave some boost. </li>\n<li>During each iteration  weighted ensemble  of  Non Pseudo labelled kernel  predictions fold 0 giving  highest public lb score of 90.6 (0.55)  and pseudo labelled kernel with public score of 88 (0.45). </li>\n<li>Took mean of albumentation augmentations ran through 4 times and affine based predictions (dihedral)</li>\n</ul>",
  "messages": [
    {
      "id": 946225,
      "postDate": "2020-07-26T12:52:49.053Z",
      "content": "<p>&gt;  In real world i suppose selections for productions are not made like this.</p>\n\n<p>Except the private data is supposed to simulate the real world scenario where you don't have access to the new, unseen data's labels. You will have to build a model that generalizes the best from what you have locally e.g training + public test set in Kaggle's case. \nEven that, you're still lucky that you even get to see the private score like in a ML competition. In the real world you will have to rely on A/B testing and other observable metrics (like how many users screamed YOUR RECOMMENDER SUCKS and voted the app 1 star in the last week).\nDon't get me started on a critical application like this case, you can't just deploy 241 models and see which one made the least people died.</p>",
      "rawMarkdown": "&gt;  In real world i suppose selections for productions are not made like this.\n\nExcept the private data is supposed to simulate the real world scenario where you don't have access to the new, unseen data's labels. You will have to build a model that generalizes the best from what you have locally e.g training + public test set in Kaggle's case. \nEven that, you're still lucky that you even get to see the private score like in a ML competition. In the real world you will have to rely on A/B testing and other observable metrics (like how many users screamed YOUR RECOMMENDER SUCKS and voted the app 1 star in the last week).\nDon't get me started on a critical application like this case, you can't just deploy 241 models and see which one made the least people died.",
      "votes": 13,
      "replies": [
        {
          "id": 946246,
          "postDate": "2020-07-26T13:06:32.080Z",
          "content": "<p>You   have real production data in test environment that user uses for final testing (eq to private set in kaggle)   . Test all your models ,put the one that give best outcome as per user acceptance criteria . Then once you deploy   rest of process follow. </p>",
          "rawMarkdown": "You   have real production data in test environment that user uses for final testing (eq to private set in kaggle)   . Test all your models ,put the one that give best outcome as per user acceptance criteria . Then once you deploy   rest of process follow. "
        },
        {
          "id": 946450,
          "postDate": "2020-07-26T15:31:51.137Z",
          "content": "<p>Here's a simple example.</p>\n\n<ul>\n<li><strong>Host</strong>: \"Can you give me a model to predict Tomorrow's Weather?\"\n<ul><li>In this sentence, <code>Tomorrow's Weather</code>=<code>Unseen data</code>=<code>Private LB</code></li></ul></li>\n<li><strong>Competition participants</strong>: “Here is my model\"\n<ul><li>What we chose by seeing local CV and Public LB and so on</li></ul></li>\n<li><strong>Your claim</strong>:  \"I don't know. I have lots of models to choose, so choose the one that best fits Tomorrow's Weather\"</li>\n<li><strong>Host</strong>: \"If we already know the Tomorrow's Weather, we're not going to have a competition!\"</li>\n</ul>",
          "rawMarkdown": "Here's a simple example.\n\n- **Host**: \"Can you give me a model to predict Tomorrow's Weather?\"\n   -  In this sentence, `Tomorrow's Weather`=`Unseen data`=`Private LB`\n- **Competition participants**: “Here is my model\"\n    - What we chose by seeing local CV and Public LB and so on\n- **Your claim**:  \"I don't know. I have lots of models to choose, so choose the one that best fits Tomorrow's Weather\"\n- **Host**: \"If we already know the Tomorrow's Weather, we're not going to have a competition!\"",
          "votes": 20
        },
        {
          "id": 953179,
          "postDate": "2020-07-31T15:27:10.813Z",
          "content": "<p><strong>\" don't know. I have lots of models to choose\"</strong>\n@famtaro first of all congrats to you and team for getting 1 st rank. \nIn reply to your comment ,No op this is definitely not the case here. We made over 100 submission , but confusion  always remains with those models whose scores close by   and  we dont know which could fare well over the private set ,there would be hardly any participant who dsnt knows which out of more than 5 to 10 models would give him best ,every one is able to fairly narrow down the list of their submission which they think would give best to below 10 or even less. In real world you are going to push  more than one models that are presumably best ones ,this is based on way  it happens in one of my colleagues company . Results from all are analysed to see which one is fitting best  . </p>",
          "rawMarkdown": "**\" don't know. I have lots of models to choose\"**\n@famtaro first of all congrats to you and team for getting 1 st rank. \nIn reply to your comment ,No op this is definitely not the case here. We made over 100 submission , but confusion  always remains with those models whose scores close by   and  we dont know which could fare well over the private set ,there would be hardly any participant who dsnt knows which out of more than 5 to 10 models would give him best ,every one is able to fairly narrow down the list of their submission which they think would give best to below 10 or even less. In real world you are going to push  more than one models that are presumably best ones ,this is based on way  it happens in one of my colleagues company . Results from all are analysed to see which one is fitting best  . \n\n",
          "votes": -1
        }
      ]
    },
    {
      "id": 946178,
      "postDate": "2020-07-26T12:19:45.093Z",
      "content": "<p>Thanks for your contribution, your questions and thoughts during the whole challenge were a great source of information. \nSo to sum-up you did not select the model which handles noisy labels but the model which got the best public LB ?</p>",
      "rawMarkdown": "Thanks for your contribution, your questions and thoughts during the whole challenge were a great source of information. \nSo to sum-up you did not select the model which handles noisy labels but the model which got the best public LB ?",
      "votes": 1
    },
    {
      "id": 945356,
      "postDate": "2020-07-25T18:50:37.543Z",
      "content": "<p>Congratulations to those who were able to retain their top lb positions based on the best selections for private data set . All those like me who couldnt make it to private lb even though with good private lb score unselected,its always a better next time. </p>\n\n<p>I suggest kaggle team &amp; host team to enable automatic selection of  best submissions based on Highest Private lb score to avoid such shake up in future which should also ensure best solution kernel  are not left out because of them not getting selected by participant. In real world i suppose selections for productions are not made like this. </p>\n\n<h1>Tile Extraction</h1>\n\n<p>This is base method published generously by <a href=\"/iafoss\">@iafoss</a>  kudos \nfurther refined that using below kernel\n<a href=\"https://www.kaggle.com/akensert/panda-optimized-tiling-tf-data-dataset\">https://www.kaggle.com/akensert/panda-optimized-tiling-tf-data-dataset</a> \noptimized compute parameter  as below to cover up more n more tissues regions</p>\n\n<p><code>def compute_coords(image,\n                   patch_size=256,\n                   precompute=False,\n                   min_patch_info=0.15,\n                   min_axis_info=0.15,\n                   min_consec_axis_info=0.35,\n                   min_decimal_keep=0.7):\n</code></p>\n\n<h1>Loss</h1>\n\n<p>Multilabel Multiclass as explaned in <a href=\"/haqishen\">@haqishen</a>  kernel </p>\n\n<h1>Highest Image Size during Training</h1>\n\n<p>6 X6 X256</p>\n\n<h1>Sampling</h1>\n\n<p>Performed balanced sampling  of Ra and Ka sources. This helped in stablizing the loss</p>\n\n<h1>Training Methodology</h1>\n\n<ul>\n<li>FastAI framework</li>\n<li>Progressively increased the  number of tiles  like 9,16,25,36 as regularization method to prevent overfitting</li>\n<li>Albumentation augmentations  comprising  of Shift scale rotation, optical distortion,One of (Median Blur,Motion Blur,Blur),  RandomGamma, Clahe  and dihedral rotations.</li>\n<li><p>Since no of tiles for many WSI were more than 36 so i used grouping methodology to select first N slides and last N slides from Tile Pool of each WSI during training iteration at random. Eg if number of tiles of WSI is 48,grouping would select either 0 to N  or  Total Slide cnt -N to Total slides  at random. to cover up entire slide. This systematic grouping method gave some boost.</p></li>\n<li><p>Pseudo Labelling  : </p>\n\n<ol><li>Found out 359 WSI (139 ka,220 Ra) which were  having poor predictions . I concluded it by keeping them in validation   and found very poor qwk ,by excluding all these 359 slides cv QWK was always  above 91</li>\n<li>Replaced for subset of 359 WSI their actual predictions with model predictions such that   difference between actual and predicted   gt than 2</li>\n<li>Kept increasing the pseudo labelling progressively till found QWK reaching 93-94</li></ol>\n\n<p>This approach gave only 88  in public leader board but 93 in private leader board perhaps this \n was reason i dint select it :( . I wish i had invested some more time on  it ,probably would have landed up  further higher rankings. I suppose one of my post suggesting this became basis for top private lb positions during last days. I m glad it helped some.</p></li>\n</ul>\n\n<h1>Inference</h1>\n\n<ul>\n<li>Used image of size 7 X 7 X 280  but tile extracted of size 256 . This gave suddenly a big boost in score when i used it for first time</li>\n<li>During inference i ran two iteration of inference one for Group 0  and one for Group 1. Gave more weightage to Group 1 prediction because i took stats of cancerous tissue mean location using masks provided,mean location happened to lie at centre of image ,assumed same for test also. This also gave some boost. </li>\n<li>During each iteration  weighted ensemble  of  Non Pseudo labelled kernel  predictions fold 0 giving  highest public lb score of 90.6 (0.55)  and pseudo labelled kernel with public score of 88 (0.45). </li>\n<li>Took mean of albumentation augmentations ran through 4 times and affine based predictions (dihedral)</li>\n</ul>",
      "rawMarkdown": "Congratulations to those who were able to retain their top lb positions based on the best selections for private data set . All those like me who couldnt make it to private lb even though with good private lb score unselected,its always a better next time. \n\nI suggest kaggle team &amp; host team to enable automatic selection of  best submissions based on Highest Private lb score to avoid such shake up in future which should also ensure best solution kernel  are not left out because of them not getting selected by participant. In real world i suppose selections for productions are not made like this. \n\n# Tile Extraction\nThis is base method published generously by @iafoss  kudos \nfurther refined that using below kernel\nhttps://www.kaggle.com/akensert/panda-optimized-tiling-tf-data-dataset \noptimized compute parameter  as below to cover up more n more tissues regions\n\n```def compute_coords(image,\n                   patch_size=256,\n                   precompute=False,\n                   min_patch_info=0.15,\n                   min_axis_info=0.15,\n                   min_consec_axis_info=0.35,\n                   min_decimal_keep=0.7):\n```\n\n \n# Loss\nMultilabel Multiclass as explaned in @haqishen  kernel \n\n\n \n# Highest Image Size during Training\n6 X6 X256\n\n\n# Sampling\nPerformed balanced sampling  of Ra and Ka sources. This helped in stablizing the loss\n\n\n#  Training Methodology\n\n- FastAI framework\n- Progressively increased the  number of tiles  like 9,16,25,36 as regularization method to prevent overfitting\n- Albumentation augmentations  comprising  of Shift scale rotation, optical distortion,One of (Median Blur,Motion Blur,Blur),  RandomGamma, Clahe  and dihedral rotations.\n- Since no of tiles for many WSI were more than 36 so i used grouping methodology to select first N slides and last N slides from Tile Pool of each WSI during training iteration at random. Eg if number of tiles of WSI is 48,grouping would select either 0 to N  or  Total Slide cnt -N to Total slides  at random. to cover up entire slide. This systematic grouping method gave some boost.\n\n- Pseudo Labelling  : \n1. Found out 359 WSI (139 ka,220 Ra) which were  having poor predictions . I concluded it by keeping them in validation   and found very poor qwk ,by excluding all these 359 slides cv QWK was always  above 91\n2.  Replaced for subset of 359 WSI their actual predictions with model predictions such that   difference between actual and predicted   gt than 2\n3.  Kept increasing the pseudo labelling progressively till found QWK reaching 93-94\n\n    This approach gave only 88  in public leader board but 93 in private leader board perhaps this \n     was reason i dint select it :( . I wish i had invested some more time on  it ,probably would have landed up  further higher rankings. I suppose one of my post suggesting this became basis for top private lb positions during last days. I m glad it helped some.\n\n# Inference\n- Used image of size 7 X 7 X 280  but tile extracted of size 256 . This gave suddenly a big boost in score when i used it for first time\n- During inference i ran two iteration of inference one for Group 0  and one for Group 1. Gave more weightage to Group 1 prediction because i took stats of cancerous tissue mean location using masks provided,mean location happened to lie at centre of image ,assumed same for test also. This also gave some boost. \n- During each iteration  weighted ensemble  of  Non Pseudo labelled kernel  predictions fold 0 giving  highest public lb score of 90.6 (0.55)  and pseudo labelled kernel with public score of 88 (0.45). \n- Took mean of albumentation augmentations ran through 4 times and affine based predictions (dihedral)\n",
      "votes": -2
    }
  ],
  "comments": [
    {
      "id": 946225,
      "author_name": "Khoi Nguyen",
      "author_url": "",
      "post_date": "2020-07-26T12:52:49.053000",
      "content": "<p>&gt;  In real world i suppose selections for productions are not made like this.</p>\n\n<p>Except the private data is supposed to simulate the real world scenario where you don't have access to the new, unseen data's labels. You will have to build a model that generalizes the best from what you have locally e.g training + public test set in Kaggle's case. \nEven that, you're still lucky that you even get to see the private score like in a ML competition. In the real world you will have to rely on A/B testing and other observable metrics (like how many users screamed YOUR RECOMMENDER SUCKS and voted the app 1 star in the last week).\nDon't get me started on a critical application like this case, you can't just deploy 241 models and see which one made the least people died.</p>",
      "votes": 13,
      "replies": [
        {
          "id": 946246,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-07-26T13:06:32.080000",
          "content": "<p>You   have real production data in test environment that user uses for final testing (eq to private set in kaggle)   . Test all your models ,put the one that give best outcome as per user acceptance criteria . Then once you deploy   rest of process follow. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 946450,
          "author_name": "fam_taro",
          "author_url": "",
          "post_date": "2020-07-26T15:31:51.137000",
          "content": "<p>Here's a simple example.</p>\n\n<ul>\n<li><strong>Host</strong>: \"Can you give me a model to predict Tomorrow's Weather?\"\n<ul><li>In this sentence, <code>Tomorrow's Weather</code>=<code>Unseen data</code>=<code>Private LB</code></li></ul></li>\n<li><strong>Competition participants</strong>: “Here is my model\"\n<ul><li>What we chose by seeing local CV and Public LB and so on</li></ul></li>\n<li><strong>Your claim</strong>:  \"I don't know. I have lots of models to choose, so choose the one that best fits Tomorrow's Weather\"</li>\n<li><strong>Host</strong>: \"If we already know the Tomorrow's Weather, we're not going to have a competition!\"</li>\n</ul>",
          "votes": 20,
          "replies": []
        },
        {
          "id": 953179,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-07-31T15:27:10.813000",
          "content": "<p><strong>\" don't know. I have lots of models to choose\"</strong>\n@famtaro first of all congrats to you and team for getting 1 st rank. \nIn reply to your comment ,No op this is definitely not the case here. We made over 100 submission , but confusion  always remains with those models whose scores close by   and  we dont know which could fare well over the private set ,there would be hardly any participant who dsnt knows which out of more than 5 to 10 models would give him best ,every one is able to fairly narrow down the list of their submission which they think would give best to below 10 or even less. In real world you are going to push  more than one models that are presumably best ones ,this is based on way  it happens in one of my colleagues company . Results from all are analysed to see which one is fitting best  . </p>",
          "votes": -1,
          "replies": []
        }
      ]
    },
    {
      "id": 946178,
      "author_name": "Alex",
      "author_url": "",
      "post_date": "2020-07-26T12:19:45.093000",
      "content": "<p>Thanks for your contribution, your questions and thoughts during the whole challenge were a great source of information. \nSo to sum-up you did not select the model which handles noisy labels but the model which got the best public LB ?</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "946225": "&gt;  In real world i suppose selections for productions are not made like this.\n\nExcept the private data is supposed to simulate the real world scenario where you don't have access to the new, unseen data's labels. You will have to build a model that generalizes the best from what you have locally e.g training + public test set in Kaggle's case. \nEven that, you're still lucky that you even get to see the private score like in a ML competition. In the real world you will have to rely on A/B testing and other observable metrics (like how many users screamed YOUR RECOMMENDER SUCKS and voted the app 1 star in the last week).\nDon't get me started on a critical application like this case, you can't just deploy 241 models and see which one made the least people died.",
    "946178": "Thanks for your contribution, your questions and thoughts during the whole challenge were a great source of information. \nSo to sum-up you did not select the model which handles noisy labels but the model which got the best public LB ?",
    "945356": "Congratulations to those who were able to retain their top lb positions based on the best selections for private data set . All those like me who couldnt make it to private lb even though with good private lb score unselected,its always a better next time. \n\nI suggest kaggle team &amp; host team to enable automatic selection of  best submissions based on Highest Private lb score to avoid such shake up in future which should also ensure best solution kernel  are not left out because of them not getting selected by participant. In real world i suppose selections for productions are not made like this. \n\n# Tile Extraction\nThis is base method published generously by @iafoss  kudos \nfurther refined that using below kernel\nhttps://www.kaggle.com/akensert/panda-optimized-tiling-tf-data-dataset \noptimized compute parameter  as below to cover up more n more tissues regions\n\n```def compute_coords(image,\n                   patch_size=256,\n                   precompute=False,\n                   min_patch_info=0.15,\n                   min_axis_info=0.15,\n                   min_consec_axis_info=0.35,\n                   min_decimal_keep=0.7):\n```\n\n \n# Loss\nMultilabel Multiclass as explaned in @haqishen  kernel \n\n\n \n# Highest Image Size during Training\n6 X6 X256\n\n\n# Sampling\nPerformed balanced sampling  of Ra and Ka sources. This helped in stablizing the loss\n\n\n#  Training Methodology\n\n- FastAI framework\n- Progressively increased the  number of tiles  like 9,16,25,36 as regularization method to prevent overfitting\n- Albumentation augmentations  comprising  of Shift scale rotation, optical distortion,One of (Median Blur,Motion Blur,Blur),  RandomGamma, Clahe  and dihedral rotations.\n- Since no of tiles for many WSI were more than 36 so i used grouping methodology to select first N slides and last N slides from Tile Pool of each WSI during training iteration at random. Eg if number of tiles of WSI is 48,grouping would select either 0 to N  or  Total Slide cnt -N to Total slides  at random. to cover up entire slide. This systematic grouping method gave some boost.\n\n- Pseudo Labelling  : \n1. Found out 359 WSI (139 ka,220 Ra) which were  having poor predictions . I concluded it by keeping them in validation   and found very poor qwk ,by excluding all these 359 slides cv QWK was always  above 91\n2.  Replaced for subset of 359 WSI their actual predictions with model predictions such that   difference between actual and predicted   gt than 2\n3.  Kept increasing the pseudo labelling progressively till found QWK reaching 93-94\n\n    This approach gave only 88  in public leader board but 93 in private leader board perhaps this \n     was reason i dint select it :( . I wish i had invested some more time on  it ,probably would have landed up  further higher rankings. I suppose one of my post suggesting this became basis for top private lb positions during last days. I m glad it helped some.\n\n# Inference\n- Used image of size 7 X 7 X 280  but tile extracted of size 256 . This gave suddenly a big boost in score when i used it for first time\n- During inference i ran two iteration of inference one for Group 0  and one for Group 1. Gave more weightage to Group 1 prediction because i took stats of cancerous tissue mean location using masks provided,mean location happened to lie at centre of image ,assumed same for test also. This also gave some boost. \n- During each iteration  weighted ensemble  of  Non Pseudo labelled kernel  predictions fold 0 giving  highest public lb score of 90.6 (0.55)  and pseudo labelled kernel with public score of 88 (0.45). \n- Took mean of albumentation augmentations ran through 4 times and affine based predictions (dihedral)\n"
  }
}