{
  "id": 169100,
  "title": "Lightgmb crop-wise solution (public 0.9)",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/169100",
  "author_name": "Artur Fattakhov (MIPT DIHT)",
  "post_date": "2020-07-23T00:02:38.519000",
  "votes": 7,
  "comment_count": 0,
  "views": 0,
  "content": "<h2>Crop-wise model</h2>\n\n<p>I think my approach is different from approaches presented in public kernels, because I did not use an end-to-end model that makes predictions throughout the slide.\nInstead I trained network on single 512x512 crops from level 0 to predict ISUP-grade. Of course, it is impossible to accurately predict the class using only one crop from slide, but this way you can get a model that gives out good crop embeddings without troubles with batch-size/resolution. For training I use effnetb5 with softmax activation and crops with cancer area more then 10%(any type of cancer).</p>\n\n<h2>Aggregation part</h2>\n\n<p>After training I split the entire slide into crops(left only those in which the proportion of tissue was above 10%) and separately predict embeddings and probabilities.\nI use embeddings and probabilities to aggregate some features for lightgbm model:\n- element-wise statistics on embeddings(min, max, std, median)\n- statistics on probas\n- the number of crops with a certain class\n- attention-like features: obviously, not all crops are equally important for predicting the ISUP. So, instead of mean, i use weighted sum where weight is sum of ISUP probabilities(I have it from my effb5 crop-wise model). Therefore, crops containing only healthy tissue will have a low weight, while crops containing a lot of cancerous tissue will have a high weight.</p>\n\n<p>I used out of the box lightgbm classification model to predict final slide-level probabilities.\nSince I trained the crop-wise model in whole train, there was a leak in the features, but since strong augmentations were used, it was insignificant.\nInstead of regression, the final prediction was sum(i*p_i) with [0.5, 1.5, 2.5, 3.5, 4.5] threasholds(i - the i'th ISUP class, p_i - the lightgbm probability of this class)\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F744554%2F25836385b85525212c3acd86ba8e8459%2Farch.png?generation=1595457060891999&amp;alt=media\" alt=\"\"></p>\n\n<h2>Training dataset modifications</h2>\n\n<p>As I know, there was some mistakes in training dataset, so i tried to fix them:\nIf I have, for example slide with 4+5, I can add embeddings from another (4+5) slide, and new sub-slide will have the same 4+5 class.\nFurthermore 4+5 means that class 4 have highter area then class 5, but pathologists could be wrong in determining areas, so i can add 4+4 class from another slide to make my current slide more specific.\nAlso I can add 0+0 tissue to every slide and it will not be harmful\nI used this augmentation-like techniques to expand dataset for 200000 rows(instead of default 10000)\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F744554%2F489fd9a4595a8973800f1ba367aaf59d%2Fdata_aug.png?generation=1595457693350028&amp;alt=media\" alt=\"\"></p>\n\n<h2>Embedding TTA</h2>\n\n<p>During inference I can make predictions 100 times(get for example random 90% of embeddings) and average probabilities.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F744554%2F3de0730d0fff89bac6f719e4094f6cab%2Ftta.png?generation=1595458209610605&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 940417,
      "postDate": "2020-07-23T00:02:38.520Z",
      "content": "<h2>Crop-wise model</h2>\n\n<p>I think my approach is different from approaches presented in public kernels, because I did not use an end-to-end model that makes predictions throughout the slide.\nInstead I trained network on single 512x512 crops from level 0 to predict ISUP-grade. Of course, it is impossible to accurately predict the class using only one crop from slide, but this way you can get a model that gives out good crop embeddings without troubles with batch-size/resolution. For training I use effnetb5 with softmax activation and crops with cancer area more then 10%(any type of cancer).</p>\n\n<h2>Aggregation part</h2>\n\n<p>After training I split the entire slide into crops(left only those in which the proportion of tissue was above 10%) and separately predict embeddings and probabilities.\nI use embeddings and probabilities to aggregate some features for lightgbm model:\n- element-wise statistics on embeddings(min, max, std, median)\n- statistics on probas\n- the number of crops with a certain class\n- attention-like features: obviously, not all crops are equally important for predicting the ISUP. So, instead of mean, i use weighted sum where weight is sum of ISUP probabilities(I have it from my effb5 crop-wise model). Therefore, crops containing only healthy tissue will have a low weight, while crops containing a lot of cancerous tissue will have a high weight.</p>\n\n<p>I used out of the box lightgbm classification model to predict final slide-level probabilities.\nSince I trained the crop-wise model in whole train, there was a leak in the features, but since strong augmentations were used, it was insignificant.\nInstead of regression, the final prediction was sum(i*p_i) with [0.5, 1.5, 2.5, 3.5, 4.5] threasholds(i - the i'th ISUP class, p_i - the lightgbm probability of this class)\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F744554%2F25836385b85525212c3acd86ba8e8459%2Farch.png?generation=1595457060891999&amp;alt=media\" alt=\"\"></p>\n\n<h2>Training dataset modifications</h2>\n\n<p>As I know, there was some mistakes in training dataset, so i tried to fix them:\nIf I have, for example slide with 4+5, I can add embeddings from another (4+5) slide, and new sub-slide will have the same 4+5 class.\nFurthermore 4+5 means that class 4 have highter area then class 5, but pathologists could be wrong in determining areas, so i can add 4+4 class from another slide to make my current slide more specific.\nAlso I can add 0+0 tissue to every slide and it will not be harmful\nI used this augmentation-like techniques to expand dataset for 200000 rows(instead of default 10000)\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F744554%2F489fd9a4595a8973800f1ba367aaf59d%2Fdata_aug.png?generation=1595457693350028&amp;alt=media\" alt=\"\"></p>\n\n<h2>Embedding TTA</h2>\n\n<p>During inference I can make predictions 100 times(get for example random 90% of embeddings) and average probabilities.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F744554%2F3de0730d0fff89bac6f719e4094f6cab%2Ftta.png?generation=1595458209610605&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "##Crop-wise model\nI think my approach is different from approaches presented in public kernels, because I did not use an end-to-end model that makes predictions throughout the slide.\nInstead I trained network on single 512x512 crops from level 0 to predict ISUP-grade. Of course, it is impossible to accurately predict the class using only one crop from slide, but this way you can get a model that gives out good crop embeddings without troubles with batch-size/resolution. For training I use effnetb5 with softmax activation and crops with cancer area more then 10%(any type of cancer).\n##Aggregation part\nAfter training I split the entire slide into crops(left only those in which the proportion of tissue was above 10%) and separately predict embeddings and probabilities.\nI use embeddings and probabilities to aggregate some features for lightgbm model:\n- element-wise statistics on embeddings(min, max, std, median)\n- statistics on probas\n- the number of crops with a certain class\n- attention-like features: obviously, not all crops are equally important for predicting the ISUP. So, instead of mean, i use weighted sum where weight is sum of ISUP probabilities(I have it from my effb5 crop-wise model). Therefore, crops containing only healthy tissue will have a low weight, while crops containing a lot of cancerous tissue will have a high weight.\n\nI used out of the box lightgbm classification model to predict final slide-level probabilities.\nSince I trained the crop-wise model in whole train, there was a leak in the features, but since strong augmentations were used, it was insignificant.\nInstead of regression, the final prediction was sum(i*p_i) with [0.5, 1.5, 2.5, 3.5, 4.5] threasholds(i - the i'th ISUP class, p_i - the lightgbm probability of this class)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F744554%2F25836385b85525212c3acd86ba8e8459%2Farch.png?generation=1595457060891999&amp;alt=media)\n\n\n##Training dataset modifications\nAs I know, there was some mistakes in training dataset, so i tried to fix them:\nIf I have, for example slide with 4+5, I can add embeddings from another (4+5) slide, and new sub-slide will have the same 4+5 class.\nFurthermore 4+5 means that class 4 have highter area then class 5, but pathologists could be wrong in determining areas, so i can add 4+4 class from another slide to make my current slide more specific.\nAlso I can add 0+0 tissue to every slide and it will not be harmful\nI used this augmentation-like techniques to expand dataset for 200000 rows(instead of default 10000)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F744554%2F489fd9a4595a8973800f1ba367aaf59d%2Fdata_aug.png?generation=1595457693350028&amp;alt=media)\n\n##Embedding TTA\nDuring inference I can make predictions 100 times(get for example random 90% of embeddings) and average probabilities.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F744554%2F3de0730d0fff89bac6f719e4094f6cab%2Ftta.png?generation=1595458209610605&amp;alt=media)\n",
      "votes": 7
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "940417": "##Crop-wise model\nI think my approach is different from approaches presented in public kernels, because I did not use an end-to-end model that makes predictions throughout the slide.\nInstead I trained network on single 512x512 crops from level 0 to predict ISUP-grade. Of course, it is impossible to accurately predict the class using only one crop from slide, but this way you can get a model that gives out good crop embeddings without troubles with batch-size/resolution. For training I use effnetb5 with softmax activation and crops with cancer area more then 10%(any type of cancer).\n##Aggregation part\nAfter training I split the entire slide into crops(left only those in which the proportion of tissue was above 10%) and separately predict embeddings and probabilities.\nI use embeddings and probabilities to aggregate some features for lightgbm model:\n- element-wise statistics on embeddings(min, max, std, median)\n- statistics on probas\n- the number of crops with a certain class\n- attention-like features: obviously, not all crops are equally important for predicting the ISUP. So, instead of mean, i use weighted sum where weight is sum of ISUP probabilities(I have it from my effb5 crop-wise model). Therefore, crops containing only healthy tissue will have a low weight, while crops containing a lot of cancerous tissue will have a high weight.\n\nI used out of the box lightgbm classification model to predict final slide-level probabilities.\nSince I trained the crop-wise model in whole train, there was a leak in the features, but since strong augmentations were used, it was insignificant.\nInstead of regression, the final prediction was sum(i*p_i) with [0.5, 1.5, 2.5, 3.5, 4.5] threasholds(i - the i'th ISUP class, p_i - the lightgbm probability of this class)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F744554%2F25836385b85525212c3acd86ba8e8459%2Farch.png?generation=1595457060891999&amp;alt=media)\n\n\n##Training dataset modifications\nAs I know, there was some mistakes in training dataset, so i tried to fix them:\nIf I have, for example slide with 4+5, I can add embeddings from another (4+5) slide, and new sub-slide will have the same 4+5 class.\nFurthermore 4+5 means that class 4 have highter area then class 5, but pathologists could be wrong in determining areas, so i can add 4+4 class from another slide to make my current slide more specific.\nAlso I can add 0+0 tissue to every slide and it will not be harmful\nI used this augmentation-like techniques to expand dataset for 200000 rows(instead of default 10000)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F744554%2F489fd9a4595a8973800f1ba367aaf59d%2Fdata_aug.png?generation=1595457693350028&amp;alt=media)\n\n##Embedding TTA\nDuring inference I can make predictions 100 times(get for example random 90% of embeddings) and average probabilities.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F744554%2F3de0730d0fff89bac6f719e4094f6cab%2Ftta.png?generation=1595458209610605&amp;alt=media)\n"
  }
}