{
  "id": 465719,
  "title": "A Bronze Solution",
  "url": "/competitions/UBC-OCEAN/discussion/465719",
  "author_name": "Manuel K",
  "post_date": "2024-01-05T11:26:20.767000",
  "votes": 11,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hey.<br>\nFirst, I want to thank the competition hosts for this challenge - I learnt a lot and it was nice to use actual histological data. Event though I \"only\" have a bronze medal, I thought it might be interesting for some people to see my approach.<br>\nI have written a more detailled report on my website if somebody wants to have all the details (it's in German but Google Translate does a god job): <a href=\"https://manuelk-net.translate.goog/portfolio/UBC_KI.html?_x_tr_sl=de&amp;_x_tr_tl=en&amp;_x_tr_hl=en-US&amp;_x_tr_pto=wapp\" target=\"_blank\">Complete case study</a></p>\n<h2>Overview of the approach: The 2 Stage Model</h2>\n<h3>Data extraction</h3>\n<p>I tried a lot of different things, but in the end my own prerequisites I set for the solution were as follows:</p>\n<ul>\n<li>The classification can only work on a cellular scale. A macroscopic image of the whole slide won't provide nearly as much information as needed.</li>\n<li>There are cancerous and non-cancerous areas in most of the slides. I have to extract relevant patches and train a model only on these.</li>\n</ul>\n<p>I guess most of the other competitors would agree when I say, that a big challenge was the data preprocessing. The method I used is based on the masks provided, so I wasn't able to use all the images. The process is as follows:</p>\n<ul>\n<li>The original image is divided into N x M patches of size 512px. This is done without any spacing, meaning the patches are right next to each other.</li>\n<li>The tissue type is read from the mask and the patch is sorted into a corresponding folder</li>\n</ul>\n<p>The procedure becomes clear if you look at the following example:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14931205%2F420d56830587ac5694d68037acb741c2%2Fimage11.png?generation=1704451956255218&amp;alt=media\" alt=\"\"></p>\n<h3>Detector and Classifier</h3>\n<p>Now that we have clear information about which patches show healthy tissue and which show malignant tissue, my approach is a 2 stage model. So two stages that carry out the following classifications one after the other:</p>\n<ol>\n<li>Detector: A neural network detects patches that show malignant tissue. This is then a binary classification of all patches.</li>\n<li>Classifier: A second neural network determines the subtype (HGSC, MC, …) of all patches detected as malignant.</li>\n</ol>\n<p>Number 1 is of course the easier task, but the above requirement should be met here: The classifier only sees actually malignant tissue and not healthy patches.</p>\n<p>EdgeNext is used for both stages, as it already delivered very good in earlier experiments.</p>\n<p>Approximately 75% accuracy was achieved in the validation data set. Interestingly, this is only slightly better than compared to my first model (~72%) which used randomly selected patches.</p>\n<p>The following procedure was then used for inference:</p>\n<ol>\n<li>M x N patches with a size of 512px are extracted from each slice image.</li>\n<li>Completely dark or light patches are sorted out</li>\n<li>Detection: The detection model determines whether all relevant patches are cancerous tissue.</li>\n<li>Classification: The Classifier Model uses all cancerous patches and determines the subtypes.</li>\n<li>The result is a list of predictions of all malignant patches. The average value is determined from this and written into the output table as the final result.</li>\n</ol>\n<p>The final result on the private test set is <strong>0.44</strong>. Unfortunately, the best model I choose on the public LB wasn't the best on the private. Theoretically, the best score was 0.48.</p>\n<h2>What didn't work</h2>\n<ul>\n<li><p>I tried Multi Instance Learning with a neural network as feature extractor and Boosting, SVM or a RNN. All of these experiments gave good results but were always a few percent worse compared to the 2 Stage approach.</p></li>\n<li><p>I used external dataset (e.g. the <a href=\"https://wiki.cancerimagingarchive.net/pages/viewpage.action?pageId=83593077\" target=\"_blank\">Bevacizumab Response Dataset</a> ) ,but I never got any improvement on the score. Maybe these images differ too much from the ones we had here.</p></li>\n<li><p>Simple novelty detection. Based on the feature extractor, I tried a One-Class SVM and Isolation Forest to detect the \"Other\" class but there was no improvement.</p></li>\n</ul>\n<h2>How to improve</h2>\n<p>Two things: Outlier detection and generalization. I didn't spent much time on the outlier detection part, but from the solutions I've seen so far it can improve the results a lot.</p>\n<p>In all my experiments the difference between train, validation and LB score was quite high. So it was always a mild form of overfitting and with more (labeled and segmented) data this could be reduced.</p>\n<p>In the end, I am very happy that I won my first bronze medal and congratulations to all other participants! The solutions I looked over from you guys were amazing.</p>",
  "messages": [
    {
      "id": 2588244,
      "postDate": "2024-01-05T11:26:20.767Z",
      "content": "<p>Hey.<br>\nFirst, I want to thank the competition hosts for this challenge - I learnt a lot and it was nice to use actual histological data. Event though I \"only\" have a bronze medal, I thought it might be interesting for some people to see my approach.<br>\nI have written a more detailled report on my website if somebody wants to have all the details (it's in German but Google Translate does a god job): <a href=\"https://manuelk-net.translate.goog/portfolio/UBC_KI.html?_x_tr_sl=de&amp;_x_tr_tl=en&amp;_x_tr_hl=en-US&amp;_x_tr_pto=wapp\" target=\"_blank\">Complete case study</a></p>\n<h2>Overview of the approach: The 2 Stage Model</h2>\n<h3>Data extraction</h3>\n<p>I tried a lot of different things, but in the end my own prerequisites I set for the solution were as follows:</p>\n<ul>\n<li>The classification can only work on a cellular scale. A macroscopic image of the whole slide won't provide nearly as much information as needed.</li>\n<li>There are cancerous and non-cancerous areas in most of the slides. I have to extract relevant patches and train a model only on these.</li>\n</ul>\n<p>I guess most of the other competitors would agree when I say, that a big challenge was the data preprocessing. The method I used is based on the masks provided, so I wasn't able to use all the images. The process is as follows:</p>\n<ul>\n<li>The original image is divided into N x M patches of size 512px. This is done without any spacing, meaning the patches are right next to each other.</li>\n<li>The tissue type is read from the mask and the patch is sorted into a corresponding folder</li>\n</ul>\n<p>The procedure becomes clear if you look at the following example:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14931205%2F420d56830587ac5694d68037acb741c2%2Fimage11.png?generation=1704451956255218&amp;alt=media\" alt=\"\"></p>\n<h3>Detector and Classifier</h3>\n<p>Now that we have clear information about which patches show healthy tissue and which show malignant tissue, my approach is a 2 stage model. So two stages that carry out the following classifications one after the other:</p>\n<ol>\n<li>Detector: A neural network detects patches that show malignant tissue. This is then a binary classification of all patches.</li>\n<li>Classifier: A second neural network determines the subtype (HGSC, MC, …) of all patches detected as malignant.</li>\n</ol>\n<p>Number 1 is of course the easier task, but the above requirement should be met here: The classifier only sees actually malignant tissue and not healthy patches.</p>\n<p>EdgeNext is used for both stages, as it already delivered very good in earlier experiments.</p>\n<p>Approximately 75% accuracy was achieved in the validation data set. Interestingly, this is only slightly better than compared to my first model (~72%) which used randomly selected patches.</p>\n<p>The following procedure was then used for inference:</p>\n<ol>\n<li>M x N patches with a size of 512px are extracted from each slice image.</li>\n<li>Completely dark or light patches are sorted out</li>\n<li>Detection: The detection model determines whether all relevant patches are cancerous tissue.</li>\n<li>Classification: The Classifier Model uses all cancerous patches and determines the subtypes.</li>\n<li>The result is a list of predictions of all malignant patches. The average value is determined from this and written into the output table as the final result.</li>\n</ol>\n<p>The final result on the private test set is <strong>0.44</strong>. Unfortunately, the best model I choose on the public LB wasn't the best on the private. Theoretically, the best score was 0.48.</p>\n<h2>What didn't work</h2>\n<ul>\n<li><p>I tried Multi Instance Learning with a neural network as feature extractor and Boosting, SVM or a RNN. All of these experiments gave good results but were always a few percent worse compared to the 2 Stage approach.</p></li>\n<li><p>I used external dataset (e.g. the <a href=\"https://wiki.cancerimagingarchive.net/pages/viewpage.action?pageId=83593077\" target=\"_blank\">Bevacizumab Response Dataset</a> ) ,but I never got any improvement on the score. Maybe these images differ too much from the ones we had here.</p></li>\n<li><p>Simple novelty detection. Based on the feature extractor, I tried a One-Class SVM and Isolation Forest to detect the \"Other\" class but there was no improvement.</p></li>\n</ul>\n<h2>How to improve</h2>\n<p>Two things: Outlier detection and generalization. I didn't spent much time on the outlier detection part, but from the solutions I've seen so far it can improve the results a lot.</p>\n<p>In all my experiments the difference between train, validation and LB score was quite high. So it was always a mild form of overfitting and with more (labeled and segmented) data this could be reduced.</p>\n<p>In the end, I am very happy that I won my first bronze medal and congratulations to all other participants! The solutions I looked over from you guys were amazing.</p>",
      "rawMarkdown": "Hey.\nFirst, I want to thank the competition hosts for this challenge - I learnt a lot and it was nice to use actual histological data. Event though I \"only\" have a bronze medal, I thought it might be interesting for some people to see my approach.\nI have written a more detailled report on my website if somebody wants to have all the details (it's in German but Google Translate does a god job): [Complete case study](https://manuelk-net.translate.goog/portfolio/UBC_KI.html?_x_tr_sl=de&_x_tr_tl=en&_x_tr_hl=en-US&_x_tr_pto=wapp)\n\n## Overview of the approach: The 2 Stage Model\n\n### Data extraction\n\nI tried a lot of different things, but in the end my own prerequisites I set for the solution were as follows:\n\n- The classification can only work on a cellular scale. A macroscopic image of the whole slide won't provide nearly as much information as needed.\n- There are cancerous and non-cancerous areas in most of the slides. I have to extract relevant patches and train a model only on these.\n\nI guess most of the other competitors would agree when I say, that a big challenge was the data preprocessing. The method I used is based on the masks provided, so I wasn't able to use all the images. The process is as follows:\n\n- The original image is divided into N x M patches of size 512px. This is done without any spacing, meaning the patches are right next to each other.\n- The tissue type is read from the mask and the patch is sorted into a corresponding folder\n\nThe procedure becomes clear if you look at the following example:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14931205%2F420d56830587ac5694d68037acb741c2%2Fimage11.png?generation=1704451956255218&alt=media)\n\n### Detector and Classifier\n\nNow that we have clear information about which patches show healthy tissue and which show malignant tissue, my approach is a 2 stage model. So two stages that carry out the following classifications one after the other:\n\n1. Detector: A neural network detects patches that show malignant tissue. This is then a binary classification of all patches.\n2. Classifier: A second neural network determines the subtype (HGSC, MC, ...) of all patches detected as malignant.\n\nNumber 1 is of course the easier task, but the above requirement should be met here: The classifier only sees actually malignant tissue and not healthy patches.\n\nEdgeNext is used for both stages, as it already delivered very good in earlier experiments.\n\nApproximately 75% accuracy was achieved in the validation data set. Interestingly, this is only slightly better than compared to my first model (~72%) which used randomly selected patches.\n\nThe following procedure was then used for inference:\n\n1. M x N patches with a size of 512px are extracted from each slice image.\n2. Completely dark or light patches are sorted out\n3. Detection: The detection model determines whether all relevant patches are cancerous tissue.\n4. Classification: The Classifier Model uses all cancerous patches and determines the subtypes.\n5. The result is a list of predictions of all malignant patches. The average value is determined from this and written into the output table as the final result.\n\nThe final result on the private test set is **0.44**. Unfortunately, the best model I choose on the public LB wasn't the best on the private. Theoretically, the best score was 0.48.\n\n## What didn't work\n\n- I tried Multi Instance Learning with a neural network as feature extractor and Boosting, SVM or a RNN. All of these experiments gave good results but were always a few percent worse compared to the 2 Stage approach.\n\n- I used external dataset (e.g. the [Bevacizumab Response Dataset](https://wiki.cancerimagingarchive.net/pages/viewpage.action?pageId=83593077) ) ,but I never got any improvement on the score. Maybe these images differ too much from the ones we had here.\n\n- Simple novelty detection. Based on the feature extractor, I tried a One-Class SVM and Isolation Forest to detect the \"Other\" class but there was no improvement.\n\n## How to improve\n\nTwo things: Outlier detection and generalization. I didn't spent much time on the outlier detection part, but from the solutions I've seen so far it can improve the results a lot.\n\nIn all my experiments the difference between train, validation and LB score was quite high. So it was always a mild form of overfitting and with more (labeled and segmented) data this could be reduced.\n\nIn the end, I am very happy that I won my first bronze medal and congratulations to all other participants! The solutions I looked over from you guys were amazing.",
      "votes": 10
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2588244": "Hey.\nFirst, I want to thank the competition hosts for this challenge - I learnt a lot and it was nice to use actual histological data. Event though I \"only\" have a bronze medal, I thought it might be interesting for some people to see my approach.\nI have written a more detailled report on my website if somebody wants to have all the details (it's in German but Google Translate does a god job): [Complete case study](https://manuelk-net.translate.goog/portfolio/UBC_KI.html?_x_tr_sl=de&_x_tr_tl=en&_x_tr_hl=en-US&_x_tr_pto=wapp)\n\n## Overview of the approach: The 2 Stage Model\n\n### Data extraction\n\nI tried a lot of different things, but in the end my own prerequisites I set for the solution were as follows:\n\n- The classification can only work on a cellular scale. A macroscopic image of the whole slide won't provide nearly as much information as needed.\n- There are cancerous and non-cancerous areas in most of the slides. I have to extract relevant patches and train a model only on these.\n\nI guess most of the other competitors would agree when I say, that a big challenge was the data preprocessing. The method I used is based on the masks provided, so I wasn't able to use all the images. The process is as follows:\n\n- The original image is divided into N x M patches of size 512px. This is done without any spacing, meaning the patches are right next to each other.\n- The tissue type is read from the mask and the patch is sorted into a corresponding folder\n\nThe procedure becomes clear if you look at the following example:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14931205%2F420d56830587ac5694d68037acb741c2%2Fimage11.png?generation=1704451956255218&alt=media)\n\n### Detector and Classifier\n\nNow that we have clear information about which patches show healthy tissue and which show malignant tissue, my approach is a 2 stage model. So two stages that carry out the following classifications one after the other:\n\n1. Detector: A neural network detects patches that show malignant tissue. This is then a binary classification of all patches.\n2. Classifier: A second neural network determines the subtype (HGSC, MC, ...) of all patches detected as malignant.\n\nNumber 1 is of course the easier task, but the above requirement should be met here: The classifier only sees actually malignant tissue and not healthy patches.\n\nEdgeNext is used for both stages, as it already delivered very good in earlier experiments.\n\nApproximately 75% accuracy was achieved in the validation data set. Interestingly, this is only slightly better than compared to my first model (~72%) which used randomly selected patches.\n\nThe following procedure was then used for inference:\n\n1. M x N patches with a size of 512px are extracted from each slice image.\n2. Completely dark or light patches are sorted out\n3. Detection: The detection model determines whether all relevant patches are cancerous tissue.\n4. Classification: The Classifier Model uses all cancerous patches and determines the subtypes.\n5. The result is a list of predictions of all malignant patches. The average value is determined from this and written into the output table as the final result.\n\nThe final result on the private test set is **0.44**. Unfortunately, the best model I choose on the public LB wasn't the best on the private. Theoretically, the best score was 0.48.\n\n## What didn't work\n\n- I tried Multi Instance Learning with a neural network as feature extractor and Boosting, SVM or a RNN. All of these experiments gave good results but were always a few percent worse compared to the 2 Stage approach.\n\n- I used external dataset (e.g. the [Bevacizumab Response Dataset](https://wiki.cancerimagingarchive.net/pages/viewpage.action?pageId=83593077) ) ,but I never got any improvement on the score. Maybe these images differ too much from the ones we had here.\n\n- Simple novelty detection. Based on the feature extractor, I tried a One-Class SVM and Isolation Forest to detect the \"Other\" class but there was no improvement.\n\n## How to improve\n\nTwo things: Outlier detection and generalization. I didn't spent much time on the outlier detection part, but from the solutions I've seen so far it can improve the results a lot.\n\nIn all my experiments the difference between train, validation and LB score was quite high. So it was always a mild form of overfitting and with more (labeled and segmented) data this could be reduced.\n\nIn the end, I am very happy that I won my first bronze medal and congratulations to all other participants! The solutions I looked over from you guys were amazing."
  }
}