{
  "id": 147076,
  "title": "Tissue Detection and Size Optimization ~80% Shrink",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/147076",
  "author_name": "Zac Dannelly",
  "post_date": "2020-04-29T12:05:23.270000",
  "votes": 5,
  "comment_count": 0,
  "views": 0,
  "content": "<p>I have made a <a href=\"https://www.kaggle.com/dannellyz/tissue-detection-and-size-optimization-70-shrink\">kernel</a> that implements the research completed in the pre-processing step of <a href=\"https://arxiv.org/pdf/1606.05718.pdf\">Deep Learning for Identifying Metastatic Breast Cancer</a>. This work identifies regions of the slides that have tissue and optimally trims away those sections with information relating to the target variable. I further augmented this with optimized bounding boxes and space trimming as seen below to reach up to 90% size reduction on the test cases.</p>\n\n<p><strong>Example</strong>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1342122%2Ff9fd5b4492dbaed9bd305c8a322d8955%2FScreen%20Shot%202020-04-29%20at%207.57.25%20AM.png?generation=1588161794118623&amp;alt=media\" alt=\"\"></p>\n\n<p>By using this the hope would be high fidelity images could be used with no increase of storage capacity, but rather just more densely packed information. Since orientation is not a factor in predicting the target variable this work should only serve to better inform models.</p>\n\n<p>My next steps are:\n1. Test methodology on masks and check for information retention\n2. Get baseline memory reduction across entire corpus\n3. Speed tests and efficiency increase\n4. Generating a dataset from lower downsampling for API access.</p>\n\n<p>I would greatly welcome recomendations or other next steps as they would be useful to the work being done on this challenge. Just let me know!</p>",
  "messages": [
    {
      "id": 826037,
      "postDate": "2020-04-29T12:05:23.270Z",
      "content": "<p>I have made a <a href=\"https://www.kaggle.com/dannellyz/tissue-detection-and-size-optimization-70-shrink\">kernel</a> that implements the research completed in the pre-processing step of <a href=\"https://arxiv.org/pdf/1606.05718.pdf\">Deep Learning for Identifying Metastatic Breast Cancer</a>. This work identifies regions of the slides that have tissue and optimally trims away those sections with information relating to the target variable. I further augmented this with optimized bounding boxes and space trimming as seen below to reach up to 90% size reduction on the test cases.</p>\n\n<p><strong>Example</strong>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1342122%2Ff9fd5b4492dbaed9bd305c8a322d8955%2FScreen%20Shot%202020-04-29%20at%207.57.25%20AM.png?generation=1588161794118623&amp;alt=media\" alt=\"\"></p>\n\n<p>By using this the hope would be high fidelity images could be used with no increase of storage capacity, but rather just more densely packed information. Since orientation is not a factor in predicting the target variable this work should only serve to better inform models.</p>\n\n<p>My next steps are:\n1. Test methodology on masks and check for information retention\n2. Get baseline memory reduction across entire corpus\n3. Speed tests and efficiency increase\n4. Generating a dataset from lower downsampling for API access.</p>\n\n<p>I would greatly welcome recomendations or other next steps as they would be useful to the work being done on this challenge. Just let me know!</p>",
      "rawMarkdown": "I have made a [kernel](https://www.kaggle.com/dannellyz/tissue-detection-and-size-optimization-70-shrink) that implements the research completed in the pre-processing step of [Deep Learning for Identifying Metastatic Breast Cancer](https://arxiv.org/pdf/1606.05718.pdf). This work identifies regions of the slides that have tissue and optimally trims away those sections with information relating to the target variable. I further augmented this with optimized bounding boxes and space trimming as seen below to reach up to 90% size reduction on the test cases.\n\n**Example**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1342122%2Ff9fd5b4492dbaed9bd305c8a322d8955%2FScreen%20Shot%202020-04-29%20at%207.57.25%20AM.png?generation=1588161794118623&amp;alt=media)\n\n\nBy using this the hope would be high fidelity images could be used with no increase of storage capacity, but rather just more densely packed information. Since orientation is not a factor in predicting the target variable this work should only serve to better inform models.\n\nMy next steps are:\n1. Test methodology on masks and check for information retention\n2. Get baseline memory reduction across entire corpus\n3. Speed tests and efficiency increase\n4. Generating a dataset from lower downsampling for API access.\n\nI would greatly welcome recomendations or other next steps as they would be useful to the work being done on this challenge. Just let me know!\n",
      "votes": 5
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "826037": "I have made a [kernel](https://www.kaggle.com/dannellyz/tissue-detection-and-size-optimization-70-shrink) that implements the research completed in the pre-processing step of [Deep Learning for Identifying Metastatic Breast Cancer](https://arxiv.org/pdf/1606.05718.pdf). This work identifies regions of the slides that have tissue and optimally trims away those sections with information relating to the target variable. I further augmented this with optimized bounding boxes and space trimming as seen below to reach up to 90% size reduction on the test cases.\n\n**Example**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1342122%2Ff9fd5b4492dbaed9bd305c8a322d8955%2FScreen%20Shot%202020-04-29%20at%207.57.25%20AM.png?generation=1588161794118623&amp;alt=media)\n\n\nBy using this the hope would be high fidelity images could be used with no increase of storage capacity, but rather just more densely packed information. Since orientation is not a factor in predicting the target variable this work should only serve to better inform models.\n\nMy next steps are:\n1. Test methodology on masks and check for information retention\n2. Get baseline memory reduction across entire corpus\n3. Speed tests and efficiency increase\n4. Generating a dataset from lower downsampling for API access.\n\nI would greatly welcome recomendations or other next steps as they would be useful to the work being done on this challenge. Just let me know!\n"
  }
}