{
  "id": 359562,
  "title": "6th place solution",
  "url": "/competitions/mayo-clinic-strip-ai/discussion/359562",
  "author_name": "Kodai",
  "post_date": "2022-10-12T15:49:12.527000",
  "votes": 5,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hi, this competition is my first challenge at kaggle. I really enjoyed this competition and thought hard for about three months. I don't know why I could obtain a gold medal, but I'm sure it's thanks to everyone who have shared the notebooks or suggestions. <br>\nHere I describe my solution in brief.</p>\n<p>Submission kernel:<br>\n<a href=\"https://www.kaggle.com/code/kodaihatayama/mayo02-submission\" target=\"_blank\">https://www.kaggle.com/code/kodaihatayama/mayo02-submission</a></p>\n<h3>Preprocessing</h3>\n<p>I created tiles with the following steps. <strong>I didn't resize the image</strong> in all preprocessing steps.</p>\n<ul>\n<li>Cut 6 rectangles of width 512 at equal intervals from every image <br>\n(The purpose of this step is to create tiles quickly and avoid memory crashes without resizing.)</li>\n<li>Make tiles of size 512x512 from the above rectangles and pick up the top 8 dark tiles per image</li>\n<li>Split each tile to instances of size 32x32 for training</li>\n</ul>\n<p>For using pyvips, I refered to the following kernel. Thank you, <a href=\"https://www.kaggle.com/analokamus\" target=\"_blank\">@analokamus</a> !<br>\n<a href=\"https://www.kaggle.com/code/analokamus/how-to-use-pyvips-offline\" target=\"_blank\">https://www.kaggle.com/code/analokamus/how-to-use-pyvips-offline</a></p>\n<h3>Validation and Undersampling</h3>\n<p>I used the hold-out method with undersampling. </p>\n<ul>\n<li>Split patient id to train and validation (80:20) which was stratified by class (CE, LAA)</li>\n<li>Perform undersampling to be CE:LAA = 1:1 on the train dataset only</li>\n</ul>\n<h3>Model</h3>\n<p>I created the ensemble model of the following 2 methods.</p>\n<ul>\n<li>CNN </li>\n<li>lightGBM</li>\n</ul>\n<p>The difference of private LB between CNN only model and the ensemble model was a bit (approximately 0.676 → 0.668).<br>\nAs a side note, I created a model per institution (center id). (This approach may not be effective against test dataset because the institutions of test may be different from them of train.)</p>\n<p>For CNN, I refered to the following kernel. Thank you, <a href=\"https://www.kaggle.com/vbookshelf\" target=\"_blank\">@vbookshelf</a> !<br>\n<a href=\"https://www.kaggle.com/code/vbookshelf/cnn-how-to-use-160-000-images-without-crashing\" target=\"_blank\">https://www.kaggle.com/code/vbookshelf/cnn-how-to-use-160-000-images-without-crashing</a></p>\n<h3>What didn't work for me</h3>\n<ul>\n<li>grayscale</li>\n<li>histogram equalization</li>\n</ul>\n<p>Thanks for reading.</p>",
  "messages": [
    {
      "id": 1984324,
      "postDate": "2022-10-12T15:49:12.527Z",
      "content": "<p>Hi, this competition is my first challenge at kaggle. I really enjoyed this competition and thought hard for about three months. I don't know why I could obtain a gold medal, but I'm sure it's thanks to everyone who have shared the notebooks or suggestions. <br>\nHere I describe my solution in brief.</p>\n<p>Submission kernel:<br>\n<a href=\"https://www.kaggle.com/code/kodaihatayama/mayo02-submission\" target=\"_blank\">https://www.kaggle.com/code/kodaihatayama/mayo02-submission</a></p>\n<h3>Preprocessing</h3>\n<p>I created tiles with the following steps. <strong>I didn't resize the image</strong> in all preprocessing steps.</p>\n<ul>\n<li>Cut 6 rectangles of width 512 at equal intervals from every image <br>\n(The purpose of this step is to create tiles quickly and avoid memory crashes without resizing.)</li>\n<li>Make tiles of size 512x512 from the above rectangles and pick up the top 8 dark tiles per image</li>\n<li>Split each tile to instances of size 32x32 for training</li>\n</ul>\n<p>For using pyvips, I refered to the following kernel. Thank you, <a href=\"https://www.kaggle.com/analokamus\" target=\"_blank\">@analokamus</a> !<br>\n<a href=\"https://www.kaggle.com/code/analokamus/how-to-use-pyvips-offline\" target=\"_blank\">https://www.kaggle.com/code/analokamus/how-to-use-pyvips-offline</a></p>\n<h3>Validation and Undersampling</h3>\n<p>I used the hold-out method with undersampling. </p>\n<ul>\n<li>Split patient id to train and validation (80:20) which was stratified by class (CE, LAA)</li>\n<li>Perform undersampling to be CE:LAA = 1:1 on the train dataset only</li>\n</ul>\n<h3>Model</h3>\n<p>I created the ensemble model of the following 2 methods.</p>\n<ul>\n<li>CNN </li>\n<li>lightGBM</li>\n</ul>\n<p>The difference of private LB between CNN only model and the ensemble model was a bit (approximately 0.676 → 0.668).<br>\nAs a side note, I created a model per institution (center id). (This approach may not be effective against test dataset because the institutions of test may be different from them of train.)</p>\n<p>For CNN, I refered to the following kernel. Thank you, <a href=\"https://www.kaggle.com/vbookshelf\" target=\"_blank\">@vbookshelf</a> !<br>\n<a href=\"https://www.kaggle.com/code/vbookshelf/cnn-how-to-use-160-000-images-without-crashing\" target=\"_blank\">https://www.kaggle.com/code/vbookshelf/cnn-how-to-use-160-000-images-without-crashing</a></p>\n<h3>What didn't work for me</h3>\n<ul>\n<li>grayscale</li>\n<li>histogram equalization</li>\n</ul>\n<p>Thanks for reading.</p>",
      "rawMarkdown": "Hi, this competition is my first challenge at kaggle. I really enjoyed this competition and thought hard for about three months. I don't know why I could obtain a gold medal, but I'm sure it's thanks to everyone who have shared the notebooks or suggestions. \nHere I describe my solution in brief.\n\nSubmission kernel:\nhttps://www.kaggle.com/code/kodaihatayama/mayo02-submission\n\n### Preprocessing\nI created tiles with the following steps. **I didn't resize the image** in all preprocessing steps.\n- Cut 6 rectangles of width 512 at equal intervals from every image \n(The purpose of this step is to create tiles quickly and avoid memory crashes without resizing.)\n- Make tiles of size 512x512 from the above rectangles and pick up the top 8 dark tiles per image\n- Split each tile to instances of size 32x32 for training\n\nFor using pyvips, I refered to the following kernel. Thank you, @analokamus !\nhttps://www.kaggle.com/code/analokamus/how-to-use-pyvips-offline\n\n### Validation and Undersampling\nI used the hold-out method with undersampling. \n- Split patient id to train and validation (80:20) which was stratified by class (CE, LAA)\n- Perform undersampling to be CE:LAA = 1:1 on the train dataset only\n\n### Model\nI created the ensemble model of the following 2 methods.\n  - CNN \n  - lightGBM\n\nThe difference of private LB between CNN only model and the ensemble model was a bit (approximately 0.676 → 0.668).\nAs a side note, I created a model per institution (center id). (This approach may not be effective against test dataset because the institutions of test may be different from them of train.)\n\nFor CNN, I refered to the following kernel. Thank you, @vbookshelf !\nhttps://www.kaggle.com/code/vbookshelf/cnn-how-to-use-160-000-images-without-crashing\n\n### What didn't work for me\n- grayscale\n- histogram equalization\n\nThanks for reading.",
      "votes": 5
    },
    {
      "id": 1985062,
      "postDate": "2022-10-13T04:38:34.837Z",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/kodaihatayama\" target=\"_blank\">@kodaihatayama</a>. Solo gold in your first competition is a fantastic result!</p>",
      "rawMarkdown": "Congratulations @kodaihatayama. Solo gold in your first competition is a fantastic result!",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 1985062,
      "author_name": "vbookshelf",
      "author_url": "",
      "post_date": "2022-10-13T04:38:34.837000",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/kodaihatayama\" target=\"_blank\">@kodaihatayama</a>. Solo gold in your first competition is a fantastic result!</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1984324": "Hi, this competition is my first challenge at kaggle. I really enjoyed this competition and thought hard for about three months. I don't know why I could obtain a gold medal, but I'm sure it's thanks to everyone who have shared the notebooks or suggestions. \nHere I describe my solution in brief.\n\nSubmission kernel:\nhttps://www.kaggle.com/code/kodaihatayama/mayo02-submission\n\n### Preprocessing\nI created tiles with the following steps. **I didn't resize the image** in all preprocessing steps.\n- Cut 6 rectangles of width 512 at equal intervals from every image \n(The purpose of this step is to create tiles quickly and avoid memory crashes without resizing.)\n- Make tiles of size 512x512 from the above rectangles and pick up the top 8 dark tiles per image\n- Split each tile to instances of size 32x32 for training\n\nFor using pyvips, I refered to the following kernel. Thank you, @analokamus !\nhttps://www.kaggle.com/code/analokamus/how-to-use-pyvips-offline\n\n### Validation and Undersampling\nI used the hold-out method with undersampling. \n- Split patient id to train and validation (80:20) which was stratified by class (CE, LAA)\n- Perform undersampling to be CE:LAA = 1:1 on the train dataset only\n\n### Model\nI created the ensemble model of the following 2 methods.\n  - CNN \n  - lightGBM\n\nThe difference of private LB between CNN only model and the ensemble model was a bit (approximately 0.676 → 0.668).\nAs a side note, I created a model per institution (center id). (This approach may not be effective against test dataset because the institutions of test may be different from them of train.)\n\nFor CNN, I refered to the following kernel. Thank you, @vbookshelf !\nhttps://www.kaggle.com/code/vbookshelf/cnn-how-to-use-160-000-images-without-crashing\n\n### What didn't work for me\n- grayscale\n- histogram equalization\n\nThanks for reading.",
    "1985062": "Congratulations @kodaihatayama. Solo gold in your first competition is a fantastic result!"
  }
}