{
  "id": 465461,
  "title": "83rd solusion and this is my firstintroduction of solution !",
  "url": "/competitions/UBC-OCEAN/discussion/465461",
  "author_name": "Taro Kuroda",
  "post_date": "2024-01-04T11:05:39.375000",
  "votes": 12,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Thank you for all people who have involved this competition.</p>\n<p>I'm very surprised to see the result because my rank was below 600 in public LB.</p>\n<p>I have devoted myself to this competition during the competition period, and I want to share my solution.</p>\n<p>As you suspect, I'm still a begginer in Kaggle, so please advise me for my improvement.</p>\n<p><strong>Data</strong></p>\n<ul>\n<li>As you know, in this competition, we had to predict the diagnosis of tissue microarray (TMA) images using whole slide images (WSI). </li>\n<li>In train data, there are many WSI images and a little number of TMA images(only 25 images!). </li>\n<li>In addition, the maginification of TMA and WSI images is 40× and 20x respectively, so I thought that preprocess shuold be applied.<br>\nEventually, I used the method which was introduced in <a href=\"https://www.kaggle.com/competitions/mayo-clinic-strip-ai/discussion/357892\" target=\"_blank\">STRIP-AI competition 1st solution</a></li>\n<li>By applying this process, I could create TMA like images from WSI images and increase the number of train data.</li>\n</ul>\n<p><strong>Model</strong></p>\n<ul>\n<li>I tried over a lot of models in timm library, and I chose maxvit-tiny and convnext_base. </li>\n<li>I added nn.MultiheadAttention layer after forward_features with num_head=4.</li>\n<li>Both of them are pre-trained and fine-tuned with max lr=1e-4, OneCycleLR, pct_start=0.2, epoch=20. These parameters were determined based on the experiment results.</li>\n</ul>\n<p><strong>loss function</strong></p>\n<ul>\n<li>Simple crossentropy-loss was used with weights.</li>\n</ul>\n<p><strong>CV Strategy</strong></p>\n<ul>\n<li>5 fold StratifiedGroupKFold using labels and is_tma == True or False.</li>\n<li>At first time, I used 25 TMA images only for validation. However, including these images improved the CV score.</li>\n</ul>\n<p><strong>cut off to determine 'Others' label</strong></p>\n<ul>\n<li>Another difficulty of this competition is that we have to diagnose labels which were not in train data but presented in test data as 'Others'. </li>\n<li>I determined this cut off label based on the results of my models.</li>\n<li>My submission model gave 'Others' when the probability was below 0.3.</li>\n<li>Unfortunately, the model with this value 0.4 got higher score and I could get first silver medal if I submitted this one…</li>\n</ul>\n<p><strong>What Works</strong></p>\n<ul>\n<li><p><strong>applying nn.MultiheadAttention layer after forward_features</strong><br>\nI also multiplied attention outputs and weights, which made strong features much stronger. I believe the reason why this method worked is that images derived from WSI had a lot of noise, and restricted features were important. </p></li>\n<li><p><strong>strong regularization</strong><br>\nI adapted RandomResizedCrop, all direction flips, CoarseDropout, label-smoothing, and mixup. I annoied whether color changes should be adapted, but I didn't because I thought color information is important in pathology.</p></li>\n<li><p><strong>delete noise-like data from training</strong><br>\nI checked all data which was diagnosed incorrectly, and deleted ones which I considered invalid. I think this process included a lot of biases, but revealed to be effective from the final results.</p></li>\n</ul>\n<p><strong>What didn't Work</strong></p>\n<ul>\n<li><p><strong>using segmentation masks</strong><br>\nDuring competition period, segmentation masks of some train data became open. I created resnet-based segmentation models to get only tumor areas in WSI, and adapted the preprocess mentioned above. However, this method made CV socre and LB score worse.</p></li>\n<li><p><strong>TTA(flip horizontally or vertically)</strong></p></li>\n</ul>\n<p>Thank you for reading !!!</p>",
  "messages": [
    {
      "id": 2586713,
      "postDate": "2024-01-04T11:05:39.377Z",
      "content": "<p>Thank you for all people who have involved this competition.</p>\n<p>I'm very surprised to see the result because my rank was below 600 in public LB.</p>\n<p>I have devoted myself to this competition during the competition period, and I want to share my solution.</p>\n<p>As you suspect, I'm still a begginer in Kaggle, so please advise me for my improvement.</p>\n<p><strong>Data</strong></p>\n<ul>\n<li>As you know, in this competition, we had to predict the diagnosis of tissue microarray (TMA) images using whole slide images (WSI). </li>\n<li>In train data, there are many WSI images and a little number of TMA images(only 25 images!). </li>\n<li>In addition, the maginification of TMA and WSI images is 40× and 20x respectively, so I thought that preprocess shuold be applied.<br>\nEventually, I used the method which was introduced in <a href=\"https://www.kaggle.com/competitions/mayo-clinic-strip-ai/discussion/357892\" target=\"_blank\">STRIP-AI competition 1st solution</a></li>\n<li>By applying this process, I could create TMA like images from WSI images and increase the number of train data.</li>\n</ul>\n<p><strong>Model</strong></p>\n<ul>\n<li>I tried over a lot of models in timm library, and I chose maxvit-tiny and convnext_base. </li>\n<li>I added nn.MultiheadAttention layer after forward_features with num_head=4.</li>\n<li>Both of them are pre-trained and fine-tuned with max lr=1e-4, OneCycleLR, pct_start=0.2, epoch=20. These parameters were determined based on the experiment results.</li>\n</ul>\n<p><strong>loss function</strong></p>\n<ul>\n<li>Simple crossentropy-loss was used with weights.</li>\n</ul>\n<p><strong>CV Strategy</strong></p>\n<ul>\n<li>5 fold StratifiedGroupKFold using labels and is_tma == True or False.</li>\n<li>At first time, I used 25 TMA images only for validation. However, including these images improved the CV score.</li>\n</ul>\n<p><strong>cut off to determine 'Others' label</strong></p>\n<ul>\n<li>Another difficulty of this competition is that we have to diagnose labels which were not in train data but presented in test data as 'Others'. </li>\n<li>I determined this cut off label based on the results of my models.</li>\n<li>My submission model gave 'Others' when the probability was below 0.3.</li>\n<li>Unfortunately, the model with this value 0.4 got higher score and I could get first silver medal if I submitted this one…</li>\n</ul>\n<p><strong>What Works</strong></p>\n<ul>\n<li><p><strong>applying nn.MultiheadAttention layer after forward_features</strong><br>\nI also multiplied attention outputs and weights, which made strong features much stronger. I believe the reason why this method worked is that images derived from WSI had a lot of noise, and restricted features were important. </p></li>\n<li><p><strong>strong regularization</strong><br>\nI adapted RandomResizedCrop, all direction flips, CoarseDropout, label-smoothing, and mixup. I annoied whether color changes should be adapted, but I didn't because I thought color information is important in pathology.</p></li>\n<li><p><strong>delete noise-like data from training</strong><br>\nI checked all data which was diagnosed incorrectly, and deleted ones which I considered invalid. I think this process included a lot of biases, but revealed to be effective from the final results.</p></li>\n</ul>\n<p><strong>What didn't Work</strong></p>\n<ul>\n<li><p><strong>using segmentation masks</strong><br>\nDuring competition period, segmentation masks of some train data became open. I created resnet-based segmentation models to get only tumor areas in WSI, and adapted the preprocess mentioned above. However, this method made CV socre and LB score worse.</p></li>\n<li><p><strong>TTA(flip horizontally or vertically)</strong></p></li>\n</ul>\n<p>Thank you for reading !!!</p>",
      "rawMarkdown": "Thank you for all people who have involved this competition.\n\nI'm very surprised to see the result because my rank was below 600 in public LB.\n\nI have devoted myself to this competition during the competition period, and I want to share my solution.\n\nAs you suspect, I'm still a begginer in Kaggle, so please advise me for my improvement.\n\n **Data**\n- As you know, in this competition, we had to predict the diagnosis of tissue microarray (TMA) images using whole slide images (WSI). \n- In train data, there are many WSI images and a little number of TMA images(only 25 images!). \n- In addition, the maginification of TMA and WSI images is 40× and 20x respectively, so I thought that preprocess shuold be applied.\nEventually, I used the method which was introduced in [STRIP-AI competition 1st solution](https://www.kaggle.com/competitions/mayo-clinic-strip-ai/discussion/357892)\n- By applying this process, I could create TMA like images from WSI images and increase the number of train data.\n\n**Model**\n- I tried over a lot of models in timm library, and I chose maxvit-tiny and convnext_base. \n- I added nn.MultiheadAttention layer after forward_features with num_head=4.\n- Both of them are pre-trained and fine-tuned with max lr=1e-4, OneCycleLR, pct_start=0.2, epoch=20. These parameters were determined based on the experiment results.\n\n**loss function**\n- Simple crossentropy-loss was used with weights.\n\n**CV Strategy**\n- 5 fold StratifiedGroupKFold using labels and is_tma == True or False.\n- At first time, I used 25 TMA images only for validation. However, including these images improved the CV score.\n\n**cut off to determine 'Others' label**\n- Another difficulty of this competition is that we have to diagnose labels which were not in train data but presented in test data as 'Others'. \n- I determined this cut off label based on the results of my models.\n- My submission model gave 'Others' when the probability was below 0.3.\n- Unfortunately, the model with this value 0.4 got higher score and I could get first silver medal if I submitted this one...\n\n**What Works**\n- **applying nn.MultiheadAttention layer after forward_features**\nI also multiplied attention outputs and weights, which made strong features much stronger. I believe the reason why this method worked is that images derived from WSI had a lot of noise, and restricted features were important. \n\n- **strong regularization**\nI adapted RandomResizedCrop, all direction flips, CoarseDropout, label-smoothing, and mixup. I annoied whether color changes should be adapted, but I didn't because I thought color information is important in pathology.\n\n- **delete noise-like data from training**\nI checked all data which was diagnosed incorrectly, and deleted ones which I considered invalid. I think this process included a lot of biases, but revealed to be effective from the final results.\n\n**What didn't Work**\n- **using segmentation masks**\nDuring competition period, segmentation masks of some train data became open. I created resnet-based segmentation models to get only tumor areas in WSI, and adapted the preprocess mentioned above. However, this method made CV socre and LB score worse.\n\n- **TTA(flip horizontally or vertically)**\n\nThank you for reading !!!\n\n\n\n",
      "votes": 12
    },
    {
      "id": 2586947,
      "postDate": "2024-01-04T13:43:14.563Z",
      "content": "<p>nice tried！</p>",
      "rawMarkdown": "nice tried！",
      "votes": 1,
      "replies": [
        {
          "id": 2587603,
          "postDate": "2024-01-04T21:34:36.133Z",
          "content": "<p>Thank you !😂</p>",
          "rawMarkdown": "Thank you !😂"
        }
      ]
    },
    {
      "id": 2586844,
      "postDate": "2024-01-04T12:37:46.667Z",
      "content": "<p>Congratulations for the medal!</p>\n<p>I have two questions. What does 'accuracy was below 0.3' mean?<br>\nAnd approximately how many did you delete? (100% unsure but with a sample showing something wrong and a sample that was definitely misdiagnosed)</p>",
      "rawMarkdown": "Congratulations for the medal!\n\nI have two questions. What does 'accuracy was below 0.3' mean?\nAnd approximately how many did you delete? (100% unsure but with a sample showing something wrong and a sample that was definitely misdiagnosed)",
      "votes": 1,
      "replies": [
        {
          "id": 2587616,
          "postDate": "2024-01-04T22:03:11.907Z",
          "content": "<p>I'm sorry for misleading. I updated.<br>\nAccurately, 'brobability was below 0.3', so if model gave probability under 30%, the label becomes 'Others'.</p>\n<p>Tha total number of images are8633(538*16+25), but 8206 images were used in training.<br>\nTherefore, 427 images were removed.<br>\nI'm not a pathologist, so I didn't have confidence in my diagnosis, so the exclusion criteria was simple; </p>\n<ol>\n<li><strong>50% and above area is blank or hematoma</strong></li>\n</ol>\n<p>2.<strong>clearly different from other images</strong> </p>\n<p>For example, below image represents tiled images from ID26950.<br>\n26950_0, 2, 14 were mis-diagnosed, and different from other 13 images.<br>\nTherefore, I removed these from training !</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8370847%2Fff848758c735290c60b99acbaf791068%2F2024-01-05%20065932.png?generation=1704405779657231&amp;alt=media\" alt=\"image\"></p>",
          "rawMarkdown": "I'm sorry for misleading. I updated.\nAccurately, 'brobability was below 0.3', so if model gave probability under 30%, the label becomes 'Others'.\n\nTha total number of images are8633(538*16+25), but 8206 images were used in training.\nTherefore, 427 images were removed.\nI'm not a pathologist, so I didn't have confidence in my diagnosis, so the exclusion criteria was simple; \n\n1. **50% and above area is blank or hematoma**\n\n2.**clearly different from other images** \n\nFor example, below image represents tiled images from ID26950.\n26950_0, 2, 14 were mis-diagnosed, and different from other 13 images.\nTherefore, I removed these from training !\n\n![image](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8370847%2Fff848758c735290c60b99acbaf791068%2F2024-01-05%20065932.png?generation=1704405779657231&alt=media)\n\n\n",
          "replies": [
            {
              "id": 2587644,
              "postDate": "2024-01-04T22:53:40.240Z",
              "content": "<p>Thank you very much for your detailed explanation. <br>\nI hope you win more than a silver medal in the next competition.</p>",
              "rawMarkdown": "Thank you very much for your detailed explanation. \nI hope you win more than a silver medal in the next competition."
            },
            {
              "id": 2587818,
              "postDate": "2024-01-05T03:52:31.770Z",
              "content": "<p>Thank you!<br>\nI hope your success, too!</p>",
              "rawMarkdown": "Thank you!\nI hope your success, too!"
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2586947,
      "author_name": "Seeing Times",
      "author_url": "",
      "post_date": "2024-01-04T13:43:14.563000",
      "content": "<p>nice tried！</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2587603,
          "author_name": "Taro Kuroda",
          "author_url": "",
          "post_date": "2024-01-04T21:34:36.133000",
          "content": "<p>Thank you !😂</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2586844,
      "author_name": "HB",
      "author_url": "",
      "post_date": "2024-01-04T12:37:46.667000",
      "content": "<p>Congratulations for the medal!</p>\n<p>I have two questions. What does 'accuracy was below 0.3' mean?<br>\nAnd approximately how many did you delete? (100% unsure but with a sample showing something wrong and a sample that was definitely misdiagnosed)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2587616,
          "author_name": "Taro Kuroda",
          "author_url": "",
          "post_date": "2024-01-04T22:03:11.907000",
          "content": "<p>I'm sorry for misleading. I updated.<br>\nAccurately, 'brobability was below 0.3', so if model gave probability under 30%, the label becomes 'Others'.</p>\n<p>Tha total number of images are8633(538*16+25), but 8206 images were used in training.<br>\nTherefore, 427 images were removed.<br>\nI'm not a pathologist, so I didn't have confidence in my diagnosis, so the exclusion criteria was simple; </p>\n<ol>\n<li><strong>50% and above area is blank or hematoma</strong></li>\n</ol>\n<p>2.<strong>clearly different from other images</strong> </p>\n<p>For example, below image represents tiled images from ID26950.<br>\n26950_0, 2, 14 were mis-diagnosed, and different from other 13 images.<br>\nTherefore, I removed these from training !</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8370847%2Fff848758c735290c60b99acbaf791068%2F2024-01-05%20065932.png?generation=1704405779657231&amp;alt=media\" alt=\"image\"></p>",
          "votes": 0,
          "replies": [
            {
              "id": 2587644,
              "author_name": "HB",
              "author_url": "",
              "post_date": "2024-01-04T22:53:40.240000",
              "content": "<p>Thank you very much for your detailed explanation. <br>\nI hope you win more than a silver medal in the next competition.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2587818,
              "author_name": "Taro Kuroda",
              "author_url": "",
              "post_date": "2024-01-05T03:52:31.770000",
              "content": "<p>Thank you!<br>\nI hope your success, too!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2586713": "Thank you for all people who have involved this competition.\n\nI'm very surprised to see the result because my rank was below 600 in public LB.\n\nI have devoted myself to this competition during the competition period, and I want to share my solution.\n\nAs you suspect, I'm still a begginer in Kaggle, so please advise me for my improvement.\n\n **Data**\n- As you know, in this competition, we had to predict the diagnosis of tissue microarray (TMA) images using whole slide images (WSI). \n- In train data, there are many WSI images and a little number of TMA images(only 25 images!). \n- In addition, the maginification of TMA and WSI images is 40× and 20x respectively, so I thought that preprocess shuold be applied.\nEventually, I used the method which was introduced in [STRIP-AI competition 1st solution](https://www.kaggle.com/competitions/mayo-clinic-strip-ai/discussion/357892)\n- By applying this process, I could create TMA like images from WSI images and increase the number of train data.\n\n**Model**\n- I tried over a lot of models in timm library, and I chose maxvit-tiny and convnext_base. \n- I added nn.MultiheadAttention layer after forward_features with num_head=4.\n- Both of them are pre-trained and fine-tuned with max lr=1e-4, OneCycleLR, pct_start=0.2, epoch=20. These parameters were determined based on the experiment results.\n\n**loss function**\n- Simple crossentropy-loss was used with weights.\n\n**CV Strategy**\n- 5 fold StratifiedGroupKFold using labels and is_tma == True or False.\n- At first time, I used 25 TMA images only for validation. However, including these images improved the CV score.\n\n**cut off to determine 'Others' label**\n- Another difficulty of this competition is that we have to diagnose labels which were not in train data but presented in test data as 'Others'. \n- I determined this cut off label based on the results of my models.\n- My submission model gave 'Others' when the probability was below 0.3.\n- Unfortunately, the model with this value 0.4 got higher score and I could get first silver medal if I submitted this one...\n\n**What Works**\n- **applying nn.MultiheadAttention layer after forward_features**\nI also multiplied attention outputs and weights, which made strong features much stronger. I believe the reason why this method worked is that images derived from WSI had a lot of noise, and restricted features were important. \n\n- **strong regularization**\nI adapted RandomResizedCrop, all direction flips, CoarseDropout, label-smoothing, and mixup. I annoied whether color changes should be adapted, but I didn't because I thought color information is important in pathology.\n\n- **delete noise-like data from training**\nI checked all data which was diagnosed incorrectly, and deleted ones which I considered invalid. I think this process included a lot of biases, but revealed to be effective from the final results.\n\n**What didn't Work**\n- **using segmentation masks**\nDuring competition period, segmentation masks of some train data became open. I created resnet-based segmentation models to get only tumor areas in WSI, and adapted the preprocess mentioned above. However, this method made CV socre and LB score worse.\n\n- **TTA(flip horizontally or vertically)**\n\nThank you for reading !!!\n\n\n\n",
    "2586947": "nice tried！",
    "2586844": "Congratulations for the medal!\n\nI have two questions. What does 'accuracy was below 0.3' mean?\nAnd approximately how many did you delete? (100% unsure but with a sample showing something wrong and a sample that was definitely misdiagnosed)"
  }
}