{
  "id": 152111,
  "title": "Why I gave up on segmentation task, some experiment results share",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/152111",
  "author_name": "Tsai29",
  "post_date": "2020-05-18T13:42:43.405000",
  "votes": 19,
  "comment_count": 8,
  "views": 0,
  "content": "<h2>The reasons I started this competition with segmentation model :</h2>\n\n<p><br>\n- More papers to reference\n- With segmentation results, the doctors can double check the segmented region faster since the model already labeled the cancer region.\n- I'm more interested in segmentation task.\n<br></p>\n\n<h2>My best experiment result so far:</h2>\n\n<p><br>\n- CV:0.66, LB: 0.68\n<br>\n- Three stage models. The first model is trained as binary classifier, which will classify the image is totally healthy or not. The second model is trained as segmentation model and it only trained with data labeled as unhealthy. The third stage model is trained as classifier, which will classify the isup_grade by the percentage of healthy/unhealthy pixels.\n<br>\n- Stage 1 model was train with 512x1024 resolution which resized from level 2 images. \nAnd the stage 2 model was trained with 1024x1024 resolution, which the images/masks are divided from level 1 resolution. Which means I trained with original image details of level 1 resolution without any resize. \nStage 3 model was trained with percentage of ground truth mask pixels from radboud.\n<br>\n- Both stage 1 and stage 2 models are using efficiententB0 backbone.\nStage 1 is binary classifier.\nStage 2 is U-Net.\nStage 3 is lgbm clssifier.\n<br>\n- Sigmoid loss for stage 1 model.\nDiceLoss for stage 2 model loss function. At first, I implemented the DiceLoss which will calculate the loss per-image. But the stage 2 model just can't recognize the mask is belong to gleasion 3, gleason 4 or gleason 5. Then I changed the DiceLoss to calculate the loss per-batch, then the stage 2 model start to converge well.\n<br>\n- Rotate/Scale/Shear/Shift/vertical flip/horizontal flip/blur data augmentation.\n<br></p>\n\n<h2>Why I gave up on segmentation model:</h2>\n\n<p><br>\n- I believe that if I collect the data with level 0 resolution can boost the score, but I don't have enough hardware resource to train that. \n<br>\n- Only half of data has ground truth mask with detail gleason(gleason 3/4/5), which is data from radbound. Maybe we can re-label the data from other data provider with high confidence prediction, but there are up to 5000 images, it is quite easy to get serious overfitting.\n<br>\n- The hardware requirement of classification task is lower.</p>\n\n<h2><br></h2>\n\n<p>But, if you are also interested in treating this as segmentation task, I'm quite welcome to team up with you, I will divide half of my resource to figure out how to get better results on segmentation task. Thanks!</p>",
  "messages": [
    {
      "id": 852504,
      "postDate": "2020-05-18T13:42:43.407Z",
      "content": "<h2>The reasons I started this competition with segmentation model :</h2>\n\n<p><br>\n- More papers to reference\n- With segmentation results, the doctors can double check the segmented region faster since the model already labeled the cancer region.\n- I'm more interested in segmentation task.\n<br></p>\n\n<h2>My best experiment result so far:</h2>\n\n<p><br>\n- CV:0.66, LB: 0.68\n<br>\n- Three stage models. The first model is trained as binary classifier, which will classify the image is totally healthy or not. The second model is trained as segmentation model and it only trained with data labeled as unhealthy. The third stage model is trained as classifier, which will classify the isup_grade by the percentage of healthy/unhealthy pixels.\n<br>\n- Stage 1 model was train with 512x1024 resolution which resized from level 2 images. \nAnd the stage 2 model was trained with 1024x1024 resolution, which the images/masks are divided from level 1 resolution. Which means I trained with original image details of level 1 resolution without any resize. \nStage 3 model was trained with percentage of ground truth mask pixels from radboud.\n<br>\n- Both stage 1 and stage 2 models are using efficiententB0 backbone.\nStage 1 is binary classifier.\nStage 2 is U-Net.\nStage 3 is lgbm clssifier.\n<br>\n- Sigmoid loss for stage 1 model.\nDiceLoss for stage 2 model loss function. At first, I implemented the DiceLoss which will calculate the loss per-image. But the stage 2 model just can't recognize the mask is belong to gleasion 3, gleason 4 or gleason 5. Then I changed the DiceLoss to calculate the loss per-batch, then the stage 2 model start to converge well.\n<br>\n- Rotate/Scale/Shear/Shift/vertical flip/horizontal flip/blur data augmentation.\n<br></p>\n\n<h2>Why I gave up on segmentation model:</h2>\n\n<p><br>\n- I believe that if I collect the data with level 0 resolution can boost the score, but I don't have enough hardware resource to train that. \n<br>\n- Only half of data has ground truth mask with detail gleason(gleason 3/4/5), which is data from radbound. Maybe we can re-label the data from other data provider with high confidence prediction, but there are up to 5000 images, it is quite easy to get serious overfitting.\n<br>\n- The hardware requirement of classification task is lower.</p>\n\n<h2><br></h2>\n\n<p>But, if you are also interested in treating this as segmentation task, I'm quite welcome to team up with you, I will divide half of my resource to figure out how to get better results on segmentation task. Thanks!</p>",
      "rawMarkdown": "## The reasons I started this competition with segmentation model :\n<br>\n- More papers to reference\n- With segmentation results, the doctors can double check the segmented region faster since the model already labeled the cancer region.\n- I'm more interested in segmentation task.\n<br>\n## My best experiment result so far:\n<br>\n- CV:0.66, LB: 0.68\n<br>\n- Three stage models. The first model is trained as binary classifier, which will classify the image is totally healthy or not. The second model is trained as segmentation model and it only trained with data labeled as unhealthy. The third stage model is trained as classifier, which will classify the isup_grade by the percentage of healthy/unhealthy pixels.\n<br>\n- Stage 1 model was train with 512x1024 resolution which resized from level 2 images. \nAnd the stage 2 model was trained with 1024x1024 resolution, which the images/masks are divided from level 1 resolution. Which means I trained with original image details of level 1 resolution without any resize. \nStage 3 model was trained with percentage of ground truth mask pixels from radboud.\n<br>\n- Both stage 1 and stage 2 models are using efficiententB0 backbone.\nStage 1 is binary classifier.\nStage 2 is U-Net.\nStage 3 is lgbm clssifier.\n<br>\n- Sigmoid loss for stage 1 model.\nDiceLoss for stage 2 model loss function. At first, I implemented the DiceLoss which will calculate the loss per-image. But the stage 2 model just can't recognize the mask is belong to gleasion 3, gleason 4 or gleason 5. Then I changed the DiceLoss to calculate the loss per-batch, then the stage 2 model start to converge well.\n<br>\n- Rotate/Scale/Shear/Shift/vertical flip/horizontal flip/blur data augmentation.\n<br>\n## Why I gave up on segmentation model:\n<br>\n- I believe that if I collect the data with level 0 resolution can boost the score, but I don't have enough hardware resource to train that. \n<br>\n- Only half of data has ground truth mask with detail gleason(gleason 3/4/5), which is data from radbound. Maybe we can re-label the data from other data provider with high confidence prediction, but there are up to 5000 images, it is quite easy to get serious overfitting.\n<br>\n- The hardware requirement of classification task is lower.\n<br>\n--------------------------------------------------------------\n\nBut, if you are also interested in treating this as segmentation task, I'm quite welcome to team up with you, I will divide half of my resource to figure out how to get better results on segmentation task. Thanks!",
      "votes": 18
    },
    {
      "id": 853034,
      "postDate": "2020-05-18T23:06:50.050Z",
      "content": "<p>You can just have a segmentation aux serving to focus the attention to the particular parts of the picture when train a classifier based model. Though, when I tried it in the past, it gave me only a little boost, and I ran other experiments without it.</p>",
      "rawMarkdown": "You can just have a segmentation aux serving to focus the attention to the particular parts of the picture when train a classifier based model. Though, when I tried it in the past, it gave me only a little boost, and I ran other experiments without it.",
      "votes": 1,
      "replies": [
        {
          "id": 853067,
          "postDate": "2020-05-19T00:04:21.413Z",
          "content": "<p>Thanks for the kind sharing. I did considered the same model arrangement for my experiment. BTW, congrats for the grandmaster title.</p>",
          "rawMarkdown": "Thanks for the kind sharing. I did considered the same model arrangement for my experiment. BTW, congrats for the grandmaster title."
        },
        {
          "id": 853094,
          "postDate": "2020-05-19T00:47:02.710Z",
          "content": "<p>Thank you so much.</p>",
          "rawMarkdown": "Thank you so much."
        }
      ]
    },
    {
      "id": 852756,
      "postDate": "2020-05-18T16:56:22.063Z",
      "content": "<p>The segmented model is worth a try. It can be divided into two categories first, excluding healthy samples, and then more categories for unhealthy samples.</p>",
      "rawMarkdown": "The segmented model is worth a try. It can be divided into two categories first, excluding healthy samples, and then more categories for unhealthy samples.",
      "votes": 1,
      "replies": [
        {
          "id": 853058,
          "postDate": "2020-05-18T23:58:55.303Z",
          "content": "<p>Yes, this seems like what I did if I understand your words right. But I might need to evaluate the tradeoff between training resource requirement and scores of both classification major method and segmentation major method. From my current result, classification major method is quite friendly to train and the results seems okay.</p>",
          "rawMarkdown": "Yes, this seems like what I did if I understand your words right. But I might need to evaluate the tradeoff between training resource requirement and scores of both classification major method and segmentation major method. From my current result, classification major method is quite friendly to train and the results seems okay.",
          "votes": 1
        }
      ]
    },
    {
      "id": 852560,
      "postDate": "2020-05-18T14:30:14.150Z",
      "content": "<p><a href=\"/xiejialun\">@xiejialun</a> We are willing to work you. I was attempting to follow the same road map. My email is dskswu@gmail.com</p>",
      "rawMarkdown": "@xiejialun We are willing to work you. I was attempting to follow the same road map. My email is dskswu@gmail.com",
      "votes": 1,
      "replies": [
        {
          "id": 853055,
          "postDate": "2020-05-18T23:51:28.680Z",
          "content": "<p><a href=\"/dskswu\">@dskswu</a> Thanks a lot for the kind invitation. But one of my friend reached me yesterday and said that he will evaluate the situation and might join me in this competition. So I might need to wait for his response. But if he decide not to join in competition, I will reach you if you still need a another person to work with :).</p>",
          "rawMarkdown": "@dskswu Thanks a lot for the kind invitation. But one of my friend reached me yesterday and said that he will evaluate the situation and might join me in this competition. So I might need to wait for his response. But if he decide not to join in competition, I will reach you if you still need a another person to work with :)."
        },
        {
          "id": 853151,
          "postDate": "2020-05-19T02:11:53.310Z",
          "content": "<p><a href=\"/xiejialun\">@xiejialun</a> Welcome. :)</p>",
          "rawMarkdown": "@xiejialun Welcome. :)"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 853034,
      "author_name": "Iafoss",
      "author_url": "",
      "post_date": "2020-05-18T23:06:50.050000",
      "content": "<p>You can just have a segmentation aux serving to focus the attention to the particular parts of the picture when train a classifier based model. Though, when I tried it in the past, it gave me only a little boost, and I ran other experiments without it.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 853067,
          "author_name": "Tsai29",
          "author_url": "",
          "post_date": "2020-05-19T00:04:21.413000",
          "content": "<p>Thanks for the kind sharing. I did considered the same model arrangement for my experiment. BTW, congrats for the grandmaster title.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 853094,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2020-05-19T00:47:02.710000",
          "content": "<p>Thank you so much.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 852756,
      "author_name": "zhangeng",
      "author_url": "",
      "post_date": "2020-05-18T16:56:22.063000",
      "content": "<p>The segmented model is worth a try. It can be divided into two categories first, excluding healthy samples, and then more categories for unhealthy samples.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 853058,
          "author_name": "Tsai29",
          "author_url": "",
          "post_date": "2020-05-18T23:58:55.303000",
          "content": "<p>Yes, this seems like what I did if I understand your words right. But I might need to evaluate the tradeoff between training resource requirement and scores of both classification major method and segmentation major method. From my current result, classification major method is quite friendly to train and the results seems okay.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 852560,
      "author_name": "William Green",
      "author_url": "",
      "post_date": "2020-05-18T14:30:14.150000",
      "content": "<p><a href=\"/xiejialun\">@xiejialun</a> We are willing to work you. I was attempting to follow the same road map. My email is dskswu@gmail.com</p>",
      "votes": 1,
      "replies": [
        {
          "id": 853055,
          "author_name": "Tsai29",
          "author_url": "",
          "post_date": "2020-05-18T23:51:28.680000",
          "content": "<p><a href=\"/dskswu\">@dskswu</a> Thanks a lot for the kind invitation. But one of my friend reached me yesterday and said that he will evaluate the situation and might join me in this competition. So I might need to wait for his response. But if he decide not to join in competition, I will reach you if you still need a another person to work with :).</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 853151,
          "author_name": "William Green",
          "author_url": "",
          "post_date": "2020-05-19T02:11:53.310000",
          "content": "<p><a href=\"/xiejialun\">@xiejialun</a> Welcome. :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "852504": "## The reasons I started this competition with segmentation model :\n<br>\n- More papers to reference\n- With segmentation results, the doctors can double check the segmented region faster since the model already labeled the cancer region.\n- I'm more interested in segmentation task.\n<br>\n## My best experiment result so far:\n<br>\n- CV:0.66, LB: 0.68\n<br>\n- Three stage models. The first model is trained as binary classifier, which will classify the image is totally healthy or not. The second model is trained as segmentation model and it only trained with data labeled as unhealthy. The third stage model is trained as classifier, which will classify the isup_grade by the percentage of healthy/unhealthy pixels.\n<br>\n- Stage 1 model was train with 512x1024 resolution which resized from level 2 images. \nAnd the stage 2 model was trained with 1024x1024 resolution, which the images/masks are divided from level 1 resolution. Which means I trained with original image details of level 1 resolution without any resize. \nStage 3 model was trained with percentage of ground truth mask pixels from radboud.\n<br>\n- Both stage 1 and stage 2 models are using efficiententB0 backbone.\nStage 1 is binary classifier.\nStage 2 is U-Net.\nStage 3 is lgbm clssifier.\n<br>\n- Sigmoid loss for stage 1 model.\nDiceLoss for stage 2 model loss function. At first, I implemented the DiceLoss which will calculate the loss per-image. But the stage 2 model just can't recognize the mask is belong to gleasion 3, gleason 4 or gleason 5. Then I changed the DiceLoss to calculate the loss per-batch, then the stage 2 model start to converge well.\n<br>\n- Rotate/Scale/Shear/Shift/vertical flip/horizontal flip/blur data augmentation.\n<br>\n## Why I gave up on segmentation model:\n<br>\n- I believe that if I collect the data with level 0 resolution can boost the score, but I don't have enough hardware resource to train that. \n<br>\n- Only half of data has ground truth mask with detail gleason(gleason 3/4/5), which is data from radbound. Maybe we can re-label the data from other data provider with high confidence prediction, but there are up to 5000 images, it is quite easy to get serious overfitting.\n<br>\n- The hardware requirement of classification task is lower.\n<br>\n--------------------------------------------------------------\n\nBut, if you are also interested in treating this as segmentation task, I'm quite welcome to team up with you, I will divide half of my resource to figure out how to get better results on segmentation task. Thanks!",
    "853034": "You can just have a segmentation aux serving to focus the attention to the particular parts of the picture when train a classifier based model. Though, when I tried it in the past, it gave me only a little boost, and I ran other experiments without it.",
    "852756": "The segmented model is worth a try. It can be divided into two categories first, excluding healthy samples, and then more categories for unhealthy samples.",
    "852560": "@xiejialun We are willing to work you. I was attempting to follow the same road map. My email is dskswu@gmail.com"
  }
}