{
  "id": 169553,
  "title": "Two items that worked that I haven't seen mentioned...",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/169553",
  "author_name": "fergusoci",
  "post_date": "2020-07-24T08:00:22.194000",
  "votes": 11,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Hi all,</p>\n\n<p>Congratulations to the winners! </p>\n\n<p>There were two items that worked that I haven't seen mentioned thus far:</p>\n\n<ol>\n<li><p>Mixed slide augmentations: I got my highest CV scores when using these (and they also generated some of the highest private LB scores). When an image was selected for training I randomly selected a second image from the same data provider, same isup grade, tiled both images, then randomly selected 50% of the tiles from each (and reordered them as you would if you were selecting tiles from a single image). This lead to a +0.005 in CV score. I tried a version whereby , as well as the same data provider + isup, the mix image was of the same tissue colour, but this didn't add anything additional.</p></li>\n<li><p>Sampling according to the Karolinska distribution: before each epoch the Radboud images were sampled according to the Karolinska distribution of isup grades (i.e. sample proportion of grades 0 - 5 as Karolinska). The idea was that the Karolinska distribution was more representative of the test set. </p></li>\n</ol>\n\n<p>Cheers</p>",
  "messages": [
    {
      "id": 943220,
      "postDate": "2020-07-24T08:00:22.193Z",
      "content": "<p>Hi all,</p>\n\n<p>Congratulations to the winners! </p>\n\n<p>There were two items that worked that I haven't seen mentioned thus far:</p>\n\n<ol>\n<li><p>Mixed slide augmentations: I got my highest CV scores when using these (and they also generated some of the highest private LB scores). When an image was selected for training I randomly selected a second image from the same data provider, same isup grade, tiled both images, then randomly selected 50% of the tiles from each (and reordered them as you would if you were selecting tiles from a single image). This lead to a +0.005 in CV score. I tried a version whereby , as well as the same data provider + isup, the mix image was of the same tissue colour, but this didn't add anything additional.</p></li>\n<li><p>Sampling according to the Karolinska distribution: before each epoch the Radboud images were sampled according to the Karolinska distribution of isup grades (i.e. sample proportion of grades 0 - 5 as Karolinska). The idea was that the Karolinska distribution was more representative of the test set. </p></li>\n</ol>\n\n<p>Cheers</p>",
      "rawMarkdown": "Hi all,\n\nCongratulations to the winners! \n\nThere were two items that worked that I haven't seen mentioned thus far:\n\n1. Mixed slide augmentations: I got my highest CV scores when using these (and they also generated some of the highest private LB scores). When an image was selected for training I randomly selected a second image from the same data provider, same isup grade, tiled both images, then randomly selected 50% of the tiles from each (and reordered them as you would if you were selecting tiles from a single image). This lead to a +0.005 in CV score. I tried a version whereby , as well as the same data provider + isup, the mix image was of the same tissue colour, but this didn't add anything additional.\n\n2. Sampling according to the Karolinska distribution: before each epoch the Radboud images were sampled according to the Karolinska distribution of isup grades (i.e. sample proportion of grades 0 - 5 as Karolinska). The idea was that the Karolinska distribution was more representative of the test set. \n\nCheers",
      "votes": 11
    },
    {
      "id": 943854,
      "postDate": "2020-07-24T16:16:58.547Z",
      "content": "<p>Hi thanks for sharing.\nI tried a similar cutmix augmentation but more \"permissive\" : my second image was randomly selected among all the images i.e. not always the same data provider/isup grade. I then changed the label based on the proportion of image 1 and 2 (I randomly select the number of tiles). So labels could be 3.x, 0.x ...  I've observed a small decrease in CV so I did not push forward the idea</p>",
      "rawMarkdown": "Hi thanks for sharing.\nI tried a similar cutmix augmentation but more \"permissive\" : my second image was randomly selected among all the images i.e. not always the same data provider/isup grade. I then changed the label based on the proportion of image 1 and 2 (I randomly select the number of tiles). So labels could be 3.x, 0.x ...  I've observed a small decrease in CV so I did not push forward the idea",
      "votes": 1,
      "replies": [
        {
          "id": 943868,
          "postDate": "2020-07-24T16:30:52.593Z",
          "content": "<p>So you mixed the images and then tiled the mixed image? Or tiled them individually and then mixed the tiles like I did?</p>",
          "rawMarkdown": "So you mixed the images and then tiled the mixed image? Or tiled them individually and then mixed the tiles like I did?"
        },
        {
          "id": 943913,
          "postDate": "2020-07-24T17:04:53.610Z",
          "content": "<p>Like you did</p>",
          "rawMarkdown": "Like you did",
          "votes": 1
        }
      ]
    },
    {
      "id": 947866,
      "postDate": "2020-07-27T14:25:51.147Z",
      "content": "<p>Thanks for sharing. How did you get the idea to do it?</p>",
      "rawMarkdown": "Thanks for sharing. How did you get the idea to do it?",
      "replies": [
        {
          "id": 947877,
          "postDate": "2020-07-27T14:31:39.070Z",
          "content": "<p>Just thinking about how best to generalise given the limited data!</p>",
          "rawMarkdown": "Just thinking about how best to generalise given the limited data!",
          "votes": 1
        },
        {
          "id": 947897,
          "postDate": "2020-07-27T14:43:24.827Z",
          "content": "<p>Is it like applying MixUp to individual tiles?</p>",
          "rawMarkdown": "Is it like applying MixUp to individual tiles?"
        },
        {
          "id": 947960,
          "postDate": "2020-07-27T15:17:06.447Z",
          "content": "<p>It would be a similar to mixup, but we're keeping the tiles intact (i.e. the individual tiles are not a mix from two images), but then mixing the tile selections from two images</p>",
          "rawMarkdown": "It would be a similar to mixup, but we're keeping the tiles intact (i.e. the individual tiles are not a mix from two images), but then mixing the tile selections from two images"
        }
      ]
    },
    {
      "id": 943327,
      "postDate": "2020-07-24T09:48:20.163Z",
      "content": "<p>Thanks for sharing <a href=\"/fergusoci\">@fergusoci</a> </p>",
      "rawMarkdown": "Thanks for sharing @fergusoci ",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 943854,
      "author_name": "Alex",
      "author_url": "",
      "post_date": "2020-07-24T16:16:58.547000",
      "content": "<p>Hi thanks for sharing.\nI tried a similar cutmix augmentation but more \"permissive\" : my second image was randomly selected among all the images i.e. not always the same data provider/isup grade. I then changed the label based on the proportion of image 1 and 2 (I randomly select the number of tiles). So labels could be 3.x, 0.x ...  I've observed a small decrease in CV so I did not push forward the idea</p>",
      "votes": 1,
      "replies": [
        {
          "id": 943868,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "2020-07-24T16:30:52.593000",
          "content": "<p>So you mixed the images and then tiled the mixed image? Or tiled them individually and then mixed the tiles like I did?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 943913,
          "author_name": "Alex",
          "author_url": "",
          "post_date": "2020-07-24T17:04:53.610000",
          "content": "<p>Like you did</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 947866,
      "author_name": "Son of Anton v3.0",
      "author_url": "",
      "post_date": "2020-07-27T14:25:51.147000",
      "content": "<p>Thanks for sharing. How did you get the idea to do it?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 947877,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "2020-07-27T14:31:39.070000",
          "content": "<p>Just thinking about how best to generalise given the limited data!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 947897,
          "author_name": "Son of Anton v3.0",
          "author_url": "",
          "post_date": "2020-07-27T14:43:24.827000",
          "content": "<p>Is it like applying MixUp to individual tiles?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 947960,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "2020-07-27T15:17:06.447000",
          "content": "<p>It would be a similar to mixup, but we're keeping the tiles intact (i.e. the individual tiles are not a mix from two images), but then mixing the tile selections from two images</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 943327,
      "author_name": "Gryffindor",
      "author_url": "",
      "post_date": "2020-07-24T09:48:20.163000",
      "content": "<p>Thanks for sharing <a href=\"/fergusoci\">@fergusoci</a> </p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "943220": "Hi all,\n\nCongratulations to the winners! \n\nThere were two items that worked that I haven't seen mentioned thus far:\n\n1. Mixed slide augmentations: I got my highest CV scores when using these (and they also generated some of the highest private LB scores). When an image was selected for training I randomly selected a second image from the same data provider, same isup grade, tiled both images, then randomly selected 50% of the tiles from each (and reordered them as you would if you were selecting tiles from a single image). This lead to a +0.005 in CV score. I tried a version whereby , as well as the same data provider + isup, the mix image was of the same tissue colour, but this didn't add anything additional.\n\n2. Sampling according to the Karolinska distribution: before each epoch the Radboud images were sampled according to the Karolinska distribution of isup grades (i.e. sample proportion of grades 0 - 5 as Karolinska). The idea was that the Karolinska distribution was more representative of the test set. \n\nCheers",
    "943854": "Hi thanks for sharing.\nI tried a similar cutmix augmentation but more \"permissive\" : my second image was randomly selected among all the images i.e. not always the same data provider/isup grade. I then changed the label based on the proportion of image 1 and 2 (I randomly select the number of tiles). So labels could be 3.x, 0.x ...  I've observed a small decrease in CV so I did not push forward the idea",
    "947866": "Thanks for sharing. How did you get the idea to do it?",
    "943327": "Thanks for sharing @fergusoci "
  }
}