{
  "id": 145280,
  "title": "Are Radboud and Karolinska using different process for imaging ? ",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/145280",
  "author_name": "Benjamin Dubreu",
  "post_date": "2020-04-22T14:39:10.538000",
  "votes": 2,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi everyone,</p>\n\n<p>I quickly put together <a href=\"https://www.kaggle.com/bdubreu/differences-between-radboud-and-karolinska-data\">this kernel</a> , and it seems like the data coming from Radboud is very different from that coming from karolinska. </p>\n\n<p>I worry we might be feeding our models pictures of whole dogs on the one hand, and a zoom on dog faces on the other hand (or something along those lines).</p>\n\n<p>Any thoughts ? </p>\n\n<p>Organizing team: can we have more info on how the two different institutions process their data ? </p>\n\n<p>Thanks !</p>",
  "messages": [
    {
      "id": 816766,
      "postDate": "2020-04-22T15:12:38.870Z",
      "content": "<p>Hi Benjamin,</p>\n\n<p>Karolinska has used two scanners (Hamamatsu and Aperio) and Radboud has used one (3DHistech). There can be some optical differences such slight color shifts. Each of the three scanners has a unique pixel spacing, so it is easy to tell them apart (look for example at the introduction notebook; some 5-10% more or less tissue fit in one pixel depending on the scanner used). Feel free to either resize your training data so all images/patches align in pixel spacing or use as is. All images (Karolinska and Radboud) have been anonymized by striping some meta data, and in this process been converted to tiff format. So, the images are comparable between both sites, but there can be slight differences in the lab procedure (e.g. Radboud biopsies are sometime colored arond the edges of the tissue) and scanner. In all cases each image contain one biopsy needle of tissue, but there may be some images from Karolinska that contain two sections of the same biopsy needle (i.e. two thin slices of the tissue).</p>\n\n<p>Best,\nPeter</p>",
      "rawMarkdown": "Hi Benjamin,\n\nKarolinska has used two scanners (Hamamatsu and Aperio) and Radboud has used one (3DHistech). There can be some optical differences such slight color shifts. Each of the three scanners has a unique pixel spacing, so it is easy to tell them apart (look for example at the introduction notebook; some 5-10% more or less tissue fit in one pixel depending on the scanner used). Feel free to either resize your training data so all images/patches align in pixel spacing or use as is. All images (Karolinska and Radboud) have been anonymized by striping some meta data, and in this process been converted to tiff format. So, the images are comparable between both sites, but there can be slight differences in the lab procedure (e.g. Radboud biopsies are sometime colored arond the edges of the tissue) and scanner. In all cases each image contain one biopsy needle of tissue, but there may be some images from Karolinska that contain two sections of the same biopsy needle (i.e. two thin slices of the tissue).\n\nBest,\nPeter\n",
      "votes": 7,
      "replies": [
        {
          "id": 816779,
          "postDate": "2020-04-22T15:33:22.043Z",
          "content": "<p>To add a point for the Radboud data: these images are from biopsy procedures that were taken over a time span of around 5 years. You will see that there can also be a lot of variance within data from a single center. Different staining protocols, the temperature in the lab, etc., all these things can influence the final appearance. This variance is one of the big challenges in AI applications for pathology: how can we build algorithms today that also work in a different hospital tomorrow? We are very excited to see your methods on how to address this!</p>",
          "rawMarkdown": "To add a point for the Radboud data: these images are from biopsy procedures that were taken over a time span of around 5 years. You will see that there can also be a lot of variance within data from a single center. Different staining protocols, the temperature in the lab, etc., all these things can influence the final appearance. This variance is one of the big challenges in AI applications for pathology: how can we build algorithms today that also work in a different hospital tomorrow? We are very excited to see your methods on how to address this!",
          "votes": 7
        },
        {
          "id": 816799,
          "postDate": "2020-04-22T15:46:25.577Z",
          "content": "<p>Makes sense, thanks for your answers guys !</p>",
          "rawMarkdown": "Makes sense, thanks for your answers guys !"
        }
      ]
    },
    {
      "id": 816730,
      "postDate": "2020-04-22T14:39:10.537Z",
      "content": "<p>Hi everyone,</p>\n\n<p>I quickly put together <a href=\"https://www.kaggle.com/bdubreu/differences-between-radboud-and-karolinska-data\">this kernel</a> , and it seems like the data coming from Radboud is very different from that coming from karolinska. </p>\n\n<p>I worry we might be feeding our models pictures of whole dogs on the one hand, and a zoom on dog faces on the other hand (or something along those lines).</p>\n\n<p>Any thoughts ? </p>\n\n<p>Organizing team: can we have more info on how the two different institutions process their data ? </p>\n\n<p>Thanks !</p>",
      "rawMarkdown": "Hi everyone,\n\nI quickly put together [this kernel](https://www.kaggle.com/bdubreu/differences-between-radboud-and-karolinska-data) , and it seems like the data coming from Radboud is very different from that coming from karolinska. \n\nI worry we might be feeding our models pictures of whole dogs on the one hand, and a zoom on dog faces on the other hand (or something along those lines).\n\nAny thoughts ? \n\nOrganizing team: can we have more info on how the two different institutions process their data ? \n\nThanks !",
      "votes": 2
    },
    {
      "id": 816955,
      "postDate": "2020-04-22T18:22:52.817Z",
      "content": "<p>Thank a lot</p>",
      "rawMarkdown": "Thank a lot"
    }
  ],
  "comments": [
    {
      "id": 816766,
      "author_name": "PeterStröm",
      "author_url": "",
      "post_date": "2020-04-22T15:12:38.870000",
      "content": "<p>Hi Benjamin,</p>\n\n<p>Karolinska has used two scanners (Hamamatsu and Aperio) and Radboud has used one (3DHistech). There can be some optical differences such slight color shifts. Each of the three scanners has a unique pixel spacing, so it is easy to tell them apart (look for example at the introduction notebook; some 5-10% more or less tissue fit in one pixel depending on the scanner used). Feel free to either resize your training data so all images/patches align in pixel spacing or use as is. All images (Karolinska and Radboud) have been anonymized by striping some meta data, and in this process been converted to tiff format. So, the images are comparable between both sites, but there can be slight differences in the lab procedure (e.g. Radboud biopsies are sometime colored arond the edges of the tissue) and scanner. In all cases each image contain one biopsy needle of tissue, but there may be some images from Karolinska that contain two sections of the same biopsy needle (i.e. two thin slices of the tissue).</p>\n\n<p>Best,\nPeter</p>",
      "votes": 7,
      "replies": [
        {
          "id": 816779,
          "author_name": "Wouter Bulten",
          "author_url": "",
          "post_date": "2020-04-22T15:33:22.043000",
          "content": "<p>To add a point for the Radboud data: these images are from biopsy procedures that were taken over a time span of around 5 years. You will see that there can also be a lot of variance within data from a single center. Different staining protocols, the temperature in the lab, etc., all these things can influence the final appearance. This variance is one of the big challenges in AI applications for pathology: how can we build algorithms today that also work in a different hospital tomorrow? We are very excited to see your methods on how to address this!</p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 816799,
          "author_name": "Benjamin Dubreu",
          "author_url": "",
          "post_date": "2020-04-22T15:46:25.577000",
          "content": "<p>Makes sense, thanks for your answers guys !</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 816955,
      "author_name": "Ashish Pandey",
      "author_url": "",
      "post_date": "2020-04-22T18:22:52.817000",
      "content": "<p>Thank a lot</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "816766": "Hi Benjamin,\n\nKarolinska has used two scanners (Hamamatsu and Aperio) and Radboud has used one (3DHistech). There can be some optical differences such slight color shifts. Each of the three scanners has a unique pixel spacing, so it is easy to tell them apart (look for example at the introduction notebook; some 5-10% more or less tissue fit in one pixel depending on the scanner used). Feel free to either resize your training data so all images/patches align in pixel spacing or use as is. All images (Karolinska and Radboud) have been anonymized by striping some meta data, and in this process been converted to tiff format. So, the images are comparable between both sites, but there can be slight differences in the lab procedure (e.g. Radboud biopsies are sometime colored arond the edges of the tissue) and scanner. In all cases each image contain one biopsy needle of tissue, but there may be some images from Karolinska that contain two sections of the same biopsy needle (i.e. two thin slices of the tissue).\n\nBest,\nPeter\n",
    "816730": "Hi everyone,\n\nI quickly put together [this kernel](https://www.kaggle.com/bdubreu/differences-between-radboud-and-karolinska-data) , and it seems like the data coming from Radboud is very different from that coming from karolinska. \n\nI worry we might be feeding our models pictures of whole dogs on the one hand, and a zoom on dog faces on the other hand (or something along those lines).\n\nAny thoughts ? \n\nOrganizing team: can we have more info on how the two different institutions process their data ? \n\nThanks !",
    "816955": "Thank a lot"
  }
}