{
  "id": 445575,
  "title": "Any reason behind why micro tissue array count is just 25? Is there any use of is_tma if its only available in training?",
  "url": "/competitions/UBC-OCEAN/discussion/445575",
  "author_name": "Prithviraj",
  "post_date": "2023-10-07T16:41:02.973000",
  "votes": 2,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Does the tissue being micro tissue have any significance on the label? What would be the use of is_tma if it is only available for training data?</p>",
  "messages": [
    {
      "id": 2472916,
      "postDate": "2023-10-07T17:57:51.820Z",
      "content": "<p>Hi Prithviraj!<br>\nI am not sure what you mean by a count of 5. According to <code>train.csv</code>, we have 513 whole-slide images and 25 tissue micro-arrays.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14817755%2Ffa0d9e2286f66a5ae676d36984a488ce%2Ftemp.jpg?generation=1696700749351612&amp;alt=media\" alt=\"\"></p>\n<p>I think we should just take care of pre-processing to render both WSLs and TMAs close enough for the model to focus on the important details of the images (i.e. cellular structures) and ignore insignificant differences between the two types of images.</p>\n<p>Also, take note that in the training images, the majority of the images are WSLs, while in the testing data, the majority of the images will be TMAs. So, our models have to generalize well from WSLs to TMAs, and to images from other hospitals (other hospitals will likely have different treatments to the tissues that will results in slightly different staining (colors) and detail (sharpness &amp; blur)). In my opinion, pre-processing (including image augmentations) will be a crucial part of solving these problems (probably even more important than the model itself).</p>",
      "rawMarkdown": "Hi Prithviraj!\nI am not sure what you mean by a count of 5. According to `train.csv`, we have 513 whole-slide images and 25 tissue micro-arrays.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14817755%2Ffa0d9e2286f66a5ae676d36984a488ce%2Ftemp.jpg?generation=1696700749351612&alt=media)\n\nI think we should just take care of pre-processing to render both WSLs and TMAs close enough for the model to focus on the important details of the images (i.e. cellular structures) and ignore insignificant differences between the two types of images.\n\nAlso, take note that in the training images, the majority of the images are WSLs, while in the testing data, the majority of the images will be TMAs. So, our models have to generalize well from WSLs to TMAs, and to images from other hospitals (other hospitals will likely have different treatments to the tissues that will results in slightly different staining (colors) and detail (sharpness & blur)). In my opinion, pre-processing (including image augmentations) will be a crucial part of solving these problems (probably even more important than the model itself).",
      "votes": 6
    },
    {
      "id": 2482634,
      "postDate": "2023-10-15T07:38:25.277Z",
      "content": "<p>The reason is enforcing participants to build models that could generalize across different image types. Domain adaptation is a challenging problem in medical imaging.</p>",
      "rawMarkdown": "The reason is enforcing participants to build models that could generalize across different image types. Domain adaptation is a challenging problem in medical imaging.",
      "votes": 1
    },
    {
      "id": 2477538,
      "postDate": "2023-10-11T11:11:44.990Z",
      "content": "<p>Although the tissue microarrays are smaller images, after viewing them, I think they provide excellent image data of the tumor class without too much extra unneeded pixels.</p>",
      "rawMarkdown": "Although the tissue microarrays are smaller images, after viewing them, I think they provide excellent image data of the tumor class without too much extra unneeded pixels.",
      "votes": 1,
      "replies": [
        {
          "id": 2478092,
          "postDate": "2023-10-11T18:40:00.380Z",
          "content": "<p>Hi Noli!<br>\nI agree. They are. However, more data is almost always better, my friend :)… The main challenge in the competition is to be able to generalize well to out-of-distribution data, and it would be a lot harder with only 25 images.<br>\nI suggest that you can use the TMAs initially to build a performant model, and then use the WSI when you need more data (and you probably will). I also suggest that you use the thumbnails as a mask that you can use to read in certain tiles off the full-sized images. I mean, you can turn the thumbnail into a mask that determines which part(s) of the image is blank and which part contains actual cell data, and then you can only read in tiles that 95% belongs within the data part off the original images. This way, you will be able to read in small parts of the WSIs at a time, will get more quality data out of them, and won't waste your resources reading in useless emptiness.</p>",
          "rawMarkdown": "Hi Noli!\nI agree. They are. However, more data is almost always better, my friend :)... The main challenge in the competition is to be able to generalize well to out-of-distribution data, and it would be a lot harder with only 25 images.\nI suggest that you can use the TMAs initially to build a performant model, and then use the WSI when you need more data (and you probably will). I also suggest that you use the thumbnails as a mask that you can use to read in certain tiles off the full-sized images. I mean, you can turn the thumbnail into a mask that determines which part(s) of the image is blank and which part contains actual cell data, and then you can only read in tiles that 95% belongs within the data part off the original images. This way, you will be able to read in small parts of the WSIs at a time, will get more quality data out of them, and won't waste your resources reading in useless emptiness.",
          "votes": 5,
          "replies": [
            {
              "id": 2538736,
              "postDate": "2023-11-26T12:54:20.017Z",
              "content": "<p>Thanks! This idea sounds pretty nice. </p>",
              "rawMarkdown": "Thanks! This idea sounds pretty nice. "
            }
          ]
        }
      ]
    },
    {
      "id": 2472849,
      "postDate": "2023-10-07T16:41:02.973Z",
      "content": "<p>Does the tissue being micro tissue have any significance on the label? What would be the use of is_tma if it is only available for training data?</p>",
      "rawMarkdown": "Does the tissue being micro tissue have any significance on the label? What would be the use of is_tma if it is only available for training data?",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 2472916,
      "author_name": "Mohamed AbdulMaksoud",
      "author_url": "",
      "post_date": "2023-10-07T17:57:51.820000",
      "content": "<p>Hi Prithviraj!<br>\nI am not sure what you mean by a count of 5. According to <code>train.csv</code>, we have 513 whole-slide images and 25 tissue micro-arrays.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14817755%2Ffa0d9e2286f66a5ae676d36984a488ce%2Ftemp.jpg?generation=1696700749351612&amp;alt=media\" alt=\"\"></p>\n<p>I think we should just take care of pre-processing to render both WSLs and TMAs close enough for the model to focus on the important details of the images (i.e. cellular structures) and ignore insignificant differences between the two types of images.</p>\n<p>Also, take note that in the training images, the majority of the images are WSLs, while in the testing data, the majority of the images will be TMAs. So, our models have to generalize well from WSLs to TMAs, and to images from other hospitals (other hospitals will likely have different treatments to the tissues that will results in slightly different staining (colors) and detail (sharpness &amp; blur)). In my opinion, pre-processing (including image augmentations) will be a crucial part of solving these problems (probably even more important than the model itself).</p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 2482634,
      "author_name": "Gunes Evitan",
      "author_url": "",
      "post_date": "2023-10-15T07:38:25.277000",
      "content": "<p>The reason is enforcing participants to build models that could generalize across different image types. Domain adaptation is a challenging problem in medical imaging.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2477538,
      "author_name": "Noli Alonso",
      "author_url": "",
      "post_date": "2023-10-11T11:11:44.990000",
      "content": "<p>Although the tissue microarrays are smaller images, after viewing them, I think they provide excellent image data of the tumor class without too much extra unneeded pixels.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2478092,
          "author_name": "Mohamed AbdulMaksoud",
          "author_url": "",
          "post_date": "2023-10-11T18:40:00.380000",
          "content": "<p>Hi Noli!<br>\nI agree. They are. However, more data is almost always better, my friend :)… The main challenge in the competition is to be able to generalize well to out-of-distribution data, and it would be a lot harder with only 25 images.<br>\nI suggest that you can use the TMAs initially to build a performant model, and then use the WSI when you need more data (and you probably will). I also suggest that you use the thumbnails as a mask that you can use to read in certain tiles off the full-sized images. I mean, you can turn the thumbnail into a mask that determines which part(s) of the image is blank and which part contains actual cell data, and then you can only read in tiles that 95% belongs within the data part off the original images. This way, you will be able to read in small parts of the WSIs at a time, will get more quality data out of them, and won't waste your resources reading in useless emptiness.</p>",
          "votes": 5,
          "replies": [
            {
              "id": 2538736,
              "author_name": "Akash Gupta",
              "author_url": "",
              "post_date": "2023-11-26T12:54:20.017000",
              "content": "<p>Thanks! This idea sounds pretty nice. </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2472916": "Hi Prithviraj!\nI am not sure what you mean by a count of 5. According to `train.csv`, we have 513 whole-slide images and 25 tissue micro-arrays.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14817755%2Ffa0d9e2286f66a5ae676d36984a488ce%2Ftemp.jpg?generation=1696700749351612&alt=media)\n\nI think we should just take care of pre-processing to render both WSLs and TMAs close enough for the model to focus on the important details of the images (i.e. cellular structures) and ignore insignificant differences between the two types of images.\n\nAlso, take note that in the training images, the majority of the images are WSLs, while in the testing data, the majority of the images will be TMAs. So, our models have to generalize well from WSLs to TMAs, and to images from other hospitals (other hospitals will likely have different treatments to the tissues that will results in slightly different staining (colors) and detail (sharpness & blur)). In my opinion, pre-processing (including image augmentations) will be a crucial part of solving these problems (probably even more important than the model itself).",
    "2482634": "The reason is enforcing participants to build models that could generalize across different image types. Domain adaptation is a challenging problem in medical imaging.",
    "2477538": "Although the tissue microarrays are smaller images, after viewing them, I think they provide excellent image data of the tumor class without too much extra unneeded pixels.",
    "2472849": "Does the tissue being micro tissue have any significance on the label? What would be the use of is_tma if it is only available for training data?"
  }
}