{
  "id": 340960,
  "title": "Image preprocessing idea, cutting similar but duplicate objects into multiple slides",
  "url": "/competitions/mayo-clinic-strip-ai/discussion/340960",
  "author_name": "yqz",
  "post_date": "2022-07-31T17:36:02.709000",
  "votes": 6,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I think many of us by now have noticed that there are many slides where there are images of what seem to be the same object, duplicated 2x or 3x times, although they do have minor differences between them. In many cases there will also be a good deal of empty space separating these objects.</p>\n<p>Is there a way to perhaps crop these duplicate objects into separate images while giving them the same target label? Or perhaps it would make sense to treat them as the same bag in the case of MIL (i.e., we tile the image, potentially ignoring empty areas, and treat all the tiles as belonging to one image with one label)? I'm worried the former maybe is too much of a task to try and automate.</p>",
  "messages": [
    {
      "id": 1878880,
      "postDate": "2022-07-31T17:36:02.710Z",
      "content": "<p>I think many of us by now have noticed that there are many slides where there are images of what seem to be the same object, duplicated 2x or 3x times, although they do have minor differences between them. In many cases there will also be a good deal of empty space separating these objects.</p>\n<p>Is there a way to perhaps crop these duplicate objects into separate images while giving them the same target label? Or perhaps it would make sense to treat them as the same bag in the case of MIL (i.e., we tile the image, potentially ignoring empty areas, and treat all the tiles as belonging to one image with one label)? I'm worried the former maybe is too much of a task to try and automate.</p>",
      "rawMarkdown": "I think many of us by now have noticed that there are many slides where there are images of what seem to be the same object, duplicated 2x or 3x times, although they do have minor differences between them. In many cases there will also be a good deal of empty space separating these objects.\n\nIs there a way to perhaps crop these duplicate objects into separate images while giving them the same target label? Or perhaps it would make sense to treat them as the same bag in the case of MIL (i.e., we tile the image, potentially ignoring empty areas, and treat all the tiles as belonging to one image with one label)? I'm worried the former maybe is too much of a task to try and automate.",
      "votes": 6
    },
    {
      "id": 1880309,
      "postDate": "2022-08-01T15:59:04.680Z",
      "content": "<p>I'm not sure how the minor differences between the slices translate to actual signal corresponding to the labels, but in both cases a similar amount of effort for data engineering is required I think. You might as well the separate image method first since it's more straightforward with tf.data or torch data loader API. </p>",
      "rawMarkdown": "I'm not sure how the minor differences between the slices translate to actual signal corresponding to the labels, but in both cases a similar amount of effort for data engineering is required I think. You might as well the separate image method first since it's more straightforward with tf.data or torch data loader API. ",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 1880309,
      "author_name": "Phan Nguyen",
      "author_url": "",
      "post_date": "2022-08-01T15:59:04.680000",
      "content": "<p>I'm not sure how the minor differences between the slices translate to actual signal corresponding to the labels, but in both cases a similar amount of effort for data engineering is required I think. You might as well the separate image method first since it's more straightforward with tf.data or torch data loader API. </p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1878880": "I think many of us by now have noticed that there are many slides where there are images of what seem to be the same object, duplicated 2x or 3x times, although they do have minor differences between them. In many cases there will also be a good deal of empty space separating these objects.\n\nIs there a way to perhaps crop these duplicate objects into separate images while giving them the same target label? Or perhaps it would make sense to treat them as the same bag in the case of MIL (i.e., we tile the image, potentially ignoring empty areas, and treat all the tiles as belonging to one image with one label)? I'm worried the former maybe is too much of a task to try and automate.",
    "1880309": "I'm not sure how the minor differences between the slices translate to actual signal corresponding to the labels, but in both cases a similar amount of effort for data engineering is required I think. You might as well the separate image method first since it's more straightforward with tf.data or torch data loader API. "
  }
}