{
  "id": 369749,
  "title": "⭐️⭐️ Breast Cancer - ROI (brest) extractor ⭐️⭐️",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/369749",
  "author_name": "Remek Kinas",
  "post_date": "2022-12-01T09:28:18.864000",
  "votes": 24,
  "comment_count": 7,
  "views": 0,
  "content": "<p>During the visual data analysis I noticed that there is a large variation in the arrangement of the object in the images. In addition, some objects occupy only a small part of the image. By converting (resizing) without ROI extraction we have a very inefficient use of the reduced image. Most of the picture is blank.</p>\n<p><strong>Solution:</strong></p>\n<ul>\n<li>annotate data - I annotated about 500 images in a human in the loop technique (3 models were created - I started from 300 images and ended up about 500)</li>\n<li>train object detector - I used yolov5 (small - balance between accuracy and speed)</li>\n</ul>\n<p><strong>Result on train DS:</strong></p>\n<ul>\n<li>54601 images processed successfully</li>\n<li>105 images - detection failed</li>\n</ul>\n<p>Model performance (on my validation DS):  mAP@50 -&gt; 0.995, mAP50-95 -&gt; 0.914</p>\n<p><img src=\"https://i.ibb.co/LdrhsWn/br001.jpg\" alt=\"ROI extractor\"></p>\n<p>Notebook: <a href=\"https://www.kaggle.com/code/remekkinas/breast-cancer-roi-brest-extractor\" target=\"_blank\">https://www.kaggle.com/code/remekkinas/breast-cancer-roi-brest-extractor</a></p>",
  "messages": [
    {
      "id": 2051208,
      "postDate": "2022-12-01T09:28:18.863Z",
      "content": "<p>During the visual data analysis I noticed that there is a large variation in the arrangement of the object in the images. In addition, some objects occupy only a small part of the image. By converting (resizing) without ROI extraction we have a very inefficient use of the reduced image. Most of the picture is blank.</p>\n<p><strong>Solution:</strong></p>\n<ul>\n<li>annotate data - I annotated about 500 images in a human in the loop technique (3 models were created - I started from 300 images and ended up about 500)</li>\n<li>train object detector - I used yolov5 (small - balance between accuracy and speed)</li>\n</ul>\n<p><strong>Result on train DS:</strong></p>\n<ul>\n<li>54601 images processed successfully</li>\n<li>105 images - detection failed</li>\n</ul>\n<p>Model performance (on my validation DS):  mAP@50 -&gt; 0.995, mAP50-95 -&gt; 0.914</p>\n<p><img src=\"https://i.ibb.co/LdrhsWn/br001.jpg\" alt=\"ROI extractor\"></p>\n<p>Notebook: <a href=\"https://www.kaggle.com/code/remekkinas/breast-cancer-roi-brest-extractor\" target=\"_blank\">https://www.kaggle.com/code/remekkinas/breast-cancer-roi-brest-extractor</a></p>",
      "rawMarkdown": "During the visual data analysis I noticed that there is a large variation in the arrangement of the object in the images. In addition, some objects occupy only a small part of the image. By converting (resizing) without ROI extraction we have a very inefficient use of the reduced image. Most of the picture is blank.\n\n**Solution:**\n* annotate data - I annotated about 500 images in a human in the loop technique (3 models were created - I started from 300 images and ended up about 500)\n* train object detector - I used yolov5 (small - balance between accuracy and speed)\n\n**Result on train DS:**\n* 54601 images processed successfully\n* 105 images - detection failed\n\nModel performance (on my validation DS):  mAP@50 -> 0.995, mAP50-95 -> 0.914\n\n![ROI extractor](https://i.ibb.co/LdrhsWn/br001.jpg)\n\nNotebook: https://www.kaggle.com/code/remekkinas/breast-cancer-roi-brest-extractor",
      "votes": 24
    },
    {
      "id": 2051338,
      "postDate": "2022-12-01T10:52:30.633Z",
      "content": "<p>Thanks, <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> for sharing. This is great!</p>",
      "rawMarkdown": "Thanks, @remekkinas for sharing. This is great!",
      "votes": 1,
      "replies": [
        {
          "id": 2051346,
          "postDate": "2022-12-01T11:02:14.947Z",
          "content": "<p>Thank you very much :) </p>",
          "rawMarkdown": "Thank you very much :) "
        }
      ]
    },
    {
      "id": 2117043,
      "postDate": "2023-01-26T23:23:39.180Z",
      "content": "<p>With respect, I believe this kind of task does not require the complexity of annotating and then training a model for it. It can be done using a simple function like this: </p>\n<p>def crop_func(img):<br>\n    thresholded = cv2.threshold(np.array(img), 0, 255, cv2.THRESH_OTSU)<br>\n    bbox = cv2.boundingRect(thresholded[1])<br>\n    x, y, w, h = bbox<br>\n    cr_img = img[y:y+h, x:x+w]<br>\n    return cr_img</p>",
      "rawMarkdown": "With respect, I believe this kind of task does not require the complexity of annotating and then training a model for it. It can be done using a simple function like this: \n\ndef crop_func(img):\n    thresholded = cv2.threshold(np.array(img), 0, 255, cv2.THRESH_OTSU)\n    bbox = cv2.boundingRect(thresholded[1])\n    x, y, w, h = bbox\n    cr_img = img[y:y+h, x:x+w]\n    return cr_img",
      "replies": [
        {
          "id": 2117130,
          "postDate": "2023-01-27T03:29:21.377Z",
          "content": "<p>There are always outliers and a detection model is able to generalize much better (almost 100% coverage) from my tests.</p>",
          "rawMarkdown": "There are always outliers and a detection model is able to generalize much better (almost 100% coverage) from my tests."
        },
        {
          "id": 2117352,
          "postDate": "2023-01-27T08:01:36.007Z",
          "content": "<p>The way you presented is totally different. </p>\n<ul>\n<li>you do not crop ROI - you crop contour</li>\n<li>my way is to crop breast (center part)</li>\n</ul>\n<p><img src=\"https://i.ibb.co/d7YkpqF/roi.jpg\" alt=\"\"></p>\n<p>AI/ML (art of details)</p>",
          "rawMarkdown": "The way you presented is totally different. \n- you do not crop ROI - you crop contour\n- my way is to crop breast (center part)\n\n![](https://i.ibb.co/d7YkpqF/roi.jpg)\n\nAI/ML (art of details)",
          "votes": 1
        }
      ]
    },
    {
      "id": 2053236,
      "postDate": "2022-12-03T02:01:54.177Z",
      "content": "<p>Thanks for sharing, really good dataset. I seem to be getting a slightly lower CV score with ROI cropping. Did this work for anyone else? <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> How did this go for you?</p>",
      "rawMarkdown": "Thanks for sharing, really good dataset. I seem to be getting a slightly lower CV score with ROI cropping. Did this work for anyone else? @remekkinas How did this go for you?",
      "replies": [
        {
          "id": 2053255,
          "postDate": "2022-12-03T03:21:26.327Z",
          "content": "<p>Ok, it started doing better at later epochs. Good stuff 👍</p>",
          "rawMarkdown": "Ok, it started doing better at later epochs. Good stuff 👍",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2051338,
      "author_name": "Mohammad Dehghanmanshadi",
      "author_url": "",
      "post_date": "2022-12-01T10:52:30.633000",
      "content": "<p>Thanks, <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> for sharing. This is great!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2051346,
          "author_name": "Remek Kinas",
          "author_url": "",
          "post_date": "2022-12-01T11:02:14.947000",
          "content": "<p>Thank you very much :) </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2117043,
      "author_name": "Sina Mehdinia",
      "author_url": "",
      "post_date": "2023-01-26T23:23:39.180000",
      "content": "<p>With respect, I believe this kind of task does not require the complexity of annotating and then training a model for it. It can be done using a simple function like this: </p>\n<p>def crop_func(img):<br>\n    thresholded = cv2.threshold(np.array(img), 0, 255, cv2.THRESH_OTSU)<br>\n    bbox = cv2.boundingRect(thresholded[1])<br>\n    x, y, w, h = bbox<br>\n    cr_img = img[y:y+h, x:x+w]<br>\n    return cr_img</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2117130,
          "author_name": "outwrest",
          "author_url": "",
          "post_date": "2023-01-27T03:29:21.377000",
          "content": "<p>There are always outliers and a detection model is able to generalize much better (almost 100% coverage) from my tests.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2117352,
          "author_name": "Remek Kinas",
          "author_url": "",
          "post_date": "2023-01-27T08:01:36.007000",
          "content": "<p>The way you presented is totally different. </p>\n<ul>\n<li>you do not crop ROI - you crop contour</li>\n<li>my way is to crop breast (center part)</li>\n</ul>\n<p><img src=\"https://i.ibb.co/d7YkpqF/roi.jpg\" alt=\"\"></p>\n<p>AI/ML (art of details)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2053236,
      "author_name": "Ivan Aerlic",
      "author_url": "",
      "post_date": "2022-12-03T02:01:54.177000",
      "content": "<p>Thanks for sharing, really good dataset. I seem to be getting a slightly lower CV score with ROI cropping. Did this work for anyone else? <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> How did this go for you?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2053255,
          "author_name": "Ivan Aerlic",
          "author_url": "",
          "post_date": "2022-12-03T03:21:26.327000",
          "content": "<p>Ok, it started doing better at later epochs. Good stuff 👍</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2051208": "During the visual data analysis I noticed that there is a large variation in the arrangement of the object in the images. In addition, some objects occupy only a small part of the image. By converting (resizing) without ROI extraction we have a very inefficient use of the reduced image. Most of the picture is blank.\n\n**Solution:**\n* annotate data - I annotated about 500 images in a human in the loop technique (3 models were created - I started from 300 images and ended up about 500)\n* train object detector - I used yolov5 (small - balance between accuracy and speed)\n\n**Result on train DS:**\n* 54601 images processed successfully\n* 105 images - detection failed\n\nModel performance (on my validation DS):  mAP@50 -> 0.995, mAP50-95 -> 0.914\n\n![ROI extractor](https://i.ibb.co/LdrhsWn/br001.jpg)\n\nNotebook: https://www.kaggle.com/code/remekkinas/breast-cancer-roi-brest-extractor",
    "2051338": "Thanks, @remekkinas for sharing. This is great!",
    "2117043": "With respect, I believe this kind of task does not require the complexity of annotating and then training a model for it. It can be done using a simple function like this: \n\ndef crop_func(img):\n    thresholded = cv2.threshold(np.array(img), 0, 255, cv2.THRESH_OTSU)\n    bbox = cv2.boundingRect(thresholded[1])\n    x, y, w, h = bbox\n    cr_img = img[y:y+h, x:x+w]\n    return cr_img",
    "2053236": "Thanks for sharing, really good dataset. I seem to be getting a slightly lower CV score with ROI cropping. Did this work for anyone else? @remekkinas How did this go for you?"
  }
}