{
  "id": 383277,
  "title": "Images with text!",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/383277",
  "author_name": "Yacine Bouaouni",
  "post_date": "2023-02-03T00:25:10.226000",
  "votes": 1,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Hello Kagglers,</p>\n<p>Did you notice that some images have text written on them, did that disturb your models?</p>",
  "messages": [
    {
      "id": 2128701,
      "postDate": "2023-02-04T00:18:57.403Z",
      "content": "<p>I usually try to get rid of texts that have been burned into the radiological images. I have read that those can disturb CNN's, but on the other hand that might have been only an issue related to a small size of the dataset.</p>",
      "rawMarkdown": "I usually try to get rid of texts that have been burned into the radiological images. I have read that those can disturb CNN's, but on the other hand that might have been only an issue related to a small size of the dataset.",
      "votes": 1
    },
    {
      "id": 2128255,
      "postDate": "2023-02-03T15:14:46.447Z",
      "content": "<p>Someone could use GradCam (<a href=\"https://github.com/jacobgil/pytorch-grad-cam\" target=\"_blank\">https://github.com/jacobgil/pytorch-grad-cam</a>) to see if the text disturbs the model to recognize the right class.</p>",
      "rawMarkdown": "Someone could use GradCam (https://github.com/jacobgil/pytorch-grad-cam) to see if the text disturbs the model to recognize the right class.",
      "votes": 1
    },
    {
      "id": 2127525,
      "postDate": "2023-02-03T00:25:10.227Z",
      "content": "<p>Hello Kagglers,</p>\n<p>Did you notice that some images have text written on them, did that disturb your models?</p>",
      "rawMarkdown": "Hello Kagglers,\n\nDid you notice that some images have text written on them, did that disturb your models?",
      "votes": 1
    },
    {
      "id": 2128815,
      "postDate": "2023-02-04T05:30:18.653Z",
      "content": "<p>It looks like those watermarks are at maximum pixel value, then you could remove it with temp=np.where(pixel&lt;max(),pixel,0). As I used this temporal image to place landmarks to crop, it doesn't hurt the original maximum pixels.</p>",
      "rawMarkdown": "It looks like those watermarks are at maximum pixel value, then you could remove it with temp=np.where(pixel<max(),pixel,0). As I used this temporal image to place landmarks to crop, it doesn't hurt the original maximum pixels."
    },
    {
      "id": 2128057,
      "postDate": "2023-02-03T12:46:10.703Z",
      "content": "<p>Hi, <a href=\"https://www.kaggle.com/jarvisai7\" target=\"_blank\">@jarvisai7</a> any idea how can i remove text(watermark) from image in <a href=\"https://www.kaggle.com/datasets/shubhamgoel27/dermnet\" target=\"_blank\">this dataset</a>. I traid cv2.inpaint to remove watermark but after removing watermark, watermark area got blur.<br>\nhelp me to remove text!<br>\nThanks</p>",
      "rawMarkdown": "Hi, @jarvisai7 any idea how can i remove text(watermark) from image in [this dataset](https://www.kaggle.com/datasets/shubhamgoel27/dermnet). I traid cv2.inpaint to remove watermark but after removing watermark, watermark area got blur.\nhelp me to remove text!\nThanks",
      "replies": [
        {
          "id": 2128519,
          "postDate": "2023-02-03T19:01:46.900Z",
          "content": "<p><a href=\"https://www.kaggle.com/vipin20\" target=\"_blank\">@vipin20</a> Perhaps by looking for the largest connected component?</p>",
          "rawMarkdown": "@vipin20 Perhaps by looking for the largest connected component?",
          "votes": 1
        },
        {
          "id": 2128553,
          "postDate": "2023-02-03T19:52:25.097Z",
          "content": "<p>I would use a deep learning SOTA model for inpainting from Hugging face maybe. Check this one (<a href=\"https://github.com/Sanster/lama-cleaner)\" target=\"_blank\">https://github.com/Sanster/lama-cleaner)</a>. I believe it will give you better results than opencv inpaint. </p>",
          "rawMarkdown": "I would use a deep learning SOTA model for inpainting from Hugging face maybe. Check this one (https://github.com/Sanster/lama-cleaner). I believe it will give you better results than opencv inpaint. ",
          "votes": 1,
          "replies": [
            {
              "id": 2129076,
              "postDate": "2023-02-04T09:49:18.490Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        },
        {
          "id": 2129077,
          "postDate": "2023-02-04T09:49:38.283Z",
          "content": "<p><a href=\"https://www.kaggle.com/davidbroberts\" target=\"_blank\">@davidbroberts</a> wrote <a href=\"https://www.kaggle.com/code/davidbroberts/mammography-remove-letter-markers\" target=\"_blank\">a notebook</a> to remove those text markers</p>",
          "rawMarkdown": "@davidbroberts wrote [a notebook](https://www.kaggle.com/code/davidbroberts/mammography-remove-letter-markers) to remove those text markers\n"
        }
      ]
    },
    {
      "id": 2127857,
      "postDate": "2023-02-03T08:55:40.667Z",
      "content": "<p>Those texts explain which angles were used to accomplish the mammography. like the CC/MLO view</p>",
      "rawMarkdown": "Those texts explain which angles were used to accomplish the mammography. like the CC/MLO view",
      "replies": [
        {
          "id": 2127861,
          "postDate": "2023-02-03T09:19:39.170Z",
          "content": "<p>Yes thank you, but did you find that it somehow affects the models? </p>",
          "rawMarkdown": "Yes thank you, but did you find that it somehow affects the models? ",
          "replies": [
            {
              "id": 2127907,
              "postDate": "2023-02-03T10:26:16.470Z",
              "content": "<p>I think by cropping ROI you able to exclude this text signs from training. See eg <a href=\"https://www.kaggle.com/code/salmanahmedtamu/faster-dicom-loading-and-cropping\" target=\"_blank\">https://www.kaggle.com/code/salmanahmedtamu/faster-dicom-loading-and-cropping</a></p>",
              "rawMarkdown": "I think by cropping ROI you able to exclude this text signs from training. See eg https://www.kaggle.com/code/salmanahmedtamu/faster-dicom-loading-and-cropping",
              "votes": 1
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2128701,
      "author_name": "Antti Isosalo",
      "author_url": "",
      "post_date": "2023-02-04T00:18:57.403000",
      "content": "<p>I usually try to get rid of texts that have been burned into the radiological images. I have read that those can disturb CNN's, but on the other hand that might have been only an issue related to a small size of the dataset.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2128255,
      "author_name": "Ali Abdin",
      "author_url": "",
      "post_date": "2023-02-03T15:14:46.447000",
      "content": "<p>Someone could use GradCam (<a href=\"https://github.com/jacobgil/pytorch-grad-cam\" target=\"_blank\">https://github.com/jacobgil/pytorch-grad-cam</a>) to see if the text disturbs the model to recognize the right class.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2128815,
      "author_name": "Alfredo Maussa",
      "author_url": "",
      "post_date": "2023-02-04T05:30:18.653000",
      "content": "<p>It looks like those watermarks are at maximum pixel value, then you could remove it with temp=np.where(pixel&lt;max(),pixel,0). As I used this temporal image to place landmarks to crop, it doesn't hurt the original maximum pixels.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2128057,
      "author_name": "Vipin Kumar",
      "author_url": "",
      "post_date": "2023-02-03T12:46:10.703000",
      "content": "<p>Hi, <a href=\"https://www.kaggle.com/jarvisai7\" target=\"_blank\">@jarvisai7</a> any idea how can i remove text(watermark) from image in <a href=\"https://www.kaggle.com/datasets/shubhamgoel27/dermnet\" target=\"_blank\">this dataset</a>. I traid cv2.inpaint to remove watermark but after removing watermark, watermark area got blur.<br>\nhelp me to remove text!<br>\nThanks</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2128519,
          "author_name": "Antti Isosalo",
          "author_url": "",
          "post_date": "2023-02-03T19:01:46.900000",
          "content": "<p><a href=\"https://www.kaggle.com/vipin20\" target=\"_blank\">@vipin20</a> Perhaps by looking for the largest connected component?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2128553,
          "author_name": "Yacine Bouaouni",
          "author_url": "",
          "post_date": "2023-02-03T19:52:25.097000",
          "content": "<p>I would use a deep learning SOTA model for inpainting from Hugging face maybe. Check this one (<a href=\"https://github.com/Sanster/lama-cleaner)\" target=\"_blank\">https://github.com/Sanster/lama-cleaner)</a>. I believe it will give you better results than opencv inpaint. </p>",
          "votes": 1,
          "replies": [
            {
              "id": 2129076,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-02-04T09:49:18.490000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2129077,
          "author_name": "FabienDaniel",
          "author_url": "",
          "post_date": "2023-02-04T09:49:38.283000",
          "content": "<p><a href=\"https://www.kaggle.com/davidbroberts\" target=\"_blank\">@davidbroberts</a> wrote <a href=\"https://www.kaggle.com/code/davidbroberts/mammography-remove-letter-markers\" target=\"_blank\">a notebook</a> to remove those text markers</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2127857,
      "author_name": "Rahel Sarif",
      "author_url": "",
      "post_date": "2023-02-03T08:55:40.667000",
      "content": "<p>Those texts explain which angles were used to accomplish the mammography. like the CC/MLO view</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2127861,
          "author_name": "Yacine Bouaouni",
          "author_url": "",
          "post_date": "2023-02-03T09:19:39.170000",
          "content": "<p>Yes thank you, but did you find that it somehow affects the models? </p>",
          "votes": 0,
          "replies": [
            {
              "id": 2127907,
              "author_name": "A.P.",
              "author_url": "",
              "post_date": "2023-02-03T10:26:16.470000",
              "content": "<p>I think by cropping ROI you able to exclude this text signs from training. See eg <a href=\"https://www.kaggle.com/code/salmanahmedtamu/faster-dicom-loading-and-cropping\" target=\"_blank\">https://www.kaggle.com/code/salmanahmedtamu/faster-dicom-loading-and-cropping</a></p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2128701": "I usually try to get rid of texts that have been burned into the radiological images. I have read that those can disturb CNN's, but on the other hand that might have been only an issue related to a small size of the dataset.",
    "2128255": "Someone could use GradCam (https://github.com/jacobgil/pytorch-grad-cam) to see if the text disturbs the model to recognize the right class.",
    "2127525": "Hello Kagglers,\n\nDid you notice that some images have text written on them, did that disturb your models?",
    "2128815": "It looks like those watermarks are at maximum pixel value, then you could remove it with temp=np.where(pixel<max(),pixel,0). As I used this temporal image to place landmarks to crop, it doesn't hurt the original maximum pixels.",
    "2128057": "Hi, @jarvisai7 any idea how can i remove text(watermark) from image in [this dataset](https://www.kaggle.com/datasets/shubhamgoel27/dermnet). I traid cv2.inpaint to remove watermark but after removing watermark, watermark area got blur.\nhelp me to remove text!\nThanks",
    "2127857": "Those texts explain which angles were used to accomplish the mammography. like the CC/MLO view"
  }
}