{
  "id": 347095,
  "title": "Texture Analysis",
  "url": "/competitions/mayo-clinic-strip-ai/discussion/347095",
  "author_name": "Barbaros",
  "post_date": "2022-08-22T20:14:39.209000",
  "votes": 26,
  "comment_count": 18,
  "views": 0,
  "content": "<p>Greetings Kagglers,</p>\n<p>It is great to see the amazing work and effort put forward to attack this clinical problem. We are very optimistic that these efforts will result in good results. </p>\n<p>Even though we cannot take part in this competition, we are still working on the problem as well, and upon completion, we are planning on the dissemination of our findings.</p>\n<p>We saw some of the competitors were having issues with file sizes, and we also have seen great comradery and aid for those in need to help tackle their problems. There are several issues with this dataset: 1) The individual file sizes are too large, which can lead to memory issues if not handled properly. 2) The number of data points is relatively small. While one may have 100K+ images for ImageNet, here, we only have 1K+. 3) If we were to transfer learn, a lot of custom code may be needed to adapt input/output to known algorithms such as Inception, etc.  4) There are issues with WSIs inherited from scanning and staining procedures (data collected from 18 institutions). Therefore, stain normalization may be needed to prevent algorithms from overfitting on color schemes. </p>\n<p>At this time, it might also be worth investigating traditional techniques such as texture analysis instead of pure deep-learning-based approaches. They have shown promising results in the past, and they can be considered as either the primary techniques or pre-processing steps for deep-learning-based methods. If we were to segment, normalize and tile these images, the things left for investigation are the shapes and distribution patterns of the cells in each tile (probably not their color). Many texture algorithms, such as GLCM and its derivatives, can work on 8-bit gray-scale images (which can potentially provide a dramatic input size reduction). They usually produce decent results that can lead to shape and pattern detection.   They are already part of packages such as scikit-image ( <a href=\"https://scikit-image.org/docs/stable/auto_examples\" target=\"_blank\">https://scikit-image.org/docs/stable/auto_examples</a> ). A potential flow for pre-processing can be:  <br>\nSegment, Tile (512x512, etc.), Convert to Gray Scale, Use Textures to either classify or Create Image input for Deep Learning based algorithms. </p>\n<p>We hope we can spark some ideas.</p>",
  "messages": [
    {
      "id": 1909670,
      "postDate": "2022-08-22T20:14:39.210Z",
      "content": "<p>Greetings Kagglers,</p>\n<p>It is great to see the amazing work and effort put forward to attack this clinical problem. We are very optimistic that these efforts will result in good results. </p>\n<p>Even though we cannot take part in this competition, we are still working on the problem as well, and upon completion, we are planning on the dissemination of our findings.</p>\n<p>We saw some of the competitors were having issues with file sizes, and we also have seen great comradery and aid for those in need to help tackle their problems. There are several issues with this dataset: 1) The individual file sizes are too large, which can lead to memory issues if not handled properly. 2) The number of data points is relatively small. While one may have 100K+ images for ImageNet, here, we only have 1K+. 3) If we were to transfer learn, a lot of custom code may be needed to adapt input/output to known algorithms such as Inception, etc.  4) There are issues with WSIs inherited from scanning and staining procedures (data collected from 18 institutions). Therefore, stain normalization may be needed to prevent algorithms from overfitting on color schemes. </p>\n<p>At this time, it might also be worth investigating traditional techniques such as texture analysis instead of pure deep-learning-based approaches. They have shown promising results in the past, and they can be considered as either the primary techniques or pre-processing steps for deep-learning-based methods. If we were to segment, normalize and tile these images, the things left for investigation are the shapes and distribution patterns of the cells in each tile (probably not their color). Many texture algorithms, such as GLCM and its derivatives, can work on 8-bit gray-scale images (which can potentially provide a dramatic input size reduction). They usually produce decent results that can lead to shape and pattern detection.   They are already part of packages such as scikit-image ( <a href=\"https://scikit-image.org/docs/stable/auto_examples\" target=\"_blank\">https://scikit-image.org/docs/stable/auto_examples</a> ). A potential flow for pre-processing can be:  <br>\nSegment, Tile (512x512, etc.), Convert to Gray Scale, Use Textures to either classify or Create Image input for Deep Learning based algorithms. </p>\n<p>We hope we can spark some ideas.</p>",
      "rawMarkdown": "Greetings Kagglers,\n\nIt is great to see the amazing work and effort put forward to attack this clinical problem. We are very optimistic that these efforts will result in good results. \n\nEven though we cannot take part in this competition, we are still working on the problem as well, and upon completion, we are planning on the dissemination of our findings.\n\nWe saw some of the competitors were having issues with file sizes, and we also have seen great comradery and aid for those in need to help tackle their problems. There are several issues with this dataset: 1) The individual file sizes are too large, which can lead to memory issues if not handled properly. 2) The number of data points is relatively small. While one may have 100K+ images for ImageNet, here, we only have 1K+. 3) If we were to transfer learn, a lot of custom code may be needed to adapt input/output to known algorithms such as Inception, etc.  4) There are issues with WSIs inherited from scanning and staining procedures (data collected from 18 institutions). Therefore, stain normalization may be needed to prevent algorithms from overfitting on color schemes. \n\nAt this time, it might also be worth investigating traditional techniques such as texture analysis instead of pure deep-learning-based approaches. They have shown promising results in the past, and they can be considered as either the primary techniques or pre-processing steps for deep-learning-based methods. If we were to segment, normalize and tile these images, the things left for investigation are the shapes and distribution patterns of the cells in each tile (probably not their color). Many texture algorithms, such as GLCM and its derivatives, can work on 8-bit gray-scale images (which can potentially provide a dramatic input size reduction). They usually produce decent results that can lead to shape and pattern detection.   They are already part of packages such as scikit-image ( https://scikit-image.org/docs/stable/auto_examples ). A potential flow for pre-processing can be:  \nSegment, Tile (512x512, etc.), Convert to Gray Scale, Use Textures to either classify or Create Image input for Deep Learning based algorithms. \n\nWe hope we can spark some ideas.\n",
      "votes": 25
    },
    {
      "id": 1909762,
      "postDate": "2022-08-22T22:45:26.680Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/barbaroserdal\" target=\"_blank\">@barbaroserdal</a> !</p>\n<p>It's great to have input from the hosts on such a challenging problem.</p>\n<p>I have spent some time on the competition and so far I have come to the conclusion that there is not enough signal in the data for deep learning models to work. I am using crop-based models, my intuition tells me that this is the approach that should work the best - I do some color normalization but nothing too fancy, as Deep Learning models should be robust to changes in hue/saturation/value.</p>\n<p>I'm having deja-vu with the <a href=\"https://www.kaggle.com/competitions/rsna-miccai-brain-tumor-radiogenomic-classification\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-miccai-brain-tumor-radiogenomic-classification</a> competition where final results were not really better than random predictions, which is why I have stopped working on the topic. </p>\n<p>Although some people have achieved 0.4 and even 0.3 public LB, I have the feeling that those results are overfitted and do not translate to private LB. </p>\n<p>So here are a few questions : </p>\n<ul>\n<li><p>What is the state of current scores on the private leaderboard ? Do some of the approaches Kagglers have been using so far actually work, i.e. are scores significantly better than the biased/random baselines ? <br>\nI am not actually sure whether you have access to such information but surely Kaggle does. This information could help avoiding the disaster of no one providing an actual useful solution to the problem, as it has happened before.</p></li>\n<li><p>You've mentioned several times segmenting images as a track to explore. As far as I understand this means segmenting red blood cells, white blood cells, platelets &amp; fibrin. But how are we supposed to do this ? In the literature I've read, they do quantifications manually which is not really possible for us. This sounds like an impossible challenge with the data we have.</p></li>\n</ul>",
      "rawMarkdown": "Hi @barbaroserdal !\n\nIt's great to have input from the hosts on such a challenging problem.\n\nI have spent some time on the competition and so far I have come to the conclusion that there is not enough signal in the data for deep learning models to work. I am using crop-based models, my intuition tells me that this is the approach that should work the best - I do some color normalization but nothing too fancy, as Deep Learning models should be robust to changes in hue/saturation/value.\n\nI'm having deja-vu with the https://www.kaggle.com/competitions/rsna-miccai-brain-tumor-radiogenomic-classification competition where final results were not really better than random predictions, which is why I have stopped working on the topic. \n\nAlthough some people have achieved 0.4 and even 0.3 public LB, I have the feeling that those results are overfitted and do not translate to private LB. \n\nSo here are a few questions : \n\n- What is the state of current scores on the private leaderboard ? Do some of the approaches Kagglers have been using so far actually work, i.e. are scores significantly better than the biased/random baselines ? \nI am not actually sure whether you have access to such information but surely Kaggle does. This information could help avoiding the disaster of no one providing an actual useful solution to the problem, as it has happened before.\n\n- You've mentioned several times segmenting images as a track to explore. As far as I understand this means segmenting red blood cells, white blood cells, platelets & fibrin. But how are we supposed to do this ? In the literature I've read, they do quantifications manually which is not really possible for us. This sounds like an impossible challenge with the data we have.",
      "votes": 11,
      "replies": [
        {
          "id": 1909858,
          "postDate": "2022-08-23T02:09:11.177Z",
          "content": "<blockquote>\n  <ul>\n  <li>You've mentioned several times segmenting images as a track to explore. As far as I understand this means segmenting red blood cells, white blood cells, platelets &amp; fibrin. But how are we supposed to do this ? In the literature I've read, they do quantifications manually which is not really possible for us. This sounds like an impossible challenge with the data we have.</li>\n  </ul>\n</blockquote>\n<p>I believe he meant segmenting the background out of images so the patches are only of the blood clot.</p>",
          "rawMarkdown": "> - You've mentioned several times segmenting images as a track to explore. As far as I understand this means segmenting red blood cells, white blood cells, platelets & fibrin. But how are we supposed to do this ? In the literature I've read, they do quantifications manually which is not really possible for us. This sounds like an impossible challenge with the data we have.\n\nI believe he meant segmenting the background out of images so the patches are only of the blood clot."
        },
        {
          "id": 1910609,
          "postDate": "2022-08-23T14:51:04.580Z",
          "content": "<p>I have the same impression as <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> </p>\n<p>I also participated in the RSNA-MICCAI Brain Tumor Radiogenomic Classification competition. I tried to create a good model by <a href=\"https://www.kaggle.com/code/ren4yu/normalized-voxels-align-planes-and-crop\" target=\"_blank\">working hard on data preprocessing</a> but the validation score was the same as random guess (and left the competition).<br>\nAgain, in this competition, my model doesn't work at all with validation data.<br>\nI really want to know whether anyone is getting good scores on the validation data (not public LB).</p>",
          "rawMarkdown": "I have the same impression as @theoviel \n\nI also participated in the RSNA-MICCAI Brain Tumor Radiogenomic Classification competition. I tried to create a good model by [working hard on data preprocessing](https://www.kaggle.com/code/ren4yu/normalized-voxels-align-planes-and-crop) but the validation score was the same as random guess (and left the competition).\nAgain, in this competition, my model doesn't work at all with validation data.\nI really want to know whether anyone is getting good scores on the validation data (not public LB).",
          "votes": 2
        },
        {
          "id": 1910652,
          "postDate": "2022-08-23T15:24:01.873Z",
          "content": "<p><a href=\"https://www.kaggle.com/ren4yu\" target=\"_blank\">@ren4yu</a>, I have the same impression at the moment too. My model does not learn enough information from the training data to detect the correct labels in the validation data. (I tried using different subsets of the images for making train and validation datasets, but this pattern still persists.)</p>",
          "rawMarkdown": "@ren4yu, I have the same impression at the moment too. My model does not learn enough information from the training data to detect the correct labels in the validation data. (I tried using different subsets of the images for making train and validation datasets, but this pattern still persists.)"
        },
        {
          "id": 1910677,
          "postDate": "2022-08-23T15:36:48.830Z",
          "content": "<p>Could some of you let us know what general techniques you have used and not found success? There are a lot of methods for such problems that generally do well, but may not be working for this data. But it hard to believe that most of those methods have been already tried, tuned and improvised on and fully rejected. Would be good to know what all has not worked.</p>",
          "rawMarkdown": "Could some of you let us know what general techniques you have used and not found success? There are a lot of methods for such problems that generally do well, but may not be working for this data. But it hard to believe that most of those methods have been already tried, tuned and improvised on and fully rejected. Would be good to know what all has not worked.",
          "votes": 2
        },
        {
          "id": 1910773,
          "postDate": "2022-08-23T17:03:06.077Z",
          "content": "<p>Kaggle is monitoring the private LB.  By segmentation, I meant segmenting the background initially (in earlier posts); however, by utilizing filters like Hessian matrix ( <a href=\"https://scikit-image.org/docs/stable/api/skimage.feature.html\" target=\"_blank\">https://scikit-image.org/docs/stable/api/skimage.feature.html</a> ) individual cells can be segmented as well</p>",
          "rawMarkdown": "Kaggle is monitoring the private LB.  By segmentation, I meant segmenting the background initially (in earlier posts); however, by utilizing filters like Hessian matrix ( https://scikit-image.org/docs/stable/api/skimage.feature.html ) individual cells can be segmented as well",
          "votes": 1
        },
        {
          "id": 1911636,
          "postDate": "2022-08-24T08:01:32.657Z",
          "content": "<p>We have approximately 750 images. Even if we suppose the images were filled with relevant data, it would be difficult to create a model that will generalise well. It is like detecting cat or dog with less than 1000 images. But it is even more difficult since only a small amount of information for each image should be extracted by the model to classify. So to rephrase the problem, it is like detecting if we have a dog or a cat in a small set of high resolution landscape photographs taken with different cameras at different places and on different weathers. The dataset is a little bit too small, especially without clear indications of how the data acquisition occurred and what are the suspected physical phenomenons showing there is at least a little bit of causality between the diseases and what is observed in the WSI.</p>",
          "rawMarkdown": "We have approximately 750 images. Even if we suppose the images were filled with relevant data, it would be difficult to create a model that will generalise well. It is like detecting cat or dog with less than 1000 images. But it is even more difficult since only a small amount of information for each image should be extracted by the model to classify. So to rephrase the problem, it is like detecting if we have a dog or a cat in a small set of high resolution landscape photographs taken with different cameras at different places and on different weathers. The dataset is a little bit too small, especially without clear indications of how the data acquisition occurred and what are the suspected physical phenomenons showing there is at least a little bit of causality between the diseases and what is observed in the WSI.",
          "votes": 1
        },
        {
          "id": 1912167,
          "postDate": "2022-08-24T14:54:44.567Z",
          "content": "<p><a href=\"https://www.kaggle.com/barbaroserdal\" target=\"_blank\">@barbaroserdal</a> If you meant segmenting the background then I am (to some extent) doing that :)</p>\n<p>Another thing I have in mind :</p>\n<p>I don't think the dataset is too small, I work with medical data and have trained models with even less data than that. However when data is small there needs to be a significant amount of signal for models to learn. <br>\nIs it the case here ? I don't know. <br>\nA good criterion I like to keep in mind is \"Can an expert eye predict the target from the images ?\". If so then there's hope models learn the same thing even with &lt;1000 samples.</p>",
          "rawMarkdown": "@barbaroserdal If you meant segmenting the background then I am (to some extent) doing that :)\n\nAnother thing I have in mind :\n\nI don't think the dataset is too small, I work with medical data and have trained models with even less data than that. However when data is small there needs to be a significant amount of signal for models to learn. \nIs it the case here ? I don't know. \nA good criterion I like to keep in mind is \"Can an expert eye predict the target from the images ?\". If so then there's hope models learn the same thing even with <1000 samples.\n",
          "votes": 4
        },
        {
          "id": 1915828,
          "postDate": "2022-08-27T12:16:04.623Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1932026,
          "postDate": "2022-09-09T07:57:28.530Z",
          "content": "<p>\"Can an expert eye predict the target from the images ?\".<br>\nWell the answer to that question is no. Pathologists have tried, and they can't figure it out. Keep in mind that mechanical thrombectomies are a relatively new thing in Medicine. It didn't use to be the case that physicians had access to ischemic stroke emboli like we do now. This is why they are trying to see if they can get somewhere with Deep Learning. Specifically, it is why they have put the problem on Kaggle. The hope, I'm sure, is that enough people will try different ideas and something that works will surface somehow.</p>",
          "rawMarkdown": "\"Can an expert eye predict the target from the images ?\".\nWell the answer to that question is no. Pathologists have tried, and they can't figure it out. Keep in mind that mechanical thrombectomies are a relatively new thing in Medicine. It didn't use to be the case that physicians had access to ischemic stroke emboli like we do now. This is why they are trying to see if they can get somewhere with Deep Learning. Specifically, it is why they have put the problem on Kaggle. The hope, I'm sure, is that enough people will try different ideas and something that works will surface somehow.",
          "votes": 2
        },
        {
          "id": 1938844,
          "postDate": "2022-09-14T11:48:47.567Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1944071,
          "postDate": "2022-09-18T03:42:09.597Z",
          "content": "<p>Do you have an idea at what scale there can be the most characteristic elements?</p>",
          "rawMarkdown": "Do you have an idea at what scale there can be the most characteristic elements?\n"
        },
        {
          "id": 1957812,
          "postDate": "2022-09-27T07:04:12.837Z",
          "content": "<p><a href=\"https://www.kaggle.com/ren4yu\" target=\"_blank\">@ren4yu</a>   We are able to achieve about 75% accuracy on the validation set, while accuracy on the training set is about 95%. The validation dataset is 6% of training dataset. However, LB penalizes big mistakes a lot. While training we are able to achieve 0.1 and 0.5 WLL on the training and validation sets respectively, but it gives the score of 1.6 on the LB.</p>",
          "rawMarkdown": "@ren4yu   We are able to achieve about 75% accuracy on the validation set, while accuracy on the training set is about 95%. The validation dataset is 6% of training dataset. However, LB penalizes big mistakes a lot. While training we are able to achieve 0.1 and 0.5 WLL on the training and validation sets respectively, but it gives the score of 1.6 on the LB.",
          "votes": 2
        },
        {
          "id": 1965744,
          "postDate": "2022-10-01T13:58:45.573Z",
          "content": "<p>I am able to achieve 79% accuracy at high resolution tiles, which gives 83% on validation sets (picking good tile with threshold methods). I found a lb 0.8, I did not understand anything. Have you any clues? </p>",
          "rawMarkdown": "I am able to achieve 79% accuracy at high resolution tiles, which gives 83% on validation sets (picking good tile with threshold methods). I found a lb 0.8, I did not understand anything. Have you any clues? "
        }
      ]
    },
    {
      "id": 1944078,
      "postDate": "2022-09-18T03:50:41.650Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/barbaroserdal\" target=\"_blank\">@barbaroserdal</a> . I have spent some time on the challenge and created this preprocessed dataset. </p>\n<p><a href=\"https://www.kaggle.com/datasets/tr1gg3rtrash/mayo-clinic\" target=\"_blank\">Dataset Link</a></p>\n<p>Would love to know your opinions if this could be a good way to solve the issue. </p>",
      "rawMarkdown": "Hi @barbaroserdal . I have spent some time on the challenge and created this preprocessed dataset. \n\n[Dataset Link](https://www.kaggle.com/datasets/tr1gg3rtrash/mayo-clinic)\n\nWould love to know your opinions if this could be a good way to solve the issue. ",
      "votes": 1,
      "replies": [
        {
          "id": 1944796,
          "postDate": "2022-09-18T15:38:19.847Z",
          "content": "<p>Hi, I think you made a good and useful job. I think getting lb 0.6 or 0.3 does not mean anything since there is only 3 elements.  I think there is information in low and very high resolution, but I am not a biologist.  I would like to join your team.</p>",
          "rawMarkdown": "Hi, I think you made a good and useful job. I think getting lb 0.6 or 0.3 does not mean anything since there is only 3 elements.  I think there is information in low and very high resolution, but I am not a biologist.  I would like to join your team.",
          "votes": 1
        },
        {
          "id": 1944809,
          "postDate": "2022-09-18T15:44:48.527Z",
          "content": "<p>Currently, I haven't used the processed dataset that I have uploaded to make any submissions due to resource constraints on our end. We have just open-sourced it for anyone who could use their resources to make a submission, thereby solving the blood clotting detection problem. I feel like rather than just being competitive, it should be our duty to help others if we can to solve the issue as this accounts for a major problem in the world and can save many lives. Thank you for your feedback on the dataset. </p>",
          "rawMarkdown": "Currently, I haven't used the processed dataset that I have uploaded to make any submissions due to resource constraints on our end. We have just open-sourced it for anyone who could use their resources to make a submission, thereby solving the blood clotting detection problem. I feel like rather than just being competitive, it should be our duty to help others if we can to solve the issue as this accounts for a major problem in the world and can save many lives. Thank you for your feedback on the dataset. ",
          "votes": 2
        }
      ]
    },
    {
      "id": 1909701,
      "postDate": "2022-08-22T20:56:49.283Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1909762,
      "author_name": "Theo Viel",
      "author_url": "",
      "post_date": "2022-08-22T22:45:26.680000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/barbaroserdal\" target=\"_blank\">@barbaroserdal</a> !</p>\n<p>It's great to have input from the hosts on such a challenging problem.</p>\n<p>I have spent some time on the competition and so far I have come to the conclusion that there is not enough signal in the data for deep learning models to work. I am using crop-based models, my intuition tells me that this is the approach that should work the best - I do some color normalization but nothing too fancy, as Deep Learning models should be robust to changes in hue/saturation/value.</p>\n<p>I'm having deja-vu with the <a href=\"https://www.kaggle.com/competitions/rsna-miccai-brain-tumor-radiogenomic-classification\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-miccai-brain-tumor-radiogenomic-classification</a> competition where final results were not really better than random predictions, which is why I have stopped working on the topic. </p>\n<p>Although some people have achieved 0.4 and even 0.3 public LB, I have the feeling that those results are overfitted and do not translate to private LB. </p>\n<p>So here are a few questions : </p>\n<ul>\n<li><p>What is the state of current scores on the private leaderboard ? Do some of the approaches Kagglers have been using so far actually work, i.e. are scores significantly better than the biased/random baselines ? <br>\nI am not actually sure whether you have access to such information but surely Kaggle does. This information could help avoiding the disaster of no one providing an actual useful solution to the problem, as it has happened before.</p></li>\n<li><p>You've mentioned several times segmenting images as a track to explore. As far as I understand this means segmenting red blood cells, white blood cells, platelets &amp; fibrin. But how are we supposed to do this ? In the literature I've read, they do quantifications manually which is not really possible for us. This sounds like an impossible challenge with the data we have.</p></li>\n</ul>",
      "votes": 11,
      "replies": [
        {
          "id": 1909858,
          "author_name": "Ari",
          "author_url": "",
          "post_date": "2022-08-23T02:09:11.177000",
          "content": "<blockquote>\n  <ul>\n  <li>You've mentioned several times segmenting images as a track to explore. As far as I understand this means segmenting red blood cells, white blood cells, platelets &amp; fibrin. But how are we supposed to do this ? In the literature I've read, they do quantifications manually which is not really possible for us. This sounds like an impossible challenge with the data we have.</li>\n  </ul>\n</blockquote>\n<p>I believe he meant segmenting the background out of images so the patches are only of the blood clot.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1910609,
          "author_name": "yu4u",
          "author_url": "",
          "post_date": "2022-08-23T14:51:04.580000",
          "content": "<p>I have the same impression as <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> </p>\n<p>I also participated in the RSNA-MICCAI Brain Tumor Radiogenomic Classification competition. I tried to create a good model by <a href=\"https://www.kaggle.com/code/ren4yu/normalized-voxels-align-planes-and-crop\" target=\"_blank\">working hard on data preprocessing</a> but the validation score was the same as random guess (and left the competition).<br>\nAgain, in this competition, my model doesn't work at all with validation data.<br>\nI really want to know whether anyone is getting good scores on the validation data (not public LB).</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1910652,
          "author_name": "Hassan Abedi",
          "author_url": "",
          "post_date": "2022-08-23T15:24:01.873000",
          "content": "<p><a href=\"https://www.kaggle.com/ren4yu\" target=\"_blank\">@ren4yu</a>, I have the same impression at the moment too. My model does not learn enough information from the training data to detect the correct labels in the validation data. (I tried using different subsets of the images for making train and validation datasets, but this pattern still persists.)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1910677,
          "author_name": "tdiceman",
          "author_url": "",
          "post_date": "2022-08-23T15:36:48.830000",
          "content": "<p>Could some of you let us know what general techniques you have used and not found success? There are a lot of methods for such problems that generally do well, but may not be working for this data. But it hard to believe that most of those methods have been already tried, tuned and improvised on and fully rejected. Would be good to know what all has not worked.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1910773,
          "author_name": "Barbaros",
          "author_url": "",
          "post_date": "2022-08-23T17:03:06.077000",
          "content": "<p>Kaggle is monitoring the private LB.  By segmentation, I meant segmenting the background initially (in earlier posts); however, by utilizing filters like Hessian matrix ( <a href=\"https://scikit-image.org/docs/stable/api/skimage.feature.html\" target=\"_blank\">https://scikit-image.org/docs/stable/api/skimage.feature.html</a> ) individual cells can be segmented as well</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1911636,
          "author_name": "nexus2049",
          "author_url": "",
          "post_date": "2022-08-24T08:01:32.657000",
          "content": "<p>We have approximately 750 images. Even if we suppose the images were filled with relevant data, it would be difficult to create a model that will generalise well. It is like detecting cat or dog with less than 1000 images. But it is even more difficult since only a small amount of information for each image should be extracted by the model to classify. So to rephrase the problem, it is like detecting if we have a dog or a cat in a small set of high resolution landscape photographs taken with different cameras at different places and on different weathers. The dataset is a little bit too small, especially without clear indications of how the data acquisition occurred and what are the suspected physical phenomenons showing there is at least a little bit of causality between the diseases and what is observed in the WSI.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1912167,
          "author_name": "Theo Viel",
          "author_url": "",
          "post_date": "2022-08-24T14:54:44.567000",
          "content": "<p><a href=\"https://www.kaggle.com/barbaroserdal\" target=\"_blank\">@barbaroserdal</a> If you meant segmenting the background then I am (to some extent) doing that :)</p>\n<p>Another thing I have in mind :</p>\n<p>I don't think the dataset is too small, I work with medical data and have trained models with even less data than that. However when data is small there needs to be a significant amount of signal for models to learn. <br>\nIs it the case here ? I don't know. <br>\nA good criterion I like to keep in mind is \"Can an expert eye predict the target from the images ?\". If so then there's hope models learn the same thing even with &lt;1000 samples.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1915828,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-08-27T12:16:04.623000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1932026,
          "author_name": "Daniel Macaulay",
          "author_url": "",
          "post_date": "2022-09-09T07:57:28.530000",
          "content": "<p>\"Can an expert eye predict the target from the images ?\".<br>\nWell the answer to that question is no. Pathologists have tried, and they can't figure it out. Keep in mind that mechanical thrombectomies are a relatively new thing in Medicine. It didn't use to be the case that physicians had access to ischemic stroke emboli like we do now. This is why they are trying to see if they can get somewhere with Deep Learning. Specifically, it is why they have put the problem on Kaggle. The hope, I'm sure, is that enough people will try different ideas and something that works will surface somehow.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1938844,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-09-14T11:48:47.567000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1944071,
          "author_name": "Pierre Tisseur",
          "author_url": "",
          "post_date": "2022-09-18T03:42:09.597000",
          "content": "<p>Do you have an idea at what scale there can be the most characteristic elements?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1957812,
          "author_name": "JanGlinko2",
          "author_url": "",
          "post_date": "2022-09-27T07:04:12.837000",
          "content": "<p><a href=\"https://www.kaggle.com/ren4yu\" target=\"_blank\">@ren4yu</a>   We are able to achieve about 75% accuracy on the validation set, while accuracy on the training set is about 95%. The validation dataset is 6% of training dataset. However, LB penalizes big mistakes a lot. While training we are able to achieve 0.1 and 0.5 WLL on the training and validation sets respectively, but it gives the score of 1.6 on the LB.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1965744,
          "author_name": "Pierre Tisseur",
          "author_url": "",
          "post_date": "2022-10-01T13:58:45.573000",
          "content": "<p>I am able to achieve 79% accuracy at high resolution tiles, which gives 83% on validation sets (picking good tile with threshold methods). I found a lb 0.8, I did not understand anything. Have you any clues? </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1944078,
      "author_name": "Mrinal Tyagi",
      "author_url": "",
      "post_date": "2022-09-18T03:50:41.650000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/barbaroserdal\" target=\"_blank\">@barbaroserdal</a> . I have spent some time on the challenge and created this preprocessed dataset. </p>\n<p><a href=\"https://www.kaggle.com/datasets/tr1gg3rtrash/mayo-clinic\" target=\"_blank\">Dataset Link</a></p>\n<p>Would love to know your opinions if this could be a good way to solve the issue. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1944796,
          "author_name": "Pierre Tisseur",
          "author_url": "",
          "post_date": "2022-09-18T15:38:19.847000",
          "content": "<p>Hi, I think you made a good and useful job. I think getting lb 0.6 or 0.3 does not mean anything since there is only 3 elements.  I think there is information in low and very high resolution, but I am not a biologist.  I would like to join your team.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1944809,
          "author_name": "Mrinal Tyagi",
          "author_url": "",
          "post_date": "2022-09-18T15:44:48.527000",
          "content": "<p>Currently, I haven't used the processed dataset that I have uploaded to make any submissions due to resource constraints on our end. We have just open-sourced it for anyone who could use their resources to make a submission, thereby solving the blood clotting detection problem. I feel like rather than just being competitive, it should be our duty to help others if we can to solve the issue as this accounts for a major problem in the world and can save many lives. Thank you for your feedback on the dataset. </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1909701,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-22T20:56:49.283000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1909670": "Greetings Kagglers,\n\nIt is great to see the amazing work and effort put forward to attack this clinical problem. We are very optimistic that these efforts will result in good results. \n\nEven though we cannot take part in this competition, we are still working on the problem as well, and upon completion, we are planning on the dissemination of our findings.\n\nWe saw some of the competitors were having issues with file sizes, and we also have seen great comradery and aid for those in need to help tackle their problems. There are several issues with this dataset: 1) The individual file sizes are too large, which can lead to memory issues if not handled properly. 2) The number of data points is relatively small. While one may have 100K+ images for ImageNet, here, we only have 1K+. 3) If we were to transfer learn, a lot of custom code may be needed to adapt input/output to known algorithms such as Inception, etc.  4) There are issues with WSIs inherited from scanning and staining procedures (data collected from 18 institutions). Therefore, stain normalization may be needed to prevent algorithms from overfitting on color schemes. \n\nAt this time, it might also be worth investigating traditional techniques such as texture analysis instead of pure deep-learning-based approaches. They have shown promising results in the past, and they can be considered as either the primary techniques or pre-processing steps for deep-learning-based methods. If we were to segment, normalize and tile these images, the things left for investigation are the shapes and distribution patterns of the cells in each tile (probably not their color). Many texture algorithms, such as GLCM and its derivatives, can work on 8-bit gray-scale images (which can potentially provide a dramatic input size reduction). They usually produce decent results that can lead to shape and pattern detection.   They are already part of packages such as scikit-image ( https://scikit-image.org/docs/stable/auto_examples ). A potential flow for pre-processing can be:  \nSegment, Tile (512x512, etc.), Convert to Gray Scale, Use Textures to either classify or Create Image input for Deep Learning based algorithms. \n\nWe hope we can spark some ideas.\n",
    "1909762": "Hi @barbaroserdal !\n\nIt's great to have input from the hosts on such a challenging problem.\n\nI have spent some time on the competition and so far I have come to the conclusion that there is not enough signal in the data for deep learning models to work. I am using crop-based models, my intuition tells me that this is the approach that should work the best - I do some color normalization but nothing too fancy, as Deep Learning models should be robust to changes in hue/saturation/value.\n\nI'm having deja-vu with the https://www.kaggle.com/competitions/rsna-miccai-brain-tumor-radiogenomic-classification competition where final results were not really better than random predictions, which is why I have stopped working on the topic. \n\nAlthough some people have achieved 0.4 and even 0.3 public LB, I have the feeling that those results are overfitted and do not translate to private LB. \n\nSo here are a few questions : \n\n- What is the state of current scores on the private leaderboard ? Do some of the approaches Kagglers have been using so far actually work, i.e. are scores significantly better than the biased/random baselines ? \nI am not actually sure whether you have access to such information but surely Kaggle does. This information could help avoiding the disaster of no one providing an actual useful solution to the problem, as it has happened before.\n\n- You've mentioned several times segmenting images as a track to explore. As far as I understand this means segmenting red blood cells, white blood cells, platelets & fibrin. But how are we supposed to do this ? In the literature I've read, they do quantifications manually which is not really possible for us. This sounds like an impossible challenge with the data we have.",
    "1944078": "Hi @barbaroserdal . I have spent some time on the challenge and created this preprocessed dataset. \n\n[Dataset Link](https://www.kaggle.com/datasets/tr1gg3rtrash/mayo-clinic)\n\nWould love to know your opinions if this could be a good way to solve the issue. ",
    "1909701": ""
  }
}