{
  "id": 147684,
  "title": "First step results and their analysis",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/147684",
  "author_name": "Vlad Vaduva",
  "post_date": "2020-05-01T12:16:00.636000",
  "votes": 16,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Hi everybody,</p>\n\n<p>My first step was to use the whole resized image in order to make a baseline score which later on to try to improve using different techniques (segmentation models, smart crop, other computer vision techniques). \nLately I seen some very interesting approaches using tiles, <a href=\"/iafoss\">@iafoss</a> made a really good job describing his idea and deserves our appreciation  for sharing it. <a href=\"/frlemarchand\">@frlemarchand</a> also is working on a somehow similar approach, very interesting idea and I am curious about his results.\nGetting back on analyzing the baseline results. From some tries made, the optimum resolution was 512x512 and other parameters are described in this <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/146332an\">discussion</a> and this <a href=\"https://www.kaggle.com/vladvdv/pytorch-training-customizable-kernel-with-5-folds\">kernel</a>.\nI used 5 folds, which using the mean prediction scored 0.69 on public leaderboard (0.70 with some added tricks).\nOn CV the confusion matrix looks interesting \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4005865%2Fc61b71aadf417b9891b029303c3818d4%2FconfMatrixAll.png?generation=1588333419925630&amp;alt=media\" alt=\"\"></p>\n\n<p>It seems like:\n- the model cannot separate true label 2 and it's predicting label 1 in most cases(478 vs 247)\n- the model cannot separate true label 3 and it's predicting label 1 in most cases(266 vs 238)\n- the model cannot separate true label 4 and it's predicting label 5 and 1 in most cases(255 vs 323 vs 242)</p>\n\n<p>So, there are big problems predicting labels 2,3 and 4 and a common ground is that a lot of mistakes made by the model are made by confusion with label 1.</p>\n\n<p>For a better understanding of what is going on, let's go deeper by separating the confusion matrix by the data provider</p>\n\n<p>This is the confusion matrix with data by karolinska\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4005865%2F6c1440139ad9f4bddfba2018182937be%2FconfMatrixK.png?generation=1588334130529204&amp;alt=media\" alt=\"\"></p>\n\n<p>and here by radboud\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4005865%2F20070763e5bec6a59da477aa7d35c6f8%2FconfMatrixR.png?generation=1588334056721834&amp;alt=media\" alt=\"\"></p>\n\n<p>It is easily to observe that the problems are still on the 2,3 and 4 categories on both data providers but the model is making different type of mistakes:\n- prediction of the 2 category: on karolinska the model is predicting a lot of 1 instead of 2(98vs 322) but on radboud the model is predicting also 1 (in a lower rate than k: 149 vs 156) but also 5 (149 vs 133)\n- prediction of the 3 category: on karolinska the model is predicting a lot of 1 (21 vs 126) but on radboud it seems that the errors are spread between multiple categories (and mostly on 5: 217 vs 270)\n- prediction of the 4 category:  again karolinska is predicting a lot of 1(97 vs 171) and most of the mistakes of radboud are made by predicting again by confusing it with 5 (159 vs 257)</p>\n\n<p>As a conclusion I can say that is clear that the images are different between providers and this can be felt in how the model behaves. The next step on my todo list is to design 2 different models, each model dedicated to a single provider, although unfortunately the data will be cut in half my hope is that the personalized model will better understand just one type of images.</p>",
  "messages": [
    {
      "id": 828967,
      "postDate": "2020-05-01T12:16:00.637Z",
      "content": "<p>Hi everybody,</p>\n\n<p>My first step was to use the whole resized image in order to make a baseline score which later on to try to improve using different techniques (segmentation models, smart crop, other computer vision techniques). \nLately I seen some very interesting approaches using tiles, <a href=\"/iafoss\">@iafoss</a> made a really good job describing his idea and deserves our appreciation  for sharing it. <a href=\"/frlemarchand\">@frlemarchand</a> also is working on a somehow similar approach, very interesting idea and I am curious about his results.\nGetting back on analyzing the baseline results. From some tries made, the optimum resolution was 512x512 and other parameters are described in this <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/146332an\">discussion</a> and this <a href=\"https://www.kaggle.com/vladvdv/pytorch-training-customizable-kernel-with-5-folds\">kernel</a>.\nI used 5 folds, which using the mean prediction scored 0.69 on public leaderboard (0.70 with some added tricks).\nOn CV the confusion matrix looks interesting \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4005865%2Fc61b71aadf417b9891b029303c3818d4%2FconfMatrixAll.png?generation=1588333419925630&amp;alt=media\" alt=\"\"></p>\n\n<p>It seems like:\n- the model cannot separate true label 2 and it's predicting label 1 in most cases(478 vs 247)\n- the model cannot separate true label 3 and it's predicting label 1 in most cases(266 vs 238)\n- the model cannot separate true label 4 and it's predicting label 5 and 1 in most cases(255 vs 323 vs 242)</p>\n\n<p>So, there are big problems predicting labels 2,3 and 4 and a common ground is that a lot of mistakes made by the model are made by confusion with label 1.</p>\n\n<p>For a better understanding of what is going on, let's go deeper by separating the confusion matrix by the data provider</p>\n\n<p>This is the confusion matrix with data by karolinska\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4005865%2F6c1440139ad9f4bddfba2018182937be%2FconfMatrixK.png?generation=1588334130529204&amp;alt=media\" alt=\"\"></p>\n\n<p>and here by radboud\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4005865%2F20070763e5bec6a59da477aa7d35c6f8%2FconfMatrixR.png?generation=1588334056721834&amp;alt=media\" alt=\"\"></p>\n\n<p>It is easily to observe that the problems are still on the 2,3 and 4 categories on both data providers but the model is making different type of mistakes:\n- prediction of the 2 category: on karolinska the model is predicting a lot of 1 instead of 2(98vs 322) but on radboud the model is predicting also 1 (in a lower rate than k: 149 vs 156) but also 5 (149 vs 133)\n- prediction of the 3 category: on karolinska the model is predicting a lot of 1 (21 vs 126) but on radboud it seems that the errors are spread between multiple categories (and mostly on 5: 217 vs 270)\n- prediction of the 4 category:  again karolinska is predicting a lot of 1(97 vs 171) and most of the mistakes of radboud are made by predicting again by confusing it with 5 (159 vs 257)</p>\n\n<p>As a conclusion I can say that is clear that the images are different between providers and this can be felt in how the model behaves. The next step on my todo list is to design 2 different models, each model dedicated to a single provider, although unfortunately the data will be cut in half my hope is that the personalized model will better understand just one type of images.</p>",
      "rawMarkdown": "Hi everybody,\n\nMy first step was to use the whole resized image in order to make a baseline score which later on to try to improve using different techniques (segmentation models, smart crop, other computer vision techniques). \nLately I seen some very interesting approaches using tiles, @iafoss made a really good job describing his idea and deserves our appreciation  for sharing it. @frlemarchand also is working on a somehow similar approach, very interesting idea and I am curious about his results.\nGetting back on analyzing the baseline results. From some tries made, the optimum resolution was 512x512 and other parameters are described in this [discussion](https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/146332an) and this [kernel](https://www.kaggle.com/vladvdv/pytorch-training-customizable-kernel-with-5-folds).\nI used 5 folds, which using the mean prediction scored 0.69 on public leaderboard (0.70 with some added tricks).\nOn CV the confusion matrix looks interesting \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4005865%2Fc61b71aadf417b9891b029303c3818d4%2FconfMatrixAll.png?generation=1588333419925630&amp;alt=media)\n\nIt seems like:\n- the model cannot separate true label 2 and it's predicting label 1 in most cases(478 vs 247)\n- the model cannot separate true label 3 and it's predicting label 1 in most cases(266 vs 238)\n- the model cannot separate true label 4 and it's predicting label 5 and 1 in most cases(255 vs 323 vs 242)\n\nSo, there are big problems predicting labels 2,3 and 4 and a common ground is that a lot of mistakes made by the model are made by confusion with label 1.\n\nFor a better understanding of what is going on, let's go deeper by separating the confusion matrix by the data provider\n\nThis is the confusion matrix with data by karolinska\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4005865%2F6c1440139ad9f4bddfba2018182937be%2FconfMatrixK.png?generation=1588334130529204&amp;alt=media)\n\n\nand here by radboud\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4005865%2F20070763e5bec6a59da477aa7d35c6f8%2FconfMatrixR.png?generation=1588334056721834&amp;alt=media)\n\nIt is easily to observe that the problems are still on the 2,3 and 4 categories on both data providers but the model is making different type of mistakes:\n- prediction of the 2 category: on karolinska the model is predicting a lot of 1 instead of 2(98vs 322) but on radboud the model is predicting also 1 (in a lower rate than k: 149 vs 156) but also 5 (149 vs 133)\n- prediction of the 3 category: on karolinska the model is predicting a lot of 1 (21 vs 126) but on radboud it seems that the errors are spread between multiple categories (and mostly on 5: 217 vs 270)\n- prediction of the 4 category:  again karolinska is predicting a lot of 1(97 vs 171) and most of the mistakes of radboud are made by predicting again by confusing it with 5 (159 vs 257)\n\nAs a conclusion I can say that is clear that the images are different between providers and this can be felt in how the model behaves. The next step on my todo list is to design 2 different models, each model dedicated to a single provider, although unfortunately the data will be cut in half my hope is that the personalized model will better understand just one type of images.\n",
      "votes": 16
    },
    {
      "id": 829653,
      "postDate": "2020-05-02T01:45:51.233Z",
      "content": "<p>Two models didnt seem to help for me :(</p>",
      "rawMarkdown": "Two models didnt seem to help for me :(",
      "votes": 3,
      "replies": [
        {
          "id": 830180,
          "postDate": "2020-05-02T11:42:35.603Z",
          "content": "<p>Sorry to hear that. The result was much worse or about the same as the single mode one ? </p>",
          "rawMarkdown": "Sorry to hear that. The result was much worse or about the same as the single mode one ? "
        },
        {
          "id": 830215,
          "postDate": "2020-05-02T12:11:45.113Z",
          "content": "<p>They were a little lower, maybe i need to tune them a bit more</p>",
          "rawMarkdown": "They were a little lower, maybe i need to tune them a bit more"
        }
      ]
    },
    {
      "id": 829628,
      "postDate": "2020-05-02T01:22:51.290Z",
      "content": "<p>Two models maybe more reasonable. The picture below shows the data in karolinska is unbalanced and radboud is more balanced(correct me if I was wrong), may be it also the cause of different performance of only a model to predict two different distributions?</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3149580%2Fc7eca96a293507a0a30fa8d18cc3ade7%2FDataDistribution.png?generation=1588382474836034&amp;alt=media\" alt=\"\"></p>\n\n<p>What's more, I have observed that the color in data of karolinska is darker compared with radboud in commom(I select some pictures randomly about 20 times and observe them, only crop-white and resize to 224, no augmentations)... So I place RandomBrightness/Contrast to my TODO list. As we can see, In sub-picture-3, the radboud data is hard to recognize by eyes.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3149580%2F80d0bff3db4cc90d8096be036a9f4a90%2FColorCompare.png?generation=1588383678831041&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3149580%2F1990ced2f6e6b7415cc9a9ba60817081%2FColorCompare_2.png?generation=1588384114939252&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3149580%2F3f6a1b85b87758cc868af79a4cd36b38%2FColorCompare_3.png?generation=1588384129739286&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Two models maybe more reasonable. The picture below shows the data in karolinska is unbalanced and radboud is more balanced(correct me if I was wrong), may be it also the cause of different performance of only a model to predict two different distributions?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3149580%2Fc7eca96a293507a0a30fa8d18cc3ade7%2FDataDistribution.png?generation=1588382474836034&amp;alt=media)\n\nWhat's more, I have observed that the color in data of karolinska is darker compared with radboud in commom(I select some pictures randomly about 20 times and observe them, only crop-white and resize to 224, no augmentations)... So I place RandomBrightness/Contrast to my TODO list. As we can see, In sub-picture-3, the radboud data is hard to recognize by eyes.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3149580%2F80d0bff3db4cc90d8096be036a9f4a90%2FColorCompare.png?generation=1588383678831041&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3149580%2F1990ced2f6e6b7415cc9a9ba60817081%2FColorCompare_2.png?generation=1588384114939252&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3149580%2F3f6a1b85b87758cc868af79a4cd36b38%2FColorCompare_3.png?generation=1588384129739286&amp;alt=media)\n",
      "votes": 2,
      "replies": [
        {
          "id": 830183,
          "postDate": "2020-05-02T11:43:38.163Z",
          "content": "<p>Good and useful observations <a href=\"/cnzengshiyuan\">@cnzengshiyuan</a> . I completely agree</p>",
          "rawMarkdown": "Good and useful observations @cnzengshiyuan . I completely agree",
          "votes": 1
        },
        {
          "id": 830944,
          "postDate": "2020-05-03T02:04:30.317Z",
          "content": "<p><a href=\"/vladvdv\">@vladvdv</a> haha, thank you! Since iafoss's great idea for cutting the picture into patches and get good score. I think an explainable way to modify the picure is the key to high score(Though I haven't finish implement). What do you think about it?\nPS: The RandomContrast behaves better than RandomBrighness(only on radboud). I used them with limit=0.3, always-apply=True from albumentations.</p>",
          "rawMarkdown": "@vladvdv haha, thank you! Since iafoss's great idea for cutting the picture into patches and get good score. I think an explainable way to modify the picure is the key to high score(Though I haven't finish implement). What do you think about it?\nPS: The RandomContrast behaves better than RandomBrighness(only on radboud). I used them with limit=0.3, always-apply=True from albumentations.",
          "votes": 1
        }
      ]
    },
    {
      "id": 896849,
      "postDate": "2020-06-22T13:27:37.873Z",
      "content": "<p>Could you please share how you calculated and made these really nice plots of the confusion matrix ? </p>\n\n<pre><code> Thanks :) \n</code></pre>",
      "rawMarkdown": " Could you please share how you calculated and made these really nice plots of the confusion matrix ? \n\n     Thanks :) \n",
      "replies": [
        {
          "id": 896882,
          "postDate": "2020-06-22T13:58:16.720Z",
          "content": "<p>Use seaborn:\n<code>\nimport seaborn as sn\nfrom sklearn.metrics import confusion_matrix <br>\ncm = confusion_matrix(truth, pred) <br>\ndf_cm = pd.DataFrame(cm, index=[i for i in \"012345\"], <br>\n                         columns=[i for i in \"012345\"]) <br>\nsns_plot = sn.heatmap(df_cm, annot=True)\n</code></p>",
          "rawMarkdown": "Use seaborn:\n```\nimport seaborn as sn\nfrom sklearn.metrics import confusion_matrix  \ncm = confusion_matrix(truth, pred)  \ndf_cm = pd.DataFrame(cm, index=[i for i in \"012345\"],  \n                         columns=[i for i in \"012345\"])  \nsns_plot = sn.heatmap(df_cm, annot=True)\n```",
          "votes": 3
        },
        {
          "id": 921555,
          "postDate": "2020-07-09T11:27:46.090Z",
          "content": "<p>Thanks ! I've been following your comments on the other threads. Very insightful all around !</p>",
          "rawMarkdown": "Thanks ! I've been following your comments on the other threads. Very insightful all around !"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 829653,
      "author_name": "Yann Majewski",
      "author_url": "",
      "post_date": "2020-05-02T01:45:51.233000",
      "content": "<p>Two models didnt seem to help for me :(</p>",
      "votes": 3,
      "replies": [
        {
          "id": 830180,
          "author_name": "Vlad Vaduva",
          "author_url": "",
          "post_date": "2020-05-02T11:42:35.603000",
          "content": "<p>Sorry to hear that. The result was much worse or about the same as the single mode one ? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 830215,
          "author_name": "Yann Majewski",
          "author_url": "",
          "post_date": "2020-05-02T12:11:45.113000",
          "content": "<p>They were a little lower, maybe i need to tune them a bit more</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 829628,
      "author_name": "Shiyuan Zeng",
      "author_url": "",
      "post_date": "2020-05-02T01:22:51.290000",
      "content": "<p>Two models maybe more reasonable. The picture below shows the data in karolinska is unbalanced and radboud is more balanced(correct me if I was wrong), may be it also the cause of different performance of only a model to predict two different distributions?</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3149580%2Fc7eca96a293507a0a30fa8d18cc3ade7%2FDataDistribution.png?generation=1588382474836034&amp;alt=media\" alt=\"\"></p>\n\n<p>What's more, I have observed that the color in data of karolinska is darker compared with radboud in commom(I select some pictures randomly about 20 times and observe them, only crop-white and resize to 224, no augmentations)... So I place RandomBrightness/Contrast to my TODO list. As we can see, In sub-picture-3, the radboud data is hard to recognize by eyes.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3149580%2F80d0bff3db4cc90d8096be036a9f4a90%2FColorCompare.png?generation=1588383678831041&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3149580%2F1990ced2f6e6b7415cc9a9ba60817081%2FColorCompare_2.png?generation=1588384114939252&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3149580%2F3f6a1b85b87758cc868af79a4cd36b38%2FColorCompare_3.png?generation=1588384129739286&amp;alt=media\" alt=\"\"></p>",
      "votes": 2,
      "replies": [
        {
          "id": 830183,
          "author_name": "Vlad Vaduva",
          "author_url": "",
          "post_date": "2020-05-02T11:43:38.163000",
          "content": "<p>Good and useful observations <a href=\"/cnzengshiyuan\">@cnzengshiyuan</a> . I completely agree</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 830944,
          "author_name": "Shiyuan Zeng",
          "author_url": "",
          "post_date": "2020-05-03T02:04:30.317000",
          "content": "<p><a href=\"/vladvdv\">@vladvdv</a> haha, thank you! Since iafoss's great idea for cutting the picture into patches and get good score. I think an explainable way to modify the picure is the key to high score(Though I haven't finish implement). What do you think about it?\nPS: The RandomContrast behaves better than RandomBrighness(only on radboud). I used them with limit=0.3, always-apply=True from albumentations.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 896849,
      "author_name": "jwelliav",
      "author_url": "",
      "post_date": "2020-06-22T13:27:37.873000",
      "content": "<p>Could you please share how you calculated and made these really nice plots of the confusion matrix ? </p>\n\n<pre><code> Thanks :) \n</code></pre>",
      "votes": 0,
      "replies": [
        {
          "id": 896882,
          "author_name": "Arnaud Roussel",
          "author_url": "",
          "post_date": "2020-06-22T13:58:16.720000",
          "content": "<p>Use seaborn:\n<code>\nimport seaborn as sn\nfrom sklearn.metrics import confusion_matrix <br>\ncm = confusion_matrix(truth, pred) <br>\ndf_cm = pd.DataFrame(cm, index=[i for i in \"012345\"], <br>\n                         columns=[i for i in \"012345\"]) <br>\nsns_plot = sn.heatmap(df_cm, annot=True)\n</code></p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 921555,
          "author_name": "jwelliav",
          "author_url": "",
          "post_date": "2020-07-09T11:27:46.090000",
          "content": "<p>Thanks ! I've been following your comments on the other threads. Very insightful all around !</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "828967": "Hi everybody,\n\nMy first step was to use the whole resized image in order to make a baseline score which later on to try to improve using different techniques (segmentation models, smart crop, other computer vision techniques). \nLately I seen some very interesting approaches using tiles, @iafoss made a really good job describing his idea and deserves our appreciation  for sharing it. @frlemarchand also is working on a somehow similar approach, very interesting idea and I am curious about his results.\nGetting back on analyzing the baseline results. From some tries made, the optimum resolution was 512x512 and other parameters are described in this [discussion](https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/146332an) and this [kernel](https://www.kaggle.com/vladvdv/pytorch-training-customizable-kernel-with-5-folds).\nI used 5 folds, which using the mean prediction scored 0.69 on public leaderboard (0.70 with some added tricks).\nOn CV the confusion matrix looks interesting \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4005865%2Fc61b71aadf417b9891b029303c3818d4%2FconfMatrixAll.png?generation=1588333419925630&amp;alt=media)\n\nIt seems like:\n- the model cannot separate true label 2 and it's predicting label 1 in most cases(478 vs 247)\n- the model cannot separate true label 3 and it's predicting label 1 in most cases(266 vs 238)\n- the model cannot separate true label 4 and it's predicting label 5 and 1 in most cases(255 vs 323 vs 242)\n\nSo, there are big problems predicting labels 2,3 and 4 and a common ground is that a lot of mistakes made by the model are made by confusion with label 1.\n\nFor a better understanding of what is going on, let's go deeper by separating the confusion matrix by the data provider\n\nThis is the confusion matrix with data by karolinska\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4005865%2F6c1440139ad9f4bddfba2018182937be%2FconfMatrixK.png?generation=1588334130529204&amp;alt=media)\n\n\nand here by radboud\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4005865%2F20070763e5bec6a59da477aa7d35c6f8%2FconfMatrixR.png?generation=1588334056721834&amp;alt=media)\n\nIt is easily to observe that the problems are still on the 2,3 and 4 categories on both data providers but the model is making different type of mistakes:\n- prediction of the 2 category: on karolinska the model is predicting a lot of 1 instead of 2(98vs 322) but on radboud the model is predicting also 1 (in a lower rate than k: 149 vs 156) but also 5 (149 vs 133)\n- prediction of the 3 category: on karolinska the model is predicting a lot of 1 (21 vs 126) but on radboud it seems that the errors are spread between multiple categories (and mostly on 5: 217 vs 270)\n- prediction of the 4 category:  again karolinska is predicting a lot of 1(97 vs 171) and most of the mistakes of radboud are made by predicting again by confusing it with 5 (159 vs 257)\n\nAs a conclusion I can say that is clear that the images are different between providers and this can be felt in how the model behaves. The next step on my todo list is to design 2 different models, each model dedicated to a single provider, although unfortunately the data will be cut in half my hope is that the personalized model will better understand just one type of images.\n",
    "829653": "Two models didnt seem to help for me :(",
    "829628": "Two models maybe more reasonable. The picture below shows the data in karolinska is unbalanced and radboud is more balanced(correct me if I was wrong), may be it also the cause of different performance of only a model to predict two different distributions?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3149580%2Fc7eca96a293507a0a30fa8d18cc3ade7%2FDataDistribution.png?generation=1588382474836034&amp;alt=media)\n\nWhat's more, I have observed that the color in data of karolinska is darker compared with radboud in commom(I select some pictures randomly about 20 times and observe them, only crop-white and resize to 224, no augmentations)... So I place RandomBrightness/Contrast to my TODO list. As we can see, In sub-picture-3, the radboud data is hard to recognize by eyes.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3149580%2F80d0bff3db4cc90d8096be036a9f4a90%2FColorCompare.png?generation=1588383678831041&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3149580%2F1990ced2f6e6b7415cc9a9ba60817081%2FColorCompare_2.png?generation=1588384114939252&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3149580%2F3f6a1b85b87758cc868af79a4cd36b38%2FColorCompare_3.png?generation=1588384129739286&amp;alt=media)\n",
    "896849": " Could you please share how you calculated and made these really nice plots of the confusion matrix ? \n\n     Thanks :) \n"
  }
}