{
  "id": 169749,
  "title": "Why did we shake up/down and what could be done about it?",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/169749",
  "author_name": "arutema47",
  "post_date": "2020-07-25T06:03:55.537000",
  "votes": 16,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Note that this is my views, and not my team's. This is some thoughts on the LB/PB after watching everyone's solutions.</p>\n\n<p>Lot of people shook up and down in the PB, but not much discussion on why this happened are discussed.\nI thought it would be valuable to have discussions on the shake, and not just clean it up as \"lottary\". (anyways, it was much better than M5.)</p>\n\n<p>As <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/169151\">Qishen mentioned</a>, it is fundamentally difficult to build a stable PB with such a small private (~500) + QWK.\nHowever, I think people tried their best to hang in there. </p>\n\n<p>There were 3 possibility for the PB in general:\n    1. it has similar data distribution as LB\n    2.the data is difficult than LB\n    3.the data is easier than LB</p>\n\n<h2>Pattern 1.</h2>\n\n<p>In case of pattern 1), shakes do not happen.\nOn the other hand, if pattern (2) or (3), the final scores will shake a lot.</p>\n\n<h2>Pattern 2.</h2>\n\n<p>If you had a model which was more focused on classifing hard data, it might have shook down.\nSome techniques I tried for this approach was 1) oversampling difficult data (opposite of denosising), 2) online hard example mining, 3) focal loss.\nI did not get good CV/LB results with this approach, so I let it go.</p>\n\n<h2>Pattern 3.</h2>\n\n<p>What happened in this competition was (3), where the data was much easier to classify than the public set.\nTherefore, filtering difficult data contributed alot in getting a good Private score. \n<a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/169143\">1st</a>, <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/169113\">4th</a>, <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/169230\">6th</a>, <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/169205\">11th</a> solutions utilized this idea, building models which can classify easy data well were effective.</p>\n\n<h2>Which approach was right?</h2>\n\n<p>All models are wrong, but some are useful.</p>\n\n<p>Since we had only 2 submissions, we submitted 1) Best LB model and 2) model that does well for easy data. In other words, we bet on pattern 1 and 3, and luckly got 1st because it matched the private data distribution.</p>\n\n<p>If the private had more difficult data than LB, we would have shook down a lot..(so it was kind of coin-flipping..) Since we did not have the information on private prior, we had to bet on either sides. It would have been interesting to have 3 submissions, so we could have worked on either data distributions.</p>\n\n<p>But if we were looking for good clinical models, I would want a model that will both classify easy and hard data correctly.\nSo I think real-world good models are ones that have good score on both LB and private (and also more external data), like 2nd and 4th solutions.</p>",
  "messages": [
    {
      "id": 944451,
      "postDate": "2020-07-25T06:03:55.537Z",
      "content": "<p>Note that this is my views, and not my team's. This is some thoughts on the LB/PB after watching everyone's solutions.</p>\n\n<p>Lot of people shook up and down in the PB, but not much discussion on why this happened are discussed.\nI thought it would be valuable to have discussions on the shake, and not just clean it up as \"lottary\". (anyways, it was much better than M5.)</p>\n\n<p>As <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/169151\">Qishen mentioned</a>, it is fundamentally difficult to build a stable PB with such a small private (~500) + QWK.\nHowever, I think people tried their best to hang in there. </p>\n\n<p>There were 3 possibility for the PB in general:\n    1. it has similar data distribution as LB\n    2.the data is difficult than LB\n    3.the data is easier than LB</p>\n\n<h2>Pattern 1.</h2>\n\n<p>In case of pattern 1), shakes do not happen.\nOn the other hand, if pattern (2) or (3), the final scores will shake a lot.</p>\n\n<h2>Pattern 2.</h2>\n\n<p>If you had a model which was more focused on classifing hard data, it might have shook down.\nSome techniques I tried for this approach was 1) oversampling difficult data (opposite of denosising), 2) online hard example mining, 3) focal loss.\nI did not get good CV/LB results with this approach, so I let it go.</p>\n\n<h2>Pattern 3.</h2>\n\n<p>What happened in this competition was (3), where the data was much easier to classify than the public set.\nTherefore, filtering difficult data contributed alot in getting a good Private score. \n<a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/169143\">1st</a>, <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/169113\">4th</a>, <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/169230\">6th</a>, <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/169205\">11th</a> solutions utilized this idea, building models which can classify easy data well were effective.</p>\n\n<h2>Which approach was right?</h2>\n\n<p>All models are wrong, but some are useful.</p>\n\n<p>Since we had only 2 submissions, we submitted 1) Best LB model and 2) model that does well for easy data. In other words, we bet on pattern 1 and 3, and luckly got 1st because it matched the private data distribution.</p>\n\n<p>If the private had more difficult data than LB, we would have shook down a lot..(so it was kind of coin-flipping..) Since we did not have the information on private prior, we had to bet on either sides. It would have been interesting to have 3 submissions, so we could have worked on either data distributions.</p>\n\n<p>But if we were looking for good clinical models, I would want a model that will both classify easy and hard data correctly.\nSo I think real-world good models are ones that have good score on both LB and private (and also more external data), like 2nd and 4th solutions.</p>",
      "rawMarkdown": "Note that this is my views, and not my team's. This is some thoughts on the LB/PB after watching everyone's solutions.\n\nLot of people shook up and down in the PB, but not much discussion on why this happened are discussed.\nI thought it would be valuable to have discussions on the shake, and not just clean it up as \"lottary\". (anyways, it was much better than M5.)\n\nAs [Qishen mentioned](https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/169151), it is fundamentally difficult to build a stable PB with such a small private (~500) + QWK.\nHowever, I think people tried their best to hang in there. \n\nThere were 3 possibility for the PB in general:\n\t1. it has similar data distribution as LB\n\t2.the data is difficult than LB\n\t3.the data is easier than LB\n\n## Pattern 1.\nIn case of pattern 1), shakes do not happen.\nOn the other hand, if pattern (2) or (3), the final scores will shake a lot.\n\n## Pattern 2.\nIf you had a model which was more focused on classifing hard data, it might have shook down.\nSome techniques I tried for this approach was 1) oversampling difficult data (opposite of denosising), 2) online hard example mining, 3) focal loss.\nI did not get good CV/LB results with this approach, so I let it go.\n\n## Pattern 3.\nWhat happened in this competition was (3), where the data was much easier to classify than the public set.\nTherefore, filtering difficult data contributed alot in getting a good Private score. \n[1st](https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/169143), [4th](https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/169113), [6th](https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/169230), [11th](https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/169205) solutions utilized this idea, building models which can classify easy data well were effective.\n\n## Which approach was right?\nAll models are wrong, but some are useful.\n\nSince we had only 2 submissions, we submitted 1) Best LB model and 2) model that does well for easy data. In other words, we bet on pattern 1 and 3, and luckly got 1st because it matched the private data distribution.\n\nIf the private had more difficult data than LB, we would have shook down a lot..(so it was kind of coin-flipping..) Since we did not have the information on private prior, we had to bet on either sides. It would have been interesting to have 3 submissions, so we could have worked on either data distributions.\n\nBut if we were looking for good clinical models, I would want a model that will both classify easy and hard data correctly.\nSo I think real-world good models are ones that have good score on both LB and private (and also more external data), like 2nd and 4th solutions.\n\n",
      "votes": 17
    },
    {
      "id": 944491,
      "postDate": "2020-07-25T06:44:50.363Z",
      "content": "<p>In my opinion, the data from this competition can be viewed in the following aspects.</p>\n\n<ul>\n<li>Data provider ( Karolinska or Radboud )</li>\n<li>Noisy or Not</li>\n<li>Duplicated or Not</li>\n<li>Easy or Hard example ( I didn't notice it during the competition… )\n<ul><li>Perhaps this is what made us big shakeup\n<ul><li>Private LB scores higher overall than Public LB’s</li>\n<li>Private test hard example may be less than Public test</li></ul></li>\n<li>Assumption\n<ul><li>Our model is strong to easy example, but weak to hard example\n<ul><li>Because our denoise method (using gap between pred and ground truth of oof) is easy to remove hard example</li></ul></li>\n<li>This led to a divergence between Public and Private </li></ul></li></ul></li>\n</ul>\n\n<p>&gt; But if we were looking for good clinical models, I would want a model that will both classify easy and hard data correctly.</p>\n\n<p>Ideally, it should work both ways, but it depends on the use case.</p>",
      "rawMarkdown": "In my opinion, the data from this competition can be viewed in the following aspects.\n\n- Data provider ( Karolinska or Radboud )\n- Noisy or Not\n- Duplicated or Not\n- Easy or Hard example ( I didn't notice it during the competition… )\n    - Perhaps this is what made us big shakeup\n        - Private LB scores higher overall than Public LB’s\n        - Private test hard example may be less than Public test\n    - Assumption\n        - Our model is strong to easy example, but weak to hard example\n            - Because our denoise method (using gap between pred and ground truth of oof) is easy to remove hard example\n        - This led to a divergence between Public and Private \n\n&gt; But if we were looking for good clinical models, I would want a model that will both classify easy and hard data correctly.\n\nIdeally, it should work both ways, but it depends on the use case.\n",
      "votes": 5
    },
    {
      "id": 948321,
      "postDate": "2020-07-27T20:00:27.363Z",
      "content": "<p>thank you for writing this up <a href=\"/kyoshioka47\">@kyoshioka47</a> . this is so helpful for all of us to learn from! </p>",
      "rawMarkdown": "thank you for writing this up @kyoshioka47 . this is so helpful for all of us to learn from! ",
      "votes": 1
    },
    {
      "id": 945837,
      "postDate": "2020-07-26T07:05:06.947Z",
      "content": "<p>It is an interesting question.  Not sure I have seen so many LB movements in the 100s of places!! in an image competition at least. (In tabular data there always seems to be leak or reverse engineering, etc.) \nThere have been competitions where questions on the metric are raised and changes are made - Alaska2 and Univ of Liverpool recently. Metric QWK  was maybe not the best here? \nIt was observed that images were duplicates, i.e., slices of the same. So having a patient ID or grouping in folders as other DICOM competitions do may have been useful here.  It would have impacted validation strategies, CV scores. Maybe those that identified the duplicates and handled them differently could comment. \nThere is maybe an idea that in these competitions the distribution of public and private are somewhat similar as you indicate in point 1.  Perhaps this is harder to do with medical images? Or was not the organisers' objective?</p>\n\n<p>There are always learnings to take away no matter the results.  Congrats on 1st place! </p>",
      "rawMarkdown": "It is an interesting question.  Not sure I have seen so many LB movements in the 100s of places!! in an image competition at least. (In tabular data there always seems to be leak or reverse engineering, etc.) \nThere have been competitions where questions on the metric are raised and changes are made - Alaska2 and Univ of Liverpool recently. Metric QWK  was maybe not the best here? \nIt was observed that images were duplicates, i.e., slices of the same. So having a patient ID or grouping in folders as other DICOM competitions do may have been useful here.  It would have impacted validation strategies, CV scores. Maybe those that identified the duplicates and handled them differently could comment. \nThere is maybe an idea that in these competitions the distribution of public and private are somewhat similar as you indicate in point 1.  Perhaps this is harder to do with medical images? Or was not the organisers' objective?\n\nThere are always learnings to take away no matter the results.  Congrats on 1st place! \n",
      "votes": 1,
      "replies": [
        {
          "id": 947114,
          "postDate": "2020-07-27T04:52:28.707Z",
          "content": "<p>Interesting point on better metrics than QWK.. Yes there should be better ones around and <a href=\"/hirune924\">@hirune924</a> metrics are quite interesting.</p>\n\n<p>BTW, I shook 100s of places in other image competitions (Begali), which in case the Private set was a lot more difficult than the Public (which was what the hosts wanted to do). That's where I learned to think a lot more about CV/LB/PB differences with a lot of pain.</p>",
          "rawMarkdown": "Interesting point on better metrics than QWK.. Yes there should be better ones around and @hirune924 metrics are quite interesting.\n\nBTW, I shook 100s of places in other image competitions (Begali), which in case the Private set was a lot more difficult than the Public (which was what the hosts wanted to do). That's where I learned to think a lot more about CV/LB/PB differences with a lot of pain.",
          "votes": 1
        }
      ]
    },
    {
      "id": 945965,
      "postDate": "2020-07-26T08:44:58.590Z",
      "content": "<h3>just idea (A better metric than qwk)</h3>\n\n<p>QWK is not a stable metric. And I think that's due to the Expected term, which is difficult to improve or optimize. It is very difficult to optimize the Expected term, so there is a large element of luck. We monitored the Expected and Observed terms separately during training, and the Observed scores are very stable. If the metric for this competition was Observed instead of qwk, shake could be much smaller.\nthis is the code we used in this comoetition\n```\ndef monitored_cohen_kappa_score(\n    y1, y2, labels=None, weights=None, sample_weight=None, verbose=False,\n):\n    confusion = metrics.confusion_matrix(y1, y2, labels=labels, sample_weight=sample_weight)\n    n_classes = confusion.shape[0]\n    sum0 = np.sum(confusion, axis=0)\n    sum1 = np.sum(confusion, axis=1)\n    expected = np.outer(sum0, sum1) / np.sum(sum0)</p>\n\n<pre><code>if weights is None:\n    w_mat = np.ones([n_classes, n_classes], dtype=np.int)\n    w_mat.flat[:: n_classes + 1] = 0\nelif weights == \"linear\" or weights == \"quadratic\":\n    w_mat = np.zeros([n_classes, n_classes], dtype=np.int)\n    w_mat += np.arange(n_classes)\n    if weights == \"linear\":\n        w_mat = np.abs(w_mat - w_mat.T)\n    else:\n        w_mat = (w_mat - w_mat.T) ** 2\nelse:\n    raise ValueError(\"Unknown kappa weighting type.\")\n\no = np.sum(w_mat * confusion) / (np.sum(confusion) * ((n_classes - 1) ** 2))\ne = np.sum(w_mat * expected) / (np.sum(confusion) * ((n_classes - 1) ** 2))\nk = o / e\nif verbose:\n    print(confusion)\nreturn 1 - k, o, e\n</code></pre>\n\n<p>```</p>",
      "rawMarkdown": "### just idea (A better metric than qwk)\nQWK is not a stable metric. And I think that's due to the Expected term, which is difficult to improve or optimize. It is very difficult to optimize the Expected term, so there is a large element of luck. We monitored the Expected and Observed terms separately during training, and the Observed scores are very stable. If the metric for this competition was Observed instead of qwk, shake could be much smaller.\nthis is the code we used in this comoetition\n```\ndef monitored_cohen_kappa_score(\n    y1, y2, labels=None, weights=None, sample_weight=None, verbose=False,\n):\n    confusion = metrics.confusion_matrix(y1, y2, labels=labels, sample_weight=sample_weight)\n    n_classes = confusion.shape[0]\n    sum0 = np.sum(confusion, axis=0)\n    sum1 = np.sum(confusion, axis=1)\n    expected = np.outer(sum0, sum1) / np.sum(sum0)\n\n    if weights is None:\n        w_mat = np.ones([n_classes, n_classes], dtype=np.int)\n        w_mat.flat[:: n_classes + 1] = 0\n    elif weights == \"linear\" or weights == \"quadratic\":\n        w_mat = np.zeros([n_classes, n_classes], dtype=np.int)\n        w_mat += np.arange(n_classes)\n        if weights == \"linear\":\n            w_mat = np.abs(w_mat - w_mat.T)\n        else:\n            w_mat = (w_mat - w_mat.T) ** 2\n    else:\n        raise ValueError(\"Unknown kappa weighting type.\")\n\n    o = np.sum(w_mat * confusion) / (np.sum(confusion) * ((n_classes - 1) ** 2))\n    e = np.sum(w_mat * expected) / (np.sum(confusion) * ((n_classes - 1) ** 2))\n    k = o / e\n    if verbose:\n        print(confusion)\n    return 1 - k, o, e\n\n```",
      "votes": 2
    },
    {
      "id": 945876,
      "postDate": "2020-07-26T07:37:51.590Z",
      "content": "<p><a href=\"/something4kag\">@something4kag</a> \nthx for comment. It is interesting point.</p>\n\n<p>This number of test data and QWK doesn't seem to be the best combination, but I don't know if there are any other good metrics out there...</p>\n\n<blockquote>\n  <p>It was observed that images were duplicates, i.e., slices of the same. So having a patient ID or grouping in folders as other DICOM competitions do may have been useful here. It would have impacted validation strategies, CV scores</p>\n</blockquote>\n\n<p>I agree with you. This is not a trend in the test data and I think this is a problem that could be solved if the host had provided the biopsy id.</p>\n\n<p>It seemed to me that this competition was throwing too many issues at the participants. (how to process WSI, multi data provider, different annotation process, annotation noise, duplicated images, etc...) <br>\nI know that the real issues are more complex, but as a competition, I think the hosts should have narrowed the issues a bit more.</p>",
      "rawMarkdown": "@something4kag \nthx for comment. It is interesting point.\n\nThis number of test data and QWK doesn't seem to be the best combination, but I don't know if there are any other good metrics out there...\n\n&gt; It was observed that images were duplicates, i.e., slices of the same. So having a patient ID or grouping in folders as other DICOM competitions do may have been useful here. It would have impacted validation strategies, CV scores\n\nI agree with you. This is not a trend in the test data and I think this is a problem that could be solved if the host had provided the biopsy id.\n\nIt seemed to me that this competition was throwing too many issues at the participants. (how to process WSI, multi data provider, different annotation process, annotation noise, duplicated images, etc...)  \nI know that the real issues are more complex, but as a competition, I think the hosts should have narrowed the issues a bit more.",
      "votes": 2
    },
    {
      "id": 954808,
      "postDate": "2020-08-02T05:46:47.280Z",
      "content": "<p>great</p>",
      "rawMarkdown": "great",
      "votes": -1
    },
    {
      "id": 947967,
      "postDate": "2020-07-27T15:23:42.283Z",
      "content": "<p>Interesting project.</p>",
      "rawMarkdown": "Interesting project."
    }
  ],
  "comments": [
    {
      "id": 944491,
      "author_name": "fam_taro",
      "author_url": "",
      "post_date": "2020-07-25T06:44:50.363000",
      "content": "<p>In my opinion, the data from this competition can be viewed in the following aspects.</p>\n\n<ul>\n<li>Data provider ( Karolinska or Radboud )</li>\n<li>Noisy or Not</li>\n<li>Duplicated or Not</li>\n<li>Easy or Hard example ( I didn't notice it during the competition… )\n<ul><li>Perhaps this is what made us big shakeup\n<ul><li>Private LB scores higher overall than Public LB’s</li>\n<li>Private test hard example may be less than Public test</li></ul></li>\n<li>Assumption\n<ul><li>Our model is strong to easy example, but weak to hard example\n<ul><li>Because our denoise method (using gap between pred and ground truth of oof) is easy to remove hard example</li></ul></li>\n<li>This led to a divergence between Public and Private </li></ul></li></ul></li>\n</ul>\n\n<p>&gt; But if we were looking for good clinical models, I would want a model that will both classify easy and hard data correctly.</p>\n\n<p>Ideally, it should work both ways, but it depends on the use case.</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 948321,
      "author_name": "Mehul Sampat",
      "author_url": "",
      "post_date": "2020-07-27T20:00:27.363000",
      "content": "<p>thank you for writing this up <a href=\"/kyoshioka47\">@kyoshioka47</a> . this is so helpful for all of us to learn from! </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 945837,
      "author_name": "something4kag",
      "author_url": "",
      "post_date": "2020-07-26T07:05:06.947000",
      "content": "<p>It is an interesting question.  Not sure I have seen so many LB movements in the 100s of places!! in an image competition at least. (In tabular data there always seems to be leak or reverse engineering, etc.) \nThere have been competitions where questions on the metric are raised and changes are made - Alaska2 and Univ of Liverpool recently. Metric QWK  was maybe not the best here? \nIt was observed that images were duplicates, i.e., slices of the same. So having a patient ID or grouping in folders as other DICOM competitions do may have been useful here.  It would have impacted validation strategies, CV scores. Maybe those that identified the duplicates and handled them differently could comment. \nThere is maybe an idea that in these competitions the distribution of public and private are somewhat similar as you indicate in point 1.  Perhaps this is harder to do with medical images? Or was not the organisers' objective?</p>\n\n<p>There are always learnings to take away no matter the results.  Congrats on 1st place! </p>",
      "votes": 1,
      "replies": [
        {
          "id": 947114,
          "author_name": "arutema47",
          "author_url": "",
          "post_date": "2020-07-27T04:52:28.707000",
          "content": "<p>Interesting point on better metrics than QWK.. Yes there should be better ones around and <a href=\"/hirune924\">@hirune924</a> metrics are quite interesting.</p>\n\n<p>BTW, I shook 100s of places in other image competitions (Begali), which in case the Private set was a lot more difficult than the Public (which was what the hosts wanted to do). That's where I learned to think a lot more about CV/LB/PB differences with a lot of pain.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 945965,
      "author_name": "hirune924",
      "author_url": "",
      "post_date": "2020-07-26T08:44:58.590000",
      "content": "<h3>just idea (A better metric than qwk)</h3>\n\n<p>QWK is not a stable metric. And I think that's due to the Expected term, which is difficult to improve or optimize. It is very difficult to optimize the Expected term, so there is a large element of luck. We monitored the Expected and Observed terms separately during training, and the Observed scores are very stable. If the metric for this competition was Observed instead of qwk, shake could be much smaller.\nthis is the code we used in this comoetition\n```\ndef monitored_cohen_kappa_score(\n    y1, y2, labels=None, weights=None, sample_weight=None, verbose=False,\n):\n    confusion = metrics.confusion_matrix(y1, y2, labels=labels, sample_weight=sample_weight)\n    n_classes = confusion.shape[0]\n    sum0 = np.sum(confusion, axis=0)\n    sum1 = np.sum(confusion, axis=1)\n    expected = np.outer(sum0, sum1) / np.sum(sum0)</p>\n\n<pre><code>if weights is None:\n    w_mat = np.ones([n_classes, n_classes], dtype=np.int)\n    w_mat.flat[:: n_classes + 1] = 0\nelif weights == \"linear\" or weights == \"quadratic\":\n    w_mat = np.zeros([n_classes, n_classes], dtype=np.int)\n    w_mat += np.arange(n_classes)\n    if weights == \"linear\":\n        w_mat = np.abs(w_mat - w_mat.T)\n    else:\n        w_mat = (w_mat - w_mat.T) ** 2\nelse:\n    raise ValueError(\"Unknown kappa weighting type.\")\n\no = np.sum(w_mat * confusion) / (np.sum(confusion) * ((n_classes - 1) ** 2))\ne = np.sum(w_mat * expected) / (np.sum(confusion) * ((n_classes - 1) ** 2))\nk = o / e\nif verbose:\n    print(confusion)\nreturn 1 - k, o, e\n</code></pre>\n\n<p>```</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 945876,
      "author_name": "fam_taro",
      "author_url": "",
      "post_date": "2020-07-26T07:37:51.590000",
      "content": "<p><a href=\"/something4kag\">@something4kag</a> \nthx for comment. It is interesting point.</p>\n\n<p>This number of test data and QWK doesn't seem to be the best combination, but I don't know if there are any other good metrics out there...</p>\n\n<blockquote>\n  <p>It was observed that images were duplicates, i.e., slices of the same. So having a patient ID or grouping in folders as other DICOM competitions do may have been useful here. It would have impacted validation strategies, CV scores</p>\n</blockquote>\n\n<p>I agree with you. This is not a trend in the test data and I think this is a problem that could be solved if the host had provided the biopsy id.</p>\n\n<p>It seemed to me that this competition was throwing too many issues at the participants. (how to process WSI, multi data provider, different annotation process, annotation noise, duplicated images, etc...) <br>\nI know that the real issues are more complex, but as a competition, I think the hosts should have narrowed the issues a bit more.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 954808,
      "author_name": "Faiza Islam Nahin",
      "author_url": "",
      "post_date": "2020-08-02T05:46:47.280000",
      "content": "<p>great</p>",
      "votes": -1,
      "replies": []
    },
    {
      "id": 947967,
      "author_name": "Mohammad Sakib Mahmood",
      "author_url": "",
      "post_date": "2020-07-27T15:23:42.283000",
      "content": "<p>Interesting project.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "944451": "Note that this is my views, and not my team's. This is some thoughts on the LB/PB after watching everyone's solutions.\n\nLot of people shook up and down in the PB, but not much discussion on why this happened are discussed.\nI thought it would be valuable to have discussions on the shake, and not just clean it up as \"lottary\". (anyways, it was much better than M5.)\n\nAs [Qishen mentioned](https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/169151), it is fundamentally difficult to build a stable PB with such a small private (~500) + QWK.\nHowever, I think people tried their best to hang in there. \n\nThere were 3 possibility for the PB in general:\n\t1. it has similar data distribution as LB\n\t2.the data is difficult than LB\n\t3.the data is easier than LB\n\n## Pattern 1.\nIn case of pattern 1), shakes do not happen.\nOn the other hand, if pattern (2) or (3), the final scores will shake a lot.\n\n## Pattern 2.\nIf you had a model which was more focused on classifing hard data, it might have shook down.\nSome techniques I tried for this approach was 1) oversampling difficult data (opposite of denosising), 2) online hard example mining, 3) focal loss.\nI did not get good CV/LB results with this approach, so I let it go.\n\n## Pattern 3.\nWhat happened in this competition was (3), where the data was much easier to classify than the public set.\nTherefore, filtering difficult data contributed alot in getting a good Private score. \n[1st](https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/169143), [4th](https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/169113), [6th](https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/169230), [11th](https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/169205) solutions utilized this idea, building models which can classify easy data well were effective.\n\n## Which approach was right?\nAll models are wrong, but some are useful.\n\nSince we had only 2 submissions, we submitted 1) Best LB model and 2) model that does well for easy data. In other words, we bet on pattern 1 and 3, and luckly got 1st because it matched the private data distribution.\n\nIf the private had more difficult data than LB, we would have shook down a lot..(so it was kind of coin-flipping..) Since we did not have the information on private prior, we had to bet on either sides. It would have been interesting to have 3 submissions, so we could have worked on either data distributions.\n\nBut if we were looking for good clinical models, I would want a model that will both classify easy and hard data correctly.\nSo I think real-world good models are ones that have good score on both LB and private (and also more external data), like 2nd and 4th solutions.\n\n",
    "944491": "In my opinion, the data from this competition can be viewed in the following aspects.\n\n- Data provider ( Karolinska or Radboud )\n- Noisy or Not\n- Duplicated or Not\n- Easy or Hard example ( I didn't notice it during the competition… )\n    - Perhaps this is what made us big shakeup\n        - Private LB scores higher overall than Public LB’s\n        - Private test hard example may be less than Public test\n    - Assumption\n        - Our model is strong to easy example, but weak to hard example\n            - Because our denoise method (using gap between pred and ground truth of oof) is easy to remove hard example\n        - This led to a divergence between Public and Private \n\n&gt; But if we were looking for good clinical models, I would want a model that will both classify easy and hard data correctly.\n\nIdeally, it should work both ways, but it depends on the use case.\n",
    "948321": "thank you for writing this up @kyoshioka47 . this is so helpful for all of us to learn from! ",
    "945837": "It is an interesting question.  Not sure I have seen so many LB movements in the 100s of places!! in an image competition at least. (In tabular data there always seems to be leak or reverse engineering, etc.) \nThere have been competitions where questions on the metric are raised and changes are made - Alaska2 and Univ of Liverpool recently. Metric QWK  was maybe not the best here? \nIt was observed that images were duplicates, i.e., slices of the same. So having a patient ID or grouping in folders as other DICOM competitions do may have been useful here.  It would have impacted validation strategies, CV scores. Maybe those that identified the duplicates and handled them differently could comment. \nThere is maybe an idea that in these competitions the distribution of public and private are somewhat similar as you indicate in point 1.  Perhaps this is harder to do with medical images? Or was not the organisers' objective?\n\nThere are always learnings to take away no matter the results.  Congrats on 1st place! \n",
    "945965": "### just idea (A better metric than qwk)\nQWK is not a stable metric. And I think that's due to the Expected term, which is difficult to improve or optimize. It is very difficult to optimize the Expected term, so there is a large element of luck. We monitored the Expected and Observed terms separately during training, and the Observed scores are very stable. If the metric for this competition was Observed instead of qwk, shake could be much smaller.\nthis is the code we used in this comoetition\n```\ndef monitored_cohen_kappa_score(\n    y1, y2, labels=None, weights=None, sample_weight=None, verbose=False,\n):\n    confusion = metrics.confusion_matrix(y1, y2, labels=labels, sample_weight=sample_weight)\n    n_classes = confusion.shape[0]\n    sum0 = np.sum(confusion, axis=0)\n    sum1 = np.sum(confusion, axis=1)\n    expected = np.outer(sum0, sum1) / np.sum(sum0)\n\n    if weights is None:\n        w_mat = np.ones([n_classes, n_classes], dtype=np.int)\n        w_mat.flat[:: n_classes + 1] = 0\n    elif weights == \"linear\" or weights == \"quadratic\":\n        w_mat = np.zeros([n_classes, n_classes], dtype=np.int)\n        w_mat += np.arange(n_classes)\n        if weights == \"linear\":\n            w_mat = np.abs(w_mat - w_mat.T)\n        else:\n            w_mat = (w_mat - w_mat.T) ** 2\n    else:\n        raise ValueError(\"Unknown kappa weighting type.\")\n\n    o = np.sum(w_mat * confusion) / (np.sum(confusion) * ((n_classes - 1) ** 2))\n    e = np.sum(w_mat * expected) / (np.sum(confusion) * ((n_classes - 1) ** 2))\n    k = o / e\n    if verbose:\n        print(confusion)\n    return 1 - k, o, e\n\n```",
    "945876": "@something4kag \nthx for comment. It is interesting point.\n\nThis number of test data and QWK doesn't seem to be the best combination, but I don't know if there are any other good metrics out there...\n\n&gt; It was observed that images were duplicates, i.e., slices of the same. So having a patient ID or grouping in folders as other DICOM competitions do may have been useful here. It would have impacted validation strategies, CV scores\n\nI agree with you. This is not a trend in the test data and I think this is a problem that could be solved if the host had provided the biopsy id.\n\nIt seemed to me that this competition was throwing too many issues at the participants. (how to process WSI, multi data provider, different annotation process, annotation noise, duplicated images, etc...)  \nI know that the real issues are more complex, but as a competition, I think the hosts should have narrowed the issues a bit more.",
    "954808": "great",
    "947967": "Interesting project."
  }
}