{
  "id": 168007,
  "title": "[Noise Cooking Recipe] Noise level for each provider?",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/168007",
  "author_name": "Nicholas Lyu",
  "post_date": "2020-07-18T18:10:58.834000",
  "votes": 5,
  "comment_count": 11,
  "views": 0,
  "content": "<p>In the competition report it was claimed that the students who graded Radboud train scored ~.85 on testset. This is a very valuable piece of information essential to winning this noise-cooking contest. \nI wonder if anyone would like to disclose their estimate of noise level in Karolinska? There exists systematic (asymmetric) noise for both providers. One of my models scored .965 on CV, Karolinska, and scored only .88 on the LB (made sure there was no leak).</p>\n\n<p>*On second thought, are there any duplicates or near-duplicates in Karolinska? Found a lot in Radboud</p>\n\n<p>Opening this thread to try find the second most important value (the first one is .85) to winning this competition.</p>",
  "messages": [
    {
      "id": 935998,
      "postDate": "2020-07-19T21:56:22.933Z",
      "content": "<p>The important point is that the test data noise is probably considerably less than the noise level of the training data. The test labels were determined by consensus of multiple pathologists, whereas the training data was determined by looking through the original pathology reports. </p>\n\n<p>I see this phenomenon a lot in radiology. The training data is labeled by looking through original radiology reports vs. test set is rated by multiple radiologists. Many times, test data performance actually exceeds validation performance (validation set being a subset of the training data), the intuition being that there is just greater variability that is harder fit in the training data labels (single read, many different radiologists reading). This seems to be the case here, my 0.934 model only gets ~0.91 CV. </p>",
      "rawMarkdown": "The important point is that the test data noise is probably considerably less than the noise level of the training data. The test labels were determined by consensus of multiple pathologists, whereas the training data was determined by looking through the original pathology reports. \n\nI see this phenomenon a lot in radiology. The training data is labeled by looking through original radiology reports vs. test set is rated by multiple radiologists. Many times, test data performance actually exceeds validation performance (validation set being a subset of the training data), the intuition being that there is just greater variability that is harder fit in the training data labels (single read, many different radiologists reading). This seems to be the case here, my 0.934 model only gets ~0.91 CV. ",
      "votes": 14,
      "replies": [
        {
          "id": 936146,
          "postDate": "2020-07-20T02:47:34.190Z",
          "content": "<p><a href=\"/vaillant\">@vaillant</a> Thanks. This seems like the case based on the competition report. The problem is to find the point where CV can't be trusted any more. I'm pretty confident that with the correct method, CV can be arbitrarily high (&gt;.94, with duplicates removed), but that would just be overfitting the label noise in training data.</p>\n\n<p>If you don't mind, what is the CV breakdown between the two institutions in your .934 model? Thanks in advance.</p>",
          "rawMarkdown": "@vaillant Thanks. This seems like the case based on the competition report. The problem is to find the point where CV can't be trusted any more. I'm pretty confident that with the correct method, CV can be arbitrarily high (&gt;.94, with duplicates removed), but that would just be overfitting the label noise in training data.\n\nIf you don't mind, what is the CV breakdown between the two institutions in your .934 model? Thanks in advance."
        },
        {
          "id": 936581,
          "postDate": "2020-07-20T10:30:52.370Z",
          "content": "<p>Yes, I think it is possible to overfit the noise in the training data. I think I am overfitting test data, so we will see what happens on private LB. </p>\n\n<p>Karolinska 0.91-0.92, Radboud 0.88-0.89</p>",
          "rawMarkdown": "Yes, I think it is possible to overfit the noise in the training data. I think I am overfitting test data, so we will see what happens on private LB. \n\nKarolinska 0.91-0.92, Radboud 0.88-0.89",
          "votes": 3
        },
        {
          "id": 936597,
          "postDate": "2020-07-20T10:44:04.003Z",
          "content": "<p><a href=\"/vaillant\">@vaillant</a> I personally think the public / private leaderboards will match quite well, because they were labeled similarly (It's another matter if the label distributions are different, though). About the CV score for Radboud, is that with or without removed duplicates?</p>",
          "rawMarkdown": "@vaillant I personally think the public / private leaderboards will match quite well, because they were labeled similarly (It's another matter if the label distributions are different, though). About the CV score for Radboud, is that with or without removed duplicates?",
          "votes": 1
        },
        {
          "id": 936600,
          "postDate": "2020-07-20T10:50:25.273Z",
          "content": "<p>I did not remove any duplicates.</p>",
          "rawMarkdown": "I did not remove any duplicates.",
          "votes": 3
        },
        {
          "id": 938871,
          "postDate": "2020-07-21T20:15:52.637Z",
          "content": "<p><a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a>  is tthis your single fold model ?<br>\ndid u keep training labels as it is ? </p>",
          "rawMarkdown": "@vaillant  is tthis your single fold model ?\ndid u keep training labels as it is ? "
        },
        {
          "id": 938902,
          "postDate": "2020-07-21T20:55:21.370Z",
          "content": "<p>It's 3-fold, I did not change any labels. </p>",
          "rawMarkdown": "It's 3-fold, I did not change any labels. "
        }
      ]
    },
    {
      "id": 934780,
      "postDate": "2020-07-18T18:10:58.833Z",
      "content": "<p>In the competition report it was claimed that the students who graded Radboud train scored ~.85 on testset. This is a very valuable piece of information essential to winning this noise-cooking contest. \nI wonder if anyone would like to disclose their estimate of noise level in Karolinska? There exists systematic (asymmetric) noise for both providers. One of my models scored .965 on CV, Karolinska, and scored only .88 on the LB (made sure there was no leak).</p>\n\n<p>*On second thought, are there any duplicates or near-duplicates in Karolinska? Found a lot in Radboud</p>\n\n<p>Opening this thread to try find the second most important value (the first one is .85) to winning this competition.</p>",
      "rawMarkdown": "In the competition report it was claimed that the students who graded Radboud train scored ~.85 on testset. This is a very valuable piece of information essential to winning this noise-cooking contest. \nI wonder if anyone would like to disclose their estimate of noise level in Karolinska? There exists systematic (asymmetric) noise for both providers. One of my models scored .965 on CV, Karolinska, and scored only .88 on the LB (made sure there was no leak).\n\n*On second thought, are there any duplicates or near-duplicates in Karolinska? Found a lot in Radboud\n\nOpening this thread to try find the second most important value (the first one is .85) to winning this competition.",
      "votes": 5
    },
    {
      "id": 934787,
      "postDate": "2020-07-18T18:19:00.087Z",
      "content": "<p><a href=\"/roguekk007\">@roguekk007</a> Karolinska \"duplicates\" are usually stacked together in a single slide </p>",
      "rawMarkdown": "@roguekk007 Karolinska \"duplicates\" are usually stacked together in a single slide ",
      "replies": [
        {
          "id": 935014,
          "postDate": "2020-07-19T02:42:14.910Z",
          "content": "<p>What does \"slide\" mean?</p>",
          "rawMarkdown": "What does \"slide\" mean?"
        },
        {
          "id": 935385,
          "postDate": "2020-07-19T10:51:32.843Z",
          "content": "<p>image.  images here are Whole Slide Image (WSI).</p>",
          "rawMarkdown": "image.  images here are Whole Slide Image (WSI).",
          "votes": 2
        },
        {
          "id": 935409,
          "postDate": "2020-07-19T11:12:25.387Z",
          "content": "<p>I understood =)\nthank you =) =) =) =) =)</p>",
          "rawMarkdown": "I understood =)\nthank you =) =) =) =) =)",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 935998,
      "author_name": "Ian Pan",
      "author_url": "",
      "post_date": "2020-07-19T21:56:22.933000",
      "content": "<p>The important point is that the test data noise is probably considerably less than the noise level of the training data. The test labels were determined by consensus of multiple pathologists, whereas the training data was determined by looking through the original pathology reports. </p>\n\n<p>I see this phenomenon a lot in radiology. The training data is labeled by looking through original radiology reports vs. test set is rated by multiple radiologists. Many times, test data performance actually exceeds validation performance (validation set being a subset of the training data), the intuition being that there is just greater variability that is harder fit in the training data labels (single read, many different radiologists reading). This seems to be the case here, my 0.934 model only gets ~0.91 CV. </p>",
      "votes": 14,
      "replies": [
        {
          "id": 936146,
          "author_name": "Nicholas Lyu",
          "author_url": "",
          "post_date": "2020-07-20T02:47:34.190000",
          "content": "<p><a href=\"/vaillant\">@vaillant</a> Thanks. This seems like the case based on the competition report. The problem is to find the point where CV can't be trusted any more. I'm pretty confident that with the correct method, CV can be arbitrarily high (&gt;.94, with duplicates removed), but that would just be overfitting the label noise in training data.</p>\n\n<p>If you don't mind, what is the CV breakdown between the two institutions in your .934 model? Thanks in advance.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 936581,
          "author_name": "Ian Pan",
          "author_url": "",
          "post_date": "2020-07-20T10:30:52.370000",
          "content": "<p>Yes, I think it is possible to overfit the noise in the training data. I think I am overfitting test data, so we will see what happens on private LB. </p>\n\n<p>Karolinska 0.91-0.92, Radboud 0.88-0.89</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 936597,
          "author_name": "Nicholas Lyu",
          "author_url": "",
          "post_date": "2020-07-20T10:44:04.003000",
          "content": "<p><a href=\"/vaillant\">@vaillant</a> I personally think the public / private leaderboards will match quite well, because they were labeled similarly (It's another matter if the label distributions are different, though). About the CV score for Radboud, is that with or without removed duplicates?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 936600,
          "author_name": "Ian Pan",
          "author_url": "",
          "post_date": "2020-07-20T10:50:25.273000",
          "content": "<p>I did not remove any duplicates.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 938871,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-07-21T20:15:52.637000",
          "content": "<p><a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a>  is tthis your single fold model ?<br>\ndid u keep training labels as it is ? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 938902,
          "author_name": "Ian Pan",
          "author_url": "",
          "post_date": "2020-07-21T20:55:21.370000",
          "content": "<p>It's 3-fold, I did not change any labels. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 934787,
      "author_name": "Eek The Cat",
      "author_url": "",
      "post_date": "2020-07-18T18:19:00.087000",
      "content": "<p><a href=\"/roguekk007\">@roguekk007</a> Karolinska \"duplicates\" are usually stacked together in a single slide </p>",
      "votes": 0,
      "replies": [
        {
          "id": 935014,
          "author_name": "YUYUTA",
          "author_url": "",
          "post_date": "2020-07-19T02:42:14.910000",
          "content": "<p>What does \"slide\" mean?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 935385,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-07-19T10:51:32.843000",
          "content": "<p>image.  images here are Whole Slide Image (WSI).</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 935409,
          "author_name": "YUYUTA",
          "author_url": "",
          "post_date": "2020-07-19T11:12:25.387000",
          "content": "<p>I understood =)\nthank you =) =) =) =) =)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "935998": "The important point is that the test data noise is probably considerably less than the noise level of the training data. The test labels were determined by consensus of multiple pathologists, whereas the training data was determined by looking through the original pathology reports. \n\nI see this phenomenon a lot in radiology. The training data is labeled by looking through original radiology reports vs. test set is rated by multiple radiologists. Many times, test data performance actually exceeds validation performance (validation set being a subset of the training data), the intuition being that there is just greater variability that is harder fit in the training data labels (single read, many different radiologists reading). This seems to be the case here, my 0.934 model only gets ~0.91 CV. ",
    "934780": "In the competition report it was claimed that the students who graded Radboud train scored ~.85 on testset. This is a very valuable piece of information essential to winning this noise-cooking contest. \nI wonder if anyone would like to disclose their estimate of noise level in Karolinska? There exists systematic (asymmetric) noise for both providers. One of my models scored .965 on CV, Karolinska, and scored only .88 on the LB (made sure there was no leak).\n\n*On second thought, are there any duplicates or near-duplicates in Karolinska? Found a lot in Radboud\n\nOpening this thread to try find the second most important value (the first one is .85) to winning this competition.",
    "934787": "@roguekk007 Karolinska \"duplicates\" are usually stacked together in a single slide "
  }
}