{
  "id": 165774,
  "title": "Analysis on Why Scores stuck at 89-90",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/165774",
  "author_name": "Jaideep",
  "post_date": "2020-07-11T04:34:27.115000",
  "votes": 20,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Upon closely analysing the data of two source providers  RA and KA ,i am getting a strong indication that there are certain set of slide from both data providers which have incorrect labels  . If we exclude  that  data cv qk shoots up to 94-95  but it dsnt materializes same on the Leader board which means Test set does have those problem making slides type contained in  it so your trained model on data excluding those WSI  predicts incorrect labels for those type of slides in test set so your score stucks at 89-90 . Any one who has well addressed those incorrect labels will be in top 3 .</p>\n\n<p>Let me know your opinion. </p>",
  "messages": [
    {
      "id": 923678,
      "postDate": "2020-07-11T04:34:27.117Z",
      "content": "<p>Upon closely analysing the data of two source providers  RA and KA ,i am getting a strong indication that there are certain set of slide from both data providers which have incorrect labels  . If we exclude  that  data cv qk shoots up to 94-95  but it dsnt materializes same on the Leader board which means Test set does have those problem making slides type contained in  it so your trained model on data excluding those WSI  predicts incorrect labels for those type of slides in test set so your score stucks at 89-90 . Any one who has well addressed those incorrect labels will be in top 3 .</p>\n\n<p>Let me know your opinion. </p>",
      "rawMarkdown": "Upon closely analysing the data of two source providers  RA and KA ,i am getting a strong indication that there are certain set of slide from both data providers which have incorrect labels  . If we exclude  that  data cv qk shoots up to 94-95  but it dsnt materializes same on the Leader board which means Test set does have those problem making slides type contained in  it so your trained model on data excluding those WSI  predicts incorrect labels for those type of slides in test set so your score stucks at 89-90 . Any one who has well addressed those incorrect labels will be in top 3 .\n\nLet me know your opinion. \n\n \n",
      "votes": 20
    },
    {
      "id": 925191,
      "postDate": "2020-07-11T22:03:57.917Z",
      "content": "<p>I found that in my best lb run rn (0.91), the accuracy for class 2,3,4,5 are only around 0.5. Radboud's labeled by students so the provided labels themselves only have 0.7 accuracy, so it's going to come down to whoever can handle the noisy labels in the best way possible</p>",
      "rawMarkdown": "I found that in my best lb run rn (0.91), the accuracy for class 2,3,4,5 are only around 0.5. Radboud's labeled by students so the provided labels themselves only have 0.7 accuracy, so it's going to come down to whoever can handle the noisy labels in the best way possible",
      "votes": 2,
      "replies": [
        {
          "id": 929531,
          "postDate": "2020-07-14T18:21:27.050Z",
          "content": "<p>Ka are also noisy, its difficult to relabel them up .</p>",
          "rawMarkdown": "Ka are also noisy, its difficult to relabel them up ."
        }
      ]
    },
    {
      "id": 924350,
      "postDate": "2020-07-11T11:43:32.467Z",
      "content": "<p>I agree. Tried a couple of tricks to address that issue, including online relabelling, but all that just slightly degrades local score. If you don't mind me asking, what certain set of slides you're referring to? All of my efforts were based mostly on online loss values.</p>",
      "rawMarkdown": "I agree. Tried a couple of tricks to address that issue, including online relabelling, but all that just slightly degrades local score. If you don't mind me asking, what certain set of slides you're referring to? All of my efforts were based mostly on online loss values.",
      "replies": [
        {
          "id": 924354,
          "postDate": "2020-07-11T11:49:18.817Z",
          "content": "<p>i can just tell you that they are less than 500-600 . If you exclude then only your local cv qwk gets sky rockted  not the lbs  as LB comprises of those same type of images that you exclude numbering same as in Train set not sure how many such would be there in Private set ,if they are in majority then scores are going to nose dive .</p>",
          "rawMarkdown": "i can just tell you that they are less than 500-600 . If you exclude then only your local cv qwk gets sky rockted  not the lbs  as LB comprises of those same type of images that you exclude numbering same as in Train set not sure how many such would be there in Private set ,if they are in majority then scores are going to nose dive ."
        },
        {
          "id": 924396,
          "postDate": "2020-07-11T12:08:25.773Z",
          "content": "<p>Well, if you're talking about labelling errors, organizers promised private test would be free of them. If you're referring to some particular subclass of slides, which are hard for model to classify correctly, then yes, shakeup will be funny :)</p>",
          "rawMarkdown": "Well, if you're talking about labelling errors, organizers promised private test would be free of them. If you're referring to some particular subclass of slides, which are hard for model to classify correctly, then yes, shakeup will be funny :)"
        },
        {
          "id": 924403,
          "postDate": "2020-07-11T12:10:10.480Z",
          "content": "<p>Test set will have correct labels but Train set for those supposedly hard slides are not surely . If I place them completely in Validation set then i cant get QWK  more than 75 -77. </p>",
          "rawMarkdown": "Test set will have correct labels but Train set for those supposedly hard slides are not surely . If I place them completely in Validation set then i cant get QWK  more than 75 -77. "
        },
        {
          "id": 924425,
          "postDate": "2020-07-11T12:19:31.863Z",
          "content": "<p>That's interesting. Slides in question are mostly from radboud, right?</p>",
          "rawMarkdown": "That's interesting. Slides in question are mostly from radboud, right?\n"
        },
        {
          "id": 924483,
          "postDate": "2020-07-11T13:05:04.957Z",
          "content": "<p>yup but interestingly the more problamatic ones are karloniska in those</p>",
          "rawMarkdown": "yup but interestingly the more problamatic ones are karloniska in those"
        },
        {
          "id": 924503,
          "postDate": "2020-07-11T13:21:29.137Z",
          "content": "<p>How do you define most problematic? Largest deviation of predicted class from true?</p>",
          "rawMarkdown": "How do you define most problematic? Largest deviation of predicted class from true?\n\n"
        },
        {
          "id": 924506,
          "postDate": "2020-07-11T13:24:50.597Z",
          "content": "<p>qwk on only that set of data... one gives -0.22 and one -0.44 ,so u see extent of mispredictions for that data set</p>",
          "rawMarkdown": "qwk on only that set of data... one gives -0.22 and one -0.44 ,so u see extent of mispredictions for that data set",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 925191,
      "author_name": "Shujun",
      "author_url": "",
      "post_date": "2020-07-11T22:03:57.917000",
      "content": "<p>I found that in my best lb run rn (0.91), the accuracy for class 2,3,4,5 are only around 0.5. Radboud's labeled by students so the provided labels themselves only have 0.7 accuracy, so it's going to come down to whoever can handle the noisy labels in the best way possible</p>",
      "votes": 2,
      "replies": [
        {
          "id": 929531,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-07-14T18:21:27.050000",
          "content": "<p>Ka are also noisy, its difficult to relabel them up .</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 924350,
      "author_name": "Eek The Cat",
      "author_url": "",
      "post_date": "2020-07-11T11:43:32.467000",
      "content": "<p>I agree. Tried a couple of tricks to address that issue, including online relabelling, but all that just slightly degrades local score. If you don't mind me asking, what certain set of slides you're referring to? All of my efforts were based mostly on online loss values.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 924354,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-07-11T11:49:18.817000",
          "content": "<p>i can just tell you that they are less than 500-600 . If you exclude then only your local cv qwk gets sky rockted  not the lbs  as LB comprises of those same type of images that you exclude numbering same as in Train set not sure how many such would be there in Private set ,if they are in majority then scores are going to nose dive .</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 924396,
          "author_name": "Eek The Cat",
          "author_url": "",
          "post_date": "2020-07-11T12:08:25.773000",
          "content": "<p>Well, if you're talking about labelling errors, organizers promised private test would be free of them. If you're referring to some particular subclass of slides, which are hard for model to classify correctly, then yes, shakeup will be funny :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 924403,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-07-11T12:10:10.480000",
          "content": "<p>Test set will have correct labels but Train set for those supposedly hard slides are not surely . If I place them completely in Validation set then i cant get QWK  more than 75 -77. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 924425,
          "author_name": "Eek The Cat",
          "author_url": "",
          "post_date": "2020-07-11T12:19:31.863000",
          "content": "<p>That's interesting. Slides in question are mostly from radboud, right?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 924483,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-07-11T13:05:04.957000",
          "content": "<p>yup but interestingly the more problamatic ones are karloniska in those</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 924503,
          "author_name": "Eek The Cat",
          "author_url": "",
          "post_date": "2020-07-11T13:21:29.137000",
          "content": "<p>How do you define most problematic? Largest deviation of predicted class from true?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 924506,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-07-11T13:24:50.597000",
          "content": "<p>qwk on only that set of data... one gives -0.22 and one -0.44 ,so u see extent of mispredictions for that data set</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "923678": "Upon closely analysing the data of two source providers  RA and KA ,i am getting a strong indication that there are certain set of slide from both data providers which have incorrect labels  . If we exclude  that  data cv qk shoots up to 94-95  but it dsnt materializes same on the Leader board which means Test set does have those problem making slides type contained in  it so your trained model on data excluding those WSI  predicts incorrect labels for those type of slides in test set so your score stucks at 89-90 . Any one who has well addressed those incorrect labels will be in top 3 .\n\nLet me know your opinion. \n\n \n",
    "925191": "I found that in my best lb run rn (0.91), the accuracy for class 2,3,4,5 are only around 0.5. Radboud's labeled by students so the provided labels themselves only have 0.7 accuracy, so it's going to come down to whoever can handle the noisy labels in the best way possible",
    "924350": "I agree. Tried a couple of tricks to address that issue, including online relabelling, but all that just slightly degrades local score. If you don't mind me asking, what certain set of slides you're referring to? All of my efforts were based mostly on online loss values."
  }
}