{
  "competition": "rsna-knee-abnormality-detection",
  "topic_id": "738851",
  "comments": [
    {
      "id": 3519642,
      "authorName": "PC Jimmmy",
      "votes": 2,
      "postDate": "2026-09-02T01:21:17.720000",
      "content": "<p>You identified the problem.  You need to build a model that uses the 12 categories as labels based on examination of only the dicom images.</p>\n<p>To get the training data you have to convert the reports to labels between 0.0 and 1.0 for each of the 12 classes.  Not an easy task.  The best set of labels probably wins the competition.</p>\n<p>You want to learn how to make labels - dig in.   </p>\n<p>You only want to learn the prediction model than a number of folks have shared a set of pretty decent labels.   </p>\n<p>Learning to make labels is a skill set you need in the real world.   Not one that LLM's seem to have mastered yet.   </p>"
    },
    {
      "id": 3519648,
      "authorName": "Ahmed Ali",
      "votes": 0,
      "postDate": "2026-09-02T01:29:13.503000",
      "content": "<p>Thanks for the clarification, I didn't realize that most of the training data isn't labeled. So if I understand correctly, I use the labeled training data to train a model on the reports to create labels for the rest of the data, then use those labels to train a different model on the images? Interesting.</p>"
    },
    {
      "id": 3519650,
      "authorName": "PC Jimmmy",
      "votes": 1,
      "postDate": "2026-09-02T01:33:54.337000",
      "content": "<p>You got it!</p>\n<p>Yes - only 58 of the images are labeled out of the 4700.   </p>\n<p>I been here for almost 9 years - have not participated in every competition but I think this is the first were we were not handed the labels.  Back in the day when I was working for real I had to do a lot of label creation - sadly I sucked at it.   You can find the shared labels searching in the discussion posts.</p>"
    },
    {
      "id": 3519728,
      "authorName": "Komil Parmar",
      "votes": 1,
      "postDate": "2026-09-02T05:22:08.457000",
      "content": "<p><a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a> nicely shared. I started to read works on Pseudo Labelling which is probably the single best investment anyone can make in this competition. And I was stunned by how many papers are there just regarding pseudoabelling. I also suggested the paperswithcode site to add that as a separate category and <strong>they added it</strong>. 🤯</p>\n<p>The problem though which one will realize while applying those methods is that all of those papers are regarding pseudo labelling on unlabeled data. That isn't the situation here. We already have near perfect labels for those samples achieved using llms. And if you replace them or even update them via pseudoabelling, the model will perform worse then simply training on LLM labels.</p>\n<p>So we have to learn how to update and when to update those pseudo labels. That's the key. 🔑</p>"
    }
  ],
  "messages": [],
  "raw_show": {
    "topic": {
      "id": 738851,
      "title": "No Reports for Test Data?",
      "authorName": "Ahmed Ali",
      "commentCount": 4,
      "votes": 0,
      "postDate": "2026-09-01T23:16:29.959000"
    },
    "comments": [
      {
        "id": 3519642,
        "authorName": "PC Jimmmy",
        "votes": 2,
        "postDate": "2026-09-02T01:21:17.720000",
        "content": "<p>You identified the problem.  You need to build a model that uses the 12 categories as labels based on examination of only the dicom images.</p>\n<p>To get the training data you have to convert the reports to labels between 0.0 and 1.0 for each of the 12 classes.  Not an easy task.  The best set of labels probably wins the competition.</p>\n<p>You want to learn how to make labels - dig in.   </p>\n<p>You only want to learn the prediction model than a number of folks have shared a set of pretty decent labels.   </p>\n<p>Learning to make labels is a skill set you need in the real world.   Not one that LLM's seem to have mastered yet.   </p>"
      },
      {
        "id": 3519648,
        "authorName": "Ahmed Ali",
        "votes": 0,
        "postDate": "2026-09-02T01:29:13.503000",
        "content": "<p>Thanks for the clarification, I didn't realize that most of the training data isn't labeled. So if I understand correctly, I use the labeled training data to train a model on the reports to create labels for the rest of the data, then use those labels to train a different model on the images? Interesting.</p>"
      },
      {
        "id": 3519650,
        "authorName": "PC Jimmmy",
        "votes": 1,
        "postDate": "2026-09-02T01:33:54.337000",
        "content": "<p>You got it!</p>\n<p>Yes - only 58 of the images are labeled out of the 4700.   </p>\n<p>I been here for almost 9 years - have not participated in every competition but I think this is the first were we were not handed the labels.  Back in the day when I was working for real I had to do a lot of label creation - sadly I sucked at it.   You can find the shared labels searching in the discussion posts.</p>"
      },
      {
        "id": 3519728,
        "authorName": "Komil Parmar",
        "votes": 1,
        "postDate": "2026-09-02T05:22:08.457000",
        "content": "<p><a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a> nicely shared. I started to read works on Pseudo Labelling which is probably the single best investment anyone can make in this competition. And I was stunned by how many papers are there just regarding pseudoabelling. I also suggested the paperswithcode site to add that as a separate category and <strong>they added it</strong>. 🤯</p>\n<p>The problem though which one will realize while applying those methods is that all of those papers are regarding pseudo labelling on unlabeled data. That isn't the situation here. We already have near perfect labels for those samples achieved using llms. And if you replace them or even update them via pseudoabelling, the model will perform worse then simply training on LLM labels.</p>\n<p>So we have to learn how to update and when to update those pseudo labels. That's the key. 🔑</p>"
      }
    ]
  },
  "topic": {
    "id": 738851,
    "title": "No Reports for Test Data?",
    "authorName": "Ahmed Ali",
    "commentCount": 4,
    "votes": 0,
    "postDate": "2026-09-01T23:16:29.959000"
  },
  "index": {
    "id": "738851",
    "title": "No Reports for Test Data?",
    "authorName": "",
    "commentCount": "4",
    "votes": "0",
    "postDate": "2026-09-01 23:16:29.959000"
  }
}