{
  "id": 182713,
  "title": "Submission label",
  "url": "/competitions/rsna-str-pulmonary-embolism-detection/discussion/182713",
  "author_name": "manpreet singh",
  "post_date": "2020-09-14T04:04:29.715000",
  "votes": -1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>We have total 17 labels in train csv and only one label in submission format. Can anyone clarify how to map those 17 (or less only exam level) into one label to be submitted.<br>\nThanks.</p>",
  "messages": [
    {
      "id": 1009516,
      "postDate": "2020-09-14T04:04:29.717Z",
      "content": "<p>We have total 17 labels in train csv and only one label in submission format. Can anyone clarify how to map those 17 (or less only exam level) into one label to be submitted.<br>\nThanks.</p>",
      "rawMarkdown": "We have total 17 labels in train csv and only one label in submission format. Can anyone clarify how to map those 17 (or less only exam level) into one label to be submitted.\nThanks.",
      "votes": -1
    },
    {
      "id": 1011385,
      "postDate": "2020-09-15T12:30:49.560Z",
      "content": "<p>Hi all! The submission format includes both study- and image-level labels. For labels you are predicting whether PE is present on the image. For studies you are predicting the 9 labels present in the table on the Evaluation page.</p>\n<p>Each study-level label is predicted as its own row - <code>&lt;study id&gt;_&lt;label name&gt;,&lt;prediction&gt;</code>, one for each study id / label combination. Each image-level label is also predicted as its own row - <code>&lt;image id&gt;,&lt;prediction&gt;</code>.</p>\n<p>All labels (both exam- and image-level) are scored, and then that score is averaged (and divided by the sum of the weights). The end result is that image-level labels DO get a large combined weight because there are more of them, but we've balanced the study-level weights against that so that their labels are ultimately more important, as they are clinically more important.</p>",
      "rawMarkdown": "Hi all! The submission format includes both study- and image-level labels. For labels you are predicting whether PE is present on the image. For studies you are predicting the 9 labels present in the table on the Evaluation page.\n\nEach study-level label is predicted as its own row - `<study id>_<label name>,<prediction>`, one for each study id / label combination. Each image-level label is also predicted as its own row - `<image id>,<prediction>`.\n\nAll labels (both exam- and image-level) are scored, and then that score is averaged (and divided by the sum of the weights). The end result is that image-level labels DO get a large combined weight because there are more of them, but we've balanced the study-level weights against that so that their labels are ultimately more important, as they are clinically more important."
    },
    {
      "id": 1009571,
      "postDate": "2020-09-14T05:31:05.203Z",
      "content": "<p>Here's what I know:</p>\n<p>In the submission, the main target is the <code>negative_exam_for_pe</code> label. If you check the size of the <code>sample_submission.csv</code> you would notice that the size is <code>(152703, 2)</code> which, compared to <code>test</code> is <strong>650</strong> lines more.</p>\n<p>This 650 lines comes from the number of unique <code>test.StudyInstanceUID</code>. In short, you need to predict all <code>test</code> for their <code>negative_exam_for_pe</code> and predict all <code>test.StudyInstanceUID</code> instances (which contains several DICOM images) for all the labels. These labels' <code>id</code> is formatted like:</p>\n<pre><code>\"%s_%s\" %(StudyInstanceUID, label_name)\n</code></pre>\n<p>I have <a href=\"https://www.kaggle.com/seraphwedd18/pe-detection-with-keras-model-creation?scriptVersionId=42618728\" target=\"_blank\">here</a> an unfinished notebook (due to some error) where I showed how I created an ensemble for the submission file.</p>\n<p>Hope this helps.</p>",
      "rawMarkdown": "Here's what I know:\n\nIn the submission, the main target is the `negative_exam_for_pe` label. If you check the size of the `sample_submission.csv` you would notice that the size is `(152703, 2)` which, compared to `test` is **650** lines more.\n\nThis 650 lines comes from the number of unique `test.StudyInstanceUID`. In short, you need to predict all `test` for their `negative_exam_for_pe` and predict all `test.StudyInstanceUID` instances (which contains several DICOM images) for all the labels. These labels' `id` is formatted like:\n```\n\"%s_%s\" %(StudyInstanceUID, label_name)\n```\n\nI have [here](https://www.kaggle.com/seraphwedd18/pe-detection-with-keras-model-creation?scriptVersionId=42618728) an unfinished notebook (due to some error) where I showed how I created an ensemble for the submission file.\n\nHope this helps.",
      "replies": [
        {
          "id": 1009595,
          "postDate": "2020-09-14T05:53:03.677Z",
          "content": "<p>thanks for the brief.</p>",
          "rawMarkdown": "thanks for the brief.",
          "replies": [
            {
              "id": 1009884,
              "postDate": "2020-09-14T10:47:17.640Z",
              "content": "<p>Evaluation rules state:</p>\n<p>The total loss is the average of all image- and exam-level loss, divided by the sum of the weights.</p>\n<p>So, I think that all the images combine to one score (weight 0.0736196319). So even though there are many images, the overall score per image doesn't outweigh the weighted scores for the study level targets.</p>\n<p>EDIT: Rereading the evaluation metric, I am not sure if Kaggle determines a score for each possible category (rightsided_pe, leftsided_pe, pe_present_on_image, etc) and then takes a weighted average of them, or if \"pe_present_on_image\" gets more comparative weight because there are so many more images than study/label combinations.</p>",
              "rawMarkdown": "Evaluation rules state:\n\nThe total loss is the average of all image- and exam-level loss, divided by the sum of the weights.\n\nSo, I think that all the images combine to one score (weight 0.0736196319). So even though there are many images, the overall score per image doesn't outweigh the weighted scores for the study level targets.\n\nEDIT: Rereading the evaluation metric, I am not sure if Kaggle determines a score for each possible category (rightsided_pe, leftsided_pe, pe_present_on_image, etc) and then takes a weighted average of them, or if \"pe_present_on_image\" gets more comparative weight because there are so many more images than study/label combinations."
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1011385,
      "author_name": "Phil Culliton",
      "author_url": "",
      "post_date": "2020-09-15T12:30:49.560000",
      "content": "<p>Hi all! The submission format includes both study- and image-level labels. For labels you are predicting whether PE is present on the image. For studies you are predicting the 9 labels present in the table on the Evaluation page.</p>\n<p>Each study-level label is predicted as its own row - <code>&lt;study id&gt;_&lt;label name&gt;,&lt;prediction&gt;</code>, one for each study id / label combination. Each image-level label is also predicted as its own row - <code>&lt;image id&gt;,&lt;prediction&gt;</code>.</p>\n<p>All labels (both exam- and image-level) are scored, and then that score is averaged (and divided by the sum of the weights). The end result is that image-level labels DO get a large combined weight because there are more of them, but we've balanced the study-level weights against that so that their labels are ultimately more important, as they are clinically more important.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1009571,
      "author_name": "Venturillo JE",
      "author_url": "",
      "post_date": "2020-09-14T05:31:05.203000",
      "content": "<p>Here's what I know:</p>\n<p>In the submission, the main target is the <code>negative_exam_for_pe</code> label. If you check the size of the <code>sample_submission.csv</code> you would notice that the size is <code>(152703, 2)</code> which, compared to <code>test</code> is <strong>650</strong> lines more.</p>\n<p>This 650 lines comes from the number of unique <code>test.StudyInstanceUID</code>. In short, you need to predict all <code>test</code> for their <code>negative_exam_for_pe</code> and predict all <code>test.StudyInstanceUID</code> instances (which contains several DICOM images) for all the labels. These labels' <code>id</code> is formatted like:</p>\n<pre><code>\"%s_%s\" %(StudyInstanceUID, label_name)\n</code></pre>\n<p>I have <a href=\"https://www.kaggle.com/seraphwedd18/pe-detection-with-keras-model-creation?scriptVersionId=42618728\" target=\"_blank\">here</a> an unfinished notebook (due to some error) where I showed how I created an ensemble for the submission file.</p>\n<p>Hope this helps.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1009595,
          "author_name": "manpreet singh",
          "author_url": "",
          "post_date": "2020-09-14T05:53:03.677000",
          "content": "<p>thanks for the brief.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 1009884,
              "author_name": "quadcore/Richard Epstein",
              "author_url": "",
              "post_date": "2020-09-14T10:47:17.640000",
              "content": "<p>Evaluation rules state:</p>\n<p>The total loss is the average of all image- and exam-level loss, divided by the sum of the weights.</p>\n<p>So, I think that all the images combine to one score (weight 0.0736196319). So even though there are many images, the overall score per image doesn't outweigh the weighted scores for the study level targets.</p>\n<p>EDIT: Rereading the evaluation metric, I am not sure if Kaggle determines a score for each possible category (rightsided_pe, leftsided_pe, pe_present_on_image, etc) and then takes a weighted average of them, or if \"pe_present_on_image\" gets more comparative weight because there are so many more images than study/label combinations.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1009516": "We have total 17 labels in train csv and only one label in submission format. Can anyone clarify how to map those 17 (or less only exam level) into one label to be submitted.\nThanks.",
    "1011385": "Hi all! The submission format includes both study- and image-level labels. For labels you are predicting whether PE is present on the image. For studies you are predicting the 9 labels present in the table on the Evaluation page.\n\nEach study-level label is predicted as its own row - `<study id>_<label name>,<prediction>`, one for each study id / label combination. Each image-level label is also predicted as its own row - `<image id>,<prediction>`.\n\nAll labels (both exam- and image-level) are scored, and then that score is averaged (and divided by the sum of the weights). The end result is that image-level labels DO get a large combined weight because there are more of them, but we've balanced the study-level weights against that so that their labels are ultimately more important, as they are clinically more important.",
    "1009571": "Here's what I know:\n\nIn the submission, the main target is the `negative_exam_for_pe` label. If you check the size of the `sample_submission.csv` you would notice that the size is `(152703, 2)` which, compared to `test` is **650** lines more.\n\nThis 650 lines comes from the number of unique `test.StudyInstanceUID`. In short, you need to predict all `test` for their `negative_exam_for_pe` and predict all `test.StudyInstanceUID` instances (which contains several DICOM images) for all the labels. These labels' `id` is formatted like:\n```\n\"%s_%s\" %(StudyInstanceUID, label_name)\n```\n\nI have [here](https://www.kaggle.com/seraphwedd18/pe-detection-with-keras-model-creation?scriptVersionId=42618728) an unfinished notebook (due to some error) where I showed how I created an ensemble for the submission file.\n\nHope this helps."
  }
}