{
  "id": 110599,
  "title": "Labelling Process",
  "url": "/competitions/rsna-intracranial-hemorrhage-detection/discussion/110599",
  "author_name": "Tom Aindow",
  "post_date": "2019-09-29T14:46:42.486000",
  "votes": 6,
  "comment_count": 10,
  "views": 0,
  "content": "<p>So we know that labelling is done at the slice level, but do we know if the experts would have labelled the data on a patient by patient basis for the competition?</p>\n\n<p>So for example, expert gets patient A and goes through labelling slice A1, A2, A3, .. AN?\nOr could it have been random with expert getting A5, D1, Z12, ... bla bla?</p>\n\n<p>First option seems more likely to me, in which case the order of slices at the patient level might be important. Would be interesting if Kaggle/organisers could tell us more about the process :)</p>",
  "messages": [
    {
      "id": 637520,
      "postDate": "2019-10-01T05:09:52.860Z",
      "content": "<p>The clarification that we can provide here (on behalf of the host) is that the annotators were presented with a single image at a time to label. However, they did have access to the entire stack of images of the same study as they were annotating. So images were not necessarily annotated in isolation, but it is feasible or even expected (as is it is clinically) that the adjacent slices contribute to the interpretation process.</p>",
      "rawMarkdown": "The clarification that we can provide here (on behalf of the host) is that the annotators were presented with a single image at a time to label. However, they did have access to the entire stack of images of the same study as they were annotating. So images were not necessarily annotated in isolation, but it is feasible or even expected (as is it is clinically) that the adjacent slices contribute to the interpretation process.",
      "votes": 5
    },
    {
      "id": 636445,
      "postDate": "2019-09-29T14:46:42.487Z",
      "content": "<p>So we know that labelling is done at the slice level, but do we know if the experts would have labelled the data on a patient by patient basis for the competition?</p>\n\n<p>So for example, expert gets patient A and goes through labelling slice A1, A2, A3, .. AN?\nOr could it have been random with expert getting A5, D1, Z12, ... bla bla?</p>\n\n<p>First option seems more likely to me, in which case the order of slices at the patient level might be important. Would be interesting if Kaggle/organisers could tell us more about the process :)</p>",
      "rawMarkdown": "So we know that labelling is done at the slice level, but do we know if the experts would have labelled the data on a patient by patient basis for the competition?\n\nSo for example, expert gets patient A and goes through labelling slice A1, A2, A3, .. AN?\nOr could it have been random with expert getting A5, D1, Z12, ... bla bla?\n\nFirst option seems more likely to me, in which case the order of slices at the patient level might be important. Would be interesting if Kaggle/organisers could tell us more about the process :)",
      "votes": 6
    },
    {
      "id": 636459,
      "postDate": "2019-09-29T15:06:04.987Z",
      "content": "<p>Dear <a href=\"/philculliton\">@philculliton</a>  and <a href=\"/juliaelliott\">@juliaelliott</a>  please check this post</p>",
      "rawMarkdown": "Dear @philculliton  and @juliaelliott  please check this post",
      "votes": 1
    },
    {
      "id": 1684537,
      "postDate": "2022-02-10T15:11:31.637Z",
      "content": "<p>Hi! I'm currently doing a research on different labelling tools for ML, so if anyone has an experience with using one and is willing to share it. Could you please fill this form <a href=\"https://forms.gle/KSEvduj155qrb8bw6\" target=\"_blank\">https://forms.gle/KSEvduj155qrb8bw6</a></p>",
      "rawMarkdown": "Hi! I'm currently doing a research on different labelling tools for ML, so if anyone has an experience with using one and is willing to share it. Could you please fill this form https://forms.gle/KSEvduj155qrb8bw6"
    },
    {
      "id": 637740,
      "postDate": "2019-10-01T08:19:00.363Z",
      "content": "<p>Thanks everyone for the input, that clears it up. </p>",
      "rawMarkdown": "Thanks everyone for the input, that clears it up. "
    },
    {
      "id": 636580,
      "postDate": "2019-09-29T20:34:21.743Z",
      "content": "<p>Fairly certain that cases are read at the patient level. It doesn't make sense for a radiologist to interpret standalone individual slices.</p>",
      "rawMarkdown": "Fairly certain that cases are read at the patient level. It doesn't make sense for a radiologist to interpret standalone individual slices."
    },
    {
      "id": 636514,
      "postDate": "2019-09-29T17:38:48.127Z",
      "content": "<p>Maybe you will have a slight difference where the correct labeling start and ends... but the question if this is a huge difference  ? Also for middle of slices if you have scans where labeling goes  000, 1, 1, 0, 1, 1, 000 you probably can correct it in post processing (convert middle 0 to 1)</p>",
      "rawMarkdown": "Maybe you will have a slight difference where the correct labeling start and ends... but the question if this is a huge difference  ? Also for middle of slices if you have scans where labeling goes  000, 1, 1, 0, 1, 1, 000 you probably can correct it in post processing (convert middle 0 to 1)",
      "replies": [
        {
          "id": 636518,
          "postDate": "2019-09-29T17:49:09.400Z",
          "content": "<p>Don't quite understand what you mean mate.  Can you explain? </p>\n\n<p>I am thinking that if the experts labelled slices by the patient, they might see information in previous slices that changes how they classify subsequent slices. Kinda like \"ah this can't be hemorrhage type X in this slice because I would have seen Y in the previous slice\". </p>\n\n<p>For example the following paper uses RNN at patient level after feature extraction even though it is labelled at slice level: <a href=\"https://rd.springer.com/content/pdf/10.1007%2Fs00330-019-06163-2.pdf\">https://rd.springer.com/content/pdf/10.1007%2Fs00330-019-06163-2.pdf</a></p>",
          "rawMarkdown": "Don't quite understand what you mean mate.  Can you explain? \n\nI am thinking that if the experts labelled slices by the patient, they might see information in previous slices that changes how they classify subsequent slices. Kinda like \"ah this can't be hemorrhage type X in this slice because I would have seen Y in the previous slice\". \n\nFor example the following paper uses RNN at patient level after feature extraction even though it is labelled at slice level: https://rd.springer.com/content/pdf/10.1007%2Fs00330-019-06163-2.pdf"
        },
        {
          "id": 636540,
          "postDate": "2019-09-29T18:40:42.920Z",
          "content": "<p>Please correct me if  I got something wrong =)</p>\n\n<p>So as I understand you can have two scenarios, (also for simplicity just have 2 labels (1, 0) and 10 slices, and only 1 place in the brain damaged) </p>\n\n<p>A) Where Doctor get random slices from difrent  patients\nB) Doctor get all slices from a single patient </p>\n\n<p>True Labels (3D  from top of the head to down): \n<code>0, 0, 0, 1, 1, 1, 1, 1, 0 0</code></p>\n\n<p>If we have scenario B , as you said if Doctor will see something he/she can go back and refine labeling, which will create continues label stretches (ideal scenario)</p>\n\n<p>if we have scenario A, each doctor will label differently e.g  <code>0, 0, 0, 0, 1, 1, 0, 1, 0, 0</code> You might have diffrent start and end position since you cant go back and correct, also in some cases you will not have continues labels.</p>\n\n<p>I dont know if kaggle will tell us how data was labelled, but it will be important to have good postprocessing script that will take advantage of such a sequential information.... </p>",
          "rawMarkdown": "Please correct me if  I got something wrong =)\n\n\nSo as I understand you can have two scenarios, (also for simplicity just have 2 labels (1, 0) and 10 slices, and only 1 place in the brain damaged) \n\nA) Where Doctor get random slices from difrent  patients\nB) Doctor get all slices from a single patient \n\nTrue Labels (3D  from top of the head to down): \n`0, 0, 0, 1, 1, 1, 1, 1, 0 0 `\n\nIf we have scenario B , as you said if Doctor will see something he/she can go back and refine labeling, which will create continues label stretches (ideal scenario)\n\nif we have scenario A, each doctor will label differently e.g  `0, 0, 0, 0, 1, 1, 0, 1, 0, 0` You might have diffrent start and end position since you cant go back and correct, also in some cases you will not have continues labels.\n\nI dont know if kaggle will tell us how data was labelled, but it will be important to have good postprocessing script that will take advantage of such a sequential information.... ",
          "votes": 3
        },
        {
          "id": 638685,
          "postDate": "2019-10-02T09:21:38.350Z",
          "content": "<p>I think we cannot implement this kind of postprocessing.</p>\n\n<p><a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109969637525\">https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109969637525</a></p>",
          "rawMarkdown": "I think we cannot implement this kind of postprocessing.\n\nhttps://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109969637525\n",
          "votes": 1
        },
        {
          "id": 638773,
          "postDate": "2019-10-02T12:04:25.947Z",
          "content": "<p>agree with <a href=\"/pestipeti\">@pestipeti</a> - looks like it is prohibited to use \"slices-based\" postprocessing for a single slice</p>",
          "rawMarkdown": "agree with @pestipeti - looks like it is prohibited to use \"slices-based\" postprocessing for a single slice",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 637520,
      "author_name": "Julia Elliott",
      "author_url": "",
      "post_date": "2019-10-01T05:09:52.860000",
      "content": "<p>The clarification that we can provide here (on behalf of the host) is that the annotators were presented with a single image at a time to label. However, they did have access to the entire stack of images of the same study as they were annotating. So images were not necessarily annotated in isolation, but it is feasible or even expected (as is it is clinically) that the adjacent slices contribute to the interpretation process.</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 636459,
      "author_name": "Mobassir",
      "author_url": "",
      "post_date": "2019-09-29T15:06:04.987000",
      "content": "<p>Dear <a href=\"/philculliton\">@philculliton</a>  and <a href=\"/juliaelliott\">@juliaelliott</a>  please check this post</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1684537,
      "author_name": "Anastasiia Moiseeva",
      "author_url": "",
      "post_date": "2022-02-10T15:11:31.637000",
      "content": "<p>Hi! I'm currently doing a research on different labelling tools for ML, so if anyone has an experience with using one and is willing to share it. Could you please fill this form <a href=\"https://forms.gle/KSEvduj155qrb8bw6\" target=\"_blank\">https://forms.gle/KSEvduj155qrb8bw6</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 637740,
      "author_name": "Tom Aindow",
      "author_url": "",
      "post_date": "2019-10-01T08:19:00.363000",
      "content": "<p>Thanks everyone for the input, that clears it up. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 636580,
      "author_name": "Ian Pan",
      "author_url": "",
      "post_date": "2019-09-29T20:34:21.743000",
      "content": "<p>Fairly certain that cases are read at the patient level. It doesn't make sense for a radiologist to interpret standalone individual slices.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 636514,
      "author_name": "DrHB",
      "author_url": "",
      "post_date": "2019-09-29T17:38:48.127000",
      "content": "<p>Maybe you will have a slight difference where the correct labeling start and ends... but the question if this is a huge difference  ? Also for middle of slices if you have scans where labeling goes  000, 1, 1, 0, 1, 1, 000 you probably can correct it in post processing (convert middle 0 to 1)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 636518,
          "author_name": "Tom Aindow",
          "author_url": "",
          "post_date": "2019-09-29T17:49:09.400000",
          "content": "<p>Don't quite understand what you mean mate.  Can you explain? </p>\n\n<p>I am thinking that if the experts labelled slices by the patient, they might see information in previous slices that changes how they classify subsequent slices. Kinda like \"ah this can't be hemorrhage type X in this slice because I would have seen Y in the previous slice\". </p>\n\n<p>For example the following paper uses RNN at patient level after feature extraction even though it is labelled at slice level: <a href=\"https://rd.springer.com/content/pdf/10.1007%2Fs00330-019-06163-2.pdf\">https://rd.springer.com/content/pdf/10.1007%2Fs00330-019-06163-2.pdf</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 636540,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-09-29T18:40:42.920000",
          "content": "<p>Please correct me if  I got something wrong =)</p>\n\n<p>So as I understand you can have two scenarios, (also for simplicity just have 2 labels (1, 0) and 10 slices, and only 1 place in the brain damaged) </p>\n\n<p>A) Where Doctor get random slices from difrent  patients\nB) Doctor get all slices from a single patient </p>\n\n<p>True Labels (3D  from top of the head to down): \n<code>0, 0, 0, 1, 1, 1, 1, 1, 0 0</code></p>\n\n<p>If we have scenario B , as you said if Doctor will see something he/she can go back and refine labeling, which will create continues label stretches (ideal scenario)</p>\n\n<p>if we have scenario A, each doctor will label differently e.g  <code>0, 0, 0, 0, 1, 1, 0, 1, 0, 0</code> You might have diffrent start and end position since you cant go back and correct, also in some cases you will not have continues labels.</p>\n\n<p>I dont know if kaggle will tell us how data was labelled, but it will be important to have good postprocessing script that will take advantage of such a sequential information.... </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 638685,
          "author_name": "Peter",
          "author_url": "",
          "post_date": "2019-10-02T09:21:38.350000",
          "content": "<p>I think we cannot implement this kind of postprocessing.</p>\n\n<p><a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109969637525\">https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109969637525</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 638773,
          "author_name": "Oleg Yaroshevskiy",
          "author_url": "",
          "post_date": "2019-10-02T12:04:25.947000",
          "content": "<p>agree with <a href=\"/pestipeti\">@pestipeti</a> - looks like it is prohibited to use \"slices-based\" postprocessing for a single slice</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "637520": "The clarification that we can provide here (on behalf of the host) is that the annotators were presented with a single image at a time to label. However, they did have access to the entire stack of images of the same study as they were annotating. So images were not necessarily annotated in isolation, but it is feasible or even expected (as is it is clinically) that the adjacent slices contribute to the interpretation process.",
    "636445": "So we know that labelling is done at the slice level, but do we know if the experts would have labelled the data on a patient by patient basis for the competition?\n\nSo for example, expert gets patient A and goes through labelling slice A1, A2, A3, .. AN?\nOr could it have been random with expert getting A5, D1, Z12, ... bla bla?\n\nFirst option seems more likely to me, in which case the order of slices at the patient level might be important. Would be interesting if Kaggle/organisers could tell us more about the process :)",
    "636459": "Dear @philculliton  and @juliaelliott  please check this post",
    "1684537": "Hi! I'm currently doing a research on different labelling tools for ML, so if anyone has an experience with using one and is willing to share it. Could you please fill this form https://forms.gle/KSEvduj155qrb8bw6",
    "637740": "Thanks everyone for the input, that clears it up. ",
    "636580": "Fairly certain that cases are read at the patient level. It doesn't make sense for a radiologist to interpret standalone individual slices.",
    "636514": "Maybe you will have a slight difference where the correct labeling start and ends... but the question if this is a huge difference  ? Also for middle of slices if you have scans where labeling goes  000, 1, 1, 0, 1, 1, 000 you probably can correct it in post processing (convert middle 0 to 1)"
  }
}