{
  "id": 110928,
  "title": "Be careful with postprocessing",
  "url": "/competitions/rsna-intracranial-hemorrhage-detection/discussion/110928",
  "author_name": "Peter",
  "post_date": "2019-10-02T09:36:36.134000",
  "votes": 16,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I've implemented an effortless postprocessing technique to correct labels.\n- Group the samples by <code>Study Instance UID.</code>\n- Order the slices by <code>ImagePositionPatient[2]</code>\n- Use the series of labels (0,0,0,1,1,1,<strong>0</strong>,1,1,0,0,0) to correct our predictions.</p>\n\n<p>I am not sure it would be useful. Fortunately, I read the comment below before I started experimenting.</p>\n\n<p><strong>We cannot use this kind of corrections!</strong></p>\n\n<p>Comment from <a href=\"/juliaelliott\">@juliaelliott</a>:\n&gt; <strong>metadata can be used for pre-processing of images, but they cannot be features in your model or used to change or label predictions post-modeling</strong></p>\n\n<p><a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109969#637525\">https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109969#637525</a></p>\n\n<p><strong>Update</strong>\n<a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/111325\">https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/111325</a></p>",
  "messages": [
    {
      "id": 638692,
      "postDate": "2019-10-02T09:36:36.133Z",
      "content": "<p>I've implemented an effortless postprocessing technique to correct labels.\n- Group the samples by <code>Study Instance UID.</code>\n- Order the slices by <code>ImagePositionPatient[2]</code>\n- Use the series of labels (0,0,0,1,1,1,<strong>0</strong>,1,1,0,0,0) to correct our predictions.</p>\n\n<p>I am not sure it would be useful. Fortunately, I read the comment below before I started experimenting.</p>\n\n<p><strong>We cannot use this kind of corrections!</strong></p>\n\n<p>Comment from <a href=\"/juliaelliott\">@juliaelliott</a>:\n&gt; <strong>metadata can be used for pre-processing of images, but they cannot be features in your model or used to change or label predictions post-modeling</strong></p>\n\n<p><a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109969#637525\">https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109969#637525</a></p>\n\n<p><strong>Update</strong>\n<a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/111325\">https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/111325</a></p>",
      "rawMarkdown": "I've implemented an effortless postprocessing technique to correct labels.\n- Group the samples by `Study Instance UID.`\n- Order the slices by `ImagePositionPatient[2]`\n- Use the series of labels (0,0,0,1,1,1,**0**,1,1,0,0,0) to correct our predictions.\n\n\nI am not sure it would be useful. Fortunately, I read the comment below before I started experimenting.\n\n**We cannot use this kind of corrections!**\n\nComment from @juliaelliott:\n&gt; **metadata can be used for pre-processing of images, but they cannot be features in your model or used to change or label predictions post-modeling**\n\nhttps://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109969#637525\n\n**Update**\nhttps://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/111325",
      "votes": 16
    },
    {
      "id": 638902,
      "postDate": "2019-10-02T14:36:20.560Z",
      "content": "<p>We definitely need a strong clarification. </p>\n\n<p>Can we use let's say simple rnn/cnn/fc model (or even layer) to post-process f.e this one mentioned (0,0,0,1,1,1,0,1,1,0,0,0)? For the post-processing we don't need metadata because we always can create our input batches as 3d CT representation. So the idea is that we always can pass information from pre-processing to post-processing. Clarify smbd pls</p>",
      "rawMarkdown": "We definitely need a strong clarification. \n\nCan we use let's say simple rnn/cnn/fc model (or even layer) to post-process f.e this one mentioned (0,0,0,1,1,1,0,1,1,0,0,0)? For the post-processing we don't need metadata because we always can create our input batches as 3d CT representation. So the idea is that we always can pass information from pre-processing to post-processing. Clarify smbd pls\n",
      "votes": 4
    },
    {
      "id": 639929,
      "postDate": "2019-10-03T17:45:05.340Z",
      "content": "<p>Hi Peter!</p>\n\n<p>Would you be so kind please and elaborate on how you are using this to improve predictions? What is the sequence of ones and zero that you share?</p>\n\n<p>Appreciate whatever you can share on this,</p>\n\n<p>Thank you,\nRadek</p>",
      "rawMarkdown": "Hi Peter!\n\nWould you be so kind please and elaborate on how you are using this to improve predictions? What is the sequence of ones and zero that you share?\n\nAppreciate whatever you can share on this,\n\nThank you,\nRadek",
      "votes": 1,
      "replies": [
        {
          "id": 640067,
          "postDate": "2019-10-03T19:42:26.043Z",
          "content": "<p>Hey <a href=\"/radek1\">@radek1</a> ,</p>\n\n<p>You can re-create the 3d CT scans/slices by using <code>Study Instance UID.</code> Once you have them, you can put them in the right order by using the <code>ImagePositionPatient[2]</code> metadata (the value for the z-axis). Since these hemorrhages usually appear on more than one CT slice, my idea was I use that information for correcting my predictions.</p>\n\n<p>For example, if the predictions for a patient (approx 25-30 slices (images) in train or test set) are {0, 0, 0, 0, 1, 1, 0, 1, 1, 0, 0, 0, 0, 0} it is almost certain that you could correct the '0' in the middle to '1', or use different thresholds at the sides, etc.</p>\n\n<p>Unfortunately, we cannot use any of the metadata for postprocessing.</p>",
          "rawMarkdown": "Hey @radek1 ,\n\nYou can re-create the 3d CT scans/slices by using `Study Instance UID.` Once you have them, you can put them in the right order by using the `ImagePositionPatient[2]` metadata (the value for the z-axis). Since these hemorrhages usually appear on more than one CT slice, my idea was I use that information for correcting my predictions.\n\nFor example, if the predictions for a patient (approx 25-30 slices (images) in train or test set) are {0, 0, 0, 0, 1, 1, 0, 1, 1, 0, 0, 0, 0, 0} it is almost certain that you could correct the '0' in the middle to '1', or use different thresholds at the sides, etc.\n\nUnfortunately, we cannot use any of the metadata for postprocessing.",
          "votes": 3
        },
        {
          "id": 640100,
          "postDate": "2019-10-03T20:32:56.183Z",
          "content": "<p>Ah I get it, thank you very much Peter for this additional explanation. Much appreciated!</p>",
          "rawMarkdown": "Ah I get it, thank you very much Peter for this additional explanation. Much appreciated!"
        },
        {
          "id": 649414,
          "postDate": "2019-10-15T10:08:50.797Z",
          "content": "<p>it would seem that you can actually do this now though, right:\n<a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/111325#latest-644288\">https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/111325#latest-644288</a></p>",
          "rawMarkdown": "it would seem that you can actually do this now though, right:\nhttps://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/111325#latest-644288"
        },
        {
          "id": 649430,
          "postDate": "2019-10-15T10:38:56.277Z",
          "content": "<p>Yes, I've already updated the post with the link. Thanks.</p>",
          "rawMarkdown": "Yes, I've already updated the post with the link. Thanks."
        }
      ]
    },
    {
      "id": 638904,
      "postDate": "2019-10-02T14:40:42.963Z",
      "content": "<p>Thanks for your finding!\nReally appreciate the spirit of sharing in Kaggle</p>",
      "rawMarkdown": "Thanks for your finding!\nReally appreciate the spirit of sharing in Kaggle"
    }
  ],
  "comments": [
    {
      "id": 638902,
      "author_name": "Oleg Yaroshevskiy",
      "author_url": "",
      "post_date": "2019-10-02T14:36:20.560000",
      "content": "<p>We definitely need a strong clarification. </p>\n\n<p>Can we use let's say simple rnn/cnn/fc model (or even layer) to post-process f.e this one mentioned (0,0,0,1,1,1,0,1,1,0,0,0)? For the post-processing we don't need metadata because we always can create our input batches as 3d CT representation. So the idea is that we always can pass information from pre-processing to post-processing. Clarify smbd pls</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 639929,
      "author_name": "Radek Osmulski",
      "author_url": "",
      "post_date": "2019-10-03T17:45:05.340000",
      "content": "<p>Hi Peter!</p>\n\n<p>Would you be so kind please and elaborate on how you are using this to improve predictions? What is the sequence of ones and zero that you share?</p>\n\n<p>Appreciate whatever you can share on this,</p>\n\n<p>Thank you,\nRadek</p>",
      "votes": 1,
      "replies": [
        {
          "id": 640067,
          "author_name": "Peter",
          "author_url": "",
          "post_date": "2019-10-03T19:42:26.043000",
          "content": "<p>Hey <a href=\"/radek1\">@radek1</a> ,</p>\n\n<p>You can re-create the 3d CT scans/slices by using <code>Study Instance UID.</code> Once you have them, you can put them in the right order by using the <code>ImagePositionPatient[2]</code> metadata (the value for the z-axis). Since these hemorrhages usually appear on more than one CT slice, my idea was I use that information for correcting my predictions.</p>\n\n<p>For example, if the predictions for a patient (approx 25-30 slices (images) in train or test set) are {0, 0, 0, 0, 1, 1, 0, 1, 1, 0, 0, 0, 0, 0} it is almost certain that you could correct the '0' in the middle to '1', or use different thresholds at the sides, etc.</p>\n\n<p>Unfortunately, we cannot use any of the metadata for postprocessing.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 640100,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2019-10-03T20:32:56.183000",
          "content": "<p>Ah I get it, thank you very much Peter for this additional explanation. Much appreciated!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 649414,
          "author_name": "Daithí O'Manacháin",
          "author_url": "",
          "post_date": "2019-10-15T10:08:50.797000",
          "content": "<p>it would seem that you can actually do this now though, right:\n<a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/111325#latest-644288\">https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/111325#latest-644288</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 649430,
          "author_name": "Peter",
          "author_url": "",
          "post_date": "2019-10-15T10:38:56.277000",
          "content": "<p>Yes, I've already updated the post with the link. Thanks.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 638904,
      "author_name": "Salaryman",
      "author_url": "",
      "post_date": "2019-10-02T14:40:42.963000",
      "content": "<p>Thanks for your finding!\nReally appreciate the spirit of sharing in Kaggle</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "638692": "I've implemented an effortless postprocessing technique to correct labels.\n- Group the samples by `Study Instance UID.`\n- Order the slices by `ImagePositionPatient[2]`\n- Use the series of labels (0,0,0,1,1,1,**0**,1,1,0,0,0) to correct our predictions.\n\n\nI am not sure it would be useful. Fortunately, I read the comment below before I started experimenting.\n\n**We cannot use this kind of corrections!**\n\nComment from @juliaelliott:\n&gt; **metadata can be used for pre-processing of images, but they cannot be features in your model or used to change or label predictions post-modeling**\n\nhttps://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109969#637525\n\n**Update**\nhttps://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/111325",
    "638902": "We definitely need a strong clarification. \n\nCan we use let's say simple rnn/cnn/fc model (or even layer) to post-process f.e this one mentioned (0,0,0,1,1,1,0,1,1,0,0,0)? For the post-processing we don't need metadata because we always can create our input batches as 3d CT representation. So the idea is that we always can pass information from pre-processing to post-processing. Clarify smbd pls\n",
    "639929": "Hi Peter!\n\nWould you be so kind please and elaborate on how you are using this to improve predictions? What is the sequence of ones and zero that you share?\n\nAppreciate whatever you can share on this,\n\nThank you,\nRadek",
    "638904": "Thanks for your finding!\nReally appreciate the spirit of sharing in Kaggle"
  }
}