{
  "id": 109281,
  "title": "2 step pipeline: Predict anys then the hemorrhage type",
  "url": "/competitions/rsna-intracranial-hemorrhage-detection/discussion/109281",
  "author_name": "datasaurus",
  "post_date": "2019-09-18T06:47:39.974000",
  "votes": 32,
  "comment_count": 21,
  "views": 0,
  "content": "<p>Figure 1 in this paper shows a pipeline very similar to the objective of this competition:\n<a href=\"https://rd.springer.com/content/pdf/10.1007%2Fs00330-019-06163-2.pdf\">https://rd.springer.com/content/pdf/10.1007%2Fs00330-019-06163-2.pdf</a></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F421965%2F03bcc16baae739172b513d016e26181e%2Fpipeline.png?generation=1568789507192980&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 628950,
      "postDate": "2019-09-18T06:47:39.973Z",
      "content": "<p>Figure 1 in this paper shows a pipeline very similar to the objective of this competition:\n<a href=\"https://rd.springer.com/content/pdf/10.1007%2Fs00330-019-06163-2.pdf\">https://rd.springer.com/content/pdf/10.1007%2Fs00330-019-06163-2.pdf</a></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F421965%2F03bcc16baae739172b513d016e26181e%2Fpipeline.png?generation=1568789507192980&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Figure 1 in this paper shows a pipeline very similar to the objective of this competition:\nhttps://rd.springer.com/content/pdf/10.1007%2Fs00330-019-06163-2.pdf\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F421965%2F03bcc16baae739172b513d016e26181e%2Fpipeline.png?generation=1568789507192980&amp;alt=media)\n",
      "votes": 32
    },
    {
      "id": 639509,
      "postDate": "2019-10-03T09:42:16.030Z",
      "content": "<p>So, have anyone yet tried with the 2 step pipeline yet? I am referring to the 2 step pipeline for per image, not for per study.</p>\n\n<p>Also, how do you tackle the 2nd step for classifying hemorrhage types? What is the target probability values for the test images that you define not to have any hemorrhage (any=0 or less than a threshold)? Are the probabilities for the 5 sub-types 0.0 for these images? It would be helpful if someone can give me any advice. 😊 </p>",
      "rawMarkdown": "So, have anyone yet tried with the 2 step pipeline yet? I am referring to the 2 step pipeline for per image, not for per study.\n\nAlso, how do you tackle the 2nd step for classifying hemorrhage types? What is the target probability values for the test images that you define not to have any hemorrhage (any=0 or less than a threshold)? Are the probabilities for the 5 sub-types 0.0 for these images? It would be helpful if someone can give me any advice. 😊 ",
      "votes": 1,
      "replies": [
        {
          "id": 645105,
          "postDate": "2019-10-09T19:44:12.350Z",
          "content": "<p>I'm working on a 2-step pipeline and am thinking of setting a threshold. 0.5 for the threshold is a good choice for example, but the threshold can also be tuned based on validation data.</p>\n\n<p>So for example a phase pipeline would be:</p>\n\n<p>```</p>\n\n<h1>Initialize predictions</h1>\n\n<p>preds = []</p>\n\n<h1>Load models</h1>\n\n<p>bin_model = load_model('model_trained_on_any.h5')\ntype_model = load_model('model_trained_on_types.h5')</p>\n\n<h1>Get predictions</h1>\n\n<p>threshold = 0.5\nfor img in list(img_dir):\n            img = cv2.imread(img)\n            bin_pred = bin_model.predict(img)\n            if bin_pred &lt; threshold:\n                self.predictions.append([0,0,0,0,0,0])\n            else:\n                pred = [bin_pred] + type_model.predict(img)\n                self.predictions.append(pred)\n```</p>\n\n<p>Hope this helps! Curious to hear your thoughts on this.</p>",
          "rawMarkdown": "I'm working on a 2-step pipeline and am thinking of setting a threshold. 0.5 for the threshold is a good choice for example, but the threshold can also be tuned based on validation data.\n\nSo for example a phase pipeline would be:\n\n```\n# Initialize predictions\npreds = []\n# Load models\nbin_model = load_model('model_trained_on_any.h5')\ntype_model = load_model('model_trained_on_types.h5')\n# Get predictions\nthreshold = 0.5\nfor img in list(img_dir):\n            img = cv2.imread(img)\n            bin_pred = bin_model.predict(img)\n            if bin_pred &lt; threshold:\n                self.predictions.append([0,0,0,0,0,0])\n            else:\n                pred = [bin_pred] + type_model.predict(img)\n                self.predictions.append(pred)\n```\n\nHope this helps! Curious to hear your thoughts on this.",
          "votes": 2
        }
      ]
    },
    {
      "id": 629110,
      "postDate": "2019-09-18T11:14:31.023Z",
      "content": "<p>Very nice paper! Just beware that the dataset comes as 2D slices. You can aggregate them per study to form the 3D image but some patients may have part of their slices in the test set and part in the train set. This is only for stage 1, as stated in the Data section.</p>",
      "rawMarkdown": "Very nice paper! Just beware that the dataset comes as 2D slices. You can aggregate them per study to form the 3D image but some patients may have part of their slices in the test set and part in the train set. This is only for stage 1, as stated in the Data section.",
      "votes": 2,
      "replies": [
        {
          "id": 629246,
          "postDate": "2019-09-18T14:59:15.240Z",
          "content": "<p>Hello, <a href=\"/felipekitamura\">@felipekitamura</a>. I didn't understand why part of the slices are split between train and test set. I mean, the detection is per study/pacient, right? Don't we need all the slices of each study to make the inference? My first tought was to use 3D Conv Nets, which would need the whole 3D Image with the slices in the right order.</p>",
          "rawMarkdown": "Hello, @felipekitamura. I didn't understand why part of the slices are split between train and test set. I mean, the detection is per study/pacient, right? Don't we need all the slices of each study to make the inference? My first tought was to use 3D Conv Nets, which would need the whole 3D Image with the slices in the right order.",
          "votes": 1
        },
        {
          "id": 629299,
          "postDate": "2019-09-18T16:14:21.597Z",
          "content": "<p>Good point. I didn't make myself clear. All I'm saying is we have to look at the data to understand if slices from the same study were split into the train and test sets. Maybe this is not the case and we just have different studies from the same patient in the train and test sets.\nAlso, I understood that the labels are at the image-level and not study-level. Do you agree?</p>",
          "rawMarkdown": "Good point. I didn't make myself clear. All I'm saying is we have to look at the data to understand if slices from the same study were split into the train and test sets. Maybe this is not the case and we just have different studies from the same patient in the train and test sets.\nAlso, I understood that the labels are at the image-level and not study-level. Do you agree?",
          "votes": 1
        },
        {
          "id": 629395,
          "postDate": "2019-09-18T18:18:34.670Z",
          "content": "<p>I'll have a deeper look at the data soon, but I thought the labels would be at study level.</p>",
          "rawMarkdown": "I'll have a deeper look at the data soon, but I thought the labels would be at study level."
        },
        {
          "id": 629438,
          "postDate": "2019-09-18T19:19:08.667Z",
          "content": "<p>I think aggregating per study to recreate the 3-d image violate the rules limiting us to only use pixel data (\"Submission predictions must be based entirely on the pixel data in the provided datasets\").  Or am I misreading that requirement?</p>",
          "rawMarkdown": "I think aggregating per study to recreate the 3-d image violate the rules limiting us to only use pixel data (\"Submission predictions must be based entirely on the pixel data in the provided datasets\").  Or am I misreading that requirement?",
          "votes": 4
        },
        {
          "id": 629542,
          "postDate": "2019-09-18T22:35:36.053Z",
          "content": "<p>Good point. It seems they were talking about other stuff that could potentially cause issues though. But I agree the way this is written raises the question.</p>",
          "rawMarkdown": "Good point. It seems they were talking about other stuff that could potentially cause issues though. But I agree the way this is written raises the question."
        },
        {
          "id": 633012,
          "postDate": "2019-09-24T10:41:17.250Z",
          "content": "<p>Is this question already answered?</p>",
          "rawMarkdown": "Is this question already answered?"
        },
        {
          "id": 633154,
          "postDate": "2019-09-24T13:46:50.270Z",
          "content": "<p>Please correct me if I'm wrong, but if we have slices of each patient in both the training and testing set, doesn't this present a big data leak? We could transfer the label for a patient over to the test set and have correct predictions, right? </p>",
          "rawMarkdown": "Please correct me if I'm wrong, but if we have slices of each patient in both the training and testing set, doesn't this present a big data leak? We could transfer the label for a patient over to the test set and have correct predictions, right? ",
          "votes": 1
        },
        {
          "id": 633280,
          "postDate": "2019-09-24T17:03:19.910Z",
          "content": "<p><a href=\"/carlolepelaars\">@carlolepelaars</a> yes it is a big data leak, but it won't work for stage-2 so it's quite useless\n<a href=\"/jledoux\">@jledoux</a> Phil Cullington clarified this question here: <a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109479#latest-632646\">https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109479#latest-632646</a>\nYou can use data reconstruction, however I have not seen anywhere so far whether the stage-2 data will contain all the slices for each patients or if missing slices are to be expected\nAlso, I think all slices for a given patient do not have the same labeling, that is, if the hemorrhage is not visible on a given slice, it is marked as 0 (to confirm)</p>",
          "rawMarkdown": "@carlolepelaars yes it is a big data leak, but it won't work for stage-2 so it's quite useless\n@jledoux Phil Cullington clarified this question here: https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109479#latest-632646\nYou can use data reconstruction, however I have not seen anywhere so far whether the stage-2 data will contain all the slices for each patients or if missing slices are to be expected\nAlso, I think all slices for a given patient do not have the same labeling, that is, if the hemorrhage is not visible on a given slice, it is marked as 0 (to confirm)",
          "votes": 3
        },
        {
          "id": 633282,
          "postDate": "2019-09-24T17:08:32.177Z",
          "content": "<p>Hmm, interesting. Curious how people are going to exploit this for medals and if it will become a nuisance for serious competitors.</p>",
          "rawMarkdown": "Hmm, interesting. Curious how people are going to exploit this for medals and if it will become a nuisance for serious competitors.",
          "votes": 1
        },
        {
          "id": 635537,
          "postDate": "2019-09-27T19:16:15.790Z",
          "content": "<p>Hi all - thanks for all the questions!  Some answers below:</p>\n\n<p><a href=\"/felipekitamura\">@felipekitamura</a>:\n&gt; All I'm saying is we have to look at the data to understand if slices from the same study were split into the train and test sets. Maybe this is not the case and we just have different studies from the same patient in the train and test sets.</p>\n\n<p>If one slice for a given study is in a set (train / validation / test), then ALL slices for that study are in that set. There are patients with multiple studies - and those occur across train and validation, but not in stage 2. Slices from the same study were <em>not</em> split into train and validation sets.</p>\n\n<p>&gt; Also, I understood that the labels are at the image-level and not study-level. Do you agree?</p>\n\n<p>All labels are at the image level, not the study level.</p>\n\n<p><a href=\"/carlolepelaars\">@carlolepelaars</a>:\n&gt; Curious how people are going to exploit this for medals and if it will become a nuisance for serious competitors.</p>\n\n<p>As medals are based on performance in stage 2, this should not become a nuisance for anyone, except for the people that try to exploit something that doesn't exist in stage 2, of course. :)</p>\n\n<p><a href=\"/jledoux\">@jledoux</a>:\n&gt; I think aggregating per study to recreate the 3-d image violate the rules limiting us to only use pixel data</p>\n\n<p>As clarified in the Metadata Q&amp;A, creating a 3D image of a given study using the metadata is within the bounds of the competition.</p>",
          "rawMarkdown": "Hi all - thanks for all the questions!  Some answers below:\n\n@felipekitamura:\n&gt; All I'm saying is we have to look at the data to understand if slices from the same study were split into the train and test sets. Maybe this is not the case and we just have different studies from the same patient in the train and test sets.\n\nIf one slice for a given study is in a set (train / validation / test), then ALL slices for that study are in that set. There are patients with multiple studies - and those occur across train and validation, but not in stage 2. Slices from the same study were *not* split into train and validation sets.\n\n&gt; Also, I understood that the labels are at the image-level and not study-level. Do you agree?\n\nAll labels are at the image level, not the study level.\n\n@carlolepelaars:\n&gt; Curious how people are going to exploit this for medals and if it will become a nuisance for serious competitors.\n\nAs medals are based on performance in stage 2, this should not become a nuisance for anyone, except for the people that try to exploit something that doesn't exist in stage 2, of course. :)\n\n@jledoux:\n&gt; I think aggregating per study to recreate the 3-d image violate the rules limiting us to only use pixel data\n\nAs clarified in the Metadata Q&amp;A, creating a 3D image of a given study using the metadata is within the bounds of the competition.",
          "votes": 6
        },
        {
          "id": 635553,
          "postDate": "2019-09-27T19:59:39.850Z",
          "content": "<p>Hi Phil,</p>\n\n<p>Thank you for clearing up the questions! It is much appreciated! </p>",
          "rawMarkdown": "Hi Phil,\n\nThank you for clearing up the questions! It is much appreciated! ",
          "votes": 2
        },
        {
          "id": 635561,
          "postDate": "2019-09-27T20:28:38.930Z",
          "content": "<p>Hi Carlo - you're welcome! Let me know if any more questions come up!</p>",
          "rawMarkdown": "Hi Carlo - you're welcome! Let me know if any more questions come up!",
          "votes": 1
        },
        {
          "id": 636019,
          "postDate": "2019-09-28T16:07:59.747Z",
          "content": "<p>Dear Phil,\nI just would be sure I correctly understood your statement \"creating a 3D image of a given study using the metadata is within the bounds of the competition.\"</p>\n\n<p>If I collect all the images of a patient's cranium (assume for simplicity each patient has only 1 CT), may I use the combination of those (3D image) to get the probability of the hemorrhage type ?</p>\n\n<p>In the article reported on the top of this discussion, may that approach be used without violate the competition rules ? or we are allowed just to use the CNN having in input the single slice (2D image) ? </p>\n\n<p>May you give an example about what is not allowed to do from the above mentioned article ?</p>",
          "rawMarkdown": "Dear Phil,\nI just would be sure I correctly understood your statement \"creating a 3D image of a given study using the metadata is within the bounds of the competition.\"\n\nIf I collect all the images of a patient's cranium (assume for simplicity each patient has only 1 CT), may I use the combination of those (3D image) to get the probability of the hemorrhage type ?\n\nIn the article reported on the top of this discussion, may that approach be used without violate the competition rules ? or we are allowed just to use the CNN having in input the single slice (2D image) ? \n\nMay you give an example about what is not allowed to do from the above mentioned article ?",
          "votes": 2
        },
        {
          "id": 636085,
          "postDate": "2019-09-28T19:46:39.757Z",
          "content": "<p>So it is allowed to \"reconstruct\" the whole study as a 3D example gathering relevant slices, and use them together as a single sample, right? With some kind of 3D convolution, or something?</p>",
          "rawMarkdown": "So it is allowed to \"reconstruct\" the whole study as a 3D example gathering relevant slices, and use them together as a single sample, right? With some kind of 3D convolution, or something?"
        },
        {
          "id": 636458,
          "postDate": "2019-09-29T15:05:22.993Z",
          "content": "<p>Hi <a href=\"/philculliton\">@philculliton</a> , I don't understand what's different between \"aggregating per study to recreate the 3-d image\" and \"creating a 3D image of a given study\". Could you clarify which ID or information in meta we can use to reconstruct the 3D image?</p>",
          "rawMarkdown": "Hi @philculliton , I don't understand what's different between \"aggregating per study to recreate the 3-d image\" and \"creating a 3D image of a given study\". Could you clarify which ID or information in meta we can use to reconstruct the 3D image?"
        },
        {
          "id": 645076,
          "postDate": "2019-10-09T18:54:20.420Z",
          "content": "<p>I think the rules have been amended to allow meta-data other than pixel data to be fair game! 👍 \n<a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/111325#latest-644288\">https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/111325#latest-644288</a></p>",
          "rawMarkdown": "I think the rules have been amended to allow meta-data other than pixel data to be fair game! 👍 \nhttps://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/111325#latest-644288"
        }
      ]
    },
    {
      "id": 638248,
      "postDate": "2019-10-01T16:56:07.447Z",
      "content": "<p>In response to some of the questions about use of non-pixel metadata, the host has clarified: <strong>metadata can be used for pre-processing of images, but they cannot be features in your model or used to change or label predictions post-modeling</strong>. </p>\n\n<p>This should help resolve many of the more specific questions that have been raised recently on what types of things can be done with the metadata. We won’t be advising on questions about how to use this data.</p>",
      "rawMarkdown": "In response to some of the questions about use of non-pixel metadata, the host has clarified: **metadata can be used for pre-processing of images, but they cannot be features in your model or used to change or label predictions post-modeling**. \n\nThis should help resolve many of the more specific questions that have been raised recently on what types of things can be done with the metadata. We won’t be advising on questions about how to use this data.",
      "votes": -2
    },
    {
      "id": 655553,
      "postDate": "2019-10-23T07:12:10.460Z",
      "content": "<p>This paper has no github link in it , so just give up.</p>",
      "rawMarkdown": "This paper has no github link in it , so just give up.\n"
    }
  ],
  "comments": [
    {
      "id": 639509,
      "author_name": "Mahtab Noor Shaan",
      "author_url": "",
      "post_date": "2019-10-03T09:42:16.030000",
      "content": "<p>So, have anyone yet tried with the 2 step pipeline yet? I am referring to the 2 step pipeline for per image, not for per study.</p>\n\n<p>Also, how do you tackle the 2nd step for classifying hemorrhage types? What is the target probability values for the test images that you define not to have any hemorrhage (any=0 or less than a threshold)? Are the probabilities for the 5 sub-types 0.0 for these images? It would be helpful if someone can give me any advice. 😊 </p>",
      "votes": 1,
      "replies": [
        {
          "id": 645105,
          "author_name": "Carlo Lepelaars",
          "author_url": "",
          "post_date": "2019-10-09T19:44:12.350000",
          "content": "<p>I'm working on a 2-step pipeline and am thinking of setting a threshold. 0.5 for the threshold is a good choice for example, but the threshold can also be tuned based on validation data.</p>\n\n<p>So for example a phase pipeline would be:</p>\n\n<p>```</p>\n\n<h1>Initialize predictions</h1>\n\n<p>preds = []</p>\n\n<h1>Load models</h1>\n\n<p>bin_model = load_model('model_trained_on_any.h5')\ntype_model = load_model('model_trained_on_types.h5')</p>\n\n<h1>Get predictions</h1>\n\n<p>threshold = 0.5\nfor img in list(img_dir):\n            img = cv2.imread(img)\n            bin_pred = bin_model.predict(img)\n            if bin_pred &lt; threshold:\n                self.predictions.append([0,0,0,0,0,0])\n            else:\n                pred = [bin_pred] + type_model.predict(img)\n                self.predictions.append(pred)\n```</p>\n\n<p>Hope this helps! Curious to hear your thoughts on this.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 629110,
      "author_name": "FelipeKitamura, MD, PhD",
      "author_url": "",
      "post_date": "2019-09-18T11:14:31.023000",
      "content": "<p>Very nice paper! Just beware that the dataset comes as 2D slices. You can aggregate them per study to form the 3D image but some patients may have part of their slices in the test set and part in the train set. This is only for stage 1, as stated in the Data section.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 629246,
          "author_name": "Fernando Camargo",
          "author_url": "",
          "post_date": "2019-09-18T14:59:15.240000",
          "content": "<p>Hello, <a href=\"/felipekitamura\">@felipekitamura</a>. I didn't understand why part of the slices are split between train and test set. I mean, the detection is per study/pacient, right? Don't we need all the slices of each study to make the inference? My first tought was to use 3D Conv Nets, which would need the whole 3D Image with the slices in the right order.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 629299,
          "author_name": "FelipeKitamura, MD, PhD",
          "author_url": "",
          "post_date": "2019-09-18T16:14:21.597000",
          "content": "<p>Good point. I didn't make myself clear. All I'm saying is we have to look at the data to understand if slices from the same study were split into the train and test sets. Maybe this is not the case and we just have different studies from the same patient in the train and test sets.\nAlso, I understood that the labels are at the image-level and not study-level. Do you agree?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 629395,
          "author_name": "Fernando Camargo",
          "author_url": "",
          "post_date": "2019-09-18T18:18:34.670000",
          "content": "<p>I'll have a deeper look at the data soon, but I thought the labels would be at study level.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 629438,
          "author_name": "jledoux",
          "author_url": "",
          "post_date": "2019-09-18T19:19:08.667000",
          "content": "<p>I think aggregating per study to recreate the 3-d image violate the rules limiting us to only use pixel data (\"Submission predictions must be based entirely on the pixel data in the provided datasets\").  Or am I misreading that requirement?</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 629542,
          "author_name": "FelipeKitamura, MD, PhD",
          "author_url": "",
          "post_date": "2019-09-18T22:35:36.053000",
          "content": "<p>Good point. It seems they were talking about other stuff that could potentially cause issues though. But I agree the way this is written raises the question.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 633012,
          "author_name": "Carlos Martín-Isla",
          "author_url": "",
          "post_date": "2019-09-24T10:41:17.250000",
          "content": "<p>Is this question already answered?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 633154,
          "author_name": "Carlo Lepelaars",
          "author_url": "",
          "post_date": "2019-09-24T13:46:50.270000",
          "content": "<p>Please correct me if I'm wrong, but if we have slices of each patient in both the training and testing set, doesn't this present a big data leak? We could transfer the label for a patient over to the test set and have correct predictions, right? </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 633280,
          "author_name": "Benoît Koenig",
          "author_url": "",
          "post_date": "2019-09-24T17:03:19.910000",
          "content": "<p><a href=\"/carlolepelaars\">@carlolepelaars</a> yes it is a big data leak, but it won't work for stage-2 so it's quite useless\n<a href=\"/jledoux\">@jledoux</a> Phil Cullington clarified this question here: <a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109479#latest-632646\">https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109479#latest-632646</a>\nYou can use data reconstruction, however I have not seen anywhere so far whether the stage-2 data will contain all the slices for each patients or if missing slices are to be expected\nAlso, I think all slices for a given patient do not have the same labeling, that is, if the hemorrhage is not visible on a given slice, it is marked as 0 (to confirm)</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 633282,
          "author_name": "Carlo Lepelaars",
          "author_url": "",
          "post_date": "2019-09-24T17:08:32.177000",
          "content": "<p>Hmm, interesting. Curious how people are going to exploit this for medals and if it will become a nuisance for serious competitors.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 635537,
          "author_name": "Phil Culliton",
          "author_url": "",
          "post_date": "2019-09-27T19:16:15.790000",
          "content": "<p>Hi all - thanks for all the questions!  Some answers below:</p>\n\n<p><a href=\"/felipekitamura\">@felipekitamura</a>:\n&gt; All I'm saying is we have to look at the data to understand if slices from the same study were split into the train and test sets. Maybe this is not the case and we just have different studies from the same patient in the train and test sets.</p>\n\n<p>If one slice for a given study is in a set (train / validation / test), then ALL slices for that study are in that set. There are patients with multiple studies - and those occur across train and validation, but not in stage 2. Slices from the same study were <em>not</em> split into train and validation sets.</p>\n\n<p>&gt; Also, I understood that the labels are at the image-level and not study-level. Do you agree?</p>\n\n<p>All labels are at the image level, not the study level.</p>\n\n<p><a href=\"/carlolepelaars\">@carlolepelaars</a>:\n&gt; Curious how people are going to exploit this for medals and if it will become a nuisance for serious competitors.</p>\n\n<p>As medals are based on performance in stage 2, this should not become a nuisance for anyone, except for the people that try to exploit something that doesn't exist in stage 2, of course. :)</p>\n\n<p><a href=\"/jledoux\">@jledoux</a>:\n&gt; I think aggregating per study to recreate the 3-d image violate the rules limiting us to only use pixel data</p>\n\n<p>As clarified in the Metadata Q&amp;A, creating a 3D image of a given study using the metadata is within the bounds of the competition.</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 635553,
          "author_name": "Carlo Lepelaars",
          "author_url": "",
          "post_date": "2019-09-27T19:59:39.850000",
          "content": "<p>Hi Phil,</p>\n\n<p>Thank you for clearing up the questions! It is much appreciated! </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 635561,
          "author_name": "Phil Culliton",
          "author_url": "",
          "post_date": "2019-09-27T20:28:38.930000",
          "content": "<p>Hi Carlo - you're welcome! Let me know if any more questions come up!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 636019,
          "author_name": "AlGiLa",
          "author_url": "",
          "post_date": "2019-09-28T16:07:59.747000",
          "content": "<p>Dear Phil,\nI just would be sure I correctly understood your statement \"creating a 3D image of a given study using the metadata is within the bounds of the competition.\"</p>\n\n<p>If I collect all the images of a patient's cranium (assume for simplicity each patient has only 1 CT), may I use the combination of those (3D image) to get the probability of the hemorrhage type ?</p>\n\n<p>In the article reported on the top of this discussion, may that approach be used without violate the competition rules ? or we are allowed just to use the CNN having in input the single slice (2D image) ? </p>\n\n<p>May you give an example about what is not allowed to do from the above mentioned article ?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 636085,
          "author_name": "Ilia Zaitsev",
          "author_url": "",
          "post_date": "2019-09-28T19:46:39.757000",
          "content": "<p>So it is allowed to \"reconstruct\" the whole study as a 3D example gathering relevant slices, and use them together as a single sample, right? With some kind of 3D convolution, or something?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 636458,
          "author_name": "BlackSwan",
          "author_url": "",
          "post_date": "2019-09-29T15:05:22.993000",
          "content": "<p>Hi <a href=\"/philculliton\">@philculliton</a> , I don't understand what's different between \"aggregating per study to recreate the 3-d image\" and \"creating a 3D image of a given study\". Could you clarify which ID or information in meta we can use to reconstruct the 3D image?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 645076,
          "author_name": "Lloyd Palum",
          "author_url": "",
          "post_date": "2019-10-09T18:54:20.420000",
          "content": "<p>I think the rules have been amended to allow meta-data other than pixel data to be fair game! 👍 \n<a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/111325#latest-644288\">https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/111325#latest-644288</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 638248,
      "author_name": "Julia Elliott",
      "author_url": "",
      "post_date": "2019-10-01T16:56:07.447000",
      "content": "<p>In response to some of the questions about use of non-pixel metadata, the host has clarified: <strong>metadata can be used for pre-processing of images, but they cannot be features in your model or used to change or label predictions post-modeling</strong>. </p>\n\n<p>This should help resolve many of the more specific questions that have been raised recently on what types of things can be done with the metadata. We won’t be advising on questions about how to use this data.</p>",
      "votes": -2,
      "replies": []
    },
    {
      "id": 655553,
      "author_name": "Tian Bingyang",
      "author_url": "",
      "post_date": "2019-10-23T07:12:10.460000",
      "content": "<p>This paper has no github link in it , so just give up.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "628950": "Figure 1 in this paper shows a pipeline very similar to the objective of this competition:\nhttps://rd.springer.com/content/pdf/10.1007%2Fs00330-019-06163-2.pdf\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F421965%2F03bcc16baae739172b513d016e26181e%2Fpipeline.png?generation=1568789507192980&amp;alt=media)\n",
    "639509": "So, have anyone yet tried with the 2 step pipeline yet? I am referring to the 2 step pipeline for per image, not for per study.\n\nAlso, how do you tackle the 2nd step for classifying hemorrhage types? What is the target probability values for the test images that you define not to have any hemorrhage (any=0 or less than a threshold)? Are the probabilities for the 5 sub-types 0.0 for these images? It would be helpful if someone can give me any advice. 😊 ",
    "629110": "Very nice paper! Just beware that the dataset comes as 2D slices. You can aggregate them per study to form the 3D image but some patients may have part of their slices in the test set and part in the train set. This is only for stage 1, as stated in the Data section.",
    "638248": "In response to some of the questions about use of non-pixel metadata, the host has clarified: **metadata can be used for pre-processing of images, but they cannot be features in your model or used to change or label predictions post-modeling**. \n\nThis should help resolve many of the more specific questions that have been raised recently on what types of things can be done with the metadata. We won’t be advising on questions about how to use this data.",
    "655553": "This paper has no github link in it , so just give up.\n"
  }
}