{
  "id": 109258,
  "title": "Welcome!",
  "url": "/competitions/rsna-intracranial-hemorrhage-detection/discussion/109258",
  "author_name": "Phil Culliton",
  "post_date": "2019-09-18T02:17:26.861000",
  "votes": 11,
  "comment_count": 17,
  "views": 0,
  "content": "<p>Welcome to the RSNA Intracranial Hemorrhage Detection competition!</p>\n\n<p>In this competition, we're detecting hemorrhages using de-identified CT studies. Each exam has many images, some of which may contain multiple types of hemorrhages. The evaluation metric is fairly straightforward and is <a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/overview/evaluation\">detailed here</a>. This is a two-stage competition and will have an entirely different final test set at the end.</p>\n\n<p><strong>Important Note: All data should now download via the Download All button.  Note that the API is not working correctly right now. We're working on a fix, but in the meantime please download using the Download All button.</strong></p>\n\n<p>A couple of additional notes about the dataset:</p>\n\n<p>1) There IS patient crossover between the training set and the stage 1 test set. It does NOT exist between the training set and the stage 2 test set. Please see the <a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/data\">Data</a> tab for more details.</p>\n\n<p>2) Note that the <a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/rules\">rules explicitly state</a> that ONLY pixel data can be used in creating your solution. There is metadata in the DICOM files, provided for the purposes of structuring your model input, etc., but it cannot be used directly without invalidating your model.</p>\n\n<p>Good luck!</p>",
  "messages": [
    {
      "id": 628841,
      "postDate": "2019-09-18T02:17:26.863Z",
      "content": "<p>Welcome to the RSNA Intracranial Hemorrhage Detection competition!</p>\n\n<p>In this competition, we're detecting hemorrhages using de-identified CT studies. Each exam has many images, some of which may contain multiple types of hemorrhages. The evaluation metric is fairly straightforward and is <a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/overview/evaluation\">detailed here</a>. This is a two-stage competition and will have an entirely different final test set at the end.</p>\n\n<p><strong>Important Note: All data should now download via the Download All button.  Note that the API is not working correctly right now. We're working on a fix, but in the meantime please download using the Download All button.</strong></p>\n\n<p>A couple of additional notes about the dataset:</p>\n\n<p>1) There IS patient crossover between the training set and the stage 1 test set. It does NOT exist between the training set and the stage 2 test set. Please see the <a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/data\">Data</a> tab for more details.</p>\n\n<p>2) Note that the <a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/rules\">rules explicitly state</a> that ONLY pixel data can be used in creating your solution. There is metadata in the DICOM files, provided for the purposes of structuring your model input, etc., but it cannot be used directly without invalidating your model.</p>\n\n<p>Good luck!</p>",
      "rawMarkdown": "Welcome to the RSNA Intracranial Hemorrhage Detection competition!\n\nIn this competition, we're detecting hemorrhages using de-identified CT studies. Each exam has many images, some of which may contain multiple types of hemorrhages. The evaluation metric is fairly straightforward and is [detailed here](https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/overview/evaluation). This is a two-stage competition and will have an entirely different final test set at the end.\n\n**Important Note: All data should now download via the Download All button.  Note that the API is not working correctly right now. We're working on a fix, but in the meantime please download using the Download All button.**\n\nA couple of additional notes about the dataset:\n\n1) There IS patient crossover between the training set and the stage 1 test set. It does NOT exist between the training set and the stage 2 test set. Please see the [Data](https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/data) tab for more details.\n\n2) Note that the [rules explicitly state](https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/rules) that ONLY pixel data can be used in creating your solution. There is metadata in the DICOM files, provided for the purposes of structuring your model input, etc., but it cannot be used directly without invalidating your model.\n\nGood luck!\n",
      "votes": 11
    },
    {
      "id": 635500,
      "postDate": "2019-09-27T17:38:38.690Z",
      "content": "<p>Thanks <a href=\"/philculliton\">@philculliton</a> and organizers\nCould you please provide the loss weights, so that the competitors don't have to probe the LB?\nI understand that the organizers might want to make the competition fair, and I strongly believe that hiding this information will not help.\nTop participants will eventually work out the weights and use that to fine tune their models, which will create a large unbalance on the LB and obscure the real best predictors.\nI personally believe openness will only contribute to a best outcome for all.\nThank you</p>",
      "rawMarkdown": "Thanks @philculliton and organizers\nCould you please provide the loss weights, so that the competitors don't have to probe the LB?\nI understand that the organizers might want to make the competition fair, and I strongly believe that hiding this information will not help.\nTop participants will eventually work out the weights and use that to fine tune their models, which will create a large unbalance on the LB and obscure the real best predictors.\nI personally believe openness will only contribute to a best outcome for all.\nThank you",
      "votes": 5,
      "replies": [
        {
          "id": 635685,
          "postDate": "2019-09-28T03:04:00.727Z",
          "content": "<p>Hi <a href=\"/hmendonca\">@hmendonca</a> - thanks for asking! The <code>any</code> label weight is 2.0.  All other labels are weighted 1.0.</p>",
          "rawMarkdown": "Hi @hmendonca - thanks for asking! The `any` label weight is 2.0.  All other labels are weighted 1.0.",
          "votes": 5
        }
      ]
    },
    {
      "id": 629428,
      "postDate": "2019-09-18T19:05:11.603Z",
      "content": "<p>Thanks for the challenge. I had a question about the evaluation metric: for weighted multi-label logarithmic loss, how much weight is applied for each label?</p>\n\n<p>Also, I had a question about \"ONLY pixel data can be used in creating your solution.\" So you're saying that the metadata shouldn't be used directly. What do you mean by \"directly\" here - I assume it's fine to use the metadata to organize into studies and potentially do pre-processing on images using the window values. Do you mean to not use the meta-data directly as a feature to your ML model?</p>",
      "rawMarkdown": "Thanks for the challenge. I had a question about the evaluation metric: for weighted multi-label logarithmic loss, how much weight is applied for each label?\n\nAlso, I had a question about \"ONLY pixel data can be used in creating your solution.\" So you're saying that the metadata shouldn't be used directly. What do you mean by \"directly\" here - I assume it's fine to use the metadata to organize into studies and potentially do pre-processing on images using the window values. Do you mean to not use the meta-data directly as a feature to your ML model?",
      "votes": 6,
      "replies": [
        {
          "id": 630351,
          "postDate": "2019-09-20T05:45:28.030Z",
          "content": "<p>Just in case someone might be as cautious about the rules forbidding direct usage of metadata as I am, I've shared a windowing technique based on pixel data only in <a href=\"https://www.kaggle.com/samusram/discovering-windowing-on-our-own-no-metadata\">the kernel <em>Discovering Windowing On Our Own (No Metadata)</em></a></p>",
          "rawMarkdown": "Just in case someone might be as cautious about the rules forbidding direct usage of metadata as I am, I've shared a windowing technique based on pixel data only in [the kernel *Discovering Windowing On Our Own (No Metadata)*](https://www.kaggle.com/samusram/discovering-windowing-on-our-own-no-metadata)",
          "votes": 2
        },
        {
          "id": 632622,
          "postDate": "2019-09-23T20:07:56.717Z",
          "content": "<p><a href=\"/philculliton\">@philculliton</a> I just wanted to follow up in regards to the evaluation metric (I read your response in regards to the pixel data question (<a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109479#latest-632490\">https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109479#latest-632490</a>), so I'm good at the moment.)</p>",
          "rawMarkdown": "@philculliton I just wanted to follow up in regards to the evaluation metric (I read your response in regards to the pixel data question (https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109479#latest-632490), so I'm good at the moment.)"
        }
      ]
    },
    {
      "id": 629149,
      "postDate": "2019-09-18T12:28:39.143Z",
      "content": "<p>Thanks Phil for these useful informations! I have 2 related questions:</p>\n\n<p>1) Can we aggregate the 2D images into a single serial 3D exam using the metadata to train a model and predict ? </p>\n\n<p>2) For stage 2, can we have a confirmation if all images from a single exam will completely be included in the stage 2 subset ? </p>",
      "rawMarkdown": "Thanks Phil for these useful informations! I have 2 related questions:\n\n1) Can we aggregate the 2D images into a single serial 3D exam using the metadata to train a model and predict ? \n\n2) For stage 2, can we have a confirmation if all images from a single exam will completely be included in the stage 2 subset ? ",
      "votes": 4,
      "replies": [
        {
          "id": 630055,
          "postDate": "2019-09-19T16:51:57.303Z",
          "content": "<p>Thanks for the questions <a href=\"/alexandrecc\">@alexandrecc</a>!</p>\n\n<p>1) Yes, feel free!\n2) There are no exams that cross between subsets. There is <em>patient</em> crossover between train and stage 1 test, but those are patients with multiple exams. The exams themselves are self-contained.</p>",
          "rawMarkdown": "Thanks for the questions @alexandrecc!\n\n1) Yes, feel free!\n2) There are no exams that cross between subsets. There is _patient_ crossover between train and stage 1 test, but those are patients with multiple exams. The exams themselves are self-contained.",
          "votes": 4
        }
      ]
    },
    {
      "id": 629117,
      "postDate": "2019-09-18T11:24:09.347Z",
      "content": "<p>Thank you for this interesting challenge. After looking at data page I have a few questions.</p>\n\n<blockquote>\n  <p>Submission predictions must be based entirely on the pixel data in the provided datasets</p>\n</blockquote>\n\n<p>Is it permitted to use this info for anything besides model training? E.g. for stratification.</p>\n\n<blockquote>\n  <p>In this competition, we're detecting hemorrhages using de-identified CT studies. Each exam has many images, some of which may contain multiple types of hemorrhages.</p>\n</blockquote>\n\n<p>Is labeling image-based or exam-based? E.g. if someone has hemorrhage, is it guaranteed that it will be present on each image corresponding to this patient?</p>\n\n<p>Also, maybe I missed it, but still: what is an exact meaning of <strong>Study Instance UID</strong> and <strong>Series Instance UID</strong>?</p>",
      "rawMarkdown": "Thank you for this interesting challenge. After looking at data page I have a few questions.\n\n&gt; Submission predictions must be based entirely on the pixel data in the provided datasets\n\nIs it permitted to use this info for anything besides model training? E.g. for stratification.\n\n&gt; In this competition, we're detecting hemorrhages using de-identified CT studies. Each exam has many images, some of which may contain multiple types of hemorrhages.\n\nIs labeling image-based or exam-based? E.g. if someone has hemorrhage, is it guaranteed that it will be present on each image corresponding to this patient?\n\nAlso, maybe I missed it, but still: what is an exact meaning of **Study Instance UID** and **Series Instance UID**?",
      "votes": 2,
      "replies": [
        {
          "id": 630058,
          "postDate": "2019-09-19T17:01:55.823Z",
          "content": "<p>Thanks <a href=\"/hokmund\">@hokmund</a>!</p>\n\n<p>&gt; Is it permitted to use this info for anything besides model training? E.g. for stratification.</p>\n\n<p>Do you have an example of what you're considering?  The host team is trying to work out exactly what's being proposed.</p>\n\n<p>&gt; Is labeling image-based or exam-based? E.g. if someone has hemorrhage, is it guaranteed that it will be present on each image corresponding to this patient?</p>\n\n<p>Labeling is image-based. Re: your second question, the training data has some great examples of how hemorrhages are distributed throughout a given exam - it varies.</p>\n\n<p>&gt; Also, maybe I missed it, but still: what is an exact meaning of Study Instance UID and Series Instance UID?</p>\n\n<p>The answer from the host is as follows:</p>\n\n<blockquote>\n  <p>In the real world one study can have multiple series within a study but in the case of the competition not so each one of them should be unique IDs for that specific study.</p>\n</blockquote>\n\n<p>For the purposes of this competition, both Study and Series Instance UID act as unique exam identifiers.</p>\n\n<p>Hopefully that clears some things up, let me know if you have more questions!</p>",
          "rawMarkdown": "Thanks @hokmund!\n\n&gt; Is it permitted to use this info for anything besides model training? E.g. for stratification.\n\nDo you have an example of what you're considering?  The host team is trying to work out exactly what's being proposed.\n\n&gt; Is labeling image-based or exam-based? E.g. if someone has hemorrhage, is it guaranteed that it will be present on each image corresponding to this patient?\n\nLabeling is image-based. Re: your second question, the training data has some great examples of how hemorrhages are distributed throughout a given exam - it varies.\n\n&gt; Also, maybe I missed it, but still: what is an exact meaning of Study Instance UID and Series Instance UID?\n\nThe answer from the host is as follows:\n\n&gt; In the real world one study can have multiple series within a study but in the case of the competition not so each one of them should be unique IDs for that specific study.\n\nFor the purposes of this competition, both Study and Series Instance UID act as unique exam identifiers.\n\nHopefully that clears some things up, let me know if you have more questions!",
          "votes": 3
        },
        {
          "id": 630120,
          "postDate": "2019-09-19T18:59:58.067Z",
          "content": "<blockquote>\n  <p>For the purposes of this competition, both Study and Series Instance UID act as unique exam identifiers.</p>\n</blockquote>\n\n<p>I guess it's a relief for competitors, since it simplifies things a lot 😄 </p>\n\n<blockquote>\n  <p>Do you have an example of what you're considering? The host team is trying to work out exactly what's being proposed.</p>\n</blockquote>\n\n<p>E.g. if I am doing crossvalidation, I want all data for a patient to be in a single fold (in order to prevent data leakage).</p>",
          "rawMarkdown": "&gt; For the purposes of this competition, both Study and Series Instance UID act as unique exam identifiers.\n\nI guess it's a relief for competitors, since it simplifies things a lot 😄 \n\n&gt; Do you have an example of what you're considering? The host team is trying to work out exactly what's being proposed.\n\nE.g. if I am doing crossvalidation, I want all data for a patient to be in a single fold (in order to prevent data leakage).",
          "votes": 1
        }
      ]
    },
    {
      "id": 635753,
      "postDate": "2019-09-28T06:18:49.950Z",
      "content": "<p>As data processing , converting from dcm to png/jpeg format with some pixel engineering to highlight edges , is taking  almost 3.5 hrs . And the model training time is extra.</p>\n\n<p>Could we do like uploading processed image and then run model training only.\nIs it ok form competitions rules point of view ? \nThanks.</p>",
      "rawMarkdown": "As data processing , converting from dcm to png/jpeg format with some pixel engineering to highlight edges , is taking  almost 3.5 hrs . And the model training time is extra.\n\nCould we do like uploading processed image and then run model training only.\nIs it ok form competitions rules point of view ? \nThanks."
    },
    {
      "id": 634635,
      "postDate": "2019-09-26T14:41:56.100Z",
      "content": "<p>Hi is there a time limit for a code to be run? Example: if my code is run for 30 hours in CPU (let's say 10 in GPU), will is still be accepted?</p>",
      "rawMarkdown": "Hi is there a time limit for a code to be run? Example: if my code is run for 30 hours in CPU (let's say 10 in GPU), will is still be accepted?",
      "replies": [
        {
          "id": 634670,
          "postDate": "2019-09-26T15:36:11.323Z",
          "content": "<p>I'd say that 10 hours on a single GPU sounds like a modest solution for a Kaggle CV competition :)</p>",
          "rawMarkdown": "I'd say that 10 hours on a single GPU sounds like a modest solution for a Kaggle CV competition :)"
        }
      ]
    },
    {
      "id": 632773,
      "postDate": "2019-09-24T03:38:23.857Z",
      "content": "<p><a href=\"/philculliton\">@philculliton</a> Hi Phil, I'm new to 2 stage competitions. So after stage 1, will the stage 1 test set labels be released? Also when submitting the model to stage 2, will there be any feedback on the performance before stage 2 is over?</p>\n\n<p>Thanks.</p>",
      "rawMarkdown": "@philculliton Hi Phil, I'm new to 2 stage competitions. So after stage 1, will the stage 1 test set labels be released? Also when submitting the model to stage 2, will there be any feedback on the performance before stage 2 is over?\n\nThanks."
    },
    {
      "id": 630861,
      "postDate": "2019-09-20T21:54:13.743Z",
      "content": "<blockquote>\n  <p>2) Note that the rules explicitly state that ONLY pixel data can be used in creating your solution. There is metadata in the DICOM files, provided for the purposes of structuring your model input, etc., but it cannot be used directly without invalidating your model.</p>\n</blockquote>\n\n<p>IIRC number of images per Study Instance UID  varies from 20 to 60.</p>\n\n<p>Can we use metadata for e.g. group all images by Study Instance UID  and run predictions on the set of grouped images (instead of running it from a single image)? The input to the model would be e.g. 20-60 images instead of 1.</p>",
      "rawMarkdown": "&gt; 2) Note that the rules explicitly state that ONLY pixel data can be used in creating your solution. There is metadata in the DICOM files, provided for the purposes of structuring your model input, etc., but it cannot be used directly without invalidating your model.\n\nIIRC number of images per Study Instance UID  varies from 20 to 60.\n\nCan we use metadata for e.g. group all images by Study Instance UID  and run predictions on the set of grouped images (instead of running it from a single image)? The input to the model would be e.g. 20-60 images instead of 1."
    },
    {
      "id": 628883,
      "postDate": "2019-09-18T04:40:54.540Z",
      "content": "<p>Year in timeline have been mistyped I guess...</p>",
      "rawMarkdown": "Year in timeline have been mistyped I guess...",
      "replies": [
        {
          "id": 628889,
          "postDate": "2019-09-18T04:48:16.070Z",
          "content": "<p>Thank you - great catch!</p>",
          "rawMarkdown": "Thank you - great catch!"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 635500,
      "author_name": "Henrique Mendonça",
      "author_url": "",
      "post_date": "2019-09-27T17:38:38.690000",
      "content": "<p>Thanks <a href=\"/philculliton\">@philculliton</a> and organizers\nCould you please provide the loss weights, so that the competitors don't have to probe the LB?\nI understand that the organizers might want to make the competition fair, and I strongly believe that hiding this information will not help.\nTop participants will eventually work out the weights and use that to fine tune their models, which will create a large unbalance on the LB and obscure the real best predictors.\nI personally believe openness will only contribute to a best outcome for all.\nThank you</p>",
      "votes": 5,
      "replies": [
        {
          "id": 635685,
          "author_name": "Phil Culliton",
          "author_url": "",
          "post_date": "2019-09-28T03:04:00.727000",
          "content": "<p>Hi <a href=\"/hmendonca\">@hmendonca</a> - thanks for asking! The <code>any</code> label weight is 2.0.  All other labels are weighted 1.0.</p>",
          "votes": 5,
          "replies": []
        }
      ]
    },
    {
      "id": 629428,
      "author_name": "ArvindVepa",
      "author_url": "",
      "post_date": "2019-09-18T19:05:11.603000",
      "content": "<p>Thanks for the challenge. I had a question about the evaluation metric: for weighted multi-label logarithmic loss, how much weight is applied for each label?</p>\n\n<p>Also, I had a question about \"ONLY pixel data can be used in creating your solution.\" So you're saying that the metadata shouldn't be used directly. What do you mean by \"directly\" here - I assume it's fine to use the metadata to organize into studies and potentially do pre-processing on images using the window values. Do you mean to not use the meta-data directly as a feature to your ML model?</p>",
      "votes": 6,
      "replies": [
        {
          "id": 630351,
          "author_name": "Raman",
          "author_url": "",
          "post_date": "2019-09-20T05:45:28.030000",
          "content": "<p>Just in case someone might be as cautious about the rules forbidding direct usage of metadata as I am, I've shared a windowing technique based on pixel data only in <a href=\"https://www.kaggle.com/samusram/discovering-windowing-on-our-own-no-metadata\">the kernel <em>Discovering Windowing On Our Own (No Metadata)</em></a></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 632622,
          "author_name": "ArvindVepa",
          "author_url": "",
          "post_date": "2019-09-23T20:07:56.717000",
          "content": "<p><a href=\"/philculliton\">@philculliton</a> I just wanted to follow up in regards to the evaluation metric (I read your response in regards to the pixel data question (<a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109479#latest-632490\">https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109479#latest-632490</a>), so I'm good at the moment.)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 629149,
      "author_name": "Alexandre Cadrin-Chênevert",
      "author_url": "",
      "post_date": "2019-09-18T12:28:39.143000",
      "content": "<p>Thanks Phil for these useful informations! I have 2 related questions:</p>\n\n<p>1) Can we aggregate the 2D images into a single serial 3D exam using the metadata to train a model and predict ? </p>\n\n<p>2) For stage 2, can we have a confirmation if all images from a single exam will completely be included in the stage 2 subset ? </p>",
      "votes": 4,
      "replies": [
        {
          "id": 630055,
          "author_name": "Phil Culliton",
          "author_url": "",
          "post_date": "2019-09-19T16:51:57.303000",
          "content": "<p>Thanks for the questions <a href=\"/alexandrecc\">@alexandrecc</a>!</p>\n\n<p>1) Yes, feel free!\n2) There are no exams that cross between subsets. There is <em>patient</em> crossover between train and stage 1 test, but those are patients with multiple exams. The exams themselves are self-contained.</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 629117,
      "author_name": "Dmytro Panchenko",
      "author_url": "",
      "post_date": "2019-09-18T11:24:09.347000",
      "content": "<p>Thank you for this interesting challenge. After looking at data page I have a few questions.</p>\n\n<blockquote>\n  <p>Submission predictions must be based entirely on the pixel data in the provided datasets</p>\n</blockquote>\n\n<p>Is it permitted to use this info for anything besides model training? E.g. for stratification.</p>\n\n<blockquote>\n  <p>In this competition, we're detecting hemorrhages using de-identified CT studies. Each exam has many images, some of which may contain multiple types of hemorrhages.</p>\n</blockquote>\n\n<p>Is labeling image-based or exam-based? E.g. if someone has hemorrhage, is it guaranteed that it will be present on each image corresponding to this patient?</p>\n\n<p>Also, maybe I missed it, but still: what is an exact meaning of <strong>Study Instance UID</strong> and <strong>Series Instance UID</strong>?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 630058,
          "author_name": "Phil Culliton",
          "author_url": "",
          "post_date": "2019-09-19T17:01:55.823000",
          "content": "<p>Thanks <a href=\"/hokmund\">@hokmund</a>!</p>\n\n<p>&gt; Is it permitted to use this info for anything besides model training? E.g. for stratification.</p>\n\n<p>Do you have an example of what you're considering?  The host team is trying to work out exactly what's being proposed.</p>\n\n<p>&gt; Is labeling image-based or exam-based? E.g. if someone has hemorrhage, is it guaranteed that it will be present on each image corresponding to this patient?</p>\n\n<p>Labeling is image-based. Re: your second question, the training data has some great examples of how hemorrhages are distributed throughout a given exam - it varies.</p>\n\n<p>&gt; Also, maybe I missed it, but still: what is an exact meaning of Study Instance UID and Series Instance UID?</p>\n\n<p>The answer from the host is as follows:</p>\n\n<blockquote>\n  <p>In the real world one study can have multiple series within a study but in the case of the competition not so each one of them should be unique IDs for that specific study.</p>\n</blockquote>\n\n<p>For the purposes of this competition, both Study and Series Instance UID act as unique exam identifiers.</p>\n\n<p>Hopefully that clears some things up, let me know if you have more questions!</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 630120,
          "author_name": "Dmytro Panchenko",
          "author_url": "",
          "post_date": "2019-09-19T18:59:58.067000",
          "content": "<blockquote>\n  <p>For the purposes of this competition, both Study and Series Instance UID act as unique exam identifiers.</p>\n</blockquote>\n\n<p>I guess it's a relief for competitors, since it simplifies things a lot 😄 </p>\n\n<blockquote>\n  <p>Do you have an example of what you're considering? The host team is trying to work out exactly what's being proposed.</p>\n</blockquote>\n\n<p>E.g. if I am doing crossvalidation, I want all data for a patient to be in a single fold (in order to prevent data leakage).</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 635753,
      "author_name": "Rajnish Chauhan",
      "author_url": "",
      "post_date": "2019-09-28T06:18:49.950000",
      "content": "<p>As data processing , converting from dcm to png/jpeg format with some pixel engineering to highlight edges , is taking  almost 3.5 hrs . And the model training time is extra.</p>\n\n<p>Could we do like uploading processed image and then run model training only.\nIs it ok form competitions rules point of view ? \nThanks.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 634635,
      "author_name": "Gurgel",
      "author_url": "",
      "post_date": "2019-09-26T14:41:56.100000",
      "content": "<p>Hi is there a time limit for a code to be run? Example: if my code is run for 30 hours in CPU (let's say 10 in GPU), will is still be accepted?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 634670,
          "author_name": "Dmytro Panchenko",
          "author_url": "",
          "post_date": "2019-09-26T15:36:11.323000",
          "content": "<p>I'd say that 10 hours on a single GPU sounds like a modest solution for a Kaggle CV competition :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 632773,
      "author_name": "Bo Peng",
      "author_url": "",
      "post_date": "2019-09-24T03:38:23.857000",
      "content": "<p><a href=\"/philculliton\">@philculliton</a> Hi Phil, I'm new to 2 stage competitions. So after stage 1, will the stage 1 test set labels be released? Also when submitting the model to stage 2, will there be any feedback on the performance before stage 2 is over?</p>\n\n<p>Thanks.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 630861,
      "author_name": "Andrés Miguel Torrubia Sáez",
      "author_url": "",
      "post_date": "2019-09-20T21:54:13.743000",
      "content": "<blockquote>\n  <p>2) Note that the rules explicitly state that ONLY pixel data can be used in creating your solution. There is metadata in the DICOM files, provided for the purposes of structuring your model input, etc., but it cannot be used directly without invalidating your model.</p>\n</blockquote>\n\n<p>IIRC number of images per Study Instance UID  varies from 20 to 60.</p>\n\n<p>Can we use metadata for e.g. group all images by Study Instance UID  and run predictions on the set of grouped images (instead of running it from a single image)? The input to the model would be e.g. 20-60 images instead of 1.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 628883,
      "author_name": "Broken",
      "author_url": "",
      "post_date": "2019-09-18T04:40:54.540000",
      "content": "<p>Year in timeline have been mistyped I guess...</p>",
      "votes": 0,
      "replies": [
        {
          "id": 628889,
          "author_name": "Julia Elliott",
          "author_url": "",
          "post_date": "2019-09-18T04:48:16.070000",
          "content": "<p>Thank you - great catch!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "628841": "Welcome to the RSNA Intracranial Hemorrhage Detection competition!\n\nIn this competition, we're detecting hemorrhages using de-identified CT studies. Each exam has many images, some of which may contain multiple types of hemorrhages. The evaluation metric is fairly straightforward and is [detailed here](https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/overview/evaluation). This is a two-stage competition and will have an entirely different final test set at the end.\n\n**Important Note: All data should now download via the Download All button.  Note that the API is not working correctly right now. We're working on a fix, but in the meantime please download using the Download All button.**\n\nA couple of additional notes about the dataset:\n\n1) There IS patient crossover between the training set and the stage 1 test set. It does NOT exist between the training set and the stage 2 test set. Please see the [Data](https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/data) tab for more details.\n\n2) Note that the [rules explicitly state](https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/rules) that ONLY pixel data can be used in creating your solution. There is metadata in the DICOM files, provided for the purposes of structuring your model input, etc., but it cannot be used directly without invalidating your model.\n\nGood luck!\n",
    "635500": "Thanks @philculliton and organizers\nCould you please provide the loss weights, so that the competitors don't have to probe the LB?\nI understand that the organizers might want to make the competition fair, and I strongly believe that hiding this information will not help.\nTop participants will eventually work out the weights and use that to fine tune their models, which will create a large unbalance on the LB and obscure the real best predictors.\nI personally believe openness will only contribute to a best outcome for all.\nThank you",
    "629428": "Thanks for the challenge. I had a question about the evaluation metric: for weighted multi-label logarithmic loss, how much weight is applied for each label?\n\nAlso, I had a question about \"ONLY pixel data can be used in creating your solution.\" So you're saying that the metadata shouldn't be used directly. What do you mean by \"directly\" here - I assume it's fine to use the metadata to organize into studies and potentially do pre-processing on images using the window values. Do you mean to not use the meta-data directly as a feature to your ML model?",
    "629149": "Thanks Phil for these useful informations! I have 2 related questions:\n\n1) Can we aggregate the 2D images into a single serial 3D exam using the metadata to train a model and predict ? \n\n2) For stage 2, can we have a confirmation if all images from a single exam will completely be included in the stage 2 subset ? ",
    "629117": "Thank you for this interesting challenge. After looking at data page I have a few questions.\n\n&gt; Submission predictions must be based entirely on the pixel data in the provided datasets\n\nIs it permitted to use this info for anything besides model training? E.g. for stratification.\n\n&gt; In this competition, we're detecting hemorrhages using de-identified CT studies. Each exam has many images, some of which may contain multiple types of hemorrhages.\n\nIs labeling image-based or exam-based? E.g. if someone has hemorrhage, is it guaranteed that it will be present on each image corresponding to this patient?\n\nAlso, maybe I missed it, but still: what is an exact meaning of **Study Instance UID** and **Series Instance UID**?",
    "635753": "As data processing , converting from dcm to png/jpeg format with some pixel engineering to highlight edges , is taking  almost 3.5 hrs . And the model training time is extra.\n\nCould we do like uploading processed image and then run model training only.\nIs it ok form competitions rules point of view ? \nThanks.",
    "634635": "Hi is there a time limit for a code to be run? Example: if my code is run for 30 hours in CPU (let's say 10 in GPU), will is still be accepted?",
    "632773": "@philculliton Hi Phil, I'm new to 2 stage competitions. So after stage 1, will the stage 1 test set labels be released? Also when submitting the model to stage 2, will there be any feedback on the performance before stage 2 is over?\n\nThanks.",
    "630861": "&gt; 2) Note that the rules explicitly state that ONLY pixel data can be used in creating your solution. There is metadata in the DICOM files, provided for the purposes of structuring your model input, etc., but it cannot be used directly without invalidating your model.\n\nIIRC number of images per Study Instance UID  varies from 20 to 60.\n\nCan we use metadata for e.g. group all images by Study Instance UID  and run predictions on the set of grouped images (instead of running it from a single image)? The input to the model would be e.g. 20-60 images instead of 1.",
    "628883": "Year in timeline have been mistyped I guess..."
  }
}