{
  "id": 113339,
  "title": "External data on qure.ai from Albert Einstein Brazil",
  "url": "/competitions/rsna-intracranial-hemorrhage-detection/discussion/113339",
  "author_name": "Anouk Stein, MD",
  "post_date": "2019-10-18T16:50:18.419000",
  "votes": 32,
  "comment_count": 42,
  "views": 0,
  "content": "<p>Reposted on external data thread: We are excited to see the progress that the Kaggle teams are making! Eduardo Reis, MD and his group from Hospital Israelita Albert Einstein, São Paulo, BR,  have annotated the qure.ai CQ500 dataset with bounding boxes for the different types of hemorrhage. They annotated the thick sliced series within each exam and extrapolated the boxes to the thinner sliced series to expand the available data. They are in the process of writing up their findings for publication. The dataset is made available to the Kaggle community and can be viewed on the MD.ai platform: <a href=\"https://public.md.ai/annotator/project/Y2qr6vqv\">https://public.md.ai/annotator/project/Y2qr6vqv</a> Eduardo's annotations are on labelgroup 4 - BrainHemX and the bounding boxes can be downloaded using the link below. The original images are hosted by qure.ai at <a href=\"http://headctstudy.qure.ai/dataset\">http://headctstudy.qure.ai/dataset</a> licensed under the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License</p>",
  "messages": [
    {
      "id": 652306,
      "postDate": "2019-10-18T16:50:18.420Z",
      "content": "<p>Reposted on external data thread: We are excited to see the progress that the Kaggle teams are making! Eduardo Reis, MD and his group from Hospital Israelita Albert Einstein, São Paulo, BR,  have annotated the qure.ai CQ500 dataset with bounding boxes for the different types of hemorrhage. They annotated the thick sliced series within each exam and extrapolated the boxes to the thinner sliced series to expand the available data. They are in the process of writing up their findings for publication. The dataset is made available to the Kaggle community and can be viewed on the MD.ai platform: <a href=\"https://public.md.ai/annotator/project/Y2qr6vqv\">https://public.md.ai/annotator/project/Y2qr6vqv</a> Eduardo's annotations are on labelgroup 4 - BrainHemX and the bounding boxes can be downloaded using the link below. The original images are hosted by qure.ai at <a href=\"http://headctstudy.qure.ai/dataset\">http://headctstudy.qure.ai/dataset</a> licensed under the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License</p>",
      "rawMarkdown": "Reposted on external data thread: We are excited to see the progress that the Kaggle teams are making! Eduardo Reis, MD and his group from Hospital Israelita Albert Einstein, São Paulo, BR,  have annotated the qure.ai CQ500 dataset with bounding boxes for the different types of hemorrhage. They annotated the thick sliced series within each exam and extrapolated the boxes to the thinner sliced series to expand the available data. They are in the process of writing up their findings for publication. The dataset is made available to the Kaggle community and can be viewed on the MD.ai platform: https://public.md.ai/annotator/project/Y2qr6vqv Eduardo's annotations are on labelgroup 4 - BrainHemX and the bounding boxes can be downloaded using the link below. The original images are hosted by qure.ai at http://headctstudy.qure.ai/dataset licensed under the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License",
      "votes": 32
    },
    {
      "id": 652499,
      "postDate": "2019-10-18T23:50:14.753Z",
      "content": "<p>I'm slightly concerned about releasing that big dataset for training purposes two weeks before the deadline. Am I the only one?</p>",
      "rawMarkdown": "I'm slightly concerned about releasing that big dataset for training purposes two weeks before the deadline. Am I the only one?",
      "votes": 14,
      "replies": [
        {
          "id": 652672,
          "postDate": "2019-10-19T08:03:53.520Z",
          "content": "<p>A little bit worried yes. It is becoming more of a - who has the most hardware to compute on all the datasets in time before the deadline(s) - competition ;-)</p>",
          "rawMarkdown": "A little bit worried yes. It is becoming more of a - who has the most hardware to compute on all the datasets in time before the deadline(s) - competition ;-)",
          "votes": 3
        },
        {
          "id": 652675,
          "postDate": "2019-10-19T08:10:58.287Z",
          "content": "<p>Totally agreed.\nFrom research point of view it empowers us a lot.\nBut from competitive point of view it makes competition a mess.</p>",
          "rawMarkdown": "Totally agreed.\nFrom research point of view it empowers us a lot.\nBut from competitive point of view it makes competition a mess.",
          "votes": 3
        },
        {
          "id": 652718,
          "postDate": "2019-10-19T09:44:46.860Z",
          "content": "<p>Not to mention this competition already provides quite a lot of data</p>",
          "rawMarkdown": "Not to mention this competition already provides quite a lot of data",
          "votes": 2
        }
      ]
    },
    {
      "id": 655168,
      "postDate": "2019-10-22T18:51:50.360Z",
      "content": "<p>If anyone has problems reading dicom files using pydicom, this can be a solution. \nThe embedded images can not be read by PIL/Pillow and GDCM is needed instead.</p>\n\n<p><a href=\"https://github.com/HealthplusAI/python3-gdcm\">https://github.com/HealthplusAI/python3-gdcm</a></p>\n\n<p><code>\nimport pydicom.pixel_data_handlers.gdcm_handler as gdcm_handler\npydicom.config.image_handlers = [None, gdcm_handler]\n</code></p>",
      "rawMarkdown": "If anyone has problems reading dicom files using pydicom, this can be a solution. \nThe embedded images can not be read by PIL/Pillow and GDCM is needed instead.\n\nhttps://github.com/HealthplusAI/python3-gdcm\n\n```\nimport pydicom.pixel_data_handlers.gdcm_handler as gdcm_handler\npydicom.config.image_handlers = [None, gdcm_handler]\n```\n",
      "votes": 8
    },
    {
      "id": 652448,
      "postDate": "2019-10-18T21:26:32.703Z",
      "content": "<p>I'd like to redirect a question from <a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109274#649282\">external data thread</a> here.\nAre we allowed to use BY-NC-SA-licensed data for the competition?</p>",
      "rawMarkdown": "I'd like to redirect a question from [external data thread](https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109274#649282) here.\nAre we allowed to use BY-NC-SA-licensed data for the competition?",
      "votes": 5,
      "replies": [
        {
          "id": 652941,
          "postDate": "2019-10-19T17:01:59.487Z",
          "content": "<p>Hi <a href=\"/hokmund\">@hokmund</a>! The qure.ai dataset is publicly and freely available for use and download by everyone. It is licensed for non-commercial use, which is allowable for the intended purposes of this competition. In fact, the dataset provided for this competition is also licensed for non-commercial use. </p>",
          "rawMarkdown": "Hi @hokmund! The qure.ai dataset is publicly and freely available for use and download by everyone. It is licensed for non-commercial use, which is allowable for the intended purposes of this competition. In fact, the dataset provided for this competition is also licensed for non-commercial use. ",
          "votes": 2
        }
      ]
    },
    {
      "id": 654459,
      "postDate": "2019-10-21T22:15:08.873Z",
      "content": "<p>Has anyone evaluated this dataset with a trained model? </p>",
      "rawMarkdown": "Has anyone evaluated this dataset with a trained model? ",
      "votes": 6
    },
    {
      "id": 652379,
      "postDate": "2019-10-18T19:15:08.827Z",
      "content": "<p>I almost had a heart-attack seeing Albert Einstein doing data annotation . Then after reading through i found it is Hospital Israelita Albert Einstein 😐 . Anyway, thanks for the awesome work.</p>",
      "rawMarkdown": "I almost had a heart-attack seeing Albert Einstein doing data annotation . Then after reading through i found it is Hospital Israelita Albert Einstein 😐 . Anyway, thanks for the awesome work.",
      "votes": 4
    },
    {
      "id": 652373,
      "postDate": "2019-10-18T18:56:15.763Z",
      "content": "<p>Right! Thank you, Anouk. The competitors now have image-level annotations for the qure.ai CQ500 external dataset. Perhaps it can help.</p>\n\n<p>Many thanks to everyone who also participated in the work: Felipe Nascimento, Fernando Secol, Mateus Aranha, Marcelo Felix, Edson Amaro, Birajara Machado, and Anouk for providing all technical support.</p>",
      "rawMarkdown": "Right! Thank you, Anouk. The competitors now have image-level annotations for the qure.ai CQ500 external dataset. Perhaps it can help.\n\nMany thanks to everyone who also participated in the work: Felipe Nascimento, Fernando Secol, Mateus Aranha, Marcelo Felix, Edson Amaro, Birajara Machado, and Anouk for providing all technical support.",
      "votes": 4
    },
    {
      "id": 652895,
      "postDate": "2019-10-19T15:25:16.847Z",
      "content": "<p>Some tips:</p>\n\n<p>Note that qure.ai dataset contains more than one series per exam: thick and thin slices, bone and soft-tissue filters, pre and post-contrast, etc. </p>\n\n<p>The neuroradiologists drew the bounding boxes in the <strong>soft-tissue thick slices</strong> . Then the bbx were extrapolated to the other series, which don’t always result in a perfect match. </p>\n\n<p>For the extrapolated bbx, what I would be most careful about are the images with 'no hemorrhage' which are adjacent to others with hemorrhage. They can instead, contain hemorrhage. (and vice-versa)</p>\n\n<p>We will soon add a global label indicating which series the bbxs were originally handmade @ <a href=\"https://public.md.ai/annotator/project/Y2qr6vqv\">https://public.md.ai/annotator/project/Y2qr6vqv</a></p>",
      "rawMarkdown": "Some tips:\n\nNote that qure.ai dataset contains more than one series per exam: thick and thin slices, bone and soft-tissue filters, pre and post-contrast, etc. \n\nThe neuroradiologists drew the bounding boxes in the **soft-tissue thick slices** . Then the bbx were extrapolated to the other series, which don’t always result in a perfect match. \n\nFor the extrapolated bbx, what I would be most careful about are the images with 'no hemorrhage' which are adjacent to others with hemorrhage. They can instead, contain hemorrhage. (and vice-versa)\n\nWe will soon add a global label indicating which series the bbxs were originally handmade @ https://public.md.ai/annotator/project/Y2qr6vqv",
      "votes": 3
    },
    {
      "id": 658745,
      "postDate": "2019-10-26T13:30:10.567Z",
      "content": "<p><a href=\"/anoukstein\">@anoukstein</a> 2 Questions for this additional data:\n1) From what I have seem in the CQ500 dataset, there are 5 labels. The annotations above are charming, but what type of hemorrhage does it correspond to? <code>Eduardo's annotations are on labelgroup 4 - BrainHemX</code>--so I take it is label 4?</p>\n\n<p>2) CQ500 dataset is missing important descriptions. For a single patient, there are many scans. e.g. patient 0 contains \"CT 4cc sec 150cc D3D on\", \"CT 4cc sec 150cc D3D on-2\" ... \"CT Plain\", \"CT PLAIN THIN\". How are these folders different? So data in which folder is consistent with the competition data? I found that there is a folder in each patient which contains dcm with names which includes all other folders. The additional data is very frustrating. Please explain :)</p>",
      "rawMarkdown": "@anoukstein 2 Questions for this additional data:\n1) From what I have seem in the CQ500 dataset, there are 5 labels. The annotations above are charming, but what type of hemorrhage does it correspond to? `Eduardo's annotations are on labelgroup 4 - BrainHemX`--so I take it is label 4?\n\n2) CQ500 dataset is missing important descriptions. For a single patient, there are many scans. e.g. patient 0 contains \"CT 4cc sec 150cc D3D on\", \"CT 4cc sec 150cc D3D on-2\" ... \"CT Plain\", \"CT PLAIN THIN\". How are these folders different? So data in which folder is consistent with the competition data? I found that there is a folder in each patient which contains dcm with names which includes all other folders. The additional data is very frustrating. Please explain :)",
      "votes": 3
    },
    {
      "id": 657038,
      "postDate": "2019-10-24T21:30:35.633Z",
      "content": "<p>Did this help anyone's performance?</p>",
      "rawMarkdown": "Did this help anyone's performance?",
      "votes": 3
    },
    {
      "id": 653164,
      "postDate": "2019-10-20T02:41:09.420Z",
      "content": "<p>I would be super happy if CQ500 dataset are provided in the same format as original dataset;)</p>",
      "rawMarkdown": "I would be super happy if CQ500 dataset are provided in the same format as original dataset;)",
      "votes": 3
    },
    {
      "id": 652346,
      "postDate": "2019-10-18T18:03:14.490Z",
      "content": "<p>Dr Reis' team will be supplying an updated version of the boxes with minor changes where the boxes skipped images on a few cases. They wanted to get out as much data as possible given the time constraint of this competition so decided to release this version rather than delay. The csv  contains 38940 hemorrhage bounding boxes on the 491 studies from CQ500 qure.ai.</p>",
      "rawMarkdown": "Dr Reis' team will be supplying an updated version of the boxes with minor changes where the boxes skipped images on a few cases. They wanted to get out as much data as possible given the time constraint of this competition so decided to release this version rather than delay. The csv  contains 38940 hemorrhage bounding boxes on the 491 studies from CQ500 qure.ai.",
      "votes": 4
    },
    {
      "id": 654352,
      "postDate": "2019-10-21T18:50:50.983Z",
      "content": "<p><a href=\"/anoukstein\">@anoukstein</a> or anyone else, Does this dataset have the target labels we use in our competition ? <code>epidural,intraparenchymal,intraventricular,subarachnoid,subdural</code> or is there a way to map the labels from this datset to our dataset ?</p>",
      "rawMarkdown": "@anoukstein or anyone else, Does this dataset have the target labels we use in our competition ? `epidural,intraparenchymal,intraventricular,subarachnoid,subdural` or is there a way to map the labels from this datset to our dataset ?",
      "votes": 1,
      "replies": [
        {
          "id": 654356,
          "postDate": "2019-10-21T18:56:04.600Z",
          "content": "<p>Ok, I think I got it.... ( .<a href=\"http://www.healthsciences.uci.edu/nursing/docs/stoke-conference/acute-care-of-patients-with-intracerebral-hemorrhage.pdf\">http://www.healthsciences.uci.edu/nursing/docs/stoke-conference/acute-care-of-patients-with-intracerebral-hemorrhage.pdf</a> )\n– Subarachnoid Hemorrhage (SAH)\n– Subdural Hematoma (SDH)\n– Epidural Hematoma (EDH)\n– Intraventricular Hemorrhage (IVH)\n– Intracerebral/Intraparenchymal Hemorrhage(ICH/IPH)</p>",
          "rawMarkdown": "Ok, I think I got it.... ( .http://www.healthsciences.uci.edu/nursing/docs/stoke-conference/acute-care-of-patients-with-intracerebral-hemorrhage.pdf )\n– Subarachnoid Hemorrhage (SAH)\n– Subdural Hematoma (SDH)\n– Epidural Hematoma (EDH)\n– Intraventricular Hemorrhage (IVH)\n– Intracerebral/Intraparenchymal Hemorrhage(ICH/IPH)",
          "votes": 1
        }
      ]
    },
    {
      "id": 653322,
      "postDate": "2019-10-20T08:37:06.547Z",
      "content": "<p>Hi, it seems that the images are not in standard dicom format, either they are compressed or something else (the buffer size is not the full image size). What am I missing? thanks</p>",
      "rawMarkdown": "Hi, it seems that the images are not in standard dicom format, either they are compressed or something else (the buffer size is not the full image size). What am I missing? thanks",
      "votes": 1,
      "replies": [
        {
          "id": 653334,
          "postDate": "2019-10-20T09:09:03.013Z",
          "content": "<p>So - I have found a workaround - one can read the images using SimpleITK </p>",
          "rawMarkdown": "So - I have found a workaround - one can read the images using SimpleITK ",
          "votes": 2
        },
        {
          "id": 655138,
          "postDate": "2019-10-22T18:09:14.043Z",
          "content": "<p>How do you read images with SimpleITK? \nsitk.GetArrayViewFromImage(sitk.ReadImage(file)) returns strange data type that doesnt match pydicom read</p>",
          "rawMarkdown": "How do you read images with SimpleITK? \nsitk.GetArrayViewFromImage(sitk.ReadImage(file)) returns strange data type that doesnt match pydicom read"
        },
        {
          "id": 655509,
          "postDate": "2019-10-23T05:55:28.037Z",
          "content": "<p>```\n      read_image = sitk.ReadImage(fn)\n      image = sitk.GetArrayFromImage(read_image)[0].astype(np.float32)</p>\n\n<p>```</p>",
          "rawMarkdown": "```\n      read_image = sitk.ReadImage(fn)\n      image = sitk.GetArrayFromImage(read_image)[0].astype(np.float32)\n\n```",
          "votes": 1
        }
      ]
    },
    {
      "id": 652405,
      "postDate": "2019-10-18T20:13:12.010Z",
      "content": "<p>Are these image slices labeled the same way as the competition dataset at the slice level or patient/study level? For example, in competition dataset, a patient's slices may have no hemorrhage in certain slices but hemorrhage in others.</p>",
      "rawMarkdown": "Are these image slices labeled the same way as the competition dataset at the slice level or patient/study level? For example, in competition dataset, a patient's slices may have no hemorrhage in certain slices but hemorrhage in others.",
      "votes": 1,
      "replies": [
        {
          "id": 652426,
          "postDate": "2019-10-18T20:48:46.327Z",
          "content": "<p>Only the slices with hemorrhage have bounding boxes. The images without hemorrhage were not specifically identified as such. Hope that helps.</p>",
          "rawMarkdown": "Only the slices with hemorrhage have bounding boxes. The images without hemorrhage were not specifically identified as such. Hope that helps.",
          "votes": 2
        }
      ]
    },
    {
      "id": 652357,
      "postDate": "2019-10-18T18:19:44.530Z",
      "content": "<p>Awesome work, <a href=\"/epreis\">@epreis</a>! <a href=\"/drvidurmahajan\">@drvidurmahajan</a> , your work is cross-pollinating!</p>",
      "rawMarkdown": "Awesome work, @epreis! @drvidurmahajan , your work is cross-pollinating!",
      "votes": 1,
      "replies": [
        {
          "id": 653604,
          "postDate": "2019-10-20T17:52:03.123Z",
          "content": "<p>Thanks Felipe. Didn't think someone would ever put so much effort into CQ500! </p>",
          "rawMarkdown": "Thanks Felipe. Didn't think someone would ever put so much effort into CQ500! ",
          "votes": 1
        }
      ]
    },
    {
      "id": 653376,
      "postDate": "2019-10-20T10:35:18.890Z",
      "content": "<p>Is it legal and under rules to make mask annotation based on your bbox annotation?</p>",
      "rawMarkdown": "Is it legal and under rules to make mask annotation based on your bbox annotation?",
      "votes": 2,
      "replies": [
        {
          "id": 653770,
          "postDate": "2019-10-20T23:50:26.953Z",
          "content": "<p>Nice! Soon we should also add masks (polygons on md.ai), working on it.</p>\n\n<p>You can also make it, according to the licensing rules of qure.ai, any modifications must be made publicly available under the same license, please check by yourself:\n- <a href=\"https://creativecommons.org/licenses/by-nc-sa/4.0/\">https://creativecommons.org/licenses/by-nc-sa/4.0/</a>\n- <a href=\"http://headctstudy.qure.ai/dataset\">http://headctstudy.qure.ai/dataset</a> </p>\n\n<p>To use it in the competition, you probably know <a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109274#latest-652678\">the External Data Thread</a>.</p>",
          "rawMarkdown": "Nice! Soon we should also add masks (polygons on md.ai), working on it.\n\nYou can also make it, according to the licensing rules of qure.ai, any modifications must be made publicly available under the same license, please check by yourself:\n- https://creativecommons.org/licenses/by-nc-sa/4.0/\n- http://headctstudy.qure.ai/dataset \n\nTo use it in the competition, you probably know [the External Data Thread](https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109274#latest-652678).",
          "votes": 2
        }
      ]
    },
    {
      "id": 652673,
      "postDate": "2019-10-19T08:04:34.200Z",
      "content": "<p>Will this new dataset be made available as a Kaggle Dataset?</p>",
      "rawMarkdown": "Will this new dataset be made available as a Kaggle Dataset?",
      "votes": 2
    },
    {
      "id": 652705,
      "postDate": "2019-10-19T09:05:30.267Z",
      "content": "<p>How to download it on terminal? <br>\n<code>\nwget hogehoge\n</code></p>",
      "rawMarkdown": "How to download it on terminal?  \n```\nwget hogehoge\n```\n"
    },
    {
      "id": 972123,
      "postDate": "2020-08-16T08:51:33.563Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/anoukstein\" target=\"_blank\">@anoukstein</a>, I have been trying to download the annotations from <a href=\"https://public.md.ai/annotator/project/Y2qr6vqv\" target=\"_blank\">https://public.md.ai/annotator/project/Y2qr6vqv</a><br>\nBut the export option is not enabled. Can you let me know how to download the data?</p>",
      "rawMarkdown": "Hi @anoukstein, I have been trying to download the annotations from https://public.md.ai/annotator/project/Y2qr6vqv\nBut the export option is not enabled. Can you let me know how to download the data?"
    },
    {
      "id": 664918,
      "postDate": "2019-11-04T12:03:43.667Z",
      "content": "<p>The boxes file was updated and it broke my code. Be careful :) </p>",
      "rawMarkdown": "The boxes file was updated and it broke my code. Be careful :) "
    },
    {
      "id": 652372,
      "postDate": "2019-10-18T18:55:29.280Z",
      "content": "<p>Hi, how do you export the images/labels? When I click on the export tab and the export window pops up, I can't click on the export button. Thanks for all the great work you guys are doing.</p>",
      "rawMarkdown": "Hi, how do you export the images/labels? When I click on the export tab and the export window pops up, I can't click on the export button. Thanks for all the great work you guys are doing.",
      "replies": [
        {
          "id": 652389,
          "postDate": "2019-10-18T19:43:32.633Z",
          "content": "<p><a href=\"/anoukstein\">@anoukstein</a> So I believe exporting annotations is disabled on the platform? I was reading a few discussion posts about it on md.ai.</p>",
          "rawMarkdown": "@anoukstein So I believe exporting annotations is disabled on the platform? I was reading a few discussion posts about it on md.ai."
        },
        {
          "id": 652395,
          "postDate": "2019-10-18T19:53:13.453Z",
          "content": "<p>Yes, the annotations export is disabled but the boxes from BrainHemX can be downloaded from the link in the post above (qureai-cq500-boxes.csv). You can download the images directly from qure.ai (<a href=\"http://headctstudy.qure.ai/dataset\">http://headctstudy.qure.ai/dataset</a>) who have so graciously made them available!</p>",
          "rawMarkdown": "Yes, the annotations export is disabled but the boxes from BrainHemX can be downloaded from the link in the post above (qureai-cq500-boxes.csv). You can download the images directly from qure.ai (http://headctstudy.qure.ai/dataset) who have so graciously made them available!",
          "replies": [
            {
              "id": 3267348,
              "postDate": "2025-08-11T01:55:37.127Z",
              "content": "<p>Dear Organizers, the website (<a href=\"http://headctstudy.qure.ai/dataset\" target=\"_blank\">http://headctstudy.qure.ai/dataset</a>) isn't working. Where can I download the annotated dataset instead? I'd appreciate your assistance with this.</p>",
              "rawMarkdown": "Dear Organizers, the website (http://headctstudy.qure.ai/dataset) isn't working. Where can I download the annotated dataset instead? I'd appreciate your assistance with this."
            }
          ]
        },
        {
          "id": 652402,
          "postDate": "2019-10-18T20:08:10.090Z",
          "content": "<p>oh i see. thanks.</p>",
          "rawMarkdown": "oh i see. thanks."
        },
        {
          "id": 655831,
          "postDate": "2019-10-23T14:50:28.653Z",
          "content": "<p>unable to download the data, is it still available?</p>",
          "rawMarkdown": "unable to download the data, is it still available?"
        },
        {
          "id": 658748,
          "postDate": "2019-10-26T13:31:08.633Z",
          "content": "<p><a href=\"/yangddd\">@yangddd</a> I just downloaded it today. You may need some tricks to download it if you are in China</p>",
          "rawMarkdown": "@yangddd I just downloaded it today. You may need some tricks to download it if you are in China",
          "votes": 1
        },
        {
          "id": 660000,
          "postDate": "2019-10-28T14:52:39.660Z",
          "content": "<p>thanks</p>",
          "rawMarkdown": "thanks"
        }
      ]
    },
    {
      "id": 1600883,
      "postDate": "2021-11-30T19:30:48.657Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 658861,
      "postDate": "2019-10-26T16:35:33.237Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 658743,
      "postDate": "2019-10-26T13:26:04.983Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 652499,
      "author_name": "Oleg Yaroshevskiy",
      "author_url": "",
      "post_date": "2019-10-18T23:50:14.753000",
      "content": "<p>I'm slightly concerned about releasing that big dataset for training purposes two weeks before the deadline. Am I the only one?</p>",
      "votes": 14,
      "replies": [
        {
          "id": 652672,
          "author_name": "Robin Smits",
          "author_url": "",
          "post_date": "2019-10-19T08:03:53.520000",
          "content": "<p>A little bit worried yes. It is becoming more of a - who has the most hardware to compute on all the datasets in time before the deadline(s) - competition ;-)</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 652675,
          "author_name": "Dmytro Panchenko",
          "author_url": "",
          "post_date": "2019-10-19T08:10:58.287000",
          "content": "<p>Totally agreed.\nFrom research point of view it empowers us a lot.\nBut from competitive point of view it makes competition a mess.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 652718,
          "author_name": "Eek The Cat",
          "author_url": "",
          "post_date": "2019-10-19T09:44:46.860000",
          "content": "<p>Not to mention this competition already provides quite a lot of data</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 655168,
      "author_name": "Appian",
      "author_url": "",
      "post_date": "2019-10-22T18:51:50.360000",
      "content": "<p>If anyone has problems reading dicom files using pydicom, this can be a solution. \nThe embedded images can not be read by PIL/Pillow and GDCM is needed instead.</p>\n\n<p><a href=\"https://github.com/HealthplusAI/python3-gdcm\">https://github.com/HealthplusAI/python3-gdcm</a></p>\n\n<p><code>\nimport pydicom.pixel_data_handlers.gdcm_handler as gdcm_handler\npydicom.config.image_handlers = [None, gdcm_handler]\n</code></p>",
      "votes": 8,
      "replies": []
    },
    {
      "id": 652448,
      "author_name": "Dmytro Panchenko",
      "author_url": "",
      "post_date": "2019-10-18T21:26:32.703000",
      "content": "<p>I'd like to redirect a question from <a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109274#649282\">external data thread</a> here.\nAre we allowed to use BY-NC-SA-licensed data for the competition?</p>",
      "votes": 5,
      "replies": [
        {
          "id": 652941,
          "author_name": "Anouk Stein, MD",
          "author_url": "",
          "post_date": "2019-10-19T17:01:59.487000",
          "content": "<p>Hi <a href=\"/hokmund\">@hokmund</a>! The qure.ai dataset is publicly and freely available for use and download by everyone. It is licensed for non-commercial use, which is allowable for the intended purposes of this competition. In fact, the dataset provided for this competition is also licensed for non-commercial use. </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 654459,
      "author_name": "Oleg Yaroshevskiy",
      "author_url": "",
      "post_date": "2019-10-21T22:15:08.873000",
      "content": "<p>Has anyone evaluated this dataset with a trained model? </p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 652379,
      "author_name": "Md Khairul Islam",
      "author_url": "",
      "post_date": "2019-10-18T19:15:08.827000",
      "content": "<p>I almost had a heart-attack seeing Albert Einstein doing data annotation . Then after reading through i found it is Hospital Israelita Albert Einstein 😐 . Anyway, thanks for the awesome work.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 652373,
      "author_name": "Eduardo Pontes Reis",
      "author_url": "",
      "post_date": "2019-10-18T18:56:15.763000",
      "content": "<p>Right! Thank you, Anouk. The competitors now have image-level annotations for the qure.ai CQ500 external dataset. Perhaps it can help.</p>\n\n<p>Many thanks to everyone who also participated in the work: Felipe Nascimento, Fernando Secol, Mateus Aranha, Marcelo Felix, Edson Amaro, Birajara Machado, and Anouk for providing all technical support.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 652895,
      "author_name": "Eduardo Pontes Reis",
      "author_url": "",
      "post_date": "2019-10-19T15:25:16.847000",
      "content": "<p>Some tips:</p>\n\n<p>Note that qure.ai dataset contains more than one series per exam: thick and thin slices, bone and soft-tissue filters, pre and post-contrast, etc. </p>\n\n<p>The neuroradiologists drew the bounding boxes in the <strong>soft-tissue thick slices</strong> . Then the bbx were extrapolated to the other series, which don’t always result in a perfect match. </p>\n\n<p>For the extrapolated bbx, what I would be most careful about are the images with 'no hemorrhage' which are adjacent to others with hemorrhage. They can instead, contain hemorrhage. (and vice-versa)</p>\n\n<p>We will soon add a global label indicating which series the bbxs were originally handmade @ <a href=\"https://public.md.ai/annotator/project/Y2qr6vqv\">https://public.md.ai/annotator/project/Y2qr6vqv</a></p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 658745,
      "author_name": "Nicholas Lyu",
      "author_url": "",
      "post_date": "2019-10-26T13:30:10.567000",
      "content": "<p><a href=\"/anoukstein\">@anoukstein</a> 2 Questions for this additional data:\n1) From what I have seem in the CQ500 dataset, there are 5 labels. The annotations above are charming, but what type of hemorrhage does it correspond to? <code>Eduardo's annotations are on labelgroup 4 - BrainHemX</code>--so I take it is label 4?</p>\n\n<p>2) CQ500 dataset is missing important descriptions. For a single patient, there are many scans. e.g. patient 0 contains \"CT 4cc sec 150cc D3D on\", \"CT 4cc sec 150cc D3D on-2\" ... \"CT Plain\", \"CT PLAIN THIN\". How are these folders different? So data in which folder is consistent with the competition data? I found that there is a folder in each patient which contains dcm with names which includes all other folders. The additional data is very frustrating. Please explain :)</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 657038,
      "author_name": "Bo Peng",
      "author_url": "",
      "post_date": "2019-10-24T21:30:35.633000",
      "content": "<p>Did this help anyone's performance?</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 653164,
      "author_name": "XY",
      "author_url": "",
      "post_date": "2019-10-20T02:41:09.420000",
      "content": "<p>I would be super happy if CQ500 dataset are provided in the same format as original dataset;)</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 652346,
      "author_name": "Anouk Stein, MD",
      "author_url": "",
      "post_date": "2019-10-18T18:03:14.490000",
      "content": "<p>Dr Reis' team will be supplying an updated version of the boxes with minor changes where the boxes skipped images on a few cases. They wanted to get out as much data as possible given the time constraint of this competition so decided to release this version rather than delay. The csv  contains 38940 hemorrhage bounding boxes on the 491 studies from CQ500 qure.ai.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 654352,
      "author_name": "Darragh",
      "author_url": "",
      "post_date": "2019-10-21T18:50:50.983000",
      "content": "<p><a href=\"/anoukstein\">@anoukstein</a> or anyone else, Does this dataset have the target labels we use in our competition ? <code>epidural,intraparenchymal,intraventricular,subarachnoid,subdural</code> or is there a way to map the labels from this datset to our dataset ?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 654356,
          "author_name": "Darragh",
          "author_url": "",
          "post_date": "2019-10-21T18:56:04.600000",
          "content": "<p>Ok, I think I got it.... ( .<a href=\"http://www.healthsciences.uci.edu/nursing/docs/stoke-conference/acute-care-of-patients-with-intracerebral-hemorrhage.pdf\">http://www.healthsciences.uci.edu/nursing/docs/stoke-conference/acute-care-of-patients-with-intracerebral-hemorrhage.pdf</a> )\n– Subarachnoid Hemorrhage (SAH)\n– Subdural Hematoma (SDH)\n– Epidural Hematoma (EDH)\n– Intraventricular Hemorrhage (IVH)\n– Intracerebral/Intraparenchymal Hemorrhage(ICH/IPH)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 653322,
      "author_name": "Hadar",
      "author_url": "",
      "post_date": "2019-10-20T08:37:06.547000",
      "content": "<p>Hi, it seems that the images are not in standard dicom format, either they are compressed or something else (the buffer size is not the full image size). What am I missing? thanks</p>",
      "votes": 1,
      "replies": [
        {
          "id": 653334,
          "author_name": "Hadar",
          "author_url": "",
          "post_date": "2019-10-20T09:09:03.013000",
          "content": "<p>So - I have found a workaround - one can read the images using SimpleITK </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 655138,
          "author_name": "Oleg Yaroshevskiy",
          "author_url": "",
          "post_date": "2019-10-22T18:09:14.043000",
          "content": "<p>How do you read images with SimpleITK? \nsitk.GetArrayViewFromImage(sitk.ReadImage(file)) returns strange data type that doesnt match pydicom read</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 655509,
          "author_name": "Hadar",
          "author_url": "",
          "post_date": "2019-10-23T05:55:28.037000",
          "content": "<p>```\n      read_image = sitk.ReadImage(fn)\n      image = sitk.GetArrayFromImage(read_image)[0].astype(np.float32)</p>\n\n<p>```</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 652405,
      "author_name": "Tim Yee",
      "author_url": "",
      "post_date": "2019-10-18T20:13:12.010000",
      "content": "<p>Are these image slices labeled the same way as the competition dataset at the slice level or patient/study level? For example, in competition dataset, a patient's slices may have no hemorrhage in certain slices but hemorrhage in others.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 652426,
          "author_name": "Anouk Stein, MD",
          "author_url": "",
          "post_date": "2019-10-18T20:48:46.327000",
          "content": "<p>Only the slices with hemorrhage have bounding boxes. The images without hemorrhage were not specifically identified as such. Hope that helps.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 652357,
      "author_name": "FelipeKitamura, MD, PhD",
      "author_url": "",
      "post_date": "2019-10-18T18:19:44.530000",
      "content": "<p>Awesome work, <a href=\"/epreis\">@epreis</a>! <a href=\"/drvidurmahajan\">@drvidurmahajan</a> , your work is cross-pollinating!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 653604,
          "author_name": "Vidur Mahajan",
          "author_url": "",
          "post_date": "2019-10-20T17:52:03.123000",
          "content": "<p>Thanks Felipe. Didn't think someone would ever put so much effort into CQ500! </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 653376,
      "author_name": "n01z3",
      "author_url": "",
      "post_date": "2019-10-20T10:35:18.890000",
      "content": "<p>Is it legal and under rules to make mask annotation based on your bbox annotation?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 653770,
          "author_name": "Eduardo Pontes Reis",
          "author_url": "",
          "post_date": "2019-10-20T23:50:26.953000",
          "content": "<p>Nice! Soon we should also add masks (polygons on md.ai), working on it.</p>\n\n<p>You can also make it, according to the licensing rules of qure.ai, any modifications must be made publicly available under the same license, please check by yourself:\n- <a href=\"https://creativecommons.org/licenses/by-nc-sa/4.0/\">https://creativecommons.org/licenses/by-nc-sa/4.0/</a>\n- <a href=\"http://headctstudy.qure.ai/dataset\">http://headctstudy.qure.ai/dataset</a> </p>\n\n<p>To use it in the competition, you probably know <a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109274#latest-652678\">the External Data Thread</a>.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 652673,
      "author_name": "Robin Smits",
      "author_url": "",
      "post_date": "2019-10-19T08:04:34.200000",
      "content": "<p>Will this new dataset be made available as a Kaggle Dataset?</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 652705,
      "author_name": "takuoko",
      "author_url": "",
      "post_date": "2019-10-19T09:05:30.267000",
      "content": "<p>How to download it on terminal? <br>\n<code>\nwget hogehoge\n</code></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 972123,
      "author_name": "sindhu",
      "author_url": "",
      "post_date": "2020-08-16T08:51:33.563000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/anoukstein\" target=\"_blank\">@anoukstein</a>, I have been trying to download the annotations from <a href=\"https://public.md.ai/annotator/project/Y2qr6vqv\" target=\"_blank\">https://public.md.ai/annotator/project/Y2qr6vqv</a><br>\nBut the export option is not enabled. Can you let me know how to download the data?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 664918,
      "author_name": "Oleg Yaroshevskiy",
      "author_url": "",
      "post_date": "2019-11-04T12:03:43.667000",
      "content": "<p>The boxes file was updated and it broke my code. Be careful :) </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 652372,
      "author_name": "ArvindVepa",
      "author_url": "",
      "post_date": "2019-10-18T18:55:29.280000",
      "content": "<p>Hi, how do you export the images/labels? When I click on the export tab and the export window pops up, I can't click on the export button. Thanks for all the great work you guys are doing.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 652389,
          "author_name": "ArvindVepa",
          "author_url": "",
          "post_date": "2019-10-18T19:43:32.633000",
          "content": "<p><a href=\"/anoukstein\">@anoukstein</a> So I believe exporting annotations is disabled on the platform? I was reading a few discussion posts about it on md.ai.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 652395,
          "author_name": "Anouk Stein, MD",
          "author_url": "",
          "post_date": "2019-10-18T19:53:13.453000",
          "content": "<p>Yes, the annotations export is disabled but the boxes from BrainHemX can be downloaded from the link in the post above (qureai-cq500-boxes.csv). You can download the images directly from qure.ai (<a href=\"http://headctstudy.qure.ai/dataset\">http://headctstudy.qure.ai/dataset</a>) who have so graciously made them available!</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3267348,
              "author_name": "zxc152",
              "author_url": "",
              "post_date": "2025-08-11T01:55:37.127000",
              "content": "<p>Dear Organizers, the website (<a href=\"http://headctstudy.qure.ai/dataset\" target=\"_blank\">http://headctstudy.qure.ai/dataset</a>) isn't working. Where can I download the annotated dataset instead? I'd appreciate your assistance with this.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 652402,
          "author_name": "ArvindVepa",
          "author_url": "",
          "post_date": "2019-10-18T20:08:10.090000",
          "content": "<p>oh i see. thanks.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 655831,
          "author_name": "yangDDD",
          "author_url": "",
          "post_date": "2019-10-23T14:50:28.653000",
          "content": "<p>unable to download the data, is it still available?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 658748,
          "author_name": "Nicholas Lyu",
          "author_url": "",
          "post_date": "2019-10-26T13:31:08.633000",
          "content": "<p><a href=\"/yangddd\">@yangddd</a> I just downloaded it today. You may need some tricks to download it if you are in China</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 660000,
          "author_name": "yangDDD",
          "author_url": "",
          "post_date": "2019-10-28T14:52:39.660000",
          "content": "<p>thanks</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1600883,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-11-30T19:30:48.657000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 658861,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-10-26T16:35:33.237000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 658743,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-10-26T13:26:04.983000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "652306": "Reposted on external data thread: We are excited to see the progress that the Kaggle teams are making! Eduardo Reis, MD and his group from Hospital Israelita Albert Einstein, São Paulo, BR,  have annotated the qure.ai CQ500 dataset with bounding boxes for the different types of hemorrhage. They annotated the thick sliced series within each exam and extrapolated the boxes to the thinner sliced series to expand the available data. They are in the process of writing up their findings for publication. The dataset is made available to the Kaggle community and can be viewed on the MD.ai platform: https://public.md.ai/annotator/project/Y2qr6vqv Eduardo's annotations are on labelgroup 4 - BrainHemX and the bounding boxes can be downloaded using the link below. The original images are hosted by qure.ai at http://headctstudy.qure.ai/dataset licensed under the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License",
    "652499": "I'm slightly concerned about releasing that big dataset for training purposes two weeks before the deadline. Am I the only one?",
    "655168": "If anyone has problems reading dicom files using pydicom, this can be a solution. \nThe embedded images can not be read by PIL/Pillow and GDCM is needed instead.\n\nhttps://github.com/HealthplusAI/python3-gdcm\n\n```\nimport pydicom.pixel_data_handlers.gdcm_handler as gdcm_handler\npydicom.config.image_handlers = [None, gdcm_handler]\n```\n",
    "652448": "I'd like to redirect a question from [external data thread](https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109274#649282) here.\nAre we allowed to use BY-NC-SA-licensed data for the competition?",
    "654459": "Has anyone evaluated this dataset with a trained model? ",
    "652379": "I almost had a heart-attack seeing Albert Einstein doing data annotation . Then after reading through i found it is Hospital Israelita Albert Einstein 😐 . Anyway, thanks for the awesome work.",
    "652373": "Right! Thank you, Anouk. The competitors now have image-level annotations for the qure.ai CQ500 external dataset. Perhaps it can help.\n\nMany thanks to everyone who also participated in the work: Felipe Nascimento, Fernando Secol, Mateus Aranha, Marcelo Felix, Edson Amaro, Birajara Machado, and Anouk for providing all technical support.",
    "652895": "Some tips:\n\nNote that qure.ai dataset contains more than one series per exam: thick and thin slices, bone and soft-tissue filters, pre and post-contrast, etc. \n\nThe neuroradiologists drew the bounding boxes in the **soft-tissue thick slices** . Then the bbx were extrapolated to the other series, which don’t always result in a perfect match. \n\nFor the extrapolated bbx, what I would be most careful about are the images with 'no hemorrhage' which are adjacent to others with hemorrhage. They can instead, contain hemorrhage. (and vice-versa)\n\nWe will soon add a global label indicating which series the bbxs were originally handmade @ https://public.md.ai/annotator/project/Y2qr6vqv",
    "658745": "@anoukstein 2 Questions for this additional data:\n1) From what I have seem in the CQ500 dataset, there are 5 labels. The annotations above are charming, but what type of hemorrhage does it correspond to? `Eduardo's annotations are on labelgroup 4 - BrainHemX`--so I take it is label 4?\n\n2) CQ500 dataset is missing important descriptions. For a single patient, there are many scans. e.g. patient 0 contains \"CT 4cc sec 150cc D3D on\", \"CT 4cc sec 150cc D3D on-2\" ... \"CT Plain\", \"CT PLAIN THIN\". How are these folders different? So data in which folder is consistent with the competition data? I found that there is a folder in each patient which contains dcm with names which includes all other folders. The additional data is very frustrating. Please explain :)",
    "657038": "Did this help anyone's performance?",
    "653164": "I would be super happy if CQ500 dataset are provided in the same format as original dataset;)",
    "652346": "Dr Reis' team will be supplying an updated version of the boxes with minor changes where the boxes skipped images on a few cases. They wanted to get out as much data as possible given the time constraint of this competition so decided to release this version rather than delay. The csv  contains 38940 hemorrhage bounding boxes on the 491 studies from CQ500 qure.ai.",
    "654352": "@anoukstein or anyone else, Does this dataset have the target labels we use in our competition ? `epidural,intraparenchymal,intraventricular,subarachnoid,subdural` or is there a way to map the labels from this datset to our dataset ?",
    "653322": "Hi, it seems that the images are not in standard dicom format, either they are compressed or something else (the buffer size is not the full image size). What am I missing? thanks",
    "652405": "Are these image slices labeled the same way as the competition dataset at the slice level or patient/study level? For example, in competition dataset, a patient's slices may have no hemorrhage in certain slices but hemorrhage in others.",
    "652357": "Awesome work, @epreis! @drvidurmahajan , your work is cross-pollinating!",
    "653376": "Is it legal and under rules to make mask annotation based on your bbox annotation?",
    "652673": "Will this new dataset be made available as a Kaggle Dataset?",
    "652705": "How to download it on terminal?  \n```\nwget hogehoge\n```\n",
    "972123": "Hi @anoukstein, I have been trying to download the annotations from https://public.md.ai/annotator/project/Y2qr6vqv\nBut the export option is not enabled. Can you let me know how to download the data?",
    "664918": "The boxes file was updated and it broke my code. Be careful :) ",
    "652372": "Hi, how do you export the images/labels? When I click on the export tab and the export window pops up, I can't click on the export button. Thanks for all the great work you guys are doing.",
    "1600883": "",
    "658861": "",
    "658743": ""
  }
}