{
  "id": 110840,
  "title": "Varying .png resolutions",
  "url": "/competitions/rsna-intracranial-hemorrhage-detection/discussion/110840",
  "author_name": "Tim Yee",
  "post_date": "2019-10-01T13:25:29.233000",
  "votes": 43,
  "comment_count": 39,
  "views": 0,
  "content": "<p>I am sharing my pre-processed .png files for those who may find it useful. It takes several hours to pre-process at each resolution and upload them to kaggle datasets. If you have issues with the datasets or questions about them, let me know and I will do my best to reply. All training data leave out a corrupted image \"ID_6431af929.dcm\". Pre-processing involved applying a linear transformation as discussed in <a href=\"https://www.kaggle.com/omission/eda-view-dicom-images-with-correct-windowing\">https://www.kaggle.com/omission/eda-view-dicom-images-with-correct-windowing</a>. I am not a subject matter expert on the \"windowing\" applied so I will defer to someone who is: <a href=\"https://www.kaggle.com/dcstang/see-like-a-radiologist-with-systematic-windowing\">https://www.kaggle.com/dcstang/see-like-a-radiologist-with-systematic-windowing</a>.</p>\n\n<p>Following the linear transform, resizing was done using albumentations with lanczos interpolation because that is what I felt was best (it may or may not be). Therefore, use the lower resolutions at your own discretion.</p>\n\n<p>224 x 224 <a href=\"https://www.kaggle.com/teeyee314/rsnatrain224\">train</a> <a href=\"https://www.kaggle.com/teeyee314/rsnatest224\">test</a>\n240 x 240 <a href=\"https://www.kaggle.com/teeyee314/rsnatrain240\">train</a> <a href=\"https://www.kaggle.com/teeyee314/rsnatest240\">test</a>\n260 x 260 <a href=\"https://www.kaggle.com/teeyee314/rsnatrain260\">train</a> <a href=\"https://www.kaggle.com/teeyee314/rsnatest260\">test</a>\n300 x 300 <a href=\"https://www.kaggle.com/teeyee314/rsnatrain300\">train</a> <a href=\"https://www.kaggle.com/teeyee314/rsnatest300\">test</a>\n380 x 380 <a href=\"https://www.kaggle.com/teeyee314/rsnatrain380\">train</a> <a href=\"https://www.kaggle.com/teeyee314/rsnatest380\">test</a>\n456 x 456 <a href=\"https://www.kaggle.com/teeyee314/rsnatrain456\">train</a> <a href=\"https://www.kaggle.com/teeyee314/rsnatest456\">test</a>\n512 x 512 <a href=\"https://www.kaggle.com/teeyee314/rsnatrain512\">train</a> <a href=\"https://www.kaggle.com/teeyee314/rsnatrain512b\">trainb</a> <a href=\"https://www.kaggle.com/teeyee314/rsnatest512\">test</a></p>\n\n<p><br></p>\n\n<p>_please note that 512 train is split into 2 zipped containers. this was intentional because of the hard 20GB limit on public datasets. once you've downloaded both containers, unzip both of them, and simply cd into the train512b folder and type 'mv * ../train512' and that command will move all files in train512b over to train512._</p>\n\n<p>Oct 21 update - there are &lt; 300 images in 512x512 that are not actually 512x512 and I didn't resize them to 512x512. Please make sure you resize your images to take on 512x512 when using that dataset.</p>",
  "messages": [
    {
      "id": 638037,
      "postDate": "2019-10-01T13:25:29.233Z",
      "content": "<p>I am sharing my pre-processed .png files for those who may find it useful. It takes several hours to pre-process at each resolution and upload them to kaggle datasets. If you have issues with the datasets or questions about them, let me know and I will do my best to reply. All training data leave out a corrupted image \"ID_6431af929.dcm\". Pre-processing involved applying a linear transformation as discussed in <a href=\"https://www.kaggle.com/omission/eda-view-dicom-images-with-correct-windowing\">https://www.kaggle.com/omission/eda-view-dicom-images-with-correct-windowing</a>. I am not a subject matter expert on the \"windowing\" applied so I will defer to someone who is: <a href=\"https://www.kaggle.com/dcstang/see-like-a-radiologist-with-systematic-windowing\">https://www.kaggle.com/dcstang/see-like-a-radiologist-with-systematic-windowing</a>.</p>\n\n<p>Following the linear transform, resizing was done using albumentations with lanczos interpolation because that is what I felt was best (it may or may not be). Therefore, use the lower resolutions at your own discretion.</p>\n\n<p>224 x 224 <a href=\"https://www.kaggle.com/teeyee314/rsnatrain224\">train</a> <a href=\"https://www.kaggle.com/teeyee314/rsnatest224\">test</a>\n240 x 240 <a href=\"https://www.kaggle.com/teeyee314/rsnatrain240\">train</a> <a href=\"https://www.kaggle.com/teeyee314/rsnatest240\">test</a>\n260 x 260 <a href=\"https://www.kaggle.com/teeyee314/rsnatrain260\">train</a> <a href=\"https://www.kaggle.com/teeyee314/rsnatest260\">test</a>\n300 x 300 <a href=\"https://www.kaggle.com/teeyee314/rsnatrain300\">train</a> <a href=\"https://www.kaggle.com/teeyee314/rsnatest300\">test</a>\n380 x 380 <a href=\"https://www.kaggle.com/teeyee314/rsnatrain380\">train</a> <a href=\"https://www.kaggle.com/teeyee314/rsnatest380\">test</a>\n456 x 456 <a href=\"https://www.kaggle.com/teeyee314/rsnatrain456\">train</a> <a href=\"https://www.kaggle.com/teeyee314/rsnatest456\">test</a>\n512 x 512 <a href=\"https://www.kaggle.com/teeyee314/rsnatrain512\">train</a> <a href=\"https://www.kaggle.com/teeyee314/rsnatrain512b\">trainb</a> <a href=\"https://www.kaggle.com/teeyee314/rsnatest512\">test</a></p>\n\n<p><br></p>\n\n<p>_please note that 512 train is split into 2 zipped containers. this was intentional because of the hard 20GB limit on public datasets. once you've downloaded both containers, unzip both of them, and simply cd into the train512b folder and type 'mv * ../train512' and that command will move all files in train512b over to train512._</p>\n\n<p>Oct 21 update - there are &lt; 300 images in 512x512 that are not actually 512x512 and I didn't resize them to 512x512. Please make sure you resize your images to take on 512x512 when using that dataset.</p>",
      "rawMarkdown": "I am sharing my pre-processed .png files for those who may find it useful. It takes several hours to pre-process at each resolution and upload them to kaggle datasets. If you have issues with the datasets or questions about them, let me know and I will do my best to reply. All training data leave out a corrupted image \"ID_6431af929.dcm\". Pre-processing involved applying a linear transformation as discussed in https://www.kaggle.com/omission/eda-view-dicom-images-with-correct-windowing. I am not a subject matter expert on the \"windowing\" applied so I will defer to someone who is: https://www.kaggle.com/dcstang/see-like-a-radiologist-with-systematic-windowing.\n\nFollowing the linear transform, resizing was done using albumentations with lanczos interpolation because that is what I felt was best (it may or may not be). Therefore, use the lower resolutions at your own discretion.\n\n\n224 x 224 [train](https://www.kaggle.com/teeyee314/rsnatrain224) [test](https://www.kaggle.com/teeyee314/rsnatest224)\n240 x 240 [train](https://www.kaggle.com/teeyee314/rsnatrain240) [test](https://www.kaggle.com/teeyee314/rsnatest240)\n260 x 260 [train](https://www.kaggle.com/teeyee314/rsnatrain260) [test](https://www.kaggle.com/teeyee314/rsnatest260)\n300 x 300 [train](https://www.kaggle.com/teeyee314/rsnatrain300) [test](https://www.kaggle.com/teeyee314/rsnatest300)\n380 x 380 [train](https://www.kaggle.com/teeyee314/rsnatrain380) [test](https://www.kaggle.com/teeyee314/rsnatest380)\n456 x 456 [train](https://www.kaggle.com/teeyee314/rsnatrain456) [test](https://www.kaggle.com/teeyee314/rsnatest456)\n512 x 512 [train](https://www.kaggle.com/teeyee314/rsnatrain512) [trainb](https://www.kaggle.com/teeyee314/rsnatrain512b) [test](https://www.kaggle.com/teeyee314/rsnatest512)\n\n<br>\n\n_please note that 512 train is split into 2 zipped containers. this was intentional because of the hard 20GB limit on public datasets. once you've downloaded both containers, unzip both of them, and simply cd into the train512b folder and type 'mv * ../train512' and that command will move all files in train512b over to train512._\n\nOct 21 update - there are &lt; 300 images in 512x512 that are not actually 512x512 and I didn't resize them to 512x512. Please make sure you resize your images to take on 512x512 when using that dataset.\n",
      "votes": 43
    },
    {
      "id": 647817,
      "postDate": "2019-10-13T10:38:33.170Z",
      "content": "<p>Hello Tim,</p>\n\n<p>thanks for your work, very much appreciated. I was wondering, since this is a 2 stages competition, would you mind releasing the kernel you used to create these datasets? This way people will be able to create their own Stage 2 dataset in the same way you used in Stage 1.</p>",
      "rawMarkdown": "Hello Tim,\n\nthanks for your work, very much appreciated. I was wondering, since this is a 2 stages competition, would you mind releasing the kernel you used to create these datasets? This way people will be able to create their own Stage 2 dataset in the same way you used in Stage 1.",
      "votes": 1,
      "replies": [
        {
          "id": 648326,
          "postDate": "2019-10-14T03:15:01.307Z",
          "content": "<p><a href=\"/juliencs\">@juliencs</a> <a href=\"https://www.kaggle.com/taindow/pytorch-efficientnet-b0-benchmark#634868\">https://www.kaggle.com/taindow/pytorch-efficientnet-b0-benchmark#634868</a></p>",
          "rawMarkdown": "@juliencs https://www.kaggle.com/taindow/pytorch-efficientnet-b0-benchmark#634868",
          "votes": 1
        },
        {
          "id": 648438,
          "postDate": "2019-10-14T07:14:33.093Z",
          "content": "<p>Thank you :)</p>",
          "rawMarkdown": "Thank you :)"
        },
        {
          "id": 648814,
          "postDate": "2019-10-14T16:28:56.587Z",
          "content": "<p>Just to be sure, the only line added to this script is the interpolation line to get the images to 224*224 right? And you also removed the parallel processing logic. Do you plan to share your own script? </p>",
          "rawMarkdown": "Just to be sure, the only line added to this script is the interpolation line to get the images to 224*224 right? And you also removed the parallel processing logic. Do you plan to share your own script? "
        }
      ]
    },
    {
      "id": 638127,
      "postDate": "2019-10-01T14:56:51.777Z",
      "content": "<p>thanks a ton for putting so much efforts on that and for this post <a href=\"/teeyee314\">@teeyee314</a> \nkaggle will always be a good place to learn data science as long as it will have dedicated people like you</p>",
      "rawMarkdown": "thanks a ton for putting so much efforts on that and for this post @teeyee314 \nkaggle will always be a good place to learn data science as long as it will have dedicated people like you",
      "votes": 1
    },
    {
      "id": 650419,
      "postDate": "2019-10-16T11:06:53.080Z",
      "content": "<p>could you tell me where's the 224x224 datasets with no windowing?\nMuch Thanks</p>",
      "rawMarkdown": "could you tell me where's the 224x224 datasets with no windowing?\nMuch Thanks"
    },
    {
      "id": 648425,
      "postDate": "2019-10-14T06:49:10.990Z",
      "content": "<p>Can kaggle api download your datasets？</p>",
      "rawMarkdown": "Can kaggle api download your datasets？",
      "replies": [
        {
          "id": 648712,
          "postDate": "2019-10-14T14:25:40.787Z",
          "content": "<p>yes </p>\n\n<blockquote>\n  <p>kaggle datasets download -d teeyee314/rsnatest224\n  kaggle datasets download -d teeyee314/rsnatrain224</p>\n</blockquote>",
          "rawMarkdown": "yes \n&gt; kaggle datasets download -d teeyee314/rsnatest224\nkaggle datasets download -d teeyee314/rsnatrain224"
        }
      ]
    },
    {
      "id": 647798,
      "postDate": "2019-10-13T09:44:21.870Z",
      "content": "<p>could anyone tell me what's the difference between train and trainb?</p>",
      "rawMarkdown": "could anyone tell me what's the difference between train and trainb?",
      "replies": [
        {
          "id": 648367,
          "postDate": "2019-10-14T04:40:08.300Z",
          "content": "<p>I split up the files 512 resolutions because the zipped container was too large to upload. They should be joined together into a single folder once you download both.</p>",
          "rawMarkdown": "I split up the files 512 resolutions because the zipped container was too large to upload. They should be joined together into a single folder once you download both."
        },
        {
          "id": 648411,
          "postDate": "2019-10-14T06:24:01.020Z",
          "content": "<p>Much thanks for your kind help and explanation~!</p>",
          "rawMarkdown": "Much thanks for your kind help and explanation~!"
        }
      ]
    },
    {
      "id": 647695,
      "postDate": "2019-10-13T05:03:10.567Z",
      "content": "<p><a href=\"/teeyee314\">@teeyee314</a>, thanks for this! I also want to get my hands on my own processing - did you add an additional line to the script used <a href=\"https://www.kaggle.com/taindow/generate-images-train\">here</a> on resizing with interpolation? I am referring to what you mentioned <a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109747633405\">here</a> and I would like to learn to do it myself.</p>",
      "rawMarkdown": "@teeyee314, thanks for this! I also want to get my hands on my own processing - did you add an additional line to the script used [here](https://www.kaggle.com/taindow/generate-images-train) on resizing with interpolation? I am referring to what you mentioned [here](https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109747633405) and I would like to learn to do it myself.",
      "replies": [
        {
          "id": 647951,
          "postDate": "2019-10-13T14:20:05.113Z",
          "content": "<p>Yes. I used a different interpolation than the script you are referring to.</p>",
          "rawMarkdown": "Yes. I used a different interpolation than the script you are referring to."
        }
      ]
    },
    {
      "id": 646950,
      "postDate": "2019-10-11T23:07:55.440Z",
      "content": "<p>I found some images being black or almost black, as was already discussed. But also some images appear to be corrupted and contain NaNs. Is it just me or are they really corrupted during convertation/upload?</p>",
      "rawMarkdown": "I found some images being black or almost black, as was already discussed. But also some images appear to be corrupted and contain NaNs. Is it just me or are they really corrupted during convertation/upload?",
      "replies": [
        {
          "id": 646974,
          "postDate": "2019-10-11T23:30:38.870Z",
          "content": "<p>can you please be more specific? which images in which resolution sizes are nans? screenshots?</p>",
          "rawMarkdown": "can you please be more specific? which images in which resolution sizes are nans? screenshots?"
        },
        {
          "id": 647840,
          "postDate": "2019-10-13T11:11:45.867Z",
          "content": "<p>For example. ID_1abc92d1f and ID_4349f759c are totally black in size 224. I believe there may be other IDs with black or NaN. </p>",
          "rawMarkdown": "For example. ID_1abc92d1f and ID_4349f759c are totally black in size 224. I believe there may be other IDs with black or NaN. "
        },
        {
          "id": 647964,
          "postDate": "2019-10-13T14:44:57.613Z",
          "content": "<p><a href=\"/cateek\">@cateek</a> It seems that you are onto something. Perhaps I will figure it out and post an update. I don't expect to have conclusive findings immediately. I'll have to take a look at this on a larger scale. For the record, my images are not the only one that have this \"issue\" because others have posted their images using the windowing method that I have copied. Until I figure it out, it is best to do your own data analysis as well. Hope this helps.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1894505%2Ffb23d33dba812be5d8975fcc66190941%2FID_4349f759c.PNG?generation=1570977587958976&amp;alt=media\" alt=\"ID_4349f759c\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1894505%2Fbf934fb44d9059a2a923c621d23bf8a8%2FID_1abc92d1f.PNG?generation=1570977588066890&amp;alt=media\" alt=\"ID_1abc92d1f\"></p>",
          "rawMarkdown": "@cateek It seems that you are onto something. Perhaps I will figure it out and post an update. I don't expect to have conclusive findings immediately. I'll have to take a look at this on a larger scale. For the record, my images are not the only one that have this \"issue\" because others have posted their images using the windowing method that I have copied. Until I figure it out, it is best to do your own data analysis as well. Hope this helps.\n\n![ID_4349f759c](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1894505%2Ffb23d33dba812be5d8975fcc66190941%2FID_4349f759c.PNG?generation=1570977587958976&amp;alt=media)\n![ID_1abc92d1f](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1894505%2Fbf934fb44d9059a2a923c621d23bf8a8%2FID_1abc92d1f.PNG?generation=1570977588066890&amp;alt=media)\n",
          "votes": 1
        },
        {
          "id": 647974,
          "postDate": "2019-10-13T15:00:41.050Z",
          "content": "<p><a href=\"/teeyee314\">@teeyee314</a> Thanks for answer. Yes, yours are not the only ones, It seems like there are some dicom files that result in empty image if specified window is applied. </p>",
          "rawMarkdown": "@teeyee314 Thanks for answer. Yes, yours are not the only ones, It seems like there are some dicom files that result in empty image if specified window is applied. "
        }
      ]
    },
    {
      "id": 646045,
      "postDate": "2019-10-10T19:24:19.057Z",
      "content": "<p><a href=\"/teeyee314\">@teeyee314</a> One question I have (I am new to Kaggle): are we allowed to submit and go to second phase using png data instead of the raw data given to us in the beggining?</p>",
      "rawMarkdown": "@teeyee314 One question I have (I am new to Kaggle): are we allowed to submit and go to second phase using png data instead of the raw data given to us in the beggining?",
      "replies": [
        {
          "id": 646071,
          "postDate": "2019-10-10T20:19:54.997Z",
          "content": "<p>What you need to do at the end of stage 1 is zip your model's weights and the notebooks/scripts that generate the stage 2 submission so that the model you are submitting as your final model can inference the test set for stage 2. </p>",
          "rawMarkdown": "What you need to do at the end of stage 1 is zip your model's weights and the notebooks/scripts that generate the stage 2 submission so that the model you are submitting as your final model can inference the test set for stage 2. ",
          "votes": 1
        },
        {
          "id": 646238,
          "postDate": "2019-10-11T02:48:14.887Z",
          "content": "<p><a href=\"/teeyee314\">@teeyee314</a> When you use a dataset provided by another person, do you download it and upload from your pc or do you use it directly from the other person's dataset (I don't even know if that's possible)? Thank you for answering.</p>",
          "rawMarkdown": "@teeyee314 When you use a dataset provided by another person, do you download it and upload from your pc or do you use it directly from the other person's dataset (I don't even know if that's possible)? Thank you for answering."
        },
        {
          "id": 646246,
          "postDate": "2019-10-11T03:04:39.920Z",
          "content": "<p>Using kaggle api, you can easily download a dataset to any cloud provider.</p>\n\n<p><a href=\"https://github.com/Kaggle/kaggle-api\">https://github.com/Kaggle/kaggle-api</a></p>\n\n<p>In case of kaggle kernels, you can directly add the dataset to your kernel.</p>",
          "rawMarkdown": "Using kaggle api, you can easily download a dataset to any cloud provider.\n\nhttps://github.com/Kaggle/kaggle-api\n\nIn case of kaggle kernels, you can directly add the dataset to your kernel.",
          "votes": 1
        },
        {
          "id": 646643,
          "postDate": "2019-10-11T14:03:34.653Z",
          "content": "<p><a href=\"/atikur\">@atikur</a> <a href=\"/teeyee314\">@teeyee314</a>  Thank you for your answer. When I time 'mv * ../train512', to join train with trainb, I am having this error mv: Argument list too long... Do you know how can I overcome this problem in Kaggle?</p>",
          "rawMarkdown": "@atikur @teeyee314  Thank you for your answer. When I time 'mv * ../train512', to join train with trainb, I am having this error mv: Argument list too long... Do you know how can I overcome this problem in Kaggle?"
        },
        {
          "id": 646919,
          "postDate": "2019-10-11T22:08:20.280Z",
          "content": "<p>are you using my 512 dataset locally, kaggle kernel, google colab, or gcp? If you're training on kaggle, training 512x512 may not be the most efficient way to utilize your time and gpu quota. Plenty of people at the top of LB are reaching &lt; 0.08 with 224x224 and if you're really good, you can reach 0.07 with 224x224.</p>",
          "rawMarkdown": "are you using my 512 dataset locally, kaggle kernel, google colab, or gcp? If you're training on kaggle, training 512x512 may not be the most efficient way to utilize your time and gpu quota. Plenty of people at the top of LB are reaching &lt; 0.08 with 224x224 and if you're really good, you can reach 0.07 with 224x224.",
          "votes": 1
        },
        {
          "id": 646921,
          "postDate": "2019-10-11T22:12:23.040Z",
          "content": "<p>On Kaggle. I will try to do as you say, but until now I only tried the original dataset, so I was planning to see the difference in time using each type of png size.</p>",
          "rawMarkdown": "On Kaggle. I will try to do as you say, but until now I only tried the original dataset, so I was planning to see the difference in time using each type of png size."
        },
        {
          "id": 648327,
          "postDate": "2019-10-14T03:15:22.283Z",
          "content": "<p>could you tell me what's the difference between train and trainb?\nThanks</p>",
          "rawMarkdown": "could you tell me what's the difference between train and trainb?\nThanks"
        }
      ]
    },
    {
      "id": 643725,
      "postDate": "2019-10-07T20:53:21.677Z",
      "content": "<p>Thank you for sharing. Did you normalize the images during your process? Or, is that something we need to do?</p>",
      "rawMarkdown": "Thank you for sharing. Did you normalize the images during your process? Or, is that something we need to do?",
      "replies": [
        {
          "id": 643753,
          "postDate": "2019-10-07T21:43:34.217Z",
          "content": "<p>No normalization was performed. The pixel data was extracted and windowed(transformed) per the kernel linked above. My pre-processing was similar to <a href=\"https://www.kaggle.com/taindow/generate-images?scriptVersionId=21147396\">this</a>. </p>",
          "rawMarkdown": "No normalization was performed. The pixel data was extracted and windowed(transformed) per the kernel linked above. My pre-processing was similar to [this](https://www.kaggle.com/taindow/generate-images?scriptVersionId=21147396). "
        },
        {
          "id": 643777,
          "postDate": "2019-10-07T23:00:12.140Z",
          "content": "<p>Thank you. </p>",
          "rawMarkdown": "Thank you. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 643159,
      "postDate": "2019-10-07T07:38:11.303Z",
      "content": "<p>hi <a href=\"/teeyee314\">@teeyee314</a>  i want to use your dataset in colab,will i have to download your data and upload in colab or there is any way to directly get it in colab from kaggle? asking this question because this is not original dataset and can't do api call for downloading this data in colab</p>",
      "rawMarkdown": "hi @teeyee314  i want to use your dataset in colab,will i have to download your data and upload in colab or there is any way to directly get it in colab from kaggle? asking this question because this is not original dataset and can't do api call for downloading this data in colab",
      "replies": [
        {
          "id": 643177,
          "postDate": "2019-10-07T08:00:10.627Z",
          "content": "<p>You can use kaggle api to download datasets. For example -</p>\n\n<p><code>\n!kaggle datasets download teeyee314/rsnatrain240\n</code></p>\n\n<p>You will find details about <code>Datasets</code> api here -</p>\n\n<p><a href=\"https://github.com/Kaggle/kaggle-api\">https://github.com/Kaggle/kaggle-api</a></p>",
          "rawMarkdown": "You can use kaggle api to download datasets. For example -\n\n```\n!kaggle datasets download teeyee314/rsnatrain240\n```\n\nYou will find details about `Datasets` api here -\n\nhttps://github.com/Kaggle/kaggle-api",
          "votes": 1
        },
        {
          "id": 643183,
          "postDate": "2019-10-07T08:03:47.123Z",
          "content": "<p>good to know,thanks for letting me know</p>",
          "rawMarkdown": "good to know,thanks for letting me know"
        }
      ]
    },
    {
      "id": 642540,
      "postDate": "2019-10-06T08:59:37.677Z",
      "content": "<p>512 train_a isn't available</p>",
      "rawMarkdown": "512 train_a isn't available",
      "replies": [
        {
          "id": 642660,
          "postDate": "2019-10-06T12:50:16.643Z",
          "content": "<p>I'll upload it as soon as I can. I may need to redo the split and I have my other things tying up my resources, namely limited IO on my harddrive that stores the files. I can't give you a turnaround time but you can feel free to use other resolutions until I have the resources to upload.</p>",
          "rawMarkdown": "I'll upload it as soon as I can. I may need to redo the split and I have my other things tying up my resources, namely limited IO on my harddrive that stores the files. I can't give you a turnaround time but you can feel free to use other resolutions until I have the resources to upload.",
          "votes": 1
        },
        {
          "id": 642968,
          "postDate": "2019-10-06T21:38:38.963Z",
          "content": "<p><a href=\"/max6296\">@max6296</a> 512x512 is uploaded. If you downloaded train512b before, re-download it. I created a new split.</p>",
          "rawMarkdown": "@max6296 512x512 is uploaded. If you downloaded train512b before, re-download it. I created a new split.",
          "votes": 1
        },
        {
          "id": 642984,
          "postDate": "2019-10-06T22:27:40.210Z",
          "content": "<p>Thank you so much!</p>",
          "rawMarkdown": "Thank you so much!"
        },
        {
          "id": 643128,
          "postDate": "2019-10-07T06:37:08.097Z",
          "content": "<p>Thanks so much for your effort!\nI found that train 260 isn't uploaded yet.</p>",
          "rawMarkdown": "Thanks so much for your effort!\nI found that train 260 isn't uploaded yet."
        },
        {
          "id": 643681,
          "postDate": "2019-10-07T19:21:27.690Z",
          "content": "<p><a href=\"/kenho211\">@kenho211</a> I re-uploaded test260 to keep the links uniform. train260 should be uploaded... </p>",
          "rawMarkdown": "@kenho211 I re-uploaded test260 to keep the links uniform. train260 should be uploaded... "
        }
      ]
    },
    {
      "id": 649459,
      "postDate": "2019-10-15T11:37:30.317Z",
      "content": "<p>Thanks for sharing.</p>",
      "rawMarkdown": "Thanks for sharing."
    }
  ],
  "comments": [
    {
      "id": 647817,
      "author_name": "juliencs",
      "author_url": "",
      "post_date": "2019-10-13T10:38:33.170000",
      "content": "<p>Hello Tim,</p>\n\n<p>thanks for your work, very much appreciated. I was wondering, since this is a 2 stages competition, would you mind releasing the kernel you used to create these datasets? This way people will be able to create their own Stage 2 dataset in the same way you used in Stage 1.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 648326,
          "author_name": "Tim Yee",
          "author_url": "",
          "post_date": "2019-10-14T03:15:01.307000",
          "content": "<p><a href=\"/juliencs\">@juliencs</a> <a href=\"https://www.kaggle.com/taindow/pytorch-efficientnet-b0-benchmark#634868\">https://www.kaggle.com/taindow/pytorch-efficientnet-b0-benchmark#634868</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 648438,
          "author_name": "juliencs",
          "author_url": "",
          "post_date": "2019-10-14T07:14:33.093000",
          "content": "<p>Thank you :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 648814,
          "author_name": "ChinHuiC",
          "author_url": "",
          "post_date": "2019-10-14T16:28:56.587000",
          "content": "<p>Just to be sure, the only line added to this script is the interpolation line to get the images to 224*224 right? And you also removed the parallel processing logic. Do you plan to share your own script? </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 638127,
      "author_name": "Mobassir",
      "author_url": "",
      "post_date": "2019-10-01T14:56:51.777000",
      "content": "<p>thanks a ton for putting so much efforts on that and for this post <a href=\"/teeyee314\">@teeyee314</a> \nkaggle will always be a good place to learn data science as long as it will have dedicated people like you</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 650419,
      "author_name": "pupil5",
      "author_url": "",
      "post_date": "2019-10-16T11:06:53.080000",
      "content": "<p>could you tell me where's the 224x224 datasets with no windowing?\nMuch Thanks</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 648425,
      "author_name": "pupil3",
      "author_url": "",
      "post_date": "2019-10-14T06:49:10.990000",
      "content": "<p>Can kaggle api download your datasets？</p>",
      "votes": 0,
      "replies": [
        {
          "id": 648712,
          "author_name": "Tim Yee",
          "author_url": "",
          "post_date": "2019-10-14T14:25:40.787000",
          "content": "<p>yes </p>\n\n<blockquote>\n  <p>kaggle datasets download -d teeyee314/rsnatest224\n  kaggle datasets download -d teeyee314/rsnatrain224</p>\n</blockquote>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 647798,
      "author_name": "pupil3",
      "author_url": "",
      "post_date": "2019-10-13T09:44:21.870000",
      "content": "<p>could anyone tell me what's the difference between train and trainb?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 648367,
          "author_name": "Tim Yee",
          "author_url": "",
          "post_date": "2019-10-14T04:40:08.300000",
          "content": "<p>I split up the files 512 resolutions because the zipped container was too large to upload. They should be joined together into a single folder once you download both.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 648411,
          "author_name": "pupil3",
          "author_url": "",
          "post_date": "2019-10-14T06:24:01.020000",
          "content": "<p>Much thanks for your kind help and explanation~!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 647695,
      "author_name": "ChinHuiC",
      "author_url": "",
      "post_date": "2019-10-13T05:03:10.567000",
      "content": "<p><a href=\"/teeyee314\">@teeyee314</a>, thanks for this! I also want to get my hands on my own processing - did you add an additional line to the script used <a href=\"https://www.kaggle.com/taindow/generate-images-train\">here</a> on resizing with interpolation? I am referring to what you mentioned <a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109747633405\">here</a> and I would like to learn to do it myself.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 647951,
          "author_name": "Tim Yee",
          "author_url": "",
          "post_date": "2019-10-13T14:20:05.113000",
          "content": "<p>Yes. I used a different interpolation than the script you are referring to.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 646950,
      "author_name": "Eek The Cat",
      "author_url": "",
      "post_date": "2019-10-11T23:07:55.440000",
      "content": "<p>I found some images being black or almost black, as was already discussed. But also some images appear to be corrupted and contain NaNs. Is it just me or are they really corrupted during convertation/upload?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 646974,
          "author_name": "Tim Yee",
          "author_url": "",
          "post_date": "2019-10-11T23:30:38.870000",
          "content": "<p>can you please be more specific? which images in which resolution sizes are nans? screenshots?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 647840,
          "author_name": "Eek The Cat",
          "author_url": "",
          "post_date": "2019-10-13T11:11:45.867000",
          "content": "<p>For example. ID_1abc92d1f and ID_4349f759c are totally black in size 224. I believe there may be other IDs with black or NaN. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 647964,
          "author_name": "Tim Yee",
          "author_url": "",
          "post_date": "2019-10-13T14:44:57.613000",
          "content": "<p><a href=\"/cateek\">@cateek</a> It seems that you are onto something. Perhaps I will figure it out and post an update. I don't expect to have conclusive findings immediately. I'll have to take a look at this on a larger scale. For the record, my images are not the only one that have this \"issue\" because others have posted their images using the windowing method that I have copied. Until I figure it out, it is best to do your own data analysis as well. Hope this helps.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1894505%2Ffb23d33dba812be5d8975fcc66190941%2FID_4349f759c.PNG?generation=1570977587958976&amp;alt=media\" alt=\"ID_4349f759c\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1894505%2Fbf934fb44d9059a2a923c621d23bf8a8%2FID_1abc92d1f.PNG?generation=1570977588066890&amp;alt=media\" alt=\"ID_1abc92d1f\"></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 647974,
          "author_name": "Eek The Cat",
          "author_url": "",
          "post_date": "2019-10-13T15:00:41.050000",
          "content": "<p><a href=\"/teeyee314\">@teeyee314</a> Thanks for answer. Yes, yours are not the only ones, It seems like there are some dicom files that result in empty image if specified window is applied. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 646045,
      "author_name": "Gurgel",
      "author_url": "",
      "post_date": "2019-10-10T19:24:19.057000",
      "content": "<p><a href=\"/teeyee314\">@teeyee314</a> One question I have (I am new to Kaggle): are we allowed to submit and go to second phase using png data instead of the raw data given to us in the beggining?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 646071,
          "author_name": "Tim Yee",
          "author_url": "",
          "post_date": "2019-10-10T20:19:54.997000",
          "content": "<p>What you need to do at the end of stage 1 is zip your model's weights and the notebooks/scripts that generate the stage 2 submission so that the model you are submitting as your final model can inference the test set for stage 2. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 646238,
          "author_name": "Gurgel",
          "author_url": "",
          "post_date": "2019-10-11T02:48:14.887000",
          "content": "<p><a href=\"/teeyee314\">@teeyee314</a> When you use a dataset provided by another person, do you download it and upload from your pc or do you use it directly from the other person's dataset (I don't even know if that's possible)? Thank you for answering.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 646246,
          "author_name": "Atikur Rahman",
          "author_url": "",
          "post_date": "2019-10-11T03:04:39.920000",
          "content": "<p>Using kaggle api, you can easily download a dataset to any cloud provider.</p>\n\n<p><a href=\"https://github.com/Kaggle/kaggle-api\">https://github.com/Kaggle/kaggle-api</a></p>\n\n<p>In case of kaggle kernels, you can directly add the dataset to your kernel.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 646643,
          "author_name": "Gurgel",
          "author_url": "",
          "post_date": "2019-10-11T14:03:34.653000",
          "content": "<p><a href=\"/atikur\">@atikur</a> <a href=\"/teeyee314\">@teeyee314</a>  Thank you for your answer. When I time 'mv * ../train512', to join train with trainb, I am having this error mv: Argument list too long... Do you know how can I overcome this problem in Kaggle?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 646919,
          "author_name": "Tim Yee",
          "author_url": "",
          "post_date": "2019-10-11T22:08:20.280000",
          "content": "<p>are you using my 512 dataset locally, kaggle kernel, google colab, or gcp? If you're training on kaggle, training 512x512 may not be the most efficient way to utilize your time and gpu quota. Plenty of people at the top of LB are reaching &lt; 0.08 with 224x224 and if you're really good, you can reach 0.07 with 224x224.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 646921,
          "author_name": "Gurgel",
          "author_url": "",
          "post_date": "2019-10-11T22:12:23.040000",
          "content": "<p>On Kaggle. I will try to do as you say, but until now I only tried the original dataset, so I was planning to see the difference in time using each type of png size.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 648327,
          "author_name": "pupil3",
          "author_url": "",
          "post_date": "2019-10-14T03:15:22.283000",
          "content": "<p>could you tell me what's the difference between train and trainb?\nThanks</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 643725,
      "author_name": "William Green",
      "author_url": "",
      "post_date": "2019-10-07T20:53:21.677000",
      "content": "<p>Thank you for sharing. Did you normalize the images during your process? Or, is that something we need to do?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 643753,
          "author_name": "Tim Yee",
          "author_url": "",
          "post_date": "2019-10-07T21:43:34.217000",
          "content": "<p>No normalization was performed. The pixel data was extracted and windowed(transformed) per the kernel linked above. My pre-processing was similar to <a href=\"https://www.kaggle.com/taindow/generate-images?scriptVersionId=21147396\">this</a>. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 643777,
          "author_name": "William Green",
          "author_url": "",
          "post_date": "2019-10-07T23:00:12.140000",
          "content": "<p>Thank you. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 643159,
      "author_name": "Mobassir",
      "author_url": "",
      "post_date": "2019-10-07T07:38:11.303000",
      "content": "<p>hi <a href=\"/teeyee314\">@teeyee314</a>  i want to use your dataset in colab,will i have to download your data and upload in colab or there is any way to directly get it in colab from kaggle? asking this question because this is not original dataset and can't do api call for downloading this data in colab</p>",
      "votes": 0,
      "replies": [
        {
          "id": 643177,
          "author_name": "Atikur Rahman",
          "author_url": "",
          "post_date": "2019-10-07T08:00:10.627000",
          "content": "<p>You can use kaggle api to download datasets. For example -</p>\n\n<p><code>\n!kaggle datasets download teeyee314/rsnatrain240\n</code></p>\n\n<p>You will find details about <code>Datasets</code> api here -</p>\n\n<p><a href=\"https://github.com/Kaggle/kaggle-api\">https://github.com/Kaggle/kaggle-api</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 643183,
          "author_name": "Mobassir",
          "author_url": "",
          "post_date": "2019-10-07T08:03:47.123000",
          "content": "<p>good to know,thanks for letting me know</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 642540,
      "author_name": "madmax0404",
      "author_url": "",
      "post_date": "2019-10-06T08:59:37.677000",
      "content": "<p>512 train_a isn't available</p>",
      "votes": 0,
      "replies": [
        {
          "id": 642660,
          "author_name": "Tim Yee",
          "author_url": "",
          "post_date": "2019-10-06T12:50:16.643000",
          "content": "<p>I'll upload it as soon as I can. I may need to redo the split and I have my other things tying up my resources, namely limited IO on my harddrive that stores the files. I can't give you a turnaround time but you can feel free to use other resolutions until I have the resources to upload.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 642968,
          "author_name": "Tim Yee",
          "author_url": "",
          "post_date": "2019-10-06T21:38:38.963000",
          "content": "<p><a href=\"/max6296\">@max6296</a> 512x512 is uploaded. If you downloaded train512b before, re-download it. I created a new split.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 642984,
          "author_name": "madmax0404",
          "author_url": "",
          "post_date": "2019-10-06T22:27:40.210000",
          "content": "<p>Thank you so much!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 643128,
          "author_name": "Ken Ho",
          "author_url": "",
          "post_date": "2019-10-07T06:37:08.097000",
          "content": "<p>Thanks so much for your effort!\nI found that train 260 isn't uploaded yet.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 643681,
          "author_name": "Tim Yee",
          "author_url": "",
          "post_date": "2019-10-07T19:21:27.690000",
          "content": "<p><a href=\"/kenho211\">@kenho211</a> I re-uploaded test260 to keep the links uniform. train260 should be uploaded... </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 649459,
      "author_name": "LongYin/杰少",
      "author_url": "",
      "post_date": "2019-10-15T11:37:30.317000",
      "content": "<p>Thanks for sharing.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "638037": "I am sharing my pre-processed .png files for those who may find it useful. It takes several hours to pre-process at each resolution and upload them to kaggle datasets. If you have issues with the datasets or questions about them, let me know and I will do my best to reply. All training data leave out a corrupted image \"ID_6431af929.dcm\". Pre-processing involved applying a linear transformation as discussed in https://www.kaggle.com/omission/eda-view-dicom-images-with-correct-windowing. I am not a subject matter expert on the \"windowing\" applied so I will defer to someone who is: https://www.kaggle.com/dcstang/see-like-a-radiologist-with-systematic-windowing.\n\nFollowing the linear transform, resizing was done using albumentations with lanczos interpolation because that is what I felt was best (it may or may not be). Therefore, use the lower resolutions at your own discretion.\n\n\n224 x 224 [train](https://www.kaggle.com/teeyee314/rsnatrain224) [test](https://www.kaggle.com/teeyee314/rsnatest224)\n240 x 240 [train](https://www.kaggle.com/teeyee314/rsnatrain240) [test](https://www.kaggle.com/teeyee314/rsnatest240)\n260 x 260 [train](https://www.kaggle.com/teeyee314/rsnatrain260) [test](https://www.kaggle.com/teeyee314/rsnatest260)\n300 x 300 [train](https://www.kaggle.com/teeyee314/rsnatrain300) [test](https://www.kaggle.com/teeyee314/rsnatest300)\n380 x 380 [train](https://www.kaggle.com/teeyee314/rsnatrain380) [test](https://www.kaggle.com/teeyee314/rsnatest380)\n456 x 456 [train](https://www.kaggle.com/teeyee314/rsnatrain456) [test](https://www.kaggle.com/teeyee314/rsnatest456)\n512 x 512 [train](https://www.kaggle.com/teeyee314/rsnatrain512) [trainb](https://www.kaggle.com/teeyee314/rsnatrain512b) [test](https://www.kaggle.com/teeyee314/rsnatest512)\n\n<br>\n\n_please note that 512 train is split into 2 zipped containers. this was intentional because of the hard 20GB limit on public datasets. once you've downloaded both containers, unzip both of them, and simply cd into the train512b folder and type 'mv * ../train512' and that command will move all files in train512b over to train512._\n\nOct 21 update - there are &lt; 300 images in 512x512 that are not actually 512x512 and I didn't resize them to 512x512. Please make sure you resize your images to take on 512x512 when using that dataset.\n",
    "647817": "Hello Tim,\n\nthanks for your work, very much appreciated. I was wondering, since this is a 2 stages competition, would you mind releasing the kernel you used to create these datasets? This way people will be able to create their own Stage 2 dataset in the same way you used in Stage 1.",
    "638127": "thanks a ton for putting so much efforts on that and for this post @teeyee314 \nkaggle will always be a good place to learn data science as long as it will have dedicated people like you",
    "650419": "could you tell me where's the 224x224 datasets with no windowing?\nMuch Thanks",
    "648425": "Can kaggle api download your datasets？",
    "647798": "could anyone tell me what's the difference between train and trainb?",
    "647695": "@teeyee314, thanks for this! I also want to get my hands on my own processing - did you add an additional line to the script used [here](https://www.kaggle.com/taindow/generate-images-train) on resizing with interpolation? I am referring to what you mentioned [here](https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109747633405) and I would like to learn to do it myself.",
    "646950": "I found some images being black or almost black, as was already discussed. But also some images appear to be corrupted and contain NaNs. Is it just me or are they really corrupted during convertation/upload?",
    "646045": "@teeyee314 One question I have (I am new to Kaggle): are we allowed to submit and go to second phase using png data instead of the raw data given to us in the beggining?",
    "643725": "Thank you for sharing. Did you normalize the images during your process? Or, is that something we need to do?",
    "643159": "hi @teeyee314  i want to use your dataset in colab,will i have to download your data and upload in colab or there is any way to directly get it in colab from kaggle? asking this question because this is not original dataset and can't do api call for downloading this data in colab",
    "642540": "512 train_a isn't available",
    "649459": "Thanks for sharing."
  }
}