{
  "id": 109747,
  "title": "Training dataset (png, 224x224)",
  "url": "/competitions/rsna-intracranial-hemorrhage-detection/discussion/109747",
  "author_name": "Tom Aindow",
  "post_date": "2019-09-22T00:18:18.109000",
  "votes": 76,
  "comment_count": 33,
  "views": 0,
  "content": "<p>Hi guys,</p>\n\n<p>Not sure if this will be useful for people but I uploaded the images from the training data as .png files to kaggle. This should help working with the data in kernels, since extracting larger data from .dcm on the fly can be slow and kaggle limits the amount you can write to disk. Dataset can be found here:</p>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/taindow/rsna-train-stage-1-images-png-224x\">https://www.kaggle.com/taindow/rsna-train-stage-1-images-png-224x</a></li>\n</ul>\n\n<p>All pixel data that is non-corrupted has been extracted and windowed before resizing to 224x224, using the windowing functions from the following kernel: </p>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/omission/eda-view-dicom-images-with-correct-windowing\">https://www.kaggle.com/omission/eda-view-dicom-images-with-correct-windowing</a></li>\n</ul>\n\n<p>Good luck everyone, looking forward to the competition :)</p>\n\n<hr>\n\n<p>EDIT: kernel using data can be found here:</p>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/taindow/pytorch-efficientnet-b0\">https://www.kaggle.com/taindow/pytorch-efficientnet-b0</a></li>\n</ul>\n\n<hr>\n\n<p>EDIT #2: approx. 3k images are missing from the dataset, which I think was due to issues in the initial download (files were skipped over in pre-processing). So you might want to re-run yourself or use one of the other datasets people have uploaded :)</p>",
  "messages": [
    {
      "id": 631400,
      "postDate": "2019-09-22T00:18:18.110Z",
      "content": "<p>Hi guys,</p>\n\n<p>Not sure if this will be useful for people but I uploaded the images from the training data as .png files to kaggle. This should help working with the data in kernels, since extracting larger data from .dcm on the fly can be slow and kaggle limits the amount you can write to disk. Dataset can be found here:</p>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/taindow/rsna-train-stage-1-images-png-224x\">https://www.kaggle.com/taindow/rsna-train-stage-1-images-png-224x</a></li>\n</ul>\n\n<p>All pixel data that is non-corrupted has been extracted and windowed before resizing to 224x224, using the windowing functions from the following kernel: </p>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/omission/eda-view-dicom-images-with-correct-windowing\">https://www.kaggle.com/omission/eda-view-dicom-images-with-correct-windowing</a></li>\n</ul>\n\n<p>Good luck everyone, looking forward to the competition :)</p>\n\n<hr>\n\n<p>EDIT: kernel using data can be found here:</p>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/taindow/pytorch-efficientnet-b0\">https://www.kaggle.com/taindow/pytorch-efficientnet-b0</a></li>\n</ul>\n\n<hr>\n\n<p>EDIT #2: approx. 3k images are missing from the dataset, which I think was due to issues in the initial download (files were skipped over in pre-processing). So you might want to re-run yourself or use one of the other datasets people have uploaded :)</p>",
      "rawMarkdown": "Hi guys,\n\nNot sure if this will be useful for people but I uploaded the images from the training data as .png files to kaggle. This should help working with the data in kernels, since extracting larger data from .dcm on the fly can be slow and kaggle limits the amount you can write to disk. Dataset can be found here:\n\n- https://www.kaggle.com/taindow/rsna-train-stage-1-images-png-224x\n\nAll pixel data that is non-corrupted has been extracted and windowed before resizing to 224x224, using the windowing functions from the following kernel: \n\n- https://www.kaggle.com/omission/eda-view-dicom-images-with-correct-windowing\n\nGood luck everyone, looking forward to the competition :)\n\n--------\n  \nEDIT: kernel using data can be found here:\n\n- https://www.kaggle.com/taindow/pytorch-efficientnet-b0\n\n--------\n\nEDIT #2: approx. 3k images are missing from the dataset, which I think was due to issues in the initial download (files were skipped over in pre-processing). So you might want to re-run yourself or use one of the other datasets people have uploaded :)\n",
      "votes": 75
    },
    {
      "id": 635745,
      "postDate": "2019-09-28T06:09:12.180Z",
      "content": "<p>How you uploaded files on kaggle.\nI am also trying to upload but no success till now.\nI used \"kaggle datasets ...\"  command to upload.\nThanks .</p>",
      "rawMarkdown": "How you uploaded files on kaggle.\nI am also trying to upload but no success till now.\nI used \"kaggle datasets ...\"  command to upload.\nThanks .",
      "votes": 1,
      "replies": [
        {
          "id": 647409,
          "postDate": "2019-10-12T15:26:53.840Z",
          "content": "<p><a href=\"/rajnishe\">@rajnishe</a> \n!pip install kaggle\n!pip install -U -q kaggle\n!mkdir  /root/.kaggle</p>\n\n<p>files.upload()\n!cp kaggle.json /root/.kaggle</p>\n\n<p>!chmod 600 /root/.kaggle/kaggle.json</p>\n\n<p>!kaggle datasets download taindow/rsna-train-stage-1-images-png-224x</p>\n\n<p>!kaggle datasets download taindow/rsna-test-stage-1-images-png-224x\n!kaggle datasets download mobassir/rsnatraintest</p>",
          "rawMarkdown": "@rajnishe \n!pip install kaggle\n!pip install -U -q kaggle\n!mkdir  /root/.kaggle\n\nfiles.upload()\n!cp kaggle.json /root/.kaggle\n\n!chmod 600 /root/.kaggle/kaggle.json\n\n!kaggle datasets download taindow/rsna-train-stage-1-images-png-224x\n\n!kaggle datasets download taindow/rsna-test-stage-1-images-png-224x\n!kaggle datasets download mobassir/rsnatraintest\n"
        },
        {
          "id": 647464,
          "postDate": "2019-10-12T16:49:34.480Z",
          "content": "<p>thanks <a href=\"/mobassir\">@mobassir</a> </p>",
          "rawMarkdown": "thanks @mobassir ",
          "votes": 1
        }
      ]
    },
    {
      "id": 633405,
      "postDate": "2019-09-24T21:28:31.977Z",
      "content": "<p>There are 674258 dcom files in the train set.\nBut only 671797 png files in your archive.\nDoes it mean we can't extract png from some dcoms?</p>",
      "rawMarkdown": "There are 674258 dcom files in the train set.\nBut only 671797 png files in your archive.\nDoes it mean we can't extract png from some dcoms?",
      "votes": 1,
      "replies": [
        {
          "id": 636094,
          "postDate": "2019-09-28T20:03:23.500Z",
          "content": "<p><a href=\"/taindow\">@taindow</a> I have a similar concern here. I was able to extract 674,257 train imgs. Can you go back and check your script and see why you left out 2461 images while I was able to leave out only 1 corrupted img? Maybe it is something to do with the multiprocessing you put in place whereas I opted for single core for loop through each individual file and a try/except block. If not, I'll upload my conversions when I'm done for those who want the extra 2460 images. Also I should mention, my converted unzipped 224x224 train images  are only 5.37GB using albumentations resize with interpolation set to lanczos4. I have no explanation but I'd like to figure out how you ended up with more than double the size. My guess is the interpolation affects the lossless compression algo employed by .png indirectly.</p>",
          "rawMarkdown": "@taindow I have a similar concern here. I was able to extract 674,257 train imgs. Can you go back and check your script and see why you left out 2461 images while I was able to leave out only 1 corrupted img? Maybe it is something to do with the multiprocessing you put in place whereas I opted for single core for loop through each individual file and a try/except block. If not, I'll upload my conversions when I'm done for those who want the extra 2460 images. Also I should mention, my converted unzipped 224x224 train images  are only 5.37GB using albumentations resize with interpolation set to lanczos4. I have no explanation but I'd like to figure out how you ended up with more than double the size. My guess is the interpolation affects the lossless compression algo employed by .png indirectly.",
          "votes": 1
        },
        {
          "id": 636438,
          "postDate": "2019-09-29T14:33:52.747Z",
          "content": "<p>Best I can make out there were some issues in the initial download where it was interrupted and restarted. So looks like I would have to re-download and run again. Since there are now many other people who have uploaded jpegs/pngs, I'll probably leave it, but will add a warning at the top. </p>\n\n<p>Adding yours would be great tbh .. and no idea why it ended up taking so much more space.</p>",
          "rawMarkdown": "Best I can make out there were some issues in the initial download where it was interrupted and restarted. So looks like I would have to re-download and run again. Since there are now many other people who have uploaded jpegs/pngs, I'll probably leave it, but will add a warning at the top. \n\nAdding yours would be great tbh .. and no idea why it ended up taking so much more space."
        }
      ]
    },
    {
      "id": 632752,
      "postDate": "2019-09-24T02:45:27.657Z",
      "content": "<p>Thanks for sharing! I want to know whether the size of this training dataset is consistent with original one or not?Thank you!</p>",
      "rawMarkdown": "Thanks for sharing! I want to know whether the size of this training dataset is consistent with original one or not?Thank you!",
      "votes": 1
    },
    {
      "id": 632157,
      "postDate": "2019-09-23T09:40:37.707Z",
      "content": "<p>Thanks for the dataset. Can you make a dataset of original size (512px) ?</p>",
      "rawMarkdown": "Thanks for the dataset. Can you make a dataset of original size (512px) ?",
      "votes": 1,
      "replies": [
        {
          "id": 632898,
          "postDate": "2019-09-24T07:49:06.483Z",
          "content": "<p>Hi <a href=\"/moewie94\">@moewie94</a>, keeping the original size means the dataset is &gt;20gb and cannot be hosted at Kaggle, sorry.</p>",
          "rawMarkdown": "Hi @moewie94, keeping the original size means the dataset is &gt;20gb and cannot be hosted at Kaggle, sorry.",
          "votes": 1
        },
        {
          "id": 633504,
          "postDate": "2019-09-25T03:03:19.453Z",
          "content": "<p><a href=\"/taindow\">@taindow</a> oh i see. but i downloaded your dataset and only have 671k images (it's 674k in the csv file). did my download failed or there are 3k missing images in the dataset?</p>",
          "rawMarkdown": "@taindow oh i see. but i downloaded your dataset and only have 671k images (it's 674k in the csv file). did my download failed or there are 3k missing images in the dataset?",
          "votes": 2
        }
      ]
    },
    {
      "id": 632121,
      "postDate": "2019-09-23T08:53:00.030Z",
      "content": "<p>awesome!</p>",
      "rawMarkdown": "awesome!",
      "votes": 1
    },
    {
      "id": 631954,
      "postDate": "2019-09-23T02:26:20.557Z",
      "content": "<p>This is awesome. Any chance you would do this for the test data too? </p>",
      "rawMarkdown": "This is awesome. Any chance you would do this for the test data too? ",
      "votes": 1,
      "replies": [
        {
          "id": 632050,
          "postDate": "2019-09-23T07:01:43.120Z",
          "content": "<p>Already done :) </p>\n\n<p><a href=\"https://www.kaggle.com/taindow/rsna-test-stage-1-images-png-224x\">https://www.kaggle.com/taindow/rsna-test-stage-1-images-png-224x</a></p>",
          "rawMarkdown": "Already done :) \n\nhttps://www.kaggle.com/taindow/rsna-test-stage-1-images-png-224x",
          "votes": 2
        }
      ]
    },
    {
      "id": 634862,
      "postDate": "2019-09-26T21:47:42.600Z",
      "content": "<p>Added the code used to generate the data as a kernel attached to the dataset. You can find it here:</p>\n\n<p><a href=\"https://www.kaggle.com/taindow/generate-images?scriptVersionId=21147396\">https://www.kaggle.com/taindow/generate-images?scriptVersionId=21147396</a></p>\n\n<p>With regards to the missing files, there was some issue with the pixel data that I haven't had a chance to look at yet. If you want to use this dataset just make sure to filter out the 3k images or so from training csv :)</p>",
      "rawMarkdown": "Added the code used to generate the data as a kernel attached to the dataset. You can find it here:\n\nhttps://www.kaggle.com/taindow/generate-images?scriptVersionId=21147396\n\nWith regards to the missing files, there was some issue with the pixel data that I haven't had a chance to look at yet. If you want to use this dataset just make sure to filter out the 3k images or so from training csv :)",
      "votes": 2,
      "replies": [
        {
          "id": 635687,
          "postDate": "2019-09-28T03:07:35.103Z",
          "content": "<p><a href=\"/taindow\">@taindow</a> Thanks for the share. I'm sure you had to have downloaded the .dcm files locally and created the .png files. When you compressed the files in a zip file, what settings did you use? Something other than default such as the highest compression ratio?</p>",
          "rawMarkdown": "@taindow Thanks for the share. I'm sure you had to have downloaded the .dcm files locally and created the .png files. When you compressed the files in a zip file, what settings did you use? Something other than default such as the highest compression ratio?",
          "votes": 1
        },
        {
          "id": 635980,
          "postDate": "2019-09-28T14:52:48.147Z",
          "content": "<p><a href=\"/taindow\">@taindow</a> I know I'm asking a lot of questions here, but how long did it take you to convert the images to 224x224 .png files using the multiprocessing? I haven't been able to figure out how to get your script to work with multiprocessing yet so I'm running on a single core with a for loop (I know.. smh). Also, when you resized your images, which library did you use? Albumentations? OpenCV? Finally, what interpolation did you use? I'm trying to re-do the images for different resolutions and just want to get a better idea. As for the windowing, have you found that this pre-processing step improves LB? I'm curious if you've tried the same model without the windowing preprocessing. Thanks</p>",
          "rawMarkdown": "@taindow I know I'm asking a lot of questions here, but how long did it take you to convert the images to 224x224 .png files using the multiprocessing? I haven't been able to figure out how to get your script to work with multiprocessing yet so I'm running on a single core with a for loop (I know.. smh). Also, when you resized your images, which library did you use? Albumentations? OpenCV? Finally, what interpolation did you use? I'm trying to re-do the images for different resolutions and just want to get a better idea. As for the windowing, have you found that this pre-processing step improves LB? I'm curious if you've tried the same model without the windowing preprocessing. Thanks",
          "votes": 1
        },
        {
          "id": 635988,
          "postDate": "2019-09-28T14:57:32.720Z",
          "content": "<p>The script already include multiprocessor see lines <code>15, 67-79</code>....  =) \nRegarding interpolation i think most common are here <code>bicubic</code> but you can check on the link below see which one uses yields better accuracy \n<a href=\"https://github.com/rwightman/pytorch-image-models\">https://github.com/rwightman/pytorch-image-models</a></p>\n\n<p>P.S \nuseful graph to understand what are interpolation (Yellow and green are original pixels, black is the pixel you will get if you use one of the interpolation methods )</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2F6d073c249b35bb0c86e5752b7aa8284e%2FComparison_of_1D_and_2D_interpolation.svg.png?generation=1569683284806024&amp;alt=media\" alt=\"\"></p>\n\n<p><a href=\"https://en.wikipedia.org/wiki/Bicubic_interpolation\">https://en.wikipedia.org/wiki/Bicubic_interpolation</a></p>",
          "rawMarkdown": "The script already include multiprocessor see lines `15, 67-79`....  =) \nRegarding interpolation i think most common are here ` bicubic` but you can check on the link below see which one uses yields better accuracy \nhttps://github.com/rwightman/pytorch-image-models\n\nP.S \nuseful graph to understand what are interpolation (Yellow and green are original pixels, black is the pixel you will get if you use one of the interpolation methods )\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2F6d073c249b35bb0c86e5752b7aa8284e%2FComparison_of_1D_and_2D_interpolation.svg.png?generation=1569683284806024&amp;alt=media)\n\nhttps://en.wikipedia.org/wiki/Bicubic_interpolation",
          "votes": 4
        },
        {
          "id": 636008,
          "postDate": "2019-09-28T15:33:16.583Z",
          "content": "<p><a href=\"/drhabib\">@drhabib</a>  Yes. I can see the multiprocessing line. I'm going line by line to try to get the script to work on my end. Everything except the multiprocessing works. The key to the script is actually in the multiprocessing because it is extremely slow without it. Thanks for the visualizations. I'm asking how the the files we resized so I have a better idea of what was performed. Maybe you can explain to me the windowing method as I don't understand what the purpose and benefit of such a preprocessing step does. I'm confused as to why windowing needs to be performed in the first place as it is not explained very clearly in the link to the notebook above.</p>",
          "rawMarkdown": "@drhabib  Yes. I can see the multiprocessing line. I'm going line by line to try to get the script to work on my end. Everything except the multiprocessing works. The key to the script is actually in the multiprocessing because it is extremely slow without it. Thanks for the visualizations. I'm asking how the the files we resized so I have a better idea of what was performed. Maybe you can explain to me the windowing method as I don't understand what the purpose and benefit of such a preprocessing step does. I'm confused as to why windowing needs to be performed in the first place as it is not explained very clearly in the link to the notebook above.",
          "votes": 1
        },
        {
          "id": 636018,
          "postDate": "2019-09-28T16:03:03.430Z",
          "content": "<p>I dont think I can do better job explaining than in this kernel <a href=\"https://www.kaggle.com/allunia/rsna-ih-detection-eda-baseline\">https://www.kaggle.com/allunia/rsna-ih-detection-eda-baseline</a> by @Allunia there is also a video =) And regarding the script if something is not working you can post here a full script and maybe someone can help you =) </p>",
          "rawMarkdown": "I dont think I can do better job explaining than in this kernel https://www.kaggle.com/allunia/rsna-ih-detection-eda-baseline by @Allunia there is also a video =) And regarding the script if something is not working you can post here a full script and maybe someone can help you =) ",
          "votes": 1
        },
        {
          "id": 636068,
          "postDate": "2019-09-28T18:57:31.153Z",
          "content": "<p>Images were resized using cv2 with bilinear interpolation. Sorry, I forgot this version of the script didn't include this step. As for the time it took, I have no idea since I just left it running while I was out. It was def. a fair few hours running on Ryzen 7 2700x.</p>",
          "rawMarkdown": "Images were resized using cv2 with bilinear interpolation. Sorry, I forgot this version of the script didn't include this step. As for the time it took, I have no idea since I just left it running while I was out. It was def. a fair few hours running on Ryzen 7 2700x.",
          "votes": 2
        }
      ]
    },
    {
      "id": 632186,
      "postDate": "2019-09-23T10:39:19.140Z",
      "content": "<p>i think these images are a little bit dark, compared to the reference kernel. can you check if there's something difference in codes?</p>",
      "rawMarkdown": "i think these images are a little bit dark, compared to the reference kernel. can you check if there's something difference in codes?",
      "votes": 2,
      "replies": [
        {
          "id": 632727,
          "postDate": "2019-09-24T01:28:07.457Z",
          "content": "<p>I observe the same issue. When you save an array to the image the reload it again. The information is loss slightly. For example, In my case, some of the pixels have value as 7 before saving and 6 after loading the image. </p>",
          "rawMarkdown": "I observe the same issue. When you save an array to the image the reload it again. The information is loss slightly. For example, In my case, some of the pixels have value as 7 before saving and 6 after loading the image. ",
          "votes": 1
        },
        {
          "id": 632901,
          "postDate": "2019-09-24T07:50:07.313Z",
          "content": "<p>This may just be because of the color mapping people are using when displaying images in kernels. But I will check the code and will upload the code to generate the images to the dataset when I get home today - maybe there is a mistake :)</p>",
          "rawMarkdown": "This may just be because of the color mapping people are using when displaying images in kernels. But I will check the code and will upload the code to generate the images to the dataset when I get home today - maybe there is a mistake :)",
          "votes": 2
        }
      ]
    },
    {
      "id": 666631,
      "postDate": "2019-11-06T09:45:40.270Z",
      "content": "<p>Has anyone used the png images for training and testing, which were giving a valid score, and now at the end of the stage1, show an error in the submission, such as “Failed; Evaluation Exception: Missing solution column: Id?</p>",
      "rawMarkdown": "Has anyone used the png images for training and testing, which were giving a valid score, and now at the end of the stage1, show an error in the submission, such as “Failed; Evaluation Exception: Missing solution column: Id?"
    },
    {
      "id": 647377,
      "postDate": "2019-10-12T14:33:09.733Z",
      "content": "<p>Thanks for sharing!\nI have a question, the original pixel array has only 1 channel. But your PNG files has 3. Did you use like 'gray2rgb' function? I don't see how a 1 channel array becomes 3 channels.\nThanks again for you work!!</p>",
      "rawMarkdown": "Thanks for sharing!\nI have a question, the original pixel array has only 1 channel. But your PNG files has 3. Did you use like 'gray2rgb' function? I don't see how a 1 channel array becomes 3 channels.\nThanks again for you work!!",
      "replies": [
        {
          "id": 647469,
          "postDate": "2019-10-12T16:56:36.583Z",
          "content": "<p><a href=\"/ivanwang2016\">@ivanwang2016</a> I just stacked the image 3 times, since I wanted to use it with pre-trained model that uses 3 channels by default :)</p>",
          "rawMarkdown": "@ivanwang2016 I just stacked the image 3 times, since I wanted to use it with pre-trained model that uses 3 channels by default :)"
        },
        {
          "id": 649361,
          "postDate": "2019-10-15T08:38:15.757Z",
          "content": "<p>Thanks, </p>",
          "rawMarkdown": "Thanks, "
        }
      ]
    },
    {
      "id": 644068,
      "postDate": "2019-10-08T09:16:08.380Z",
      "content": "<p>hi <a href=\"/taindow\">@taindow</a>  whenever i download your dataset in google colab and extract , my browser crashes,any solution please?</p>",
      "rawMarkdown": "hi @taindow  whenever i download your dataset in google colab and extract , my browser crashes,any solution please?",
      "replies": [
        {
          "id": 644086,
          "postDate": "2019-10-08T09:51:00.340Z",
          "content": "<p>Try adding the <code>-q</code> argument to <code>unzip</code> command (which performs the unzip operation quietly) -</p>\n\n<p><code>\n!unzip -q yourFileName.zip\n</code></p>",
          "rawMarkdown": "Try adding the `-q` argument to `unzip` command (which performs the unzip operation quietly) -\n\n```\n!unzip -q yourFileName.zip\n```",
          "votes": 2
        },
        {
          "id": 644149,
          "postDate": "2019-10-08T12:17:45.633Z",
          "content": "<p>I have found two other ways to suppress output - the first which only works on colab. The other uses cell magic and works fine elsewhere.</p>\n\n<blockquote>\n  <p>!unzip yourFileName.zip &amp;&gt; /dev/null</p>\n</blockquote>\n\n<p>and </p>\n\n<blockquote>\n  <p>%%capture\n  !unzip yourFileName.zip</p>\n</blockquote>",
          "rawMarkdown": "I have found two other ways to suppress output - the first which only works on colab. The other uses cell magic and works fine elsewhere.\n\n&gt; !unzip yourFileName.zip &amp;&gt; /dev/null\n\nand \n\n&gt; %%capture\n!unzip yourFileName.zip",
          "votes": 3
        },
        {
          "id": 644333,
          "postDate": "2019-10-08T16:28:43.893Z",
          "content": "<p><a href=\"/teeyee314\">@teeyee314</a>  i haven't tried your solutions but <a href=\"/atikur\">@atikur</a> suggested solution is working for me,thanks a lot</p>",
          "rawMarkdown": "@teeyee314  i haven't tried your solutions but @atikur suggested solution is working for me,thanks a lot"
        }
      ]
    },
    {
      "id": 631623,
      "postDate": "2019-09-22T10:59:52.747Z",
      "content": "<p>Give this man a medal. thank you Tom.</p>",
      "rawMarkdown": "Give this man a medal. thank you Tom.",
      "votes": 4
    },
    {
      "id": 649278,
      "postDate": "2019-10-15T06:59:07.007Z",
      "content": "<p>Great work Tom , thanks for the dataset </p>",
      "rawMarkdown": "Great work Tom , thanks for the dataset "
    }
  ],
  "comments": [
    {
      "id": 635745,
      "author_name": "Rajnish Chauhan",
      "author_url": "",
      "post_date": "2019-09-28T06:09:12.180000",
      "content": "<p>How you uploaded files on kaggle.\nI am also trying to upload but no success till now.\nI used \"kaggle datasets ...\"  command to upload.\nThanks .</p>",
      "votes": 1,
      "replies": [
        {
          "id": 647409,
          "author_name": "Mobassir",
          "author_url": "",
          "post_date": "2019-10-12T15:26:53.840000",
          "content": "<p><a href=\"/rajnishe\">@rajnishe</a> \n!pip install kaggle\n!pip install -U -q kaggle\n!mkdir  /root/.kaggle</p>\n\n<p>files.upload()\n!cp kaggle.json /root/.kaggle</p>\n\n<p>!chmod 600 /root/.kaggle/kaggle.json</p>\n\n<p>!kaggle datasets download taindow/rsna-train-stage-1-images-png-224x</p>\n\n<p>!kaggle datasets download taindow/rsna-test-stage-1-images-png-224x\n!kaggle datasets download mobassir/rsnatraintest</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 647464,
          "author_name": "Rajnish Chauhan",
          "author_url": "",
          "post_date": "2019-10-12T16:49:34.480000",
          "content": "<p>thanks <a href=\"/mobassir\">@mobassir</a> </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 633405,
      "author_name": "Sergey Zlobin",
      "author_url": "",
      "post_date": "2019-09-24T21:28:31.977000",
      "content": "<p>There are 674258 dcom files in the train set.\nBut only 671797 png files in your archive.\nDoes it mean we can't extract png from some dcoms?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 636094,
          "author_name": "Tim Yee",
          "author_url": "",
          "post_date": "2019-09-28T20:03:23.500000",
          "content": "<p><a href=\"/taindow\">@taindow</a> I have a similar concern here. I was able to extract 674,257 train imgs. Can you go back and check your script and see why you left out 2461 images while I was able to leave out only 1 corrupted img? Maybe it is something to do with the multiprocessing you put in place whereas I opted for single core for loop through each individual file and a try/except block. If not, I'll upload my conversions when I'm done for those who want the extra 2460 images. Also I should mention, my converted unzipped 224x224 train images  are only 5.37GB using albumentations resize with interpolation set to lanczos4. I have no explanation but I'd like to figure out how you ended up with more than double the size. My guess is the interpolation affects the lossless compression algo employed by .png indirectly.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 636438,
          "author_name": "Tom Aindow",
          "author_url": "",
          "post_date": "2019-09-29T14:33:52.747000",
          "content": "<p>Best I can make out there were some issues in the initial download where it was interrupted and restarted. So looks like I would have to re-download and run again. Since there are now many other people who have uploaded jpegs/pngs, I'll probably leave it, but will add a warning at the top. </p>\n\n<p>Adding yours would be great tbh .. and no idea why it ended up taking so much more space.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 632752,
      "author_name": "baijuguoxi",
      "author_url": "",
      "post_date": "2019-09-24T02:45:27.657000",
      "content": "<p>Thanks for sharing! I want to know whether the size of this training dataset is consistent with original one or not?Thank you!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 632157,
      "author_name": "DatNT",
      "author_url": "",
      "post_date": "2019-09-23T09:40:37.707000",
      "content": "<p>Thanks for the dataset. Can you make a dataset of original size (512px) ?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 632898,
          "author_name": "Tom Aindow",
          "author_url": "",
          "post_date": "2019-09-24T07:49:06.483000",
          "content": "<p>Hi <a href=\"/moewie94\">@moewie94</a>, keeping the original size means the dataset is &gt;20gb and cannot be hosted at Kaggle, sorry.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 633504,
          "author_name": "DatNT",
          "author_url": "",
          "post_date": "2019-09-25T03:03:19.453000",
          "content": "<p><a href=\"/taindow\">@taindow</a> oh i see. but i downloaded your dataset and only have 671k images (it's 674k in the csv file). did my download failed or there are 3k missing images in the dataset?</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 632121,
      "author_name": "Erwin John T. Carpio",
      "author_url": "",
      "post_date": "2019-09-23T08:53:00.030000",
      "content": "<p>awesome!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 631954,
      "author_name": "Alex Federation",
      "author_url": "",
      "post_date": "2019-09-23T02:26:20.557000",
      "content": "<p>This is awesome. Any chance you would do this for the test data too? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 632050,
          "author_name": "Tom Aindow",
          "author_url": "",
          "post_date": "2019-09-23T07:01:43.120000",
          "content": "<p>Already done :) </p>\n\n<p><a href=\"https://www.kaggle.com/taindow/rsna-test-stage-1-images-png-224x\">https://www.kaggle.com/taindow/rsna-test-stage-1-images-png-224x</a></p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 634862,
      "author_name": "Tom Aindow",
      "author_url": "",
      "post_date": "2019-09-26T21:47:42.600000",
      "content": "<p>Added the code used to generate the data as a kernel attached to the dataset. You can find it here:</p>\n\n<p><a href=\"https://www.kaggle.com/taindow/generate-images?scriptVersionId=21147396\">https://www.kaggle.com/taindow/generate-images?scriptVersionId=21147396</a></p>\n\n<p>With regards to the missing files, there was some issue with the pixel data that I haven't had a chance to look at yet. If you want to use this dataset just make sure to filter out the 3k images or so from training csv :)</p>",
      "votes": 2,
      "replies": [
        {
          "id": 635687,
          "author_name": "Tim Yee",
          "author_url": "",
          "post_date": "2019-09-28T03:07:35.103000",
          "content": "<p><a href=\"/taindow\">@taindow</a> Thanks for the share. I'm sure you had to have downloaded the .dcm files locally and created the .png files. When you compressed the files in a zip file, what settings did you use? Something other than default such as the highest compression ratio?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 635980,
          "author_name": "Tim Yee",
          "author_url": "",
          "post_date": "2019-09-28T14:52:48.147000",
          "content": "<p><a href=\"/taindow\">@taindow</a> I know I'm asking a lot of questions here, but how long did it take you to convert the images to 224x224 .png files using the multiprocessing? I haven't been able to figure out how to get your script to work with multiprocessing yet so I'm running on a single core with a for loop (I know.. smh). Also, when you resized your images, which library did you use? Albumentations? OpenCV? Finally, what interpolation did you use? I'm trying to re-do the images for different resolutions and just want to get a better idea. As for the windowing, have you found that this pre-processing step improves LB? I'm curious if you've tried the same model without the windowing preprocessing. Thanks</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 635988,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-09-28T14:57:32.720000",
          "content": "<p>The script already include multiprocessor see lines <code>15, 67-79</code>....  =) \nRegarding interpolation i think most common are here <code>bicubic</code> but you can check on the link below see which one uses yields better accuracy \n<a href=\"https://github.com/rwightman/pytorch-image-models\">https://github.com/rwightman/pytorch-image-models</a></p>\n\n<p>P.S \nuseful graph to understand what are interpolation (Yellow and green are original pixels, black is the pixel you will get if you use one of the interpolation methods )</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2F6d073c249b35bb0c86e5752b7aa8284e%2FComparison_of_1D_and_2D_interpolation.svg.png?generation=1569683284806024&amp;alt=media\" alt=\"\"></p>\n\n<p><a href=\"https://en.wikipedia.org/wiki/Bicubic_interpolation\">https://en.wikipedia.org/wiki/Bicubic_interpolation</a></p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 636008,
          "author_name": "Tim Yee",
          "author_url": "",
          "post_date": "2019-09-28T15:33:16.583000",
          "content": "<p><a href=\"/drhabib\">@drhabib</a>  Yes. I can see the multiprocessing line. I'm going line by line to try to get the script to work on my end. Everything except the multiprocessing works. The key to the script is actually in the multiprocessing because it is extremely slow without it. Thanks for the visualizations. I'm asking how the the files we resized so I have a better idea of what was performed. Maybe you can explain to me the windowing method as I don't understand what the purpose and benefit of such a preprocessing step does. I'm confused as to why windowing needs to be performed in the first place as it is not explained very clearly in the link to the notebook above.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 636018,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-09-28T16:03:03.430000",
          "content": "<p>I dont think I can do better job explaining than in this kernel <a href=\"https://www.kaggle.com/allunia/rsna-ih-detection-eda-baseline\">https://www.kaggle.com/allunia/rsna-ih-detection-eda-baseline</a> by @Allunia there is also a video =) And regarding the script if something is not working you can post here a full script and maybe someone can help you =) </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 636068,
          "author_name": "Tom Aindow",
          "author_url": "",
          "post_date": "2019-09-28T18:57:31.153000",
          "content": "<p>Images were resized using cv2 with bilinear interpolation. Sorry, I forgot this version of the script didn't include this step. As for the time it took, I have no idea since I just left it running while I was out. It was def. a fair few hours running on Ryzen 7 2700x.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 632186,
      "author_name": "Richul Oh",
      "author_url": "",
      "post_date": "2019-09-23T10:39:19.140000",
      "content": "<p>i think these images are a little bit dark, compared to the reference kernel. can you check if there's something difference in codes?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 632727,
          "author_name": "cab",
          "author_url": "",
          "post_date": "2019-09-24T01:28:07.457000",
          "content": "<p>I observe the same issue. When you save an array to the image the reload it again. The information is loss slightly. For example, In my case, some of the pixels have value as 7 before saving and 6 after loading the image. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 632901,
          "author_name": "Tom Aindow",
          "author_url": "",
          "post_date": "2019-09-24T07:50:07.313000",
          "content": "<p>This may just be because of the color mapping people are using when displaying images in kernels. But I will check the code and will upload the code to generate the images to the dataset when I get home today - maybe there is a mistake :)</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 666631,
      "author_name": "RosArr",
      "author_url": "",
      "post_date": "2019-11-06T09:45:40.270000",
      "content": "<p>Has anyone used the png images for training and testing, which were giving a valid score, and now at the end of the stage1, show an error in the submission, such as “Failed; Evaluation Exception: Missing solution column: Id?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 647377,
      "author_name": "ivanwang",
      "author_url": "",
      "post_date": "2019-10-12T14:33:09.733000",
      "content": "<p>Thanks for sharing!\nI have a question, the original pixel array has only 1 channel. But your PNG files has 3. Did you use like 'gray2rgb' function? I don't see how a 1 channel array becomes 3 channels.\nThanks again for you work!!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 647469,
          "author_name": "Tom Aindow",
          "author_url": "",
          "post_date": "2019-10-12T16:56:36.583000",
          "content": "<p><a href=\"/ivanwang2016\">@ivanwang2016</a> I just stacked the image 3 times, since I wanted to use it with pre-trained model that uses 3 channels by default :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 649361,
          "author_name": "LongYin/杰少",
          "author_url": "",
          "post_date": "2019-10-15T08:38:15.757000",
          "content": "<p>Thanks, </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 644068,
      "author_name": "Mobassir",
      "author_url": "",
      "post_date": "2019-10-08T09:16:08.380000",
      "content": "<p>hi <a href=\"/taindow\">@taindow</a>  whenever i download your dataset in google colab and extract , my browser crashes,any solution please?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 644086,
          "author_name": "Atikur Rahman",
          "author_url": "",
          "post_date": "2019-10-08T09:51:00.340000",
          "content": "<p>Try adding the <code>-q</code> argument to <code>unzip</code> command (which performs the unzip operation quietly) -</p>\n\n<p><code>\n!unzip -q yourFileName.zip\n</code></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 644149,
          "author_name": "Tim Yee",
          "author_url": "",
          "post_date": "2019-10-08T12:17:45.633000",
          "content": "<p>I have found two other ways to suppress output - the first which only works on colab. The other uses cell magic and works fine elsewhere.</p>\n\n<blockquote>\n  <p>!unzip yourFileName.zip &amp;&gt; /dev/null</p>\n</blockquote>\n\n<p>and </p>\n\n<blockquote>\n  <p>%%capture\n  !unzip yourFileName.zip</p>\n</blockquote>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 644333,
          "author_name": "Mobassir",
          "author_url": "",
          "post_date": "2019-10-08T16:28:43.893000",
          "content": "<p><a href=\"/teeyee314\">@teeyee314</a>  i haven't tried your solutions but <a href=\"/atikur\">@atikur</a> suggested solution is working for me,thanks a lot</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 631623,
      "author_name": "iamkhader",
      "author_url": "",
      "post_date": "2019-09-22T10:59:52.747000",
      "content": "<p>Give this man a medal. thank you Tom.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 649278,
      "author_name": "Ashik M",
      "author_url": "",
      "post_date": "2019-10-15T06:59:07.007000",
      "content": "<p>Great work Tom , thanks for the dataset </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "631400": "Hi guys,\n\nNot sure if this will be useful for people but I uploaded the images from the training data as .png files to kaggle. This should help working with the data in kernels, since extracting larger data from .dcm on the fly can be slow and kaggle limits the amount you can write to disk. Dataset can be found here:\n\n- https://www.kaggle.com/taindow/rsna-train-stage-1-images-png-224x\n\nAll pixel data that is non-corrupted has been extracted and windowed before resizing to 224x224, using the windowing functions from the following kernel: \n\n- https://www.kaggle.com/omission/eda-view-dicom-images-with-correct-windowing\n\nGood luck everyone, looking forward to the competition :)\n\n--------\n  \nEDIT: kernel using data can be found here:\n\n- https://www.kaggle.com/taindow/pytorch-efficientnet-b0\n\n--------\n\nEDIT #2: approx. 3k images are missing from the dataset, which I think was due to issues in the initial download (files were skipped over in pre-processing). So you might want to re-run yourself or use one of the other datasets people have uploaded :)\n",
    "635745": "How you uploaded files on kaggle.\nI am also trying to upload but no success till now.\nI used \"kaggle datasets ...\"  command to upload.\nThanks .",
    "633405": "There are 674258 dcom files in the train set.\nBut only 671797 png files in your archive.\nDoes it mean we can't extract png from some dcoms?",
    "632752": "Thanks for sharing! I want to know whether the size of this training dataset is consistent with original one or not?Thank you!",
    "632157": "Thanks for the dataset. Can you make a dataset of original size (512px) ?",
    "632121": "awesome!",
    "631954": "This is awesome. Any chance you would do this for the test data too? ",
    "634862": "Added the code used to generate the data as a kernel attached to the dataset. You can find it here:\n\nhttps://www.kaggle.com/taindow/generate-images?scriptVersionId=21147396\n\nWith regards to the missing files, there was some issue with the pixel data that I haven't had a chance to look at yet. If you want to use this dataset just make sure to filter out the 3k images or so from training csv :)",
    "632186": "i think these images are a little bit dark, compared to the reference kernel. can you check if there's something difference in codes?",
    "666631": "Has anyone used the png images for training and testing, which were giving a valid score, and now at the end of the stage1, show an error in the submission, such as “Failed; Evaluation Exception: Missing solution column: Id?",
    "647377": "Thanks for sharing!\nI have a question, the original pixel array has only 1 channel. But your PNG files has 3. Did you use like 'gray2rgb' function? I don't see how a 1 channel array becomes 3 channels.\nThanks again for you work!!",
    "644068": "hi @taindow  whenever i download your dataset in google colab and extract , my browser crashes,any solution please?",
    "631623": "Give this man a medal. thank you Tom.",
    "649278": "Great work Tom , thanks for the dataset "
  }
}