{
  "id": 182930,
  "title": "STARTER DATASET: Train JPEGs (256x256)",
  "url": "/competitions/rsna-str-pulmonary-embolism-detection/discussion/182930",
  "author_name": "Ian Pan",
  "post_date": "2020-09-14T23:59:57.516000",
  "votes": 193,
  "comment_count": 66,
  "views": 0,
  "content": "<p>Update: source code to produce these images has been added as attachment</p>\n<p><a href=\"https://www.kaggle.com/vaillant/rsna-str-pe-detection-jpeg-256\" target=\"_blank\">https://www.kaggle.com/vaillant/rsna-str-pe-detection-jpeg-256</a></p>\n<h1>Introduction</h1>\n<p>This competition is challenging due to the huge dataset size. In addition, people may not be familiar with medical imaging and CT scans. </p>\n<p>To lower the barrier to entry to this competition and to encourage more people to participate, I have converted all of the DICOM images into JPEGs (about 52GB), 256x256 pixels (standard is 512x512). </p>\n<h1>Windowing</h1>\n<p>The values in CT scans tend to range from -1000 to 3000; we are used to 8-bit images with pixel values ranging from 0 to 255. Radiologists use a technique called <strong>windowing</strong> to visualize CT scans. Different types of tissues are better evaluated using different windows. Windows are defined by 2 numbers: window <strong>width</strong> and window <strong>level</strong>. </p>\n<p>Here is the function I use to window an image:</p>\n<pre><code>def window(img, WL=50, WW=350):\n    upper, lower = WL+WW//2, WL-WW//2\n    X = np.clip(img.copy(), lower, upper)\n    X = X - np.min(X)\n    X = X / np.max(X)\n    X = (X*255.0).astype('uint8')\n    return X\n</code></pre>\n<p>Basically, we compute the lower and upper limits using the window level (WL) and window width (WW). Upper limit is level plus half the width and lower limit is level minus half the width. </p>\n<p>Then, all values above the upper limit are set to the upper limit; vice versa for the lower limit. We can then normalize the images to [0, 1] and multiply by 255 to convert them into an 8-bit image. </p>\n<p>Note that these images are single channel. I have provided them in 3-channel RGB format. <strong>Each channel is a different window.</strong> </p>\n<ul>\n<li><strong>RED</strong> channel / <strong>LUNG</strong> window / level=-600, width=1500</li>\n<li><strong>GREEN</strong> channel / <strong>PE</strong> window / level=100, width=700</li>\n<li><strong>BLUE</strong> channel / <strong>MEDIASTINAL</strong> window / level=40, width=400</li>\n</ul>\n<p>Please remember that <code>cv2.imread</code> by default loads images in <strong>BGR</strong> so if you are using that function make sure you're not confusing the red and blue channels. The <strong>PE specific and mediastinal windows</strong> will be most useful for evaluating whether there is blood clot. However, the lung window can show abnormalities in the lungs which may be suggestive of a clot, so I included it as well. Feel free to experiment with any single window or combination of windows.</p>\n<h3>Lung</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F281652%2Faa4bd969f80498908e8f40bade1e69b7%2Flung_window.jpg?generation=1600128037814115&amp;alt=media\" alt=\"\"></p>\n<h3>PE Specific</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F281652%2F8e5d707a2daa4606b720de0c6030aba4%2Fpe_window.jpg?generation=1600128049440405&amp;alt=media\" alt=\"\"></p>\n<h3>Mediastinal</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F281652%2F8950b63595f0b36b6d52a48907d32ffb%2Fmediastinum_window.jpg?generation=1600128059340756&amp;alt=media\" alt=\"\"></p>\n<p>This has the disadvantage that you cannot select other windows and you cannot design a learnable windowing function using this data. <strong>However, based on my experience, I do not believe these are necessary to do well in this competition.</strong></p>\n<h1>3D Reconstruction</h1>\n<p>CT scans are 3D. Each image is 2D, but taking into context the surrounding image slices can improve performance (either using a fully 3D model, hybrid 3D-2D model, or sequence model). In order to facilitate reconstructing the 3D image, <strong>each image in each series (folder) has a prefix with the slice number (increasing from bottom to top).</strong> That way, you can easily organize the slices in 3D. </p>\n<p>This is the function I used to create the 3D array:</p>\n<pre><code>def load_dicom_array(f):\n    dicom_files = glob.glob(osp.join(f, '*.dcm'))\n    dicoms = [pydicom.dcmread(d) for d in dicom_files]\n    M = float(dicoms[0].RescaleSlope)\n    B = float(dicoms[0].RescaleIntercept)\n    # Assume all images are axial\n    z_pos = [float(d.ImagePositionPatient[-1]) for d in dicoms]\n    dicoms = np.asarray([d.pixel_array for d in dicoms])\n    dicoms = dicoms[np.argsort(z_pos)]\n    dicoms = dicoms * M\n    dicoms = dicoms + B\n    return dicoms, np.asarray(dicom_files)[np.argsort(z_pos)]\n</code></pre>\n<p>Important things to note: you need to use <code>RescaleSlope</code> and <code>RescaleIntercept</code> to alter the values from <code>dicom.pixel_array</code> first. In this dataset, <code>RescaleSlope</code> is always 1 and <code>RescaleIntercept</code> is 0 or -1024 (more common). Then, you can use <code>ImagePositionPatient</code> to sort the slices. The last element <code>ImagePositionPatient[-1]</code> is used to sort in the z-axis. </p>\n<p>Please let me know if you encounter any issues with this dataset. I will try my best to fix them. It took about 36 hours to generate this dataset. I will try to produce datasets of different sizes if people request it. </p>",
  "messages": [
    {
      "id": 1010629,
      "postDate": "2020-09-14T23:59:57.517Z",
      "content": "<p>Update: source code to produce these images has been added as attachment</p>\n<p><a href=\"https://www.kaggle.com/vaillant/rsna-str-pe-detection-jpeg-256\" target=\"_blank\">https://www.kaggle.com/vaillant/rsna-str-pe-detection-jpeg-256</a></p>\n<h1>Introduction</h1>\n<p>This competition is challenging due to the huge dataset size. In addition, people may not be familiar with medical imaging and CT scans. </p>\n<p>To lower the barrier to entry to this competition and to encourage more people to participate, I have converted all of the DICOM images into JPEGs (about 52GB), 256x256 pixels (standard is 512x512). </p>\n<h1>Windowing</h1>\n<p>The values in CT scans tend to range from -1000 to 3000; we are used to 8-bit images with pixel values ranging from 0 to 255. Radiologists use a technique called <strong>windowing</strong> to visualize CT scans. Different types of tissues are better evaluated using different windows. Windows are defined by 2 numbers: window <strong>width</strong> and window <strong>level</strong>. </p>\n<p>Here is the function I use to window an image:</p>\n<pre><code>def window(img, WL=50, WW=350):\n    upper, lower = WL+WW//2, WL-WW//2\n    X = np.clip(img.copy(), lower, upper)\n    X = X - np.min(X)\n    X = X / np.max(X)\n    X = (X*255.0).astype('uint8')\n    return X\n</code></pre>\n<p>Basically, we compute the lower and upper limits using the window level (WL) and window width (WW). Upper limit is level plus half the width and lower limit is level minus half the width. </p>\n<p>Then, all values above the upper limit are set to the upper limit; vice versa for the lower limit. We can then normalize the images to [0, 1] and multiply by 255 to convert them into an 8-bit image. </p>\n<p>Note that these images are single channel. I have provided them in 3-channel RGB format. <strong>Each channel is a different window.</strong> </p>\n<ul>\n<li><strong>RED</strong> channel / <strong>LUNG</strong> window / level=-600, width=1500</li>\n<li><strong>GREEN</strong> channel / <strong>PE</strong> window / level=100, width=700</li>\n<li><strong>BLUE</strong> channel / <strong>MEDIASTINAL</strong> window / level=40, width=400</li>\n</ul>\n<p>Please remember that <code>cv2.imread</code> by default loads images in <strong>BGR</strong> so if you are using that function make sure you're not confusing the red and blue channels. The <strong>PE specific and mediastinal windows</strong> will be most useful for evaluating whether there is blood clot. However, the lung window can show abnormalities in the lungs which may be suggestive of a clot, so I included it as well. Feel free to experiment with any single window or combination of windows.</p>\n<h3>Lung</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F281652%2Faa4bd969f80498908e8f40bade1e69b7%2Flung_window.jpg?generation=1600128037814115&amp;alt=media\" alt=\"\"></p>\n<h3>PE Specific</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F281652%2F8e5d707a2daa4606b720de0c6030aba4%2Fpe_window.jpg?generation=1600128049440405&amp;alt=media\" alt=\"\"></p>\n<h3>Mediastinal</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F281652%2F8950b63595f0b36b6d52a48907d32ffb%2Fmediastinum_window.jpg?generation=1600128059340756&amp;alt=media\" alt=\"\"></p>\n<p>This has the disadvantage that you cannot select other windows and you cannot design a learnable windowing function using this data. <strong>However, based on my experience, I do not believe these are necessary to do well in this competition.</strong></p>\n<h1>3D Reconstruction</h1>\n<p>CT scans are 3D. Each image is 2D, but taking into context the surrounding image slices can improve performance (either using a fully 3D model, hybrid 3D-2D model, or sequence model). In order to facilitate reconstructing the 3D image, <strong>each image in each series (folder) has a prefix with the slice number (increasing from bottom to top).</strong> That way, you can easily organize the slices in 3D. </p>\n<p>This is the function I used to create the 3D array:</p>\n<pre><code>def load_dicom_array(f):\n    dicom_files = glob.glob(osp.join(f, '*.dcm'))\n    dicoms = [pydicom.dcmread(d) for d in dicom_files]\n    M = float(dicoms[0].RescaleSlope)\n    B = float(dicoms[0].RescaleIntercept)\n    # Assume all images are axial\n    z_pos = [float(d.ImagePositionPatient[-1]) for d in dicoms]\n    dicoms = np.asarray([d.pixel_array for d in dicoms])\n    dicoms = dicoms[np.argsort(z_pos)]\n    dicoms = dicoms * M\n    dicoms = dicoms + B\n    return dicoms, np.asarray(dicom_files)[np.argsort(z_pos)]\n</code></pre>\n<p>Important things to note: you need to use <code>RescaleSlope</code> and <code>RescaleIntercept</code> to alter the values from <code>dicom.pixel_array</code> first. In this dataset, <code>RescaleSlope</code> is always 1 and <code>RescaleIntercept</code> is 0 or -1024 (more common). Then, you can use <code>ImagePositionPatient</code> to sort the slices. The last element <code>ImagePositionPatient[-1]</code> is used to sort in the z-axis. </p>\n<p>Please let me know if you encounter any issues with this dataset. I will try my best to fix them. It took about 36 hours to generate this dataset. I will try to produce datasets of different sizes if people request it. </p>",
      "rawMarkdown": "Update: source code to produce these images has been added as attachment\n\nhttps://www.kaggle.com/vaillant/rsna-str-pe-detection-jpeg-256\n\n# Introduction\nThis competition is challenging due to the huge dataset size. In addition, people may not be familiar with medical imaging and CT scans. \n\nTo lower the barrier to entry to this competition and to encourage more people to participate, I have converted all of the DICOM images into JPEGs (about 52GB), 256x256 pixels (standard is 512x512). \n\n# Windowing\nThe values in CT scans tend to range from -1000 to 3000; we are used to 8-bit images with pixel values ranging from 0 to 255. Radiologists use a technique called **windowing** to visualize CT scans. Different types of tissues are better evaluated using different windows. Windows are defined by 2 numbers: window **width** and window **level**. \n\nHere is the function I use to window an image:\n```\ndef window(img, WL=50, WW=350):\n    upper, lower = WL+WW//2, WL-WW//2\n    X = np.clip(img.copy(), lower, upper)\n    X = X - np.min(X)\n    X = X / np.max(X)\n    X = (X*255.0).astype('uint8')\n    return X\n```\nBasically, we compute the lower and upper limits using the window level (WL) and window width (WW). Upper limit is level plus half the width and lower limit is level minus half the width. \n\nThen, all values above the upper limit are set to the upper limit; vice versa for the lower limit. We can then normalize the images to [0, 1] and multiply by 255 to convert them into an 8-bit image. \n\nNote that these images are single channel. I have provided them in 3-channel RGB format. **Each channel is a different window.** \n\n- **RED** channel / **LUNG** window / level=-600, width=1500\n- **GREEN** channel / **PE** window / level=100, width=700\n- **BLUE** channel / **MEDIASTINAL** window / level=40, width=400\n\nPlease remember that `cv2.imread` by default loads images in **BGR** so if you are using that function make sure you're not confusing the red and blue channels. The **PE specific and mediastinal windows** will be most useful for evaluating whether there is blood clot. However, the lung window can show abnormalities in the lungs which may be suggestive of a clot, so I included it as well. Feel free to experiment with any single window or combination of windows.\n\n### Lung\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F281652%2Faa4bd969f80498908e8f40bade1e69b7%2Flung_window.jpg?generation=1600128037814115&alt=media)\n### PE Specific\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F281652%2F8e5d707a2daa4606b720de0c6030aba4%2Fpe_window.jpg?generation=1600128049440405&alt=media)\n### Mediastinal\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F281652%2F8950b63595f0b36b6d52a48907d32ffb%2Fmediastinum_window.jpg?generation=1600128059340756&alt=media)\n\nThis has the disadvantage that you cannot select other windows and you cannot design a learnable windowing function using this data. **However, based on my experience, I do not believe these are necessary to do well in this competition.**\n\n# 3D Reconstruction\nCT scans are 3D. Each image is 2D, but taking into context the surrounding image slices can improve performance (either using a fully 3D model, hybrid 3D-2D model, or sequence model). In order to facilitate reconstructing the 3D image, **each image in each series (folder) has a prefix with the slice number (increasing from bottom to top).** That way, you can easily organize the slices in 3D. \n\nThis is the function I used to create the 3D array:\n```\ndef load_dicom_array(f):\n    dicom_files = glob.glob(osp.join(f, '*.dcm'))\n    dicoms = [pydicom.dcmread(d) for d in dicom_files]\n    M = float(dicoms[0].RescaleSlope)\n    B = float(dicoms[0].RescaleIntercept)\n    # Assume all images are axial\n    z_pos = [float(d.ImagePositionPatient[-1]) for d in dicoms]\n    dicoms = np.asarray([d.pixel_array for d in dicoms])\n    dicoms = dicoms[np.argsort(z_pos)]\n    dicoms = dicoms * M\n    dicoms = dicoms + B\n    return dicoms, np.asarray(dicom_files)[np.argsort(z_pos)]\n```\nImportant things to note: you need to use `RescaleSlope` and `RescaleIntercept` to alter the values from `dicom.pixel_array` first. In this dataset, `RescaleSlope` is always 1 and `RescaleIntercept` is 0 or -1024 (more common). Then, you can use `ImagePositionPatient` to sort the slices. The last element `ImagePositionPatient[-1]` is used to sort in the z-axis. \n\nPlease let me know if you encounter any issues with this dataset. I will try my best to fix them. It took about 36 hours to generate this dataset. I will try to produce datasets of different sizes if people request it. ",
      "votes": 192
    },
    {
      "id": 1011537,
      "postDate": "2020-09-15T14:42:46.857Z",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> for sharing the dataset! Without this, I would've just ignored the competition because there's no way for me to load 1TB of image data into 13GB of RAM ;)</p>\n<p>One small note: I think it's better if you change the license of your dataset to a Custom license (and refer to this competition) or a CC-BY-NC-SA 4.0 (if that's approved by the organizers). A CC0 (open domain) license might confuse people who are not familiar with Kaggle.</p>",
      "rawMarkdown": "Thanks @vaillant for sharing the dataset! Without this, I would've just ignored the competition because there's no way for me to load 1TB of image data into 13GB of RAM ;)\n\nOne small note: I think it's better if you change the license of your dataset to a Custom license (and refer to this competition) or a CC-BY-NC-SA 4.0 (if that's approved by the organizers). A CC0 (open domain) license might confuse people who are not familiar with Kaggle.",
      "votes": 6,
      "replies": [
        {
          "id": 1011902,
          "postDate": "2020-09-15T18:39:48.753Z",
          "content": "<p>Good catch, thanks! I will update it.</p>",
          "rawMarkdown": "Good catch, thanks! I will update it.",
          "votes": 4
        }
      ]
    },
    {
      "id": 1045494,
      "postDate": "2020-10-10T17:43:34.887Z",
      "content": "<p>Thanks so much for sharing this <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a>. Especially useful for those of us getting a late start to this competition.</p>\n<p>A couple of questions to you or anyone who sees this:</p>\n<p>1) What's with the <code>[5645:]</code>? I suppose you were just continuing where it left of at some point?<br>\n2) Any idea what the compression level is on the jpegs? Is it lossless?</p>",
      "rawMarkdown": "Thanks so much for sharing this @vaillant. Especially useful for those of us getting a late start to this competition.\n\nA couple of questions to you or anyone who sees this:\n\n1) What's with the `[5645:]`? I suppose you were just continuing where it left of at some point?\n2) Any idea what the compression level is on the jpegs? Is it lossless?",
      "votes": 3
    },
    {
      "id": 1016480,
      "postDate": "2020-09-19T02:33:55.003Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> </p>\n<p>Thought I would use your method to convert <code>.dcm</code> to <code>.jpg</code> files.</p>\n<p>But, in the example below, I am not sure if I can tell the difference whereas the first image has <code>LABEL=0</code> while the second image has <code>LABEL=1</code></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2116404%2Fd098a05e6c56272618c7f6c867be41ed%2FScreenshot%202020-09-19%20123222.png?generation=1600482785163541&amp;alt=media\" alt=\"\"></p>\n<p>Does this seem right?</p>",
      "rawMarkdown": "Hi @vaillant \n\nThought I would use your method to convert `.dcm` to `.jpg` files.\n\nBut, in the example below, I am not sure if I can tell the difference whereas the first image has `LABEL=0` while the second image has `LABEL=1`\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2116404%2Fd098a05e6c56272618c7f6c867be41ed%2FScreenshot%202020-09-19%20123222.png?generation=1600482785163541&alt=media)\n\nDoes this seem right?",
      "votes": 1,
      "replies": [
        {
          "id": 1017814,
          "postDate": "2020-09-19T09:27:57.243Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/aroraaman\" target=\"_blank\">@aroraaman</a> , Looking at the combined RGB channel view, I myself don't find any difference. Here there can be several possible steps in my opinion.  You can view three channels (RGB) separately and see if you can find any difference for label 1 and 0. </p>",
          "rawMarkdown": "Hi @aroraaman , Looking at the combined RGB channel view, I myself don't find any difference. Here there can be several possible steps in my opinion.  You can view three channels (RGB) separately and see if you can find any difference for label 1 and 0. \n"
        },
        {
          "id": 1017819,
          "postDate": "2020-09-19T09:30:20.703Z",
          "content": "<p><a href=\"https://www.kaggle.com/aroraaman\" target=\"_blank\">@aroraaman</a>, you can see the corresponding original DICOM and try to find out any difference. Even if you don't find anything there, you can check neighboring DICOM slices to these slices in question to check whether they have PE. … <br>\ngood luck..  </p>",
          "rawMarkdown": "@aroraaman, you can see the corresponding original DICOM and try to find out any difference. Even if you don't find anything there, you can check neighboring DICOM slices to these slices in question to check whether they have PE. ... \ngood luck..  ",
          "votes": 2
        },
        {
          "id": 1018112,
          "postDate": "2020-09-19T13:03:22.157Z",
          "content": "<p>Can you provide the filename of these images?</p>",
          "rawMarkdown": "Can you provide the filename of these images?",
          "votes": 1
        },
        {
          "id": 1047029,
          "postDate": "2020-10-12T07:27:08.613Z",
          "content": "<p><a href=\"https://www.kaggle.com/aroraaman\" target=\"_blank\">@aroraaman</a> i think you should try to visualize the image with no PE  using only PE channel.. i feel hte marked white contrasts across the encircle are clots,if a channel is not having PE then it wont show up in the  stacked image also</p>",
          "rawMarkdown": "@aroraaman i think you should try to visualize the image with no PE  using only PE channel.. i feel hte marked white contrasts across the encircle are clots,if a channel is not having PE then it wont show up in the  stacked image also"
        }
      ]
    },
    {
      "id": 1015179,
      "postDate": "2020-09-18T02:54:08.213Z",
      "content": "<p>I didn't know this PE window. Where did you get these values? Thanks</p>",
      "rawMarkdown": "I didn't know this PE window. Where did you get these values? Thanks",
      "votes": 1,
      "replies": [
        {
          "id": 1017837,
          "postDate": "2020-09-19T09:40:47.083Z",
          "content": "<p>Hi, <a href=\"https://www.kaggle.com/ronaldokun\" target=\"_blank\">@ronaldokun</a> , Different window parameters are available in different public documents. You can easily find them in the internet. Here is one of the links.. <br>\n<a href=\"https://radiopaedia.org/articles/windowing-ct\" target=\"_blank\">Windowing (CT)</a></p>",
          "rawMarkdown": "Hi, @ronaldokun , Different window parameters are available in different public documents. You can easily find them in the internet. Here is one of the links.. \n[Windowing (CT)](https://radiopaedia.org/articles/windowing-ct)",
          "votes": 2
        },
        {
          "id": 1019916,
          "postDate": "2020-09-20T18:43:50.947Z",
          "content": "<p>Thank you!</p>",
          "rawMarkdown": "Thank you!"
        }
      ]
    },
    {
      "id": 1012768,
      "postDate": "2020-09-16T09:24:04.033Z",
      "content": "<p><a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> can you provide the full source code version to produce these images? thanks in advance.</p>",
      "rawMarkdown": "@vaillant can you provide the full source code version to produce these images? thanks in advance.",
      "votes": 1,
      "replies": [
        {
          "id": 1013430,
          "postDate": "2020-09-16T17:30:50.387Z",
          "content": "<p>same question here! Thanks so much!</p>",
          "rawMarkdown": "same question here! Thanks so much!"
        },
        {
          "id": 1013486,
          "postDate": "2020-09-16T18:03:15.437Z",
          "content": "<p>Let me know if you are able to access this attachment.</p>",
          "rawMarkdown": "Let me know if you are able to access this attachment.",
          "votes": 3
        },
        {
          "id": 1013494,
          "postDate": "2020-09-16T18:07:07.757Z",
          "content": "<p>Yeah, I am able to access it. Thanks.</p>",
          "rawMarkdown": "Yeah, I am able to access it. Thanks."
        },
        {
          "id": 1014286,
          "postDate": "2020-09-17T10:16:40.897Z",
          "content": "<p><a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> thanks for the script but I think you should make use of parallelism to speed up the process :)</p>",
          "rawMarkdown": "@vaillant thanks for the script but I think you should make use of parallelism to speed up the process :)"
        },
        {
          "id": 1014678,
          "postDate": "2020-09-17T16:33:14.180Z",
          "content": "<p>Yaa, I took me 10 hrs to create 512x512x3 image with parallelism using 8 cores</p>",
          "rawMarkdown": "Yaa, I took me 10 hrs to create 512x512x3 image with parallelism using 8 cores",
          "votes": 1
        },
        {
          "id": 1015161,
          "postDate": "2020-09-18T02:25:27.117Z",
          "content": "<p>I don't have much experience with that so feel free to implement it and share it with us. </p>",
          "rawMarkdown": "I don't have much experience with that so feel free to implement it and share it with us. ",
          "votes": 2
        },
        {
          "id": 1015390,
          "postDate": "2020-09-18T06:45:46.553Z",
          "content": "<p>Thanks so much!</p>",
          "rawMarkdown": "Thanks so much!",
          "votes": 1
        },
        {
          "id": 1017811,
          "postDate": "2020-09-19T09:25:11.080Z",
          "content": "<p><a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> I have added parallelism to your data generation code. It is showing 1.5 hrs to generate data for me using 4 cores. Please have a look. <a href=\"https://gist.github.com/SumanSudhir/7f3b72510eb2b8cff080a225334f5dc4\" target=\"_blank\">https://gist.github.com/SumanSudhir/7f3b72510eb2b8cff080a225334f5dc4</a></p>",
          "rawMarkdown": "@vaillant I have added parallelism to your data generation code. It is showing 1.5 hrs to generate data for me using 4 cores. Please have a look. [https://gist.github.com/SumanSudhir/7f3b72510eb2b8cff080a225334f5dc4](https://gist.github.com/SumanSudhir/7f3b72510eb2b8cff080a225334f5dc4)",
          "votes": 16
        },
        {
          "id": 1017834,
          "postDate": "2020-09-19T09:38:30.090Z",
          "content": "<p>Thanks a lot, <a href=\"https://www.kaggle.com/sudhiriitb\" target=\"_blank\">@sudhiriitb</a> for your nice tweak.<br>\nI am stealing some of your code for my personal project if  that's ok. 😃😃</p>",
          "rawMarkdown": "Thanks a lot, @sudhiriitb for your nice tweak.\nI am stealing some of your code for my personal project if  that's ok. 😃😃"
        },
        {
          "id": 1017842,
          "postDate": "2020-09-19T09:42:48.690Z",
          "content": "<p>of course, you can, but it not mine its <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> code😹</p>",
          "rawMarkdown": "of course, you can, but it not mine its @vaillant code😹"
        },
        {
          "id": 1017897,
          "postDate": "2020-09-19T10:22:58.687Z",
          "content": "<p>🤓🤓🤓 xOOx</p>",
          "rawMarkdown": "🤓🤓🤓 xOOx"
        },
        {
          "id": 1018363,
          "postDate": "2020-09-19T16:28:47.813Z",
          "content": "<p>Excellent. Thank you!</p>",
          "rawMarkdown": "Excellent. Thank you!"
        },
        {
          "id": 1047020,
          "postDate": "2020-10-12T07:22:21.027Z",
          "content": "<p><a href=\"https://www.kaggle.com/sudhiriitb\" target=\"_blank\">@sudhiriitb</a>  are these dstacked channel images helping get score ?</p>\n<p><a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a>    Is medistenial same as Subdural ?</p>",
          "rawMarkdown": "@sudhiriitb  are these dstacked channel images helping get score ?\n\n@vaillant    Is medistenial same as Subdural ?"
        },
        {
          "id": 1050831,
          "postDate": "2020-10-15T19:06:59.643Z",
          "content": "<p><a href=\"https://www.kaggle.com/sudhiriitb\" target=\"_blank\">@sudhiriitb</a> thanks for the parallel code! I'm having one problem with it though. For some reason it's not releasing memory and I run out at about 39 iterations. Any experience with such a bug?</p>",
          "rawMarkdown": "@sudhiriitb thanks for the parallel code! I'm having one problem with it though. For some reason it's not releasing memory and I run out at about 39 iterations. Any experience with such a bug?"
        },
        {
          "id": 1050871,
          "postDate": "2020-10-15T20:06:14.100Z",
          "content": "<p><a href=\"https://www.kaggle.com/alexandersoare\" target=\"_blank\">@alexandersoare</a> are you running it as a python script or running in a notebook?</p>",
          "rawMarkdown": "@alexandersoare are you running it as a python script or running in a notebook?"
        }
      ]
    },
    {
      "id": 1011979,
      "postDate": "2020-09-15T20:00:37.910Z",
      "content": "<blockquote>\n  <p>I have provided them in 3-channel RGB format. Each channel is a different window.</p>\n</blockquote>\n<p>I am asking a silly question because of my little knowledge. Why you are using the three mentioned windows (Lung, PE and MEDIASTINAL ) as we have other windows like chest, abdomen and spine? </p>",
      "rawMarkdown": ">  I have provided them in 3-channel RGB format. Each channel is a different window.\n\nI am asking a silly question because of my little knowledge. Why you are using the three mentioned windows (Lung, PE and MEDIASTINAL ) as we have other windows like chest, abdomen and spine? ",
      "votes": 1,
      "replies": [
        {
          "id": 1012007,
          "postDate": "2020-09-15T20:45:28.243Z",
          "content": "<p>These are the windows I believe to be most useful based on domain knowledge.</p>",
          "rawMarkdown": "These are the windows I believe to be most useful based on domain knowledge.",
          "votes": 7
        }
      ]
    },
    {
      "id": 1011809,
      "postDate": "2020-09-15T17:32:25.437Z",
      "content": "<p>You guys think Colab Pro with TPU using Google Bucket for dataset would work for this? Is Pytorch models better than tensorflow? Cant use pytorch on TPU right? Thanks,</p>",
      "rawMarkdown": "You guys think Colab Pro with TPU using Google Bucket for dataset would work for this? Is Pytorch models better than tensorflow? Cant use pytorch on TPU right? Thanks,",
      "votes": 1,
      "replies": [
        {
          "id": 1011815,
          "postDate": "2020-09-15T17:36:39.120Z",
          "content": "<p>You can use pytorch-xla: <a href=\"https://github.com/pytorch/xla\" target=\"_blank\">https://github.com/pytorch/xla</a></p>\n<p>there's many great notebooks that train transformer models on TPUs</p>",
          "rawMarkdown": "You can use pytorch-xla: https://github.com/pytorch/xla\n\nthere's many great notebooks that train transformer models on TPUs",
          "votes": 3
        },
        {
          "id": 1011818,
          "postDate": "2020-09-15T17:39:25.203Z",
          "content": "<p>Thanks for the share Xing! So that is mentioning you can use pytorch with TPUs correct? Thanks,</p>",
          "rawMarkdown": "Thanks for the share Xing! So that is mentioning you can use pytorch with TPUs correct? Thanks,",
          "votes": 1
        },
        {
          "id": 1011826,
          "postDate": "2020-09-15T17:42:49.663Z",
          "content": "<p>Yep! It should work; a good starting point is the jigsaw multilingual tpu competition with a lot of pytorch tpu kernels</p>",
          "rawMarkdown": "Yep! It should work; a good starting point is the jigsaw multilingual tpu competition with a lot of pytorch tpu kernels",
          "votes": 2
        },
        {
          "id": 1011829,
          "postDate": "2020-09-15T17:45:31.867Z",
          "content": "<p>Nice, thanks Xing.</p>",
          "rawMarkdown": "Nice, thanks Xing."
        }
      ]
    },
    {
      "id": 1012346,
      "postDate": "2020-09-16T03:25:36.007Z",
      "content": "<p>Lan, Many Thanks! </p>\n<p>I'm wondering if LUNG window level should be <strong>600</strong> not <strong>-600</strong> based on the description \"Our CT techniques are shown in the ,Table. Images are displayed with three different gray scales for interpretation of lung window (window width/level [HU] = 1500/600), mediastinal window (400/40), and pulmonary embolism–specific (700/100) settings.\" from the RSNA <a href=\"https://pubs.rsna.org/doi/full/10.1148/rg.245045008\" target=\"_blank\">paper </a></p>",
      "rawMarkdown": "Lan, Many Thanks! \n\nI'm wondering if LUNG window level should be **600** not **-600** based on the description \"Our CT techniques are shown in the ,Table. Images are displayed with three different gray scales for interpretation of lung window (window width/level [HU] = 1500/600), mediastinal window (400/40), and pulmonary embolism–specific (700/100) settings.\" from the RSNA [paper ](https://pubs.rsna.org/doi/full/10.1148/rg.245045008)",
      "votes": 2,
      "replies": [
        {
          "id": 1013489,
          "postDate": "2020-09-16T18:05:02.420Z",
          "content": "<p>I used the windows from that paper, 600 is a typo. It should be -600, as that is about the HU of lung parenchyma. </p>",
          "rawMarkdown": "I used the windows from that paper, 600 is a typo. It should be -600, as that is about the HU of lung parenchyma. ",
          "votes": 7
        }
      ]
    },
    {
      "id": 1053520,
      "postDate": "2020-10-19T04:52:56.290Z",
      "content": "<p><a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> Thank You for sharing this excellent usefull information.<br>\nHow to run this file convert_to_jpeg_for_kaggle.py on Kaggle notebook?<br>\nBecasue I got error during saving the images.<br>\nAny idea please help me.</p>",
      "rawMarkdown": "@vaillant Thank You for sharing this excellent usefull information.\nHow to run this file convert_to_jpeg_for_kaggle.py on Kaggle notebook?\nBecasue I got error during saving the images.\nAny idea please help me.",
      "votes": -1
    },
    {
      "id": 1059395,
      "postDate": "2020-10-25T03:29:30.667Z",
      "content": "<p>hai all… anyone cross 0.20 with this dataset</p>",
      "rawMarkdown": "hai all... anyone cross 0.20 with this dataset",
      "replies": [
        {
          "id": 1059588,
          "postDate": "2020-10-25T08:53:18.477Z",
          "content": "<p>I think some would have.<br>\nMy current score 0.203 is with this dataset and could cross 0.2 if I have the time and energy over next 2 days.</p>",
          "rawMarkdown": "I think some would have.\nMy current score 0.203 is with this dataset and could cross 0.2 if I have the time and energy over next 2 days.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1046958,
      "postDate": "2020-10-12T06:18:19.903Z",
      "content": "<p>If anyone is facing issues like me in the extraction of images on Windows please change this function and it will work properly. </p>\n<pre><code>def edit_filenames(files):\n    dicoms = ['{:04d}_{}'.format(ind,f.split(\"\\\\\")[-1].replace(\"dcm\",\"jpg\")) for ind,f in enumerate(files)]\n    series = ['/'.join(f.split('/')[-3:-1]) for f in files]\n    return [osp.join(s,d) for s,d in zip(series, dicoms)]\n</code></pre>",
      "rawMarkdown": "If anyone is facing issues like me in the extraction of images on Windows please change this function and it will work properly. \n\n```\ndef edit_filenames(files):\n    dicoms = ['{:04d}_{}'.format(ind,f.split(\"\\\\\")[-1].replace(\"dcm\",\"jpg\")) for ind,f in enumerate(files)]\n    series = ['/'.join(f.split('/')[-3:-1]) for f in files]\n    return [osp.join(s,d) for s,d in zip(series, dicoms)]\n```"
    },
    {
      "id": 1035975,
      "postDate": "2020-10-03T09:54:19.793Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> <br>\nThank you very much for your post. I was trying to use this method, but when I execute the code I just get a lot of 'list index out of range' through the console and it finishes in a second. I'm running it in a kaggle kernel, I only changed the route to the csv and the train files, the rest is exactly as you posted. Any idea? I'm quite new on this I'm not sure if I'm missing something important.</p>\n<p>Regards</p>",
      "rawMarkdown": "Hi @vaillant \nThank you very much for your post. I was trying to use this method, but when I execute the code I just get a lot of 'list index out of range' through the console and it finishes in a second. I'm running it in a kaggle kernel, I only changed the route to the csv and the train files, the rest is exactly as you posted. Any idea? I'm quite new on this I'm not sure if I'm missing something important.\n\nRegards",
      "replies": [
        {
          "id": 1036442,
          "postDate": "2020-10-03T18:42:17.417Z",
          "content": "<p>You should post the snippet of code which returns an error.</p>",
          "rawMarkdown": "You should post the snippet of code which returns an error.",
          "votes": 1
        },
        {
          "id": 1039293,
          "postDate": "2020-10-06T13:26:47.020Z",
          "content": "<p>You're right, sorry. Anyways, the error was due to trying to write in a route Kaggle environment doesn't allow. Noob mistake, I guess 😅</p>\n<p>Thanks for your answer!</p>",
          "rawMarkdown": "You're right, sorry. Anyways, the error was due to trying to write in a route Kaggle environment doesn't allow. Noob mistake, I guess 😅\n\nThanks for your answer!"
        }
      ]
    },
    {
      "id": 1029806,
      "postDate": "2020-09-28T06:36:08.430Z",
      "content": "<p>Thank you!<br>\nI just wanted to ask can someone reduce the datasets,and then I saw your notebook</p>",
      "rawMarkdown": "Thank you!\nI just wanted to ask can someone reduce the datasets,and then I saw your notebook"
    },
    {
      "id": 1024703,
      "postDate": "2020-09-24T04:38:21.343Z",
      "content": "<p>Thanks for sharing I was actually checking this paper before seeing your post, same 3 windowing for diagnosis is used! <a href=\"https://pubs.rsna.org/doi/pdf/10.1148/rg.245045008\" target=\"_blank\">https://pubs.rsna.org/doi/pdf/10.1148/rg.245045008</a></p>",
      "rawMarkdown": "Thanks for sharing I was actually checking this paper before seeing your post, same 3 windowing for diagnosis is used! https://pubs.rsna.org/doi/pdf/10.1148/rg.245045008"
    },
    {
      "id": 1017852,
      "postDate": "2020-09-19T09:50:02.360Z",
      "content": "<p>Thanks, <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> for sharing the idea of using three channels. Previously, I was just the whole lung segment with lung window (window width/level [HU] = 1500/600). This three windows in three channels literally packing ~3x information compared to regular single channel grayscales.. </p>\n<p>Great work!!!!</p>",
      "rawMarkdown": "Thanks, @vaillant for sharing the idea of using three channels. Previously, I was just the whole lung segment with lung window (window width/level [HU] = 1500/600). This three windows in three channels literally packing ~3x information compared to regular single channel grayscales.. \n\nGreat work!!!!"
    },
    {
      "id": 1015789,
      "postDate": "2020-09-18T12:25:21.573Z",
      "content": "<p>Thank you for the effort - <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> :)</p>\n<p>Quick question please, what does the following line mean: </p>\n<blockquote>\n  <p>Assume all images are axial</p>\n</blockquote>\n<p>Also, I am new to medical imaging so spent the past two days reading about CT scans. Wish to corroborate my understanding from yourself if possible please. </p>\n<p>If I understand the CT scanner essentially takes multiple 2-D images which are indexed. The 2-D images themselves are referred as \"slices\" ? Therefore, if slice at t'th position is <code>S(t)</code> then <code>S(t-1)</code> and <code>S(t+1)</code> are slice before and after and we could leverage this information to create better models. </p>\n<p>Does that seem right? Thanks! :)</p>",
      "rawMarkdown": "Thank you for the effort - @vaillant :)\n\nQuick question please, what does the following line mean: \n> Assume all images are axial\n\nAlso, I am new to medical imaging so spent the past two days reading about CT scans. Wish to corroborate my understanding from yourself if possible please. \n\nIf I understand the CT scanner essentially takes multiple 2-D images which are indexed. The 2-D images themselves are referred as \"slices\" ? Therefore, if slice at t'th position is `S(t)` then `S(t-1)` and `S(t+1)` are slice before and after and we could leverage this information to create better models. \n\nDoes that seem right? Thanks! :)",
      "replies": [
        {
          "id": 1015870,
          "postDate": "2020-09-18T13:50:19.157Z",
          "content": "<p>Assume all images are axial - Typically a CT scan takes what we call \"axial\" images. Like slicing a loaf of bread. The CT scanners can also take the pixels from that scan and reconstruct them in Sagittal (like looking at someone from the side) or Coronal (like looking at someone from in front of them) planes. We don't have those for this competition, although you could built a 3-D model from the axial slices and create them yourself.</p>\n<p>Somethings things are easier to see on Sagittal or Coronal images, but typically the radiologist mainly looks at the Axial images.</p>\n<p>Your idea on positioning is correct. The DICOM metadata defines the orientation and relationship of each slice in a 3-dimensional space.</p>\n<p>Copying from my prior post:</p>\n<p>Patient position - FFP (Feet first prone), FFS (feet first supine), HFS (head first supine). Refers to the position of the patient in the machine. In the train dataset:</p>\n<p>FFP 242<br>\nFFS 1732839<br>\nHFS 57513</p>\n<p>In theory, this could affect the patient position. The FFP (prone) images suggest the patient is lying on their belly. Supposedly you would want to flip these images (but I haven't looked at the DICOM to see if this is true).</p>\n<p>InstanceNumber - a unique number for each image within a Series. In these series, usually starts at 1 and counts up in the physical order of the images. There are Studies that don't start at zero. I don't know if there are cases where numbers are skipped. Could count from top to bottom or bottom to top. </p>\n<p>Image Patient Position - three floating point numbers that define a relative location in the real world. Image Patient Position [2] is the \"Z-axis\" and can be used to put the images in order.</p>\n<p>Image Orientation - another parameter that defines the orientation of the patient/images in space. Might be relevant for building three dimensional models and defining top/bottom of patient. I haven't checked.</p>\n<p>You can use these parameters to reassemble the images in order and build a 3D model. Note that you are not guaranteed that the images give you a physically complete image. Nothing stops you from having gaps in the images. This would not be normal processing, but since we cannot see the real test data, we cannot check this.</p>\n<p>Also, slice thickness looks consistent within each scan, but varies between Studies.</p>\n<p>In theory, if you put them in order by InstanceNumber or Image Patient Position [2], the change in position would equal the Slice Thickness. But in practice, you could have gaps or even overlaps (5 mm slices taken every 2.5 mm). How far you go to cover the possibilities depends on your model needs. If you look at the data in the Pulmonary Fibrosis competition, the CT images have missing images, gaps in images, repeated imaging through the patient in one Study, etc.</p>\n<p>-Rich</p>",
          "rawMarkdown": "Assume all images are axial - Typically a CT scan takes what we call \"axial\" images. Like slicing a loaf of bread. The CT scanners can also take the pixels from that scan and reconstruct them in Sagittal (like looking at someone from the side) or Coronal (like looking at someone from in front of them) planes. We don't have those for this competition, although you could built a 3-D model from the axial slices and create them yourself.\n\nSomethings things are easier to see on Sagittal or Coronal images, but typically the radiologist mainly looks at the Axial images.\n\nYour idea on positioning is correct. The DICOM metadata defines the orientation and relationship of each slice in a 3-dimensional space.\n\nCopying from my prior post:\n\nPatient position - FFP (Feet first prone), FFS (feet first supine), HFS (head first supine). Refers to the position of the patient in the machine. In the train dataset:\n\nFFP 242\nFFS 1732839\nHFS 57513\n\nIn theory, this could affect the patient position. The FFP (prone) images suggest the patient is lying on their belly. Supposedly you would want to flip these images (but I haven't looked at the DICOM to see if this is true).\n\nInstanceNumber - a unique number for each image within a Series. In these series, usually starts at 1 and counts up in the physical order of the images. There are Studies that don't start at zero. I don't know if there are cases where numbers are skipped. Could count from top to bottom or bottom to top. \n\nImage Patient Position - three floating point numbers that define a relative location in the real world. Image Patient Position [2] is the \"Z-axis\" and can be used to put the images in order.\n\nImage Orientation - another parameter that defines the orientation of the patient/images in space. Might be relevant for building three dimensional models and defining top/bottom of patient. I haven't checked.\n\nYou can use these parameters to reassemble the images in order and build a 3D model. Note that you are not guaranteed that the images give you a physically complete image. Nothing stops you from having gaps in the images. This would not be normal processing, but since we cannot see the real test data, we cannot check this.\n\nAlso, slice thickness looks consistent within each scan, but varies between Studies.\n\nIn theory, if you put them in order by InstanceNumber or Image Patient Position [2], the change in position would equal the Slice Thickness. But in practice, you could have gaps or even overlaps (5 mm slices taken every 2.5 mm). How far you go to cover the possibilities depends on your model needs. If you look at the data in the Pulmonary Fibrosis competition, the CT images have missing images, gaps in images, repeated imaging through the patient in one Study, etc.\n\n-Rich",
          "votes": 6
        },
        {
          "id": 1015942,
          "postDate": "2020-09-18T14:42:15.353Z",
          "content": "<p>Thank you. So in theory, each Series could have different number of slices (SOPInstanceID) and we could then leverage 3-D CNN or sequential model to predict PE, right? :)</p>\n<p>Again, thanks a lot for being so active on the discussions and sharing your knowledge. :)</p>",
          "rawMarkdown": "Thank you. So in theory, each Series could have different number of slices (SOPInstanceID) and we could then leverage 3-D CNN or sequential model to predict PE, right? :)\n\nAgain, thanks a lot for being so active on the discussions and sharing your knowledge. :)",
          "replies": [
            {
              "id": 1015996,
              "postDate": "2020-09-18T15:34:26.977Z",
              "content": "<p>Yes. Slice thickness typically range from 1.25 mm to 5 mm. The training set has number of images per study ranging from 63 to 1083. Presumably the 63 slices are 5 mm slices. The 1083 slices are presumably 1.25 mm slices.</p>\n<p>No guarantee that a series doesn't include two passes through the chest.</p>\n<p>Also, at least one study is actually mainly an abdomen and pelvis study. It captures the bottom of the chest, but most of the study is abdomen and pelvis.</p>\n<p>There could be a chest+abdomen+pelvis study also.</p>\n<p>Most studies appear to be a standard Chest study.</p>\n<p>There is a free to download software called K-Pacs which lets you look at the studies in DICOM format.</p>",
              "rawMarkdown": "Yes. Slice thickness typically range from 1.25 mm to 5 mm. The training set has number of images per study ranging from 63 to 1083. Presumably the 63 slices are 5 mm slices. The 1083 slices are presumably 1.25 mm slices.\n\nNo guarantee that a series doesn't include two passes through the chest.\n\nAlso, at least one study is actually mainly an abdomen and pelvis study. It captures the bottom of the chest, but most of the study is abdomen and pelvis.\n\nThere could be a chest+abdomen+pelvis study also.\n\nMost studies appear to be a standard Chest study.\n\nThere is a free to download software called K-Pacs which lets you look at the studies in DICOM format.",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 1012786,
      "postDate": "2020-09-16T09:35:06.843Z",
      "content": "<p>Windowing - Exposed it. Great Sharing</p>",
      "rawMarkdown": "Windowing - Exposed it. Great Sharing"
    },
    {
      "id": 1011925,
      "postDate": "2020-09-15T19:00:03.143Z",
      "content": "<p><a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> You did not provide jpeg for the test set. I think we should need test set in jpeg too to make the submission file easily.</p>",
      "rawMarkdown": "@vaillant You did not provide jpeg for the test set. I think we should need test set in jpeg too to make the submission file easily.",
      "replies": [
        {
          "id": 1011945,
          "postDate": "2020-09-15T19:19:44.390Z",
          "content": "<p>Part of the test set is missing and will only be available when you submit your notebook.</p>",
          "rawMarkdown": "Part of the test set is missing and will only be available when you submit your notebook."
        },
        {
          "id": 1012006,
          "postDate": "2020-09-15T20:45:12.810Z",
          "content": "<p>As <a href=\"https://www.kaggle.com/xhlulu\" target=\"_blank\">@xhlulu</a> mentioned, you will need to use the functions I provided above to load in the DICOMs and convert them appropriately in your inference kernel. The private test set is computed on test data that cannot be accessed unless you submit a kernel. </p>",
          "rawMarkdown": "As @xhlulu mentioned, you will need to use the functions I provided above to load in the DICOMs and convert them appropriately in your inference kernel. The private test set is computed on test data that cannot be accessed unless you submit a kernel. ",
          "votes": 1
        },
        {
          "id": 1017849,
          "postDate": "2020-09-19T09:45:16.670Z",
          "content": "<p>The competition says that it's an <strong>inference-only</strong> competition. That means training should be done separately and use that pre-trained model on the test set given. The test set available with the data is just for demonstration. Your submitted kernel will be re-run on the hidden-kernel to produce the final submission result. </p>",
          "rawMarkdown": "The competition says that it's an **inference-only** competition. That means training should be done separately and use that pre-trained model on the test set given. The test set available with the data is just for demonstration. Your submitted kernel will be re-run on the hidden-kernel to produce the final submission result. ",
          "votes": 1,
          "replies": [
            {
              "id": 1017970,
              "postDate": "2020-09-19T11:12:03.003Z",
              "content": "<p>The public test set is about 40 percent of the real test set. So it is not just for demonstration. I note, in the Pulmonary Fibrosis competition,  the test set was just for demonstration. </p>\n<p>I am successfully doing inference on the public test data and then combining that submission file in a committed notebook with the rest of the private test data to make a submission. </p>",
              "rawMarkdown": "The public test set is about 40 percent of the real test set. So it is not just for demonstration. I note, in the Pulmonary Fibrosis competition,  the test set was just for demonstration. \n\nI am successfully doing inference on the public test data and then combining that submission file in a committed notebook with the rest of the private test data to make a submission. "
            },
            {
              "id": 1018755,
              "postDate": "2020-09-20T00:04:04.760Z",
              "content": "<p><a href=\"https://www.kaggle.com/richardepstein\" target=\"_blank\">@richardepstein</a> May I ask how you solve the GCDM problem during inference? Since even we can generate jpegs/tfrecords of testset in other kernel(with internet), but there are only 40% of testset.</p>",
              "rawMarkdown": "@richardepstein May I ask how you solve the GCDM problem during inference? Since even we can generate jpegs/tfrecords of testset in other kernel(with internet), but there are only 40% of testset."
            },
            {
              "id": 1018761,
              "postDate": "2020-09-20T00:25:35.013Z",
              "content": "<p>I haven't solved it. I use a try/except construct around my access to the pydicom pixel data. If I have an error, I use a blank image as a placeholder.</p>\n<p>I think maybe 5% or less of the images have this error, although I haven't run an actual test.</p>\n<p>Eventually, we'll need to solve this (or at least try). I suspect you could download the GDCM \"files\" into a dataset and import it during inference. I haven't done that yet. I cannot even get GDCM to install on my development machine (but it does work in kernels with Internet access.</p>",
              "rawMarkdown": "I haven't solved it. I use a try/except construct around my access to the pydicom pixel data. If I have an error, I use a blank image as a placeholder.\n\nI think maybe 5% or less of the images have this error, although I haven't run an actual test.\n\nEventually, we'll need to solve this (or at least try). I suspect you could download the GDCM \"files\" into a dataset and import it during inference. I haven't done that yet. I cannot even get GDCM to install on my development machine (but it does work in kernels with Internet access.",
              "votes": 1
            },
            {
              "id": 1018796,
              "postDate": "2020-09-20T01:40:10.853Z",
              "content": "<p>Thanks for the reply.<br>\nI think maybe kaggle should consider add GCDM in their default environment or something. </p>",
              "rawMarkdown": "Thanks for the reply.\nI think maybe kaggle should consider add GCDM in their default environment or something. "
            },
            {
              "id": 1018857,
              "postDate": "2020-09-20T03:44:23.420Z",
              "content": "<p>I think I solved this in this notebook:</p>\n<p><a href=\"https://www.kaggle.com/richardepstein/load-gdcm-in-notebook-without-internet\" target=\"_blank\">Load GDCM in Notebook without Internet</a></p>\n<p>It uses a Dataset that somebody else created that has the tar file for GDCM. </p>\n<p>It seems to work both interactively and in a Committed Notebook. I haven't yet tried it in an actual competition submission, but it runs without Internet access, so it should work.</p>\n<p>Note that once you have the \"GDCM needed\" error in a notebook, you must restart the Kernel. So it should be the first thing in your notebook.</p>\n<p>Also, you might need to wait to import Pydicom until after it is run, but I have not tested that.</p>\n<p>Note that you do not \"import GDCM\".</p>\n<p>Let me know if it works for you.</p>\n<p>-Rich</p>",
              "rawMarkdown": "I think I solved this in this notebook:\n\n[Load GDCM in Notebook without Internet](https://www.kaggle.com/richardepstein/load-gdcm-in-notebook-without-internet)\n\nIt uses a Dataset that somebody else created that has the tar file for GDCM. \n\nIt seems to work both interactively and in a Committed Notebook. I haven't yet tried it in an actual competition submission, but it runs without Internet access, so it should work.\n\nNote that once you have the \"GDCM needed\" error in a notebook, you must restart the Kernel. So it should be the first thing in your notebook.\n\nAlso, you might need to wait to import Pydicom until after it is run, but I have not tested that.\n\nNote that you do not \"import GDCM\".\n\nLet me know if it works for you.\n\n-Rich",
              "votes": 3
            },
            {
              "id": 1018973,
              "postDate": "2020-09-20T05:38:47.720Z",
              "content": "<p><a href=\"https://www.kaggle.com/richardepstein\" target=\"_blank\">@richardepstein</a> It totally works! Thank you so much for the kind sharing!</p>",
              "rawMarkdown": "@richardepstein It totally works! Thank you so much for the kind sharing!"
            }
          ]
        }
      ]
    },
    {
      "id": 1010807,
      "postDate": "2020-09-15T04:23:02.027Z",
      "content": "<p>Thanks for sharing the dataset and the insights! I seem to be unable to download it. Could anyone else download?</p>",
      "rawMarkdown": "Thanks for sharing the dataset and the insights! I seem to be unable to download it. Could anyone else download?",
      "replies": [
        {
          "id": 1010858,
          "postDate": "2020-09-15T05:26:28.130Z",
          "content": "<p>nvm, got the download option now!</p>",
          "rawMarkdown": "nvm, got the download option now!"
        }
      ]
    },
    {
      "id": 1010645,
      "postDate": "2020-09-15T00:39:29.243Z",
      "content": "<p>Thanks for sharing preprocessed images and insights Ian! I would not have known about the PE and mediastinal windows. I just used the metadata windows which just so happens to be mediastinal window.</p>",
      "rawMarkdown": "Thanks for sharing preprocessed images and insights Ian! I would not have known about the PE and mediastinal windows. I just used the metadata windows which just so happens to be mediastinal window."
    },
    {
      "id": 1031850,
      "postDate": "2020-09-29T18:18:36.603Z",
      "content": "<p>This is really helpful. Thanks.</p>",
      "rawMarkdown": "This is really helpful. Thanks."
    },
    {
      "id": 1025736,
      "postDate": "2020-09-24T18:22:39.407Z",
      "content": "<p>Thanks for sharing , very useful</p>",
      "rawMarkdown": "Thanks for sharing , very useful"
    }
  ],
  "comments": [
    {
      "id": 1011537,
      "author_name": "xhlulu",
      "author_url": "",
      "post_date": "2020-09-15T14:42:46.857000",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> for sharing the dataset! Without this, I would've just ignored the competition because there's no way for me to load 1TB of image data into 13GB of RAM ;)</p>\n<p>One small note: I think it's better if you change the license of your dataset to a Custom license (and refer to this competition) or a CC-BY-NC-SA 4.0 (if that's approved by the organizers). A CC0 (open domain) license might confuse people who are not familiar with Kaggle.</p>",
      "votes": 6,
      "replies": [
        {
          "id": 1011902,
          "author_name": "Ian Pan",
          "author_url": "",
          "post_date": "2020-09-15T18:39:48.753000",
          "content": "<p>Good catch, thanks! I will update it.</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 1045494,
      "author_name": "Alexander Soare",
      "author_url": "",
      "post_date": "2020-10-10T17:43:34.887000",
      "content": "<p>Thanks so much for sharing this <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a>. Especially useful for those of us getting a late start to this competition.</p>\n<p>A couple of questions to you or anyone who sees this:</p>\n<p>1) What's with the <code>[5645:]</code>? I suppose you were just continuing where it left of at some point?<br>\n2) Any idea what the compression level is on the jpegs? Is it lossless?</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1016480,
      "author_name": "Aman Arora",
      "author_url": "",
      "post_date": "2020-09-19T02:33:55.003000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> </p>\n<p>Thought I would use your method to convert <code>.dcm</code> to <code>.jpg</code> files.</p>\n<p>But, in the example below, I am not sure if I can tell the difference whereas the first image has <code>LABEL=0</code> while the second image has <code>LABEL=1</code></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2116404%2Fd098a05e6c56272618c7f6c867be41ed%2FScreenshot%202020-09-19%20123222.png?generation=1600482785163541&amp;alt=media\" alt=\"\"></p>\n<p>Does this seem right?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1017814,
          "author_name": "Redwan Sony",
          "author_url": "",
          "post_date": "2020-09-19T09:27:57.243000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/aroraaman\" target=\"_blank\">@aroraaman</a> , Looking at the combined RGB channel view, I myself don't find any difference. Here there can be several possible steps in my opinion.  You can view three channels (RGB) separately and see if you can find any difference for label 1 and 0. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1017819,
          "author_name": "Redwan Sony",
          "author_url": "",
          "post_date": "2020-09-19T09:30:20.703000",
          "content": "<p><a href=\"https://www.kaggle.com/aroraaman\" target=\"_blank\">@aroraaman</a>, you can see the corresponding original DICOM and try to find out any difference. Even if you don't find anything there, you can check neighboring DICOM slices to these slices in question to check whether they have PE. … <br>\ngood luck..  </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1018112,
          "author_name": "quadcore/Richard Epstein",
          "author_url": "",
          "post_date": "2020-09-19T13:03:22.157000",
          "content": "<p>Can you provide the filename of these images?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1047029,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-10-12T07:27:08.613000",
          "content": "<p><a href=\"https://www.kaggle.com/aroraaman\" target=\"_blank\">@aroraaman</a> i think you should try to visualize the image with no PE  using only PE channel.. i feel hte marked white contrasts across the encircle are clots,if a channel is not having PE then it wont show up in the  stacked image also</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1015179,
      "author_name": "Ronaldo S.A. Batista",
      "author_url": "",
      "post_date": "2020-09-18T02:54:08.213000",
      "content": "<p>I didn't know this PE window. Where did you get these values? Thanks</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1017837,
          "author_name": "Redwan Sony",
          "author_url": "",
          "post_date": "2020-09-19T09:40:47.083000",
          "content": "<p>Hi, <a href=\"https://www.kaggle.com/ronaldokun\" target=\"_blank\">@ronaldokun</a> , Different window parameters are available in different public documents. You can easily find them in the internet. Here is one of the links.. <br>\n<a href=\"https://radiopaedia.org/articles/windowing-ct\" target=\"_blank\">Windowing (CT)</a></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1019916,
          "author_name": "Ronaldo S.A. Batista",
          "author_url": "",
          "post_date": "2020-09-20T18:43:50.947000",
          "content": "<p>Thank you!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1012768,
      "author_name": "DatNT",
      "author_url": "",
      "post_date": "2020-09-16T09:24:04.033000",
      "content": "<p><a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> can you provide the full source code version to produce these images? thanks in advance.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1013430,
          "author_name": "Salaryman",
          "author_url": "",
          "post_date": "2020-09-16T17:30:50.387000",
          "content": "<p>same question here! Thanks so much!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1013486,
          "author_name": "Ian Pan",
          "author_url": "",
          "post_date": "2020-09-16T18:03:15.437000",
          "content": "<p>Let me know if you are able to access this attachment.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1013494,
          "author_name": "Abdur Rehman",
          "author_url": "",
          "post_date": "2020-09-16T18:07:07.757000",
          "content": "<p>Yeah, I am able to access it. Thanks.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1014286,
          "author_name": "DatNT",
          "author_url": "",
          "post_date": "2020-09-17T10:16:40.897000",
          "content": "<p><a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> thanks for the script but I think you should make use of parallelism to speed up the process :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1014678,
          "author_name": "SumanSudhir",
          "author_url": "",
          "post_date": "2020-09-17T16:33:14.180000",
          "content": "<p>Yaa, I took me 10 hrs to create 512x512x3 image with parallelism using 8 cores</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1015161,
          "author_name": "Ian Pan",
          "author_url": "",
          "post_date": "2020-09-18T02:25:27.117000",
          "content": "<p>I don't have much experience with that so feel free to implement it and share it with us. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1015390,
          "author_name": "Salaryman",
          "author_url": "",
          "post_date": "2020-09-18T06:45:46.553000",
          "content": "<p>Thanks so much!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1017811,
          "author_name": "SumanSudhir",
          "author_url": "",
          "post_date": "2020-09-19T09:25:11.080000",
          "content": "<p><a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> I have added parallelism to your data generation code. It is showing 1.5 hrs to generate data for me using 4 cores. Please have a look. <a href=\"https://gist.github.com/SumanSudhir/7f3b72510eb2b8cff080a225334f5dc4\" target=\"_blank\">https://gist.github.com/SumanSudhir/7f3b72510eb2b8cff080a225334f5dc4</a></p>",
          "votes": 16,
          "replies": []
        },
        {
          "id": 1017834,
          "author_name": "Redwan Sony",
          "author_url": "",
          "post_date": "2020-09-19T09:38:30.090000",
          "content": "<p>Thanks a lot, <a href=\"https://www.kaggle.com/sudhiriitb\" target=\"_blank\">@sudhiriitb</a> for your nice tweak.<br>\nI am stealing some of your code for my personal project if  that's ok. 😃😃</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1017842,
          "author_name": "SumanSudhir",
          "author_url": "",
          "post_date": "2020-09-19T09:42:48.690000",
          "content": "<p>of course, you can, but it not mine its <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> code😹</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1017897,
          "author_name": "Redwan Sony",
          "author_url": "",
          "post_date": "2020-09-19T10:22:58.687000",
          "content": "<p>🤓🤓🤓 xOOx</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1018363,
          "author_name": "Ian Pan",
          "author_url": "",
          "post_date": "2020-09-19T16:28:47.813000",
          "content": "<p>Excellent. Thank you!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1047020,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-10-12T07:22:21.027000",
          "content": "<p><a href=\"https://www.kaggle.com/sudhiriitb\" target=\"_blank\">@sudhiriitb</a>  are these dstacked channel images helping get score ?</p>\n<p><a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a>    Is medistenial same as Subdural ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1050831,
          "author_name": "Alexander Soare",
          "author_url": "",
          "post_date": "2020-10-15T19:06:59.643000",
          "content": "<p><a href=\"https://www.kaggle.com/sudhiriitb\" target=\"_blank\">@sudhiriitb</a> thanks for the parallel code! I'm having one problem with it though. For some reason it's not releasing memory and I run out at about 39 iterations. Any experience with such a bug?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1050871,
          "author_name": "SumanSudhir",
          "author_url": "",
          "post_date": "2020-10-15T20:06:14.100000",
          "content": "<p><a href=\"https://www.kaggle.com/alexandersoare\" target=\"_blank\">@alexandersoare</a> are you running it as a python script or running in a notebook?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1011979,
      "author_name": "Abdur Rehman",
      "author_url": "",
      "post_date": "2020-09-15T20:00:37.910000",
      "content": "<blockquote>\n  <p>I have provided them in 3-channel RGB format. Each channel is a different window.</p>\n</blockquote>\n<p>I am asking a silly question because of my little knowledge. Why you are using the three mentioned windows (Lung, PE and MEDIASTINAL ) as we have other windows like chest, abdomen and spine? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1012007,
          "author_name": "Ian Pan",
          "author_url": "",
          "post_date": "2020-09-15T20:45:28.243000",
          "content": "<p>These are the windows I believe to be most useful based on domain knowledge.</p>",
          "votes": 7,
          "replies": []
        }
      ]
    },
    {
      "id": 1011809,
      "author_name": "Bo Peng",
      "author_url": "",
      "post_date": "2020-09-15T17:32:25.437000",
      "content": "<p>You guys think Colab Pro with TPU using Google Bucket for dataset would work for this? Is Pytorch models better than tensorflow? Cant use pytorch on TPU right? Thanks,</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1011815,
          "author_name": "xhlulu",
          "author_url": "",
          "post_date": "2020-09-15T17:36:39.120000",
          "content": "<p>You can use pytorch-xla: <a href=\"https://github.com/pytorch/xla\" target=\"_blank\">https://github.com/pytorch/xla</a></p>\n<p>there's many great notebooks that train transformer models on TPUs</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1011818,
          "author_name": "Bo Peng",
          "author_url": "",
          "post_date": "2020-09-15T17:39:25.203000",
          "content": "<p>Thanks for the share Xing! So that is mentioning you can use pytorch with TPUs correct? Thanks,</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1011826,
          "author_name": "xhlulu",
          "author_url": "",
          "post_date": "2020-09-15T17:42:49.663000",
          "content": "<p>Yep! It should work; a good starting point is the jigsaw multilingual tpu competition with a lot of pytorch tpu kernels</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1011829,
          "author_name": "Bo Peng",
          "author_url": "",
          "post_date": "2020-09-15T17:45:31.867000",
          "content": "<p>Nice, thanks Xing.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1012346,
      "author_name": "Paul Chen",
      "author_url": "",
      "post_date": "2020-09-16T03:25:36.007000",
      "content": "<p>Lan, Many Thanks! </p>\n<p>I'm wondering if LUNG window level should be <strong>600</strong> not <strong>-600</strong> based on the description \"Our CT techniques are shown in the ,Table. Images are displayed with three different gray scales for interpretation of lung window (window width/level [HU] = 1500/600), mediastinal window (400/40), and pulmonary embolism–specific (700/100) settings.\" from the RSNA <a href=\"https://pubs.rsna.org/doi/full/10.1148/rg.245045008\" target=\"_blank\">paper </a></p>",
      "votes": 2,
      "replies": [
        {
          "id": 1013489,
          "author_name": "Ian Pan",
          "author_url": "",
          "post_date": "2020-09-16T18:05:02.420000",
          "content": "<p>I used the windows from that paper, 600 is a typo. It should be -600, as that is about the HU of lung parenchyma. </p>",
          "votes": 7,
          "replies": []
        }
      ]
    },
    {
      "id": 1053520,
      "author_name": "Shubham",
      "author_url": "",
      "post_date": "2020-10-19T04:52:56.290000",
      "content": "<p><a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> Thank You for sharing this excellent usefull information.<br>\nHow to run this file convert_to_jpeg_for_kaggle.py on Kaggle notebook?<br>\nBecasue I got error during saving the images.<br>\nAny idea please help me.</p>",
      "votes": -1,
      "replies": []
    },
    {
      "id": 1059395,
      "author_name": "Gopi Durgaprasad",
      "author_url": "",
      "post_date": "2020-10-25T03:29:30.667000",
      "content": "<p>hai all… anyone cross 0.20 with this dataset</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1059588,
          "author_name": "Vee",
          "author_url": "",
          "post_date": "2020-10-25T08:53:18.477000",
          "content": "<p>I think some would have.<br>\nMy current score 0.203 is with this dataset and could cross 0.2 if I have the time and energy over next 2 days.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1046958,
      "author_name": "Nitin Datta",
      "author_url": "",
      "post_date": "2020-10-12T06:18:19.903000",
      "content": "<p>If anyone is facing issues like me in the extraction of images on Windows please change this function and it will work properly. </p>\n<pre><code>def edit_filenames(files):\n    dicoms = ['{:04d}_{}'.format(ind,f.split(\"\\\\\")[-1].replace(\"dcm\",\"jpg\")) for ind,f in enumerate(files)]\n    series = ['/'.join(f.split('/')[-3:-1]) for f in files]\n    return [osp.join(s,d) for s,d in zip(series, dicoms)]\n</code></pre>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1035975,
      "author_name": "dgandiaga",
      "author_url": "",
      "post_date": "2020-10-03T09:54:19.793000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> <br>\nThank you very much for your post. I was trying to use this method, but when I execute the code I just get a lot of 'list index out of range' through the console and it finishes in a second. I'm running it in a kaggle kernel, I only changed the route to the csv and the train files, the rest is exactly as you posted. Any idea? I'm quite new on this I'm not sure if I'm missing something important.</p>\n<p>Regards</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1036442,
          "author_name": "Ronaldo S.A. Batista",
          "author_url": "",
          "post_date": "2020-10-03T18:42:17.417000",
          "content": "<p>You should post the snippet of code which returns an error.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1039293,
          "author_name": "dgandiaga",
          "author_url": "",
          "post_date": "2020-10-06T13:26:47.020000",
          "content": "<p>You're right, sorry. Anyways, the error was due to trying to write in a route Kaggle environment doesn't allow. Noob mistake, I guess 😅</p>\n<p>Thanks for your answer!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1029806,
      "author_name": "cxt",
      "author_url": "",
      "post_date": "2020-09-28T06:36:08.430000",
      "content": "<p>Thank you!<br>\nI just wanted to ask can someone reduce the datasets,and then I saw your notebook</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1024703,
      "author_name": "Kerem Turgutlu",
      "author_url": "",
      "post_date": "2020-09-24T04:38:21.343000",
      "content": "<p>Thanks for sharing I was actually checking this paper before seeing your post, same 3 windowing for diagnosis is used! <a href=\"https://pubs.rsna.org/doi/pdf/10.1148/rg.245045008\" target=\"_blank\">https://pubs.rsna.org/doi/pdf/10.1148/rg.245045008</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1017852,
      "author_name": "Redwan Sony",
      "author_url": "",
      "post_date": "2020-09-19T09:50:02.360000",
      "content": "<p>Thanks, <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> for sharing the idea of using three channels. Previously, I was just the whole lung segment with lung window (window width/level [HU] = 1500/600). This three windows in three channels literally packing ~3x information compared to regular single channel grayscales.. </p>\n<p>Great work!!!!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1015789,
      "author_name": "Aman Arora",
      "author_url": "",
      "post_date": "2020-09-18T12:25:21.573000",
      "content": "<p>Thank you for the effort - <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> :)</p>\n<p>Quick question please, what does the following line mean: </p>\n<blockquote>\n  <p>Assume all images are axial</p>\n</blockquote>\n<p>Also, I am new to medical imaging so spent the past two days reading about CT scans. Wish to corroborate my understanding from yourself if possible please. </p>\n<p>If I understand the CT scanner essentially takes multiple 2-D images which are indexed. The 2-D images themselves are referred as \"slices\" ? Therefore, if slice at t'th position is <code>S(t)</code> then <code>S(t-1)</code> and <code>S(t+1)</code> are slice before and after and we could leverage this information to create better models. </p>\n<p>Does that seem right? Thanks! :)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1015870,
          "author_name": "quadcore/Richard Epstein",
          "author_url": "",
          "post_date": "2020-09-18T13:50:19.157000",
          "content": "<p>Assume all images are axial - Typically a CT scan takes what we call \"axial\" images. Like slicing a loaf of bread. The CT scanners can also take the pixels from that scan and reconstruct them in Sagittal (like looking at someone from the side) or Coronal (like looking at someone from in front of them) planes. We don't have those for this competition, although you could built a 3-D model from the axial slices and create them yourself.</p>\n<p>Somethings things are easier to see on Sagittal or Coronal images, but typically the radiologist mainly looks at the Axial images.</p>\n<p>Your idea on positioning is correct. The DICOM metadata defines the orientation and relationship of each slice in a 3-dimensional space.</p>\n<p>Copying from my prior post:</p>\n<p>Patient position - FFP (Feet first prone), FFS (feet first supine), HFS (head first supine). Refers to the position of the patient in the machine. In the train dataset:</p>\n<p>FFP 242<br>\nFFS 1732839<br>\nHFS 57513</p>\n<p>In theory, this could affect the patient position. The FFP (prone) images suggest the patient is lying on their belly. Supposedly you would want to flip these images (but I haven't looked at the DICOM to see if this is true).</p>\n<p>InstanceNumber - a unique number for each image within a Series. In these series, usually starts at 1 and counts up in the physical order of the images. There are Studies that don't start at zero. I don't know if there are cases where numbers are skipped. Could count from top to bottom or bottom to top. </p>\n<p>Image Patient Position - three floating point numbers that define a relative location in the real world. Image Patient Position [2] is the \"Z-axis\" and can be used to put the images in order.</p>\n<p>Image Orientation - another parameter that defines the orientation of the patient/images in space. Might be relevant for building three dimensional models and defining top/bottom of patient. I haven't checked.</p>\n<p>You can use these parameters to reassemble the images in order and build a 3D model. Note that you are not guaranteed that the images give you a physically complete image. Nothing stops you from having gaps in the images. This would not be normal processing, but since we cannot see the real test data, we cannot check this.</p>\n<p>Also, slice thickness looks consistent within each scan, but varies between Studies.</p>\n<p>In theory, if you put them in order by InstanceNumber or Image Patient Position [2], the change in position would equal the Slice Thickness. But in practice, you could have gaps or even overlaps (5 mm slices taken every 2.5 mm). How far you go to cover the possibilities depends on your model needs. If you look at the data in the Pulmonary Fibrosis competition, the CT images have missing images, gaps in images, repeated imaging through the patient in one Study, etc.</p>\n<p>-Rich</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 1015942,
          "author_name": "Aman Arora",
          "author_url": "",
          "post_date": "2020-09-18T14:42:15.353000",
          "content": "<p>Thank you. So in theory, each Series could have different number of slices (SOPInstanceID) and we could then leverage 3-D CNN or sequential model to predict PE, right? :)</p>\n<p>Again, thanks a lot for being so active on the discussions and sharing your knowledge. :)</p>",
          "votes": 0,
          "replies": [
            {
              "id": 1015996,
              "author_name": "quadcore/Richard Epstein",
              "author_url": "",
              "post_date": "2020-09-18T15:34:26.977000",
              "content": "<p>Yes. Slice thickness typically range from 1.25 mm to 5 mm. The training set has number of images per study ranging from 63 to 1083. Presumably the 63 slices are 5 mm slices. The 1083 slices are presumably 1.25 mm slices.</p>\n<p>No guarantee that a series doesn't include two passes through the chest.</p>\n<p>Also, at least one study is actually mainly an abdomen and pelvis study. It captures the bottom of the chest, but most of the study is abdomen and pelvis.</p>\n<p>There could be a chest+abdomen+pelvis study also.</p>\n<p>Most studies appear to be a standard Chest study.</p>\n<p>There is a free to download software called K-Pacs which lets you look at the studies in DICOM format.</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 1012786,
      "author_name": "R. Joseph Manoj, PhD",
      "author_url": "",
      "post_date": "2020-09-16T09:35:06.843000",
      "content": "<p>Windowing - Exposed it. Great Sharing</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1011925,
      "author_name": "Abdur Rehman",
      "author_url": "",
      "post_date": "2020-09-15T19:00:03.143000",
      "content": "<p><a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> You did not provide jpeg for the test set. I think we should need test set in jpeg too to make the submission file easily.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1011945,
          "author_name": "xhlulu",
          "author_url": "",
          "post_date": "2020-09-15T19:19:44.390000",
          "content": "<p>Part of the test set is missing and will only be available when you submit your notebook.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1012006,
          "author_name": "Ian Pan",
          "author_url": "",
          "post_date": "2020-09-15T20:45:12.810000",
          "content": "<p>As <a href=\"https://www.kaggle.com/xhlulu\" target=\"_blank\">@xhlulu</a> mentioned, you will need to use the functions I provided above to load in the DICOMs and convert them appropriately in your inference kernel. The private test set is computed on test data that cannot be accessed unless you submit a kernel. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1017849,
          "author_name": "Redwan Sony",
          "author_url": "",
          "post_date": "2020-09-19T09:45:16.670000",
          "content": "<p>The competition says that it's an <strong>inference-only</strong> competition. That means training should be done separately and use that pre-trained model on the test set given. The test set available with the data is just for demonstration. Your submitted kernel will be re-run on the hidden-kernel to produce the final submission result. </p>",
          "votes": 1,
          "replies": [
            {
              "id": 1017970,
              "author_name": "quadcore/Richard Epstein",
              "author_url": "",
              "post_date": "2020-09-19T11:12:03.003000",
              "content": "<p>The public test set is about 40 percent of the real test set. So it is not just for demonstration. I note, in the Pulmonary Fibrosis competition,  the test set was just for demonstration. </p>\n<p>I am successfully doing inference on the public test data and then combining that submission file in a committed notebook with the rest of the private test data to make a submission. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 1018755,
              "author_name": "Tsai29",
              "author_url": "",
              "post_date": "2020-09-20T00:04:04.760000",
              "content": "<p><a href=\"https://www.kaggle.com/richardepstein\" target=\"_blank\">@richardepstein</a> May I ask how you solve the GCDM problem during inference? Since even we can generate jpegs/tfrecords of testset in other kernel(with internet), but there are only 40% of testset.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 1018761,
              "author_name": "quadcore/Richard Epstein",
              "author_url": "",
              "post_date": "2020-09-20T00:25:35.013000",
              "content": "<p>I haven't solved it. I use a try/except construct around my access to the pydicom pixel data. If I have an error, I use a blank image as a placeholder.</p>\n<p>I think maybe 5% or less of the images have this error, although I haven't run an actual test.</p>\n<p>Eventually, we'll need to solve this (or at least try). I suspect you could download the GDCM \"files\" into a dataset and import it during inference. I haven't done that yet. I cannot even get GDCM to install on my development machine (but it does work in kernels with Internet access.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 1018796,
              "author_name": "Tsai29",
              "author_url": "",
              "post_date": "2020-09-20T01:40:10.853000",
              "content": "<p>Thanks for the reply.<br>\nI think maybe kaggle should consider add GCDM in their default environment or something. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 1018857,
              "author_name": "quadcore/Richard Epstein",
              "author_url": "",
              "post_date": "2020-09-20T03:44:23.420000",
              "content": "<p>I think I solved this in this notebook:</p>\n<p><a href=\"https://www.kaggle.com/richardepstein/load-gdcm-in-notebook-without-internet\" target=\"_blank\">Load GDCM in Notebook without Internet</a></p>\n<p>It uses a Dataset that somebody else created that has the tar file for GDCM. </p>\n<p>It seems to work both interactively and in a Committed Notebook. I haven't yet tried it in an actual competition submission, but it runs without Internet access, so it should work.</p>\n<p>Note that once you have the \"GDCM needed\" error in a notebook, you must restart the Kernel. So it should be the first thing in your notebook.</p>\n<p>Also, you might need to wait to import Pydicom until after it is run, but I have not tested that.</p>\n<p>Note that you do not \"import GDCM\".</p>\n<p>Let me know if it works for you.</p>\n<p>-Rich</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 1018973,
              "author_name": "Tsai29",
              "author_url": "",
              "post_date": "2020-09-20T05:38:47.720000",
              "content": "<p><a href=\"https://www.kaggle.com/richardepstein\" target=\"_blank\">@richardepstein</a> It totally works! Thank you so much for the kind sharing!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 1010807,
      "author_name": "hirviö",
      "author_url": "",
      "post_date": "2020-09-15T04:23:02.027000",
      "content": "<p>Thanks for sharing the dataset and the insights! I seem to be unable to download it. Could anyone else download?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1010858,
          "author_name": "hirviö",
          "author_url": "",
          "post_date": "2020-09-15T05:26:28.130000",
          "content": "<p>nvm, got the download option now!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1010645,
      "author_name": "Tim Yee",
      "author_url": "",
      "post_date": "2020-09-15T00:39:29.243000",
      "content": "<p>Thanks for sharing preprocessed images and insights Ian! I would not have known about the PE and mediastinal windows. I just used the metadata windows which just so happens to be mediastinal window.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1031850,
      "author_name": "RaviJain",
      "author_url": "",
      "post_date": "2020-09-29T18:18:36.603000",
      "content": "<p>This is really helpful. Thanks.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1025736,
      "author_name": "Pinaki MIshra",
      "author_url": "",
      "post_date": "2020-09-24T18:22:39.407000",
      "content": "<p>Thanks for sharing , very useful</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1010629": "Update: source code to produce these images has been added as attachment\n\nhttps://www.kaggle.com/vaillant/rsna-str-pe-detection-jpeg-256\n\n# Introduction\nThis competition is challenging due to the huge dataset size. In addition, people may not be familiar with medical imaging and CT scans. \n\nTo lower the barrier to entry to this competition and to encourage more people to participate, I have converted all of the DICOM images into JPEGs (about 52GB), 256x256 pixels (standard is 512x512). \n\n# Windowing\nThe values in CT scans tend to range from -1000 to 3000; we are used to 8-bit images with pixel values ranging from 0 to 255. Radiologists use a technique called **windowing** to visualize CT scans. Different types of tissues are better evaluated using different windows. Windows are defined by 2 numbers: window **width** and window **level**. \n\nHere is the function I use to window an image:\n```\ndef window(img, WL=50, WW=350):\n    upper, lower = WL+WW//2, WL-WW//2\n    X = np.clip(img.copy(), lower, upper)\n    X = X - np.min(X)\n    X = X / np.max(X)\n    X = (X*255.0).astype('uint8')\n    return X\n```\nBasically, we compute the lower and upper limits using the window level (WL) and window width (WW). Upper limit is level plus half the width and lower limit is level minus half the width. \n\nThen, all values above the upper limit are set to the upper limit; vice versa for the lower limit. We can then normalize the images to [0, 1] and multiply by 255 to convert them into an 8-bit image. \n\nNote that these images are single channel. I have provided them in 3-channel RGB format. **Each channel is a different window.** \n\n- **RED** channel / **LUNG** window / level=-600, width=1500\n- **GREEN** channel / **PE** window / level=100, width=700\n- **BLUE** channel / **MEDIASTINAL** window / level=40, width=400\n\nPlease remember that `cv2.imread` by default loads images in **BGR** so if you are using that function make sure you're not confusing the red and blue channels. The **PE specific and mediastinal windows** will be most useful for evaluating whether there is blood clot. However, the lung window can show abnormalities in the lungs which may be suggestive of a clot, so I included it as well. Feel free to experiment with any single window or combination of windows.\n\n### Lung\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F281652%2Faa4bd969f80498908e8f40bade1e69b7%2Flung_window.jpg?generation=1600128037814115&alt=media)\n### PE Specific\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F281652%2F8e5d707a2daa4606b720de0c6030aba4%2Fpe_window.jpg?generation=1600128049440405&alt=media)\n### Mediastinal\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F281652%2F8950b63595f0b36b6d52a48907d32ffb%2Fmediastinum_window.jpg?generation=1600128059340756&alt=media)\n\nThis has the disadvantage that you cannot select other windows and you cannot design a learnable windowing function using this data. **However, based on my experience, I do not believe these are necessary to do well in this competition.**\n\n# 3D Reconstruction\nCT scans are 3D. Each image is 2D, but taking into context the surrounding image slices can improve performance (either using a fully 3D model, hybrid 3D-2D model, or sequence model). In order to facilitate reconstructing the 3D image, **each image in each series (folder) has a prefix with the slice number (increasing from bottom to top).** That way, you can easily organize the slices in 3D. \n\nThis is the function I used to create the 3D array:\n```\ndef load_dicom_array(f):\n    dicom_files = glob.glob(osp.join(f, '*.dcm'))\n    dicoms = [pydicom.dcmread(d) for d in dicom_files]\n    M = float(dicoms[0].RescaleSlope)\n    B = float(dicoms[0].RescaleIntercept)\n    # Assume all images are axial\n    z_pos = [float(d.ImagePositionPatient[-1]) for d in dicoms]\n    dicoms = np.asarray([d.pixel_array for d in dicoms])\n    dicoms = dicoms[np.argsort(z_pos)]\n    dicoms = dicoms * M\n    dicoms = dicoms + B\n    return dicoms, np.asarray(dicom_files)[np.argsort(z_pos)]\n```\nImportant things to note: you need to use `RescaleSlope` and `RescaleIntercept` to alter the values from `dicom.pixel_array` first. In this dataset, `RescaleSlope` is always 1 and `RescaleIntercept` is 0 or -1024 (more common). Then, you can use `ImagePositionPatient` to sort the slices. The last element `ImagePositionPatient[-1]` is used to sort in the z-axis. \n\nPlease let me know if you encounter any issues with this dataset. I will try my best to fix them. It took about 36 hours to generate this dataset. I will try to produce datasets of different sizes if people request it. ",
    "1011537": "Thanks @vaillant for sharing the dataset! Without this, I would've just ignored the competition because there's no way for me to load 1TB of image data into 13GB of RAM ;)\n\nOne small note: I think it's better if you change the license of your dataset to a Custom license (and refer to this competition) or a CC-BY-NC-SA 4.0 (if that's approved by the organizers). A CC0 (open domain) license might confuse people who are not familiar with Kaggle.",
    "1045494": "Thanks so much for sharing this @vaillant. Especially useful for those of us getting a late start to this competition.\n\nA couple of questions to you or anyone who sees this:\n\n1) What's with the `[5645:]`? I suppose you were just continuing where it left of at some point?\n2) Any idea what the compression level is on the jpegs? Is it lossless?",
    "1016480": "Hi @vaillant \n\nThought I would use your method to convert `.dcm` to `.jpg` files.\n\nBut, in the example below, I am not sure if I can tell the difference whereas the first image has `LABEL=0` while the second image has `LABEL=1`\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2116404%2Fd098a05e6c56272618c7f6c867be41ed%2FScreenshot%202020-09-19%20123222.png?generation=1600482785163541&alt=media)\n\nDoes this seem right?",
    "1015179": "I didn't know this PE window. Where did you get these values? Thanks",
    "1012768": "@vaillant can you provide the full source code version to produce these images? thanks in advance.",
    "1011979": ">  I have provided them in 3-channel RGB format. Each channel is a different window.\n\nI am asking a silly question because of my little knowledge. Why you are using the three mentioned windows (Lung, PE and MEDIASTINAL ) as we have other windows like chest, abdomen and spine? ",
    "1011809": "You guys think Colab Pro with TPU using Google Bucket for dataset would work for this? Is Pytorch models better than tensorflow? Cant use pytorch on TPU right? Thanks,",
    "1012346": "Lan, Many Thanks! \n\nI'm wondering if LUNG window level should be **600** not **-600** based on the description \"Our CT techniques are shown in the ,Table. Images are displayed with three different gray scales for interpretation of lung window (window width/level [HU] = 1500/600), mediastinal window (400/40), and pulmonary embolism–specific (700/100) settings.\" from the RSNA [paper ](https://pubs.rsna.org/doi/full/10.1148/rg.245045008)",
    "1053520": "@vaillant Thank You for sharing this excellent usefull information.\nHow to run this file convert_to_jpeg_for_kaggle.py on Kaggle notebook?\nBecasue I got error during saving the images.\nAny idea please help me.",
    "1059395": "hai all... anyone cross 0.20 with this dataset",
    "1046958": "If anyone is facing issues like me in the extraction of images on Windows please change this function and it will work properly. \n\n```\ndef edit_filenames(files):\n    dicoms = ['{:04d}_{}'.format(ind,f.split(\"\\\\\")[-1].replace(\"dcm\",\"jpg\")) for ind,f in enumerate(files)]\n    series = ['/'.join(f.split('/')[-3:-1]) for f in files]\n    return [osp.join(s,d) for s,d in zip(series, dicoms)]\n```",
    "1035975": "Hi @vaillant \nThank you very much for your post. I was trying to use this method, but when I execute the code I just get a lot of 'list index out of range' through the console and it finishes in a second. I'm running it in a kaggle kernel, I only changed the route to the csv and the train files, the rest is exactly as you posted. Any idea? I'm quite new on this I'm not sure if I'm missing something important.\n\nRegards",
    "1029806": "Thank you!\nI just wanted to ask can someone reduce the datasets,and then I saw your notebook",
    "1024703": "Thanks for sharing I was actually checking this paper before seeing your post, same 3 windowing for diagnosis is used! https://pubs.rsna.org/doi/pdf/10.1148/rg.245045008",
    "1017852": "Thanks, @vaillant for sharing the idea of using three channels. Previously, I was just the whole lung segment with lung window (window width/level [HU] = 1500/600). This three windows in three channels literally packing ~3x information compared to regular single channel grayscales.. \n\nGreat work!!!!",
    "1015789": "Thank you for the effort - @vaillant :)\n\nQuick question please, what does the following line mean: \n> Assume all images are axial\n\nAlso, I am new to medical imaging so spent the past two days reading about CT scans. Wish to corroborate my understanding from yourself if possible please. \n\nIf I understand the CT scanner essentially takes multiple 2-D images which are indexed. The 2-D images themselves are referred as \"slices\" ? Therefore, if slice at t'th position is `S(t)` then `S(t-1)` and `S(t+1)` are slice before and after and we could leverage this information to create better models. \n\nDoes that seem right? Thanks! :)",
    "1012786": "Windowing - Exposed it. Great Sharing",
    "1011925": "@vaillant You did not provide jpeg for the test set. I think we should need test set in jpeg too to make the submission file easily.",
    "1010807": "Thanks for sharing the dataset and the insights! I seem to be unable to download it. Could anyone else download?",
    "1010645": "Thanks for sharing preprocessed images and insights Ian! I would not have known about the PE and mediastinal windows. I just used the metadata windows which just so happens to be mediastinal window.",
    "1031850": "This is really helpful. Thanks.",
    "1025736": "Thanks for sharing , very useful"
  }
}