{
  "id": 110728,
  "title": "Why you need use subdural window settings",
  "url": "/competitions/rsna-intracranial-hemorrhage-detection/discussion/110728",
  "author_name": "David Tang",
  "post_date": "2019-09-30T17:11:38.724000",
  "votes": 53,
  "comment_count": 39,
  "views": 0,
  "content": "<p>Dear Kagglers,</p>\n\n<p>I hope to raise awareness about this setting. Currently I notice that most kernels are using the brain or parenchymal windows to view the hemorrhages.</p>\n\n<p>Here is a quick image from radiopedia.org, which I have annotated with arrows.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3454874%2F340f0cfb56257ff428b9e70193017731%2Fsubdural_window.png?generation=1569863200163417&amp;alt=media\" alt=\"\">\nOn the left, it is easy to miss subdural hemorrhage in the brain windows or default windows. </p>\n\n<p>I have discussed more in my <a href=\"https://www.kaggle.com/dcstang/see-like-a-radiologist-with-systematic-windowing\">kernel \"See Like A Radiologist, With Systematic Windowing\"</a> on how and why subdural windows are part of a standard workflow for a radiologist. </p>\n\n<p>This is a way to give back to the Kaggle community, from whom I have learnt so much of Python, coding, data science and machine learning. I hope this nugget of information passed down from my medical school professors will help you in devising a better model.</p>",
  "messages": [
    {
      "id": 637116,
      "postDate": "2019-09-30T17:11:38.723Z",
      "content": "<p>Dear Kagglers,</p>\n\n<p>I hope to raise awareness about this setting. Currently I notice that most kernels are using the brain or parenchymal windows to view the hemorrhages.</p>\n\n<p>Here is a quick image from radiopedia.org, which I have annotated with arrows.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3454874%2F340f0cfb56257ff428b9e70193017731%2Fsubdural_window.png?generation=1569863200163417&amp;alt=media\" alt=\"\">\nOn the left, it is easy to miss subdural hemorrhage in the brain windows or default windows. </p>\n\n<p>I have discussed more in my <a href=\"https://www.kaggle.com/dcstang/see-like-a-radiologist-with-systematic-windowing\">kernel \"See Like A Radiologist, With Systematic Windowing\"</a> on how and why subdural windows are part of a standard workflow for a radiologist. </p>\n\n<p>This is a way to give back to the Kaggle community, from whom I have learnt so much of Python, coding, data science and machine learning. I hope this nugget of information passed down from my medical school professors will help you in devising a better model.</p>",
      "rawMarkdown": "Dear Kagglers,\n\nI hope to raise awareness about this setting. Currently I notice that most kernels are using the brain or parenchymal windows to view the hemorrhages.\n\nHere is a quick image from radiopedia.org, which I have annotated with arrows.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3454874%2F340f0cfb56257ff428b9e70193017731%2Fsubdural_window.png?generation=1569863200163417&amp;alt=media)\nOn the left, it is easy to miss subdural hemorrhage in the brain windows or default windows. \n\nI have discussed more in my [kernel \"See Like A Radiologist, With Systematic Windowing\"](https://www.kaggle.com/dcstang/see-like-a-radiologist-with-systematic-windowing) on how and why subdural windows are part of a standard workflow for a radiologist. \n\nThis is a way to give back to the Kaggle community, from whom I have learnt so much of Python, coding, data science and machine learning. I hope this nugget of information passed down from my medical school professors will help you in devising a better model.",
      "votes": 53
    },
    {
      "id": 638164,
      "postDate": "2019-10-01T15:38:14.627Z",
      "content": "<p>Why would you need to use this window? Isn't windowing a solution for our \"poor\" human eye-sight? I'm feeding my network the full 16-bit data to see if it will solve this problem itself. If not I'll add a \"blood map\" (HU 50-70) in a second channel.</p>",
      "rawMarkdown": "Why would you need to use this window? Isn't windowing a solution for our \"poor\" human eye-sight? I'm feeding my network the full 16-bit data to see if it will solve this problem itself. If not I'll add a \"blood map\" (HU 50-70) in a second channel.",
      "votes": 3,
      "replies": [
        {
          "id": 638317,
          "postDate": "2019-10-01T18:48:26.367Z",
          "content": "<p>Hi <a href=\"/reflexion\">@reflexion</a>, thanks for the question. I agree about your statement on human eyes.\nBasically this is to raise awareness about how one should pre-process the data from dicom images. I think your implementation should not run into issues. For those who did save the dicom into .png format some of this info may be lost.</p>\n\n<p><a href=\"https://rd.springer.com/content/pdf/10.1007%2Fs00330-019-06163-2.pdf\">https://rd.springer.com/content/pdf/10.1007%2Fs00330-019-06163-2.pdf</a>\nI like to bring your attention to this paper on hemorrhage detection. Interestingly the authors chose to use 3 window settings for all the images, then scale them before feeding into the neural network. Do you think because of the scaling, there will be significant gains in the training speed / accuracy of the network?</p>",
          "rawMarkdown": "Hi @reflexion, thanks for the question. I agree about your statement on human eyes.\nBasically this is to raise awareness about how one should pre-process the data from dicom images. I think your implementation should not run into issues. For those who did save the dicom into .png format some of this info may be lost.\n\nhttps://rd.springer.com/content/pdf/10.1007%2Fs00330-019-06163-2.pdf\nI like to bring your attention to this paper on hemorrhage detection. Interestingly the authors chose to use 3 window settings for all the images, then scale them before feeding into the neural network. Do you think because of the scaling, there will be significant gains in the training speed / accuracy of the network?",
          "votes": 1
        },
        {
          "id": 638371,
          "postDate": "2019-10-01T20:33:01.423Z",
          "content": "<p><a href=\"/reflexion\">@reflexion</a> <a href=\"/dcstang\">@dcstang</a> This is how I see it: The images were labelled by <strong><em>our \"poor\" human eye-sight</em></strong>, and so will the test images. If we give the network these \"extra dimensions\" it will just be noise (no label will correlate with that). Tell me if I'm wrong, but I can't see how the neural net can make sense of it.</p>",
          "rawMarkdown": "@reflexion @dcstang This is how I see it: The images were labelled by ***our \"poor\" human eye-sight***, and so will the test images. If we give the network these \"extra dimensions\" it will just be noise (no label will correlate with that). Tell me if I'm wrong, but I can't see how the neural net can make sense of it.",
          "votes": 2
        },
        {
          "id": 638420,
          "postDate": "2019-10-01T22:04:38.613Z",
          "content": "<p><code>it will just be noise (no label will correlate with that).</code> \nI see your argument... but as far I can tell noise can provide good regularization.. also we dont know what kind of information NN will find useful... the best way to test will be to run small experiment and see if this improves the result. </p>",
          "rawMarkdown": "`it will just be noise (no label will correlate with that).` \nI see your argument... but as far I can tell noise can provide good regularization.. also we dont know what kind of information NN will find useful... the best way to test will be to run small experiment and see if this improves the result. ",
          "votes": 1
        },
        {
          "id": 638614,
          "postDate": "2019-10-02T06:40:06.430Z",
          "content": "<p>I've tested both now, and it doesn't seem to converge faster. I don't know if you can focus a network on certain value-ranges this way.</p>",
          "rawMarkdown": "I've tested both now, and it doesn't seem to converge faster. I don't know if you can focus a network on certain value-ranges this way.",
          "votes": 1
        },
        {
          "id": 638617,
          "postDate": "2019-10-02T06:43:23.593Z",
          "content": "<p>The network should increase weights on relevant datapoints and ignore noise. But things like skull fractures (bone window) are not noise, they correlate with epidural bleeds, and we use them to increase likelihood of diagnosis (I'm a resident radiology).</p>",
          "rawMarkdown": "The network should increase weights on relevant datapoints and ignore noise. But things like skull fractures (bone window) are not noise, they correlate with epidural bleeds, and we use them to increase likelihood of diagnosis (I'm a resident radiology).",
          "votes": 1
        },
        {
          "id": 638651,
          "postDate": "2019-10-02T08:05:16.010Z",
          "content": "<p><a href=\"/drhabib\">@drhabib</a> I agree that one should always run small experiments either way :-) I usually find my intuition wrong (way too often!). I didn't consider the regularization effect. Btw, have you experimented with the full pixel range? ;-) <a href=\"/reflexion\">@reflexion</a> I understand, and I won't argue with that :-), which one works better?</p>\n\n<p>First of all, I think this is an interesting topic! So let me clarify my point (not sure if this is necessary though):  if let's say a person labels pictures of small birds on 128x128 images; And say, the original image sizes are 1024x1024 (so it has been resized to 128x128 prior to labeling). Using the original 1024x1024 information for the neural net wouldn't be necessary: there might be birds appearing in those pictures, but the <strong><em>labeller</em></strong> wouldn't label it, and (s)he won't label those images for the test set either (because they have also been resized). So our network might fire a neuron or two seeing that spot in the sky with 1024x1024, but it won't be useful (instead creating noise)... because the label is 0.</p>",
          "rawMarkdown": "@drhabib I agree that one should always run small experiments either way :-) I usually find my intuition wrong (way too often!). I didn't consider the regularization effect. Btw, have you experimented with the full pixel range? ;-) @reflexion I understand, and I won't argue with that :-), which one works better?\n\nFirst of all, I think this is an interesting topic! So let me clarify my point (not sure if this is necessary though):  if let's say a person labels pictures of small birds on 128x128 images; And say, the original image sizes are 1024x1024 (so it has been resized to 128x128 prior to labeling). Using the original 1024x1024 information for the neural net wouldn't be necessary: there might be birds appearing in those pictures, but the ***labeller*** wouldn't label it, and (s)he won't label those images for the test set either (because they have also been resized). So our network might fire a neuron or two seeing that spot in the sky with 1024x1024, but it won't be useful (instead creating noise)... because the label is 0.",
          "votes": 2
        },
        {
          "id": 638679,
          "postDate": "2019-10-02T09:12:57.173Z",
          "content": "<p>That's true, but radiologists never just look in one window setting, so the information might contain more than noise and also have been used by the human readers.</p>",
          "rawMarkdown": "That's true, but radiologists never just look in one window setting, so the information might contain more than noise and also have been used by the human readers."
        },
        {
          "id": 638767,
          "postDate": "2019-10-02T11:56:22.547Z",
          "content": "<p>Yes that is important. Do you think they considered/based their labeling on multiple windows although only one is reported in the dicom file? </p>",
          "rawMarkdown": "Yes that is important. Do you think they considered/based their labeling on multiple windows although only one is reported in the dicom file? "
        },
        {
          "id": 638777,
          "postDate": "2019-10-02T12:10:33.567Z",
          "content": "<p>very interesting thread, subscribing :) </p>",
          "rawMarkdown": "very interesting thread, subscribing :) ",
          "votes": 1
        },
        {
          "id": 638783,
          "postDate": "2019-10-02T12:17:55.067Z",
          "content": "<blockquote>\n  <p><strong>akensert wrote:</strong></p>\n  \n  <p>Do you think they considered/based their labeling on multiple windows although only one is reported in the dicom file?</p>\n</blockquote>\n\n<p>As a neuroradiology fellow who has done these types of annotations for an ML project previously, I believe it's unlikely that any labels were generated without viewing the images in multiple windows.</p>\n\n<p>But...as has been discussed by others, I think it remains to be shown which if any window settings actually increase signal-to-noise for CNNs in any meaningful way.</p>",
          "rawMarkdown": "&gt; **akensert wrote:**\n&gt; \n&gt; Do you think they considered/based their labeling on multiple windows although only one is reported in the dicom file?\n\nAs a neuroradiology fellow who has done these types of annotations for an ML project previously, I believe it's unlikely that any labels were generated without viewing the images in multiple windows.\n\nBut...as has been discussed by others, I think it remains to be shown which if any window settings actually increase signal-to-noise for CNNs in any meaningful way.",
          "votes": 3
        },
        {
          "id": 638883,
          "postDate": "2019-10-02T14:23:21.230Z",
          "content": "<p><a href=\"/wfwiggins203\">@wfwiggins203</a> Great, thank you for the information! </p>\n\n<p>May I ask if you know why these specific window parameters of the dicom files are reported?</p>",
          "rawMarkdown": "@wfwiggins203 Great, thank you for the information! \n\nMay I ask if you know why these specific window parameters of the dicom files are reported?"
        },
        {
          "id": 638894,
          "postDate": "2019-10-02T14:28:45.133Z",
          "content": "<p>Not him but I can respond =)</p>\n\n<p><code>ranges to normalize images: -50–150, 100–300 and 250–450</code> </p>",
          "rawMarkdown": "Not him but I can respond =)\n\n`ranges to normalize images: -50–150, 100–300 and 250–450` "
        },
        {
          "id": 639095,
          "postDate": "2019-10-02T18:58:32.323Z",
          "content": "<p><a href=\"/drhabib\">@drhabib</a> These are the ranges I just started using (stacked it into 3-channeled image). From the paper above right? </p>",
          "rawMarkdown": "@drhabib These are the ranges I just started using (stacked it into 3-channeled image). From the paper above right? "
        },
        {
          "id": 639190,
          "postDate": "2019-10-02T21:35:26.037Z",
          "content": "<p>Do you guys think the brain, subdural, and bone windows are enough to predict hemorrhages or are the soft tissue and grey-white differentiation windows also needed?</p>\n\n<p>Thanks.</p>",
          "rawMarkdown": "Do you guys think the brain, subdural, and bone windows are enough to predict hemorrhages or are the soft tissue and grey-white differentiation windows also needed?\n\nThanks."
        },
        {
          "id": 639751,
          "postDate": "2019-10-03T14:05:48.297Z",
          "content": "<p>Thanks for all the valuable comments. I appreciate Kevin trying out the comparison between rescaling and none. </p>\n\n<p><a href=\"/bopengiowa\">@bopengiowa</a>, as Walter said a typical radiologist would have gone through all windows as part of a systematic workflow process. This is because having soft tissue changes (ie. a swelling on your head) also increases the probability of having an intracranial bleed. With complex models, it may be possible that a neural network will have specific neurons that also take this into account.</p>\n\n<p>In regards to grey-white differentiation, personally I think it is better suited to other applications. Age prediction, tumor detection, stroke, for instance are things that might be easier to pick up with this window. This is just my intuition, and I will be happily surprised if a neural network proves me wrong. </p>",
          "rawMarkdown": "Thanks for all the valuable comments. I appreciate Kevin trying out the comparison between rescaling and none. \n\n@bopengiowa, as Walter said a typical radiologist would have gone through all windows as part of a systematic workflow process. This is because having soft tissue changes (ie. a swelling on your head) also increases the probability of having an intracranial bleed. With complex models, it may be possible that a neural network will have specific neurons that also take this into account.\n\nIn regards to grey-white differentiation, personally I think it is better suited to other applications. Age prediction, tumor detection, stroke, for instance are things that might be easier to pick up with this window. This is just my intuition, and I will be happily surprised if a neural network proves me wrong. ",
          "votes": 2
        },
        {
          "id": 641077,
          "postDate": "2019-10-04T12:28:28.273Z",
          "content": "<blockquote>\n  <p>Do you think because of the scaling, there will be significant gains in the training speed / accuracy of the network?</p>\n</blockquote>\n\n<p>Linear rescaling has an interesting relationship with batch norm, and perhaps the rescaling has an effect on training (even for large networks) because it is done prior to batch normalization. It has the effect of centering the window at the mean and scaling the stddev according to the slope.</p>\n\n<p>Note that a network could, in principle, learn something like the rescaling function. That is, if it is actually a shortcut to a solution. The function in question is really simple (f(x)=ax+b) so this shouldn't matter if the network has a sufficient number of parameters. You could, in principle, get a smaller network to train faster and perform more accurately by including additional prior information about the input data if (as mentioned) the prior already has predictive capacity. I think this is kind of like training to find a transformation g(x) given f(x), x, a, and b in f(x) = g(ax+b) or f(x) = ax+b+g(x), which is basically an error term for a linear model. A robust network should be able to learn the linear portion of the equation (ax+b) given only x and f(x), but it can use the included ax+b as a shortcut. Doing so may also preclude the network from approximating a more complex function that, overall, will perform better than a linear model. IIRC this all matters less with network size.</p>\n\n<p>Are window width and center set by the radiologist or by the scanner? </p>",
          "rawMarkdown": "&gt; Do you think because of the scaling, there will be significant gains in the training speed / accuracy of the network?\n\nLinear rescaling has an interesting relationship with batch norm, and perhaps the rescaling has an effect on training (even for large networks) because it is done prior to batch normalization. It has the effect of centering the window at the mean and scaling the stddev according to the slope.\n\nNote that a network could, in principle, learn something like the rescaling function. That is, if it is actually a shortcut to a solution. The function in question is really simple (f(x)=ax+b) so this shouldn't matter if the network has a sufficient number of parameters. You could, in principle, get a smaller network to train faster and perform more accurately by including additional prior information about the input data if (as mentioned) the prior already has predictive capacity. I think this is kind of like training to find a transformation g(x) given f(x), x, a, and b in f(x) = g(ax+b) or f(x) = ax+b+g(x), which is basically an error term for a linear model. A robust network should be able to learn the linear portion of the equation (ax+b) given only x and f(x), but it can use the included ax+b as a shortcut. Doing so may also preclude the network from approximating a more complex function that, overall, will perform better than a linear model. IIRC this all matters less with network size.\n\nAre window width and center set by the radiologist or by the scanner? ",
          "votes": 1
        },
        {
          "id": 641390,
          "postDate": "2019-10-04T16:39:28.257Z",
          "content": "<p>Hi <a href=\"/thavlik\">@thavlik</a>, the windows are more fluid in practice, and the radiologist can \"scroll through\" a range of window values on their DICOM software. <a href=\"http://dicomviewer.booogle.net/\">http://dicomviewer.booogle.net/</a> is a sample implementation, where you can drag left and right to adjust these values. Though I suspect for this competition, the DICOM metadata has been washed clean to anonymize it and standardize most info.</p>\n\n<p>Thank you for the insight about how network size impacts this.</p>",
          "rawMarkdown": "Hi @thavlik, the windows are more fluid in practice, and the radiologist can \"scroll through\" a range of window values on their DICOM software. http://dicomviewer.booogle.net/ is a sample implementation, where you can drag left and right to adjust these values. Though I suspect for this competition, the DICOM metadata has been washed clean to anonymize it and standardize most info.\n\nThank you for the insight about how network size impacts this.",
          "votes": 2
        },
        {
          "id": 642755,
          "postDate": "2019-10-06T15:32:17.070Z",
          "content": "<p>Thanks <a href=\"/dcstang\">@dcstang</a> for all the medical info here, really appreciate it, suuper helpful!</p>",
          "rawMarkdown": "Thanks @dcstang for all the medical info here, really appreciate it, suuper helpful!",
          "votes": 1
        },
        {
          "id": 647078,
          "postDate": "2019-10-12T03:22:29.790Z",
          "content": "<p><a href=\"/drhabib\">@drhabib</a> <a href=\"/akensert\">@akensert</a> I think your argument is charming! I wonder if the images are labeled with default window only? If so, I agree with akensert since our model should see the same thing as the annotator. But are we sure of that?</p>",
          "rawMarkdown": "@drhabib @akensert I think your argument is charming! I wonder if the images are labeled with default window only? If so, I agree with akensert since our model should see the same thing as the annotator. But are we sure of that?",
          "votes": 2
        },
        {
          "id": 647246,
          "postDate": "2019-10-12T09:11:37.470Z",
          "content": "<p><a href=\"/roguekk007\">@roguekk007</a> From skimming around a bit, I think the annotators base their classification on multiple different windows (so additional windows to the default). These questions are for real annotators to answer, but I can imagine that they even have a software to check through different windows in a 'continuous way'. I still believe HUs below, let's say -100, is irrelevant. </p>",
          "rawMarkdown": "@roguekk007 From skimming around a bit, I think the annotators base their classification on multiple different windows (so additional windows to the default). These questions are for real annotators to answer, but I can imagine that they even have a software to check through different windows in a 'continuous way'. I still believe HUs below, let's say -100, is irrelevant. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 637933,
      "postDate": "2019-10-01T11:33:40.020Z",
      "content": "<p>Thanks a lot for sharing this. Would you happen to know why the fields Window Center and Window Width in the metadata have a bunch of inconsistent values? Some are floats, others are arrays with two values that are always the same except in 34 cases of Window Center: ['50', '47'].</p>",
      "rawMarkdown": "Thanks a lot for sharing this. Would you happen to know why the fields Window Center and Window Width in the metadata have a bunch of inconsistent values? Some are floats, others are arrays with two values that are always the same except in 34 cases of Window Center: ['50', '47'].",
      "votes": 3,
      "replies": [
        {
          "id": 638158,
          "postDate": "2019-10-01T15:34:01.947Z",
          "content": "<p>Hello <a href=\"/giuliasavorgnan\">@giuliasavorgnan</a>, yes I am able to answer this question. It is common for these scans to contain an array of metadata because the radiologist would have viewed the scans with more than one window setting. </p>\n\n<p>For example: \nWindow center : [‘50’, ‘47’]\nWindow width : [‘150’,’80’]</p>\n\n<p>This means two window settings were used in the clinical setting. \n1. WL: 50 WW: 150 \n2. WL: 47 WW: 80</p>",
          "rawMarkdown": "Hello @giuliasavorgnan, yes I am able to answer this question. It is common for these scans to contain an array of metadata because the radiologist would have viewed the scans with more than one window setting. \n\nFor example: \nWindow center : [‘50’, ‘47’]\nWindow width : [‘150’,’80’]\n\nThis means two window settings were used in the clinical setting. \n1. WL: 50 WW: 150 \n2. WL: 47 WW: 80",
          "votes": 2
        },
        {
          "id": 638333,
          "postDate": "2019-10-01T19:29:15.137Z",
          "content": "<p>Understood, thank you a lot.</p>",
          "rawMarkdown": "Understood, thank you a lot.",
          "votes": 2
        }
      ]
    },
    {
      "id": 637213,
      "postDate": "2019-09-30T19:12:57.047Z",
      "content": "<p>Looks excellent, can't wait to read it and \"think more like a radiologist\"! Thank you. </p>",
      "rawMarkdown": "Looks excellent, can't wait to read it and \"think more like a radiologist\"! Thank you. ",
      "votes": 3,
      "replies": [
        {
          "id": 638216,
          "postDate": "2019-10-01T16:19:27.843Z",
          "content": "<p>Thank you for the kind words.</p>",
          "rawMarkdown": "Thank you for the kind words.",
          "votes": 1
        }
      ]
    },
    {
      "id": 658319,
      "postDate": "2019-10-25T22:19:35.817Z",
      "content": "<p>I created a dataset.png 256x256 using subdural window <a href=\"https://www.kaggle.com/custodiogabriel/rsna-256-subdural-window\">Dataset Subdural Window All Channels</a> I used the same values ​​you showed in your kernel. I adapted the code by <a href=\"/guiferviz\">@guiferviz</a>  <a href=\"https://colab.research.google.com/gist/guiferviz/50912a681776d5afe012b1a9259bd637/resize-dataset.ipynb\">https://colab.research.google.com/gist/guiferviz/50912a681776d5afe012b1a9259bd637/resize-dataset.ipynb</a></p>",
      "rawMarkdown": "I created a dataset.png 256x256 using subdural window [Dataset Subdural Window All Channels](https://www.kaggle.com/custodiogabriel/rsna-256-subdural-window) I used the same values ​​you showed in your kernel. I adapted the code by @guiferviz  https://colab.research.google.com/gist/guiferviz/50912a681776d5afe012b1a9259bd637/resize-dataset.ipynb",
      "votes": 4,
      "replies": [
        {
          "id": 659011,
          "postDate": "2019-10-26T23:18:01.187Z",
          "content": "<p>Thanks <a href=\"/custodiogabriel\">@custodiogabriel</a> for fleshing out the idea! Upvoted your dataset, and I hope it helps you or others on their competition journey. </p>",
          "rawMarkdown": "Thanks @custodiogabriel for fleshing out the idea! Upvoted your dataset, and I hope it helps you or others on their competition journey. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 652029,
      "postDate": "2019-10-18T08:43:48.850Z",
      "content": "<p>Very important detail, thanks for sharing.</p>",
      "rawMarkdown": "Very important detail, thanks for sharing.",
      "votes": 1,
      "replies": [
        {
          "id": 658042,
          "postDate": "2019-10-25T16:08:13.590Z",
          "content": "<p>You are welcome and I'm sure it helps you out!</p>",
          "rawMarkdown": "You are welcome and I'm sure it helps you out!",
          "votes": 1
        }
      ]
    },
    {
      "id": 646952,
      "postDate": "2019-10-11T23:09:39.743Z",
      "content": "<p>Very important detail for further development.\nmany thanks.</p>",
      "rawMarkdown": "Very important detail for further development.\nmany thanks.",
      "votes": 2,
      "replies": [
        {
          "id": 647827,
          "postDate": "2019-10-13T11:02:01.097Z",
          "content": "<p><a href=\"/zinovadr\">@zinovadr</a> you are welcome, and hope this helps in your kernels.</p>",
          "rawMarkdown": "@zinovadr you are welcome, and hope this helps in your kernels.",
          "votes": 2
        },
        {
          "id": 647993,
          "postDate": "2019-10-13T15:29:42.553Z",
          "content": "<p>love the sharing and learning. thanks man</p>",
          "rawMarkdown": "love the sharing and learning. thanks man",
          "votes": 1
        }
      ]
    },
    {
      "id": 646938,
      "postDate": "2019-10-11T22:42:33.203Z",
      "content": "<p>David <a href=\"/dcstang\">@dcstang</a> you're still my master. Thanks.</p>",
      "rawMarkdown": "David @dcstang you're still my master. Thanks.",
      "votes": 2,
      "replies": [
        {
          "id": 647828,
          "postDate": "2019-10-13T11:02:13.777Z",
          "content": "<p>Very kind words!</p>",
          "rawMarkdown": "Very kind words!",
          "votes": 1
        }
      ]
    },
    {
      "id": 637754,
      "postDate": "2019-10-01T08:31:42.897Z",
      "content": "<p>Thank you for this example and the explanation!</p>",
      "rawMarkdown": "Thank you for this example and the explanation!",
      "votes": 2,
      "replies": [
        {
          "id": 638160,
          "postDate": "2019-10-01T15:34:37.693Z",
          "content": "<p>Glad this has helped you out! </p>",
          "rawMarkdown": "Glad this has helped you out! "
        }
      ]
    },
    {
      "id": 637600,
      "postDate": "2019-10-01T06:40:50.070Z",
      "content": "<p>For somebody who has never worked in medicine, this is extremely helpful information. Thank you!</p>",
      "rawMarkdown": "For somebody who has never worked in medicine, this is extremely helpful information. Thank you!",
      "votes": 2,
      "replies": [
        {
          "id": 638217,
          "postDate": "2019-10-01T16:19:52.743Z",
          "content": "<p>Likewise, I have learnt a lot from the Kaggle community.</p>",
          "rawMarkdown": "Likewise, I have learnt a lot from the Kaggle community.",
          "votes": 3
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 638164,
      "author_name": "Kevin",
      "author_url": "",
      "post_date": "2019-10-01T15:38:14.627000",
      "content": "<p>Why would you need to use this window? Isn't windowing a solution for our \"poor\" human eye-sight? I'm feeding my network the full 16-bit data to see if it will solve this problem itself. If not I'll add a \"blood map\" (HU 50-70) in a second channel.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 638317,
          "author_name": "David Tang",
          "author_url": "",
          "post_date": "2019-10-01T18:48:26.367000",
          "content": "<p>Hi <a href=\"/reflexion\">@reflexion</a>, thanks for the question. I agree about your statement on human eyes.\nBasically this is to raise awareness about how one should pre-process the data from dicom images. I think your implementation should not run into issues. For those who did save the dicom into .png format some of this info may be lost.</p>\n\n<p><a href=\"https://rd.springer.com/content/pdf/10.1007%2Fs00330-019-06163-2.pdf\">https://rd.springer.com/content/pdf/10.1007%2Fs00330-019-06163-2.pdf</a>\nI like to bring your attention to this paper on hemorrhage detection. Interestingly the authors chose to use 3 window settings for all the images, then scale them before feeding into the neural network. Do you think because of the scaling, there will be significant gains in the training speed / accuracy of the network?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 638371,
          "author_name": "Alex",
          "author_url": "",
          "post_date": "2019-10-01T20:33:01.423000",
          "content": "<p><a href=\"/reflexion\">@reflexion</a> <a href=\"/dcstang\">@dcstang</a> This is how I see it: The images were labelled by <strong><em>our \"poor\" human eye-sight</em></strong>, and so will the test images. If we give the network these \"extra dimensions\" it will just be noise (no label will correlate with that). Tell me if I'm wrong, but I can't see how the neural net can make sense of it.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 638420,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-10-01T22:04:38.613000",
          "content": "<p><code>it will just be noise (no label will correlate with that).</code> \nI see your argument... but as far I can tell noise can provide good regularization.. also we dont know what kind of information NN will find useful... the best way to test will be to run small experiment and see if this improves the result. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 638614,
          "author_name": "Kevin",
          "author_url": "",
          "post_date": "2019-10-02T06:40:06.430000",
          "content": "<p>I've tested both now, and it doesn't seem to converge faster. I don't know if you can focus a network on certain value-ranges this way.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 638617,
          "author_name": "Kevin",
          "author_url": "",
          "post_date": "2019-10-02T06:43:23.593000",
          "content": "<p>The network should increase weights on relevant datapoints and ignore noise. But things like skull fractures (bone window) are not noise, they correlate with epidural bleeds, and we use them to increase likelihood of diagnosis (I'm a resident radiology).</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 638651,
          "author_name": "Alex",
          "author_url": "",
          "post_date": "2019-10-02T08:05:16.010000",
          "content": "<p><a href=\"/drhabib\">@drhabib</a> I agree that one should always run small experiments either way :-) I usually find my intuition wrong (way too often!). I didn't consider the regularization effect. Btw, have you experimented with the full pixel range? ;-) <a href=\"/reflexion\">@reflexion</a> I understand, and I won't argue with that :-), which one works better?</p>\n\n<p>First of all, I think this is an interesting topic! So let me clarify my point (not sure if this is necessary though):  if let's say a person labels pictures of small birds on 128x128 images; And say, the original image sizes are 1024x1024 (so it has been resized to 128x128 prior to labeling). Using the original 1024x1024 information for the neural net wouldn't be necessary: there might be birds appearing in those pictures, but the <strong><em>labeller</em></strong> wouldn't label it, and (s)he won't label those images for the test set either (because they have also been resized). So our network might fire a neuron or two seeing that spot in the sky with 1024x1024, but it won't be useful (instead creating noise)... because the label is 0.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 638679,
          "author_name": "Kevin",
          "author_url": "",
          "post_date": "2019-10-02T09:12:57.173000",
          "content": "<p>That's true, but radiologists never just look in one window setting, so the information might contain more than noise and also have been used by the human readers.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 638767,
          "author_name": "Alex",
          "author_url": "",
          "post_date": "2019-10-02T11:56:22.547000",
          "content": "<p>Yes that is important. Do you think they considered/based their labeling on multiple windows although only one is reported in the dicom file? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 638777,
          "author_name": "Oleg Yaroshevskiy",
          "author_url": "",
          "post_date": "2019-10-02T12:10:33.567000",
          "content": "<p>very interesting thread, subscribing :) </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 638783,
          "author_name": "Walter Wiggins",
          "author_url": "",
          "post_date": "2019-10-02T12:17:55.067000",
          "content": "<blockquote>\n  <p><strong>akensert wrote:</strong></p>\n  \n  <p>Do you think they considered/based their labeling on multiple windows although only one is reported in the dicom file?</p>\n</blockquote>\n\n<p>As a neuroradiology fellow who has done these types of annotations for an ML project previously, I believe it's unlikely that any labels were generated without viewing the images in multiple windows.</p>\n\n<p>But...as has been discussed by others, I think it remains to be shown which if any window settings actually increase signal-to-noise for CNNs in any meaningful way.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 638883,
          "author_name": "Alex",
          "author_url": "",
          "post_date": "2019-10-02T14:23:21.230000",
          "content": "<p><a href=\"/wfwiggins203\">@wfwiggins203</a> Great, thank you for the information! </p>\n\n<p>May I ask if you know why these specific window parameters of the dicom files are reported?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 638894,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-10-02T14:28:45.133000",
          "content": "<p>Not him but I can respond =)</p>\n\n<p><code>ranges to normalize images: -50–150, 100–300 and 250–450</code> </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 639095,
          "author_name": "Alex",
          "author_url": "",
          "post_date": "2019-10-02T18:58:32.323000",
          "content": "<p><a href=\"/drhabib\">@drhabib</a> These are the ranges I just started using (stacked it into 3-channeled image). From the paper above right? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 639190,
          "author_name": "Bo Peng",
          "author_url": "",
          "post_date": "2019-10-02T21:35:26.037000",
          "content": "<p>Do you guys think the brain, subdural, and bone windows are enough to predict hemorrhages or are the soft tissue and grey-white differentiation windows also needed?</p>\n\n<p>Thanks.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 639751,
          "author_name": "David Tang",
          "author_url": "",
          "post_date": "2019-10-03T14:05:48.297000",
          "content": "<p>Thanks for all the valuable comments. I appreciate Kevin trying out the comparison between rescaling and none. </p>\n\n<p><a href=\"/bopengiowa\">@bopengiowa</a>, as Walter said a typical radiologist would have gone through all windows as part of a systematic workflow process. This is because having soft tissue changes (ie. a swelling on your head) also increases the probability of having an intracranial bleed. With complex models, it may be possible that a neural network will have specific neurons that also take this into account.</p>\n\n<p>In regards to grey-white differentiation, personally I think it is better suited to other applications. Age prediction, tumor detection, stroke, for instance are things that might be easier to pick up with this window. This is just my intuition, and I will be happily surprised if a neural network proves me wrong. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 641077,
          "author_name": "Tom H.",
          "author_url": "",
          "post_date": "2019-10-04T12:28:28.273000",
          "content": "<blockquote>\n  <p>Do you think because of the scaling, there will be significant gains in the training speed / accuracy of the network?</p>\n</blockquote>\n\n<p>Linear rescaling has an interesting relationship with batch norm, and perhaps the rescaling has an effect on training (even for large networks) because it is done prior to batch normalization. It has the effect of centering the window at the mean and scaling the stddev according to the slope.</p>\n\n<p>Note that a network could, in principle, learn something like the rescaling function. That is, if it is actually a shortcut to a solution. The function in question is really simple (f(x)=ax+b) so this shouldn't matter if the network has a sufficient number of parameters. You could, in principle, get a smaller network to train faster and perform more accurately by including additional prior information about the input data if (as mentioned) the prior already has predictive capacity. I think this is kind of like training to find a transformation g(x) given f(x), x, a, and b in f(x) = g(ax+b) or f(x) = ax+b+g(x), which is basically an error term for a linear model. A robust network should be able to learn the linear portion of the equation (ax+b) given only x and f(x), but it can use the included ax+b as a shortcut. Doing so may also preclude the network from approximating a more complex function that, overall, will perform better than a linear model. IIRC this all matters less with network size.</p>\n\n<p>Are window width and center set by the radiologist or by the scanner? </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 641390,
          "author_name": "David Tang",
          "author_url": "",
          "post_date": "2019-10-04T16:39:28.257000",
          "content": "<p>Hi <a href=\"/thavlik\">@thavlik</a>, the windows are more fluid in practice, and the radiologist can \"scroll through\" a range of window values on their DICOM software. <a href=\"http://dicomviewer.booogle.net/\">http://dicomviewer.booogle.net/</a> is a sample implementation, where you can drag left and right to adjust these values. Though I suspect for this competition, the DICOM metadata has been washed clean to anonymize it and standardize most info.</p>\n\n<p>Thank you for the insight about how network size impacts this.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 642755,
          "author_name": "morg",
          "author_url": "",
          "post_date": "2019-10-06T15:32:17.070000",
          "content": "<p>Thanks <a href=\"/dcstang\">@dcstang</a> for all the medical info here, really appreciate it, suuper helpful!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 647078,
          "author_name": "Nicholas Lyu",
          "author_url": "",
          "post_date": "2019-10-12T03:22:29.790000",
          "content": "<p><a href=\"/drhabib\">@drhabib</a> <a href=\"/akensert\">@akensert</a> I think your argument is charming! I wonder if the images are labeled with default window only? If so, I agree with akensert since our model should see the same thing as the annotator. But are we sure of that?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 647246,
          "author_name": "Alex",
          "author_url": "",
          "post_date": "2019-10-12T09:11:37.470000",
          "content": "<p><a href=\"/roguekk007\">@roguekk007</a> From skimming around a bit, I think the annotators base their classification on multiple different windows (so additional windows to the default). These questions are for real annotators to answer, but I can imagine that they even have a software to check through different windows in a 'continuous way'. I still believe HUs below, let's say -100, is irrelevant. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 637933,
      "author_name": "Giulia Savorgnan",
      "author_url": "",
      "post_date": "2019-10-01T11:33:40.020000",
      "content": "<p>Thanks a lot for sharing this. Would you happen to know why the fields Window Center and Window Width in the metadata have a bunch of inconsistent values? Some are floats, others are arrays with two values that are always the same except in 34 cases of Window Center: ['50', '47'].</p>",
      "votes": 3,
      "replies": [
        {
          "id": 638158,
          "author_name": "David Tang",
          "author_url": "",
          "post_date": "2019-10-01T15:34:01.947000",
          "content": "<p>Hello <a href=\"/giuliasavorgnan\">@giuliasavorgnan</a>, yes I am able to answer this question. It is common for these scans to contain an array of metadata because the radiologist would have viewed the scans with more than one window setting. </p>\n\n<p>For example: \nWindow center : [‘50’, ‘47’]\nWindow width : [‘150’,’80’]</p>\n\n<p>This means two window settings were used in the clinical setting. \n1. WL: 50 WW: 150 \n2. WL: 47 WW: 80</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 638333,
          "author_name": "Giulia Savorgnan",
          "author_url": "",
          "post_date": "2019-10-01T19:29:15.137000",
          "content": "<p>Understood, thank you a lot.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 637213,
      "author_name": "Tom Aindow",
      "author_url": "",
      "post_date": "2019-09-30T19:12:57.047000",
      "content": "<p>Looks excellent, can't wait to read it and \"think more like a radiologist\"! Thank you. </p>",
      "votes": 3,
      "replies": [
        {
          "id": 638216,
          "author_name": "David Tang",
          "author_url": "",
          "post_date": "2019-10-01T16:19:27.843000",
          "content": "<p>Thank you for the kind words.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 658319,
      "author_name": "Gabriel",
      "author_url": "",
      "post_date": "2019-10-25T22:19:35.817000",
      "content": "<p>I created a dataset.png 256x256 using subdural window <a href=\"https://www.kaggle.com/custodiogabriel/rsna-256-subdural-window\">Dataset Subdural Window All Channels</a> I used the same values ​​you showed in your kernel. I adapted the code by <a href=\"/guiferviz\">@guiferviz</a>  <a href=\"https://colab.research.google.com/gist/guiferviz/50912a681776d5afe012b1a9259bd637/resize-dataset.ipynb\">https://colab.research.google.com/gist/guiferviz/50912a681776d5afe012b1a9259bd637/resize-dataset.ipynb</a></p>",
      "votes": 4,
      "replies": [
        {
          "id": 659011,
          "author_name": "David Tang",
          "author_url": "",
          "post_date": "2019-10-26T23:18:01.187000",
          "content": "<p>Thanks <a href=\"/custodiogabriel\">@custodiogabriel</a> for fleshing out the idea! Upvoted your dataset, and I hope it helps you or others on their competition journey. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 652029,
      "author_name": "LongYin/杰少",
      "author_url": "",
      "post_date": "2019-10-18T08:43:48.850000",
      "content": "<p>Very important detail, thanks for sharing.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 658042,
          "author_name": "David Tang",
          "author_url": "",
          "post_date": "2019-10-25T16:08:13.590000",
          "content": "<p>You are welcome and I'm sure it helps you out!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 646952,
      "author_name": "Adrian Zinovei",
      "author_url": "",
      "post_date": "2019-10-11T23:09:39.743000",
      "content": "<p>Very important detail for further development.\nmany thanks.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 647827,
          "author_name": "David Tang",
          "author_url": "",
          "post_date": "2019-10-13T11:02:01.097000",
          "content": "<p><a href=\"/zinovadr\">@zinovadr</a> you are welcome, and hope this helps in your kernels.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 647993,
          "author_name": "Adrian Zinovei",
          "author_url": "",
          "post_date": "2019-10-13T15:29:42.553000",
          "content": "<p>love the sharing and learning. thanks man</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 646938,
      "author_name": "Marília Prata",
      "author_url": "",
      "post_date": "2019-10-11T22:42:33.203000",
      "content": "<p>David <a href=\"/dcstang\">@dcstang</a> you're still my master. Thanks.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 647828,
          "author_name": "David Tang",
          "author_url": "",
          "post_date": "2019-10-13T11:02:13.777000",
          "content": "<p>Very kind words!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 637754,
      "author_name": "Sergey Zlobin",
      "author_url": "",
      "post_date": "2019-10-01T08:31:42.897000",
      "content": "<p>Thank you for this example and the explanation!</p>",
      "votes": 2,
      "replies": [
        {
          "id": 638160,
          "author_name": "David Tang",
          "author_url": "",
          "post_date": "2019-10-01T15:34:37.693000",
          "content": "<p>Glad this has helped you out! </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 637600,
      "author_name": "Alexander Abstreiter",
      "author_url": "",
      "post_date": "2019-10-01T06:40:50.070000",
      "content": "<p>For somebody who has never worked in medicine, this is extremely helpful information. Thank you!</p>",
      "votes": 2,
      "replies": [
        {
          "id": 638217,
          "author_name": "David Tang",
          "author_url": "",
          "post_date": "2019-10-01T16:19:52.743000",
          "content": "<p>Likewise, I have learnt a lot from the Kaggle community.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "637116": "Dear Kagglers,\n\nI hope to raise awareness about this setting. Currently I notice that most kernels are using the brain or parenchymal windows to view the hemorrhages.\n\nHere is a quick image from radiopedia.org, which I have annotated with arrows.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3454874%2F340f0cfb56257ff428b9e70193017731%2Fsubdural_window.png?generation=1569863200163417&amp;alt=media)\nOn the left, it is easy to miss subdural hemorrhage in the brain windows or default windows. \n\nI have discussed more in my [kernel \"See Like A Radiologist, With Systematic Windowing\"](https://www.kaggle.com/dcstang/see-like-a-radiologist-with-systematic-windowing) on how and why subdural windows are part of a standard workflow for a radiologist. \n\nThis is a way to give back to the Kaggle community, from whom I have learnt so much of Python, coding, data science and machine learning. I hope this nugget of information passed down from my medical school professors will help you in devising a better model.",
    "638164": "Why would you need to use this window? Isn't windowing a solution for our \"poor\" human eye-sight? I'm feeding my network the full 16-bit data to see if it will solve this problem itself. If not I'll add a \"blood map\" (HU 50-70) in a second channel.",
    "637933": "Thanks a lot for sharing this. Would you happen to know why the fields Window Center and Window Width in the metadata have a bunch of inconsistent values? Some are floats, others are arrays with two values that are always the same except in 34 cases of Window Center: ['50', '47'].",
    "637213": "Looks excellent, can't wait to read it and \"think more like a radiologist\"! Thank you. ",
    "658319": "I created a dataset.png 256x256 using subdural window [Dataset Subdural Window All Channels](https://www.kaggle.com/custodiogabriel/rsna-256-subdural-window) I used the same values ​​you showed in your kernel. I adapted the code by @guiferviz  https://colab.research.google.com/gist/guiferviz/50912a681776d5afe012b1a9259bd637/resize-dataset.ipynb",
    "652029": "Very important detail, thanks for sharing.",
    "646952": "Very important detail for further development.\nmany thanks.",
    "646938": "David @dcstang you're still my master. Thanks.",
    "637754": "Thank you for this example and the explanation!",
    "637600": "For somebody who has never worked in medicine, this is extremely helpful information. Thank you!"
  }
}