{
  "id": 109272,
  "title": "[lb 0.070 resnet34-512] yet another pytorch starterkit",
  "url": "/competitions/rsna-intracranial-hemorrhage-detection/discussion/109272",
  "author_name": "hengck23",
  "post_date": "2019-09-18T05:28:33.647000",
  "votes": 25,
  "comment_count": 43,
  "views": 0,
  "content": "<p>i started to join in rather late (after <a href=\"https://www.kaggle.com/c/severstal-steel-defect-detection\">https://www.kaggle.com/c/severstal-steel-defect-detection</a>).\nthe dataset was rather large. it took me a few days to setup and train the model.</p>\n\n<p>Here is the starter kit.</p>\n\n<p>you should be about to get:</p>\n\n<ul>\n<li>lb 0.070 resnet34-512 (single, tta= null, flip_lr)</li>\n<li>lb 0.080 resnet34-288 (single, tta= null, flip_lr)</li>\n<li>lb 0.070 se-resnext50-512 (single, tta= null, flip_lr)</li>\n<li>lb 0.081 efficientb4-224 (single, tta= null, flip_lr)</li>\n</ul>\n\n<p>ensemble will give about lb 0.067</p>\n\n<p>xception, efficientb0, inception3 are also implemented but not tested.</p>\n\n<p>you can use BalanceClassSampler2 in fisrt epoch to speed up training, it is important to use RandomSampler at the final training (so that the distribution of the data is not distorted)</p>",
  "messages": [
    {
      "id": 628907,
      "postDate": "2019-09-18T05:28:33.647Z",
      "content": "<p>i started to join in rather late (after <a href=\"https://www.kaggle.com/c/severstal-steel-defect-detection\">https://www.kaggle.com/c/severstal-steel-defect-detection</a>).\nthe dataset was rather large. it took me a few days to setup and train the model.</p>\n\n<p>Here is the starter kit.</p>\n\n<p>you should be about to get:</p>\n\n<ul>\n<li>lb 0.070 resnet34-512 (single, tta= null, flip_lr)</li>\n<li>lb 0.080 resnet34-288 (single, tta= null, flip_lr)</li>\n<li>lb 0.070 se-resnext50-512 (single, tta= null, flip_lr)</li>\n<li>lb 0.081 efficientb4-224 (single, tta= null, flip_lr)</li>\n</ul>\n\n<p>ensemble will give about lb 0.067</p>\n\n<p>xception, efficientb0, inception3 are also implemented but not tested.</p>\n\n<p>you can use BalanceClassSampler2 in fisrt epoch to speed up training, it is important to use RandomSampler at the final training (so that the distribution of the data is not distorted)</p>",
      "rawMarkdown": "i started to join in rather late (after https://www.kaggle.com/c/severstal-steel-defect-detection).\nthe dataset was rather large. it took me a few days to setup and train the model.\n\n\nHere is the starter kit.\n\nyou should be about to get:\n\n- lb 0.070 resnet34-512 (single, tta= null, flip_lr)\n- lb 0.080 resnet34-288 (single, tta= null, flip_lr)\n- lb 0.070 se-resnext50-512 (single, tta= null, flip_lr)\n- lb 0.081 efficientb4-224 (single, tta= null, flip_lr)\n\nensemble will give about lb 0.067\n\nxception, efficientb0, inception3 are also implemented but not tested.\n\nyou can use BalanceClassSampler2 in fisrt epoch to speed up training, it is important to use RandomSampler at the final training (so that the distribution of the data is not distorted)",
      "votes": 25
    },
    {
      "id": 661193,
      "postDate": "2019-10-30T02:07:12.430Z",
      "content": "<p>Thank you for your sharing, but I feel it is not allowed to share high score script within one weeks before deadline.  </p>",
      "rawMarkdown": "Thank you for your sharing, but I feel it is not allowed to share high score script within one weeks before deadline.  ",
      "votes": 9
    },
    {
      "id": 661699,
      "postDate": "2019-10-30T15:53:09.837Z",
      "content": "<p>i wonder if WindowCenter WindowWidth  in dicom is correlated to the target label?</p>\n\n<p>this look fishy ...</p>\n\n<p>```\ntest data:</p>\n\n<h2>only 11 unique combinations ????</h2>\n\n<pre><code>WindowCenter  WindowWidth   Freq\n</code></pre>\n\n<p>0             30           80  69272\n1             35          100   4172\n2             35          135    762\n3             36           80   2157\n4             40           80    985\n5             40           85     22\n6             40          100    136\n7             40          110     68\n8             40          120     16\n9             40          150    921\n10            47           80     34</p>\n\n<p>train data:</p>\n\n<pre><code>WindowCenter  WindowWidth    Freq\n</code></pre>\n\n<p>0             25           60      16\n1             25          100      43\n2             27           90      33\n3             28          335      34\n4             29           78      34\n5             30           80  211860\n6             30           95     344\n7             30          100    1121\n8             30          120      43\n...</p>\n\n<p>93           500         2500    531\n94           600         4000    102\n95           600         4095     32\n96           650         4000    191\n97           800         3000     32\n```</p>",
      "rawMarkdown": "i wonder if WindowCenter WindowWidth  in dicom is correlated to the target label?\n\nthis look fishy ...\n\n```\ntest data:\n## only 11 unique combinations ????\n    WindowCenter  WindowWidth   Freq\n0             30           80  69272\n1             35          100   4172\n2             35          135    762\n3             36           80   2157\n4             40           80    985\n5             40           85     22\n6             40          100    136\n7             40          110     68\n8             40          120     16\n9             40          150    921\n10            47           80     34\n\n\ntrain data:\n\n    WindowCenter  WindowWidth    Freq\n0             25           60      16\n1             25          100      43\n2             27           90      33\n3             28          335      34\n4             29           78      34\n5             30           80  211860\n6             30           95     344\n7             30          100    1121\n8             30          120      43\n...\n\n93           500         2500    531\n94           600         4000    102\n95           600         4095     32\n96           650         4000    191\n97           800         3000     32\n```",
      "votes": 3,
      "replies": [
        {
          "id": 663234,
          "postDate": "2019-11-01T16:43:01.550Z",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F29ca2ea15571b5c8333ea68f4c1c6f94%2FSelection_079.png?generation=1572626578238224&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F29ca2ea15571b5c8333ea68f4c1c6f94%2FSelection_079.png?generation=1572626578238224&amp;alt=media)\n"
        }
      ]
    },
    {
      "id": 660567,
      "postDate": "2019-10-29T10:34:58.413Z",
      "content": "<p>updated finally!</p>\n\n<p>the results is close to @Appian at</p>\n\n<p>\"Tips to get 0.066 on LB (with code on github)\"\n<a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/112819#latest-660428\">https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/112819#latest-660428</a></p>",
      "rawMarkdown": "updated finally!\n\nthe results is close to @Appian at\n\n\"Tips to get 0.066 on LB (with code on github)\"\nhttps://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/112819#latest-660428\n\n\n",
      "votes": 3,
      "replies": [
        {
          "id": 660608,
          "postDate": "2019-10-29T11:22:31.643Z",
          "content": "<p>Thank you for sharing. \nNow I see you use three windows. Have you tried without windowing?</p>\n\n<p>I have a hunch that with windowing\n- It can be trained much faster because cnn can focus on the most important range of the HU values.\n- It loses some information which can be valuable for discriminating normal brains and the others.</p>\n\n<p>but haven't tried without windowing yet.</p>",
          "rawMarkdown": "Thank you for sharing. \nNow I see you use three windows. Have you tried without windowing?\n\nI have a hunch that with windowing\n- It can be trained much faster because cnn can focus on the most important range of the HU values.\n- It loses some information which can be valuable for discriminating normal brains and the others.\n\nbut haven't tried without windowing yet."
        }
      ]
    },
    {
      "id": 630637,
      "postDate": "2019-09-20T14:30:00.447Z",
      "content": "<p>[place holder] questions regarding the details of the described steps\n... in peraparation ...</p>",
      "rawMarkdown": "[place holder] questions regarding the details of the described steps\n... in peraparation ...",
      "votes": 3
    },
    {
      "id": 628922,
      "postDate": "2019-09-18T05:49:04.840Z",
      "content": "<p>[place holder] \n… many words of appreciate in preparation …</p>",
      "rawMarkdown": "[place holder] \n… many words of appreciate in preparation …",
      "votes": 3
    },
    {
      "id": 664376,
      "postDate": "2019-11-03T15:33:57.267Z",
      "content": "<p>i only know today that the position of the CT slice is given in the dicom file.\nif you arrange the image in order, the label are continuous without break.</p>\n\n<p>this is so especially for the 'any' class.  </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F0c83f69c86b7dedfa7d6d23e726fd130%2FSelection_096.png?generation=1572798390097701&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "i only know today that the position of the CT slice is given in the dicom file.\nif you arrange the image in order, the label are continuous without break.\n\n\nthis is so especially for the 'any' class.  \n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F0c83f69c86b7dedfa7d6d23e726fd130%2FSelection_096.png?generation=1572798390097701&amp;alt=media)\n",
      "votes": 1,
      "replies": [
        {
          "id": 664436,
          "postDate": "2019-11-03T16:58:50.960Z",
          "content": "<p>it is actually mentioned in the early post here : <a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109953\">https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109953</a></p>\n\n<p>but i missed the post as i enter the competition late</p>",
          "rawMarkdown": "it is actually mentioned in the early post here : https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109953\n\nbut i missed the post as i enter the competition late"
        }
      ]
    },
    {
      "id": 663236,
      "postDate": "2019-11-01T16:48:27.710Z",
      "content": "<p>possible trick?</p>\n\n<p>input = 3 nearby slices of same patient, (if not available, just copy center slice)</p>",
      "rawMarkdown": "possible trick?\n\ninput = 3 nearby slices of same patient, (if not available, just copy center slice)",
      "votes": 1
    },
    {
      "id": 660579,
      "postDate": "2019-10-29T11:03:37.373Z",
      "content": "<p>other insight:</p>\n\n<p>initially i was using balanced sampler i.e. each train batch consist of equal images randomly draw for each class. But i got very bad results. I was very puzzled, especially after reading the forum where even simple models can give good LB score. This wasted about 2 days. But it make me realized that in order to get good Lb score, it is important:</p>\n\n<ul>\n<li><p>your probability score should be smoothed (hence ensemble is important). It should be thoroughly smoothed at each probability value</p></li>\n<li><p>your probability score must reflect the truth distribution (i.e. calibrated probability). Now the kaggle measures \"performance at each probability value\" for log loss.</p></li>\n</ul>",
      "rawMarkdown": "other insight:\n\ninitially i was using balanced sampler i.e. each train batch consist of equal images randomly draw for each class. But i got very bad results. I was very puzzled, especially after reading the forum where even simple models can give good LB score. This wasted about 2 days. But it make me realized that in order to get good Lb score, it is important:\n\n- your probability score should be smoothed (hence ensemble is important). It should be thoroughly smoothed at each probability value\n\n- your probability score must reflect the truth distribution (i.e. calibrated probability). Now the kaggle measures \"performance at each probability value\" for log loss.\n\n",
      "votes": 1,
      "replies": [
        {
          "id": 660609,
          "postDate": "2019-10-29T11:23:21.127Z",
          "content": "<p>I had exact same experience with balanced batch sampler and was not sure why. Thank you for your insight.</p>",
          "rawMarkdown": "I had exact same experience with balanced batch sampler and was not sure why. Thank you for your insight."
        },
        {
          "id": 663398,
          "postDate": "2019-11-02T00:21:37.900Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 660569,
      "postDate": "2019-10-29T10:39:39.633Z",
      "content": "<p>how a LB 0.067 ensemble results would look like:</p>\n\n<p>```\ntest submission .... @ ensemble\ncsv_file=/root/share/project/kaggle/2019/intracranial_hemorrhage/result-ensemble/xxx3/ensemble-xxx.csv</p>\n\n<h2>                          |  ratio  truth_num | pred_num |  0.0            0.2           0.4           0.6           0.8           1.0</h2>\n\n<pre><code>           any    0   |   0.138   10830   |    9518  |     66095 (0.841)  2189 (0.028)  1507 (0.019)  1623 (0.021)  7131 (0.091)\n      epidural    1   |   0.005     384   |      76  |     78164 (0.995)   252 (0.003)    85 (0.001)    36 (0.000)     8 (0.000)\n</code></pre>\n\n<p>intraparenchymal    2   |   0.045    3554   |    3001  |     74449 (0.948)   833 (0.011)   543 (0.007)   529 (0.007)  2191 (0.028)\n  intraventricular    3   |   0.031    2439   |    2174  |     75712 (0.964)   480 (0.006)   354 (0.005)   429 (0.005)  1570 (0.020)\n      subarachnoid    4   |   0.045    3553   |    2541  |     73954 (0.942)  1567 (0.020)   888 (0.011)   735 (0.009)  1401 (0.018)\n          subdural    5   |   0.059    4670   |    3577  |     72393 (0.922)  1904 (0.024)  1261 (0.016)  1208 (0.015)  1779 (0.023)</p>\n\n<p>```</p>",
      "rawMarkdown": "how a LB 0.067 ensemble results would look like:\n\n```\ntest submission .... @ ensemble\ncsv_file=/root/share/project/kaggle/2019/intracranial_hemorrhage/result-ensemble/xxx3/ensemble-xxx.csv\n\n                          |  ratio  truth_num | pred_num |  0.0            0.2           0.4           0.6           0.8           1.0\n---------------------------------------------------------------------------------------------------------------------------------------\n               any    0   |   0.138   10830   |    9518  |     66095 (0.841)  2189 (0.028)  1507 (0.019)  1623 (0.021)  7131 (0.091)\n          epidural    1   |   0.005     384   |      76  |     78164 (0.995)   252 (0.003)    85 (0.001)    36 (0.000)     8 (0.000)\n  intraparenchymal    2   |   0.045    3554   |    3001  |     74449 (0.948)   833 (0.011)   543 (0.007)   529 (0.007)  2191 (0.028)\n  intraventricular    3   |   0.031    2439   |    2174  |     75712 (0.964)   480 (0.006)   354 (0.005)   429 (0.005)  1570 (0.020)\n      subarachnoid    4   |   0.045    3553   |    2541  |     73954 (0.942)  1567 (0.020)   888 (0.011)   735 (0.009)  1401 (0.018)\n          subdural    5   |   0.059    4670   |    3577  |     72393 (0.922)  1904 (0.024)  1261 (0.016)  1208 (0.015)  1779 (0.023)\n\n```",
      "votes": 1,
      "replies": [
        {
          "id": 661229,
          "postDate": "2019-10-30T03:31:21.243Z",
          "content": "<p>i just realized that you can use the following to see how well your probability is calibrated:</p>\n\n<p>```\npredicted_score = -sum( p[i]*log(p[i]) + (1-p[i])*log(1-p[i]))</p>\n\n<p>then compare predicted_score with returned lb_score</p>\n\n<p>if your probability is well estimated, then p[i] is close to the actually value :\nnum_positive_with_predict_probe_close_to_p_i / num_all_with_predict_probe_close_to_p_i </p>\n\n<p>```</p>\n\n<p>now assume that you divide the test samples into different groups according to predicted probability, e.g. [0 to 0.2], [0.2 to 0.4] , ....</p>\n\n<p>you can make a partial submission for a single group only. you can then  find out which group is better calibrated then others.</p>\n\n<p>in fact you can repeat this probing and reverse engineer to guess the test labels with high confidence.</p>",
          "rawMarkdown": "i just realized that you can use the following to see how well your probability is calibrated:\n\n```\npredicted_score = -sum( p[i]*log(p[i]) + (1-p[i])*log(1-p[i]))\n\nthen compare predicted_score with returned lb_score\n\nif your probability is well estimated, then p[i] is close to the actually value :\nnum_positive_with_predict_probe_close_to_p_i / num_all_with_predict_probe_close_to_p_i \n\n```\n\nnow assume that you divide the test samples into different groups according to predicted probability, e.g. [0 to 0.2], [0.2 to 0.4] , ....\n\nyou can make a partial submission for a single group only. you can then  find out which group is better calibrated then others.\n\nin fact you can repeat this probing and reverse engineer to guess the test labels with high confidence.\n\n "
        }
      ]
    },
    {
      "id": 629144,
      "postDate": "2019-09-18T12:24:56.913Z",
      "content": "<p>This guy is really nice!</p>",
      "rawMarkdown": "This guy is really nice!",
      "votes": 2
    },
    {
      "id": 628928,
      "postDate": "2019-09-18T06:02:50.573Z",
      "content": "<p>人类starterkit精华</p>",
      "rawMarkdown": "人类starterkit精华",
      "votes": 1,
      "replies": [
        {
          "id": 629263,
          "postDate": "2019-09-18T15:27:26.227Z",
          "content": "<p>[Google Translate] Human starterkit essence\nDon't think itz the right translation....Could u help??</p>",
          "rawMarkdown": "[Google Translate] Human starterkit essence\nDon't think itz the right translation....Could u help??"
        },
        {
          "id": 629440,
          "postDate": "2019-09-18T19:19:30.653Z",
          "content": "<p>close...</p>\n\n<p>The crème de la crème of creating starter kits</p>",
          "rawMarkdown": "close...\n\nThe crème de la crème of creating starter kits",
          "votes": 1
        },
        {
          "id": 629652,
          "postDate": "2019-09-19T03:03:20.713Z",
          "content": "<p>Mmmmmm....human essence </p>",
          "rawMarkdown": "Mmmmmm....human essence "
        },
        {
          "id": 629831,
          "postDate": "2019-09-19T08:42:28.503Z",
          "content": "<p>The best starterkit that humans can make</p>",
          "rawMarkdown": "The best starterkit that humans can make",
          "votes": 2
        },
        {
          "id": 635994,
          "postDate": "2019-09-28T15:08:47.037Z",
          "content": "<p>Perfect translation!</p>",
          "rawMarkdown": "Perfect translation!"
        }
      ]
    },
    {
      "id": 664258,
      "postDate": "2019-11-03T12:13:15.833Z",
      "content": "<p>good online dicom viewer:\n<a href=\"https://www.imaios.com/\">https://www.imaios.com/</a></p>\n\n<p>example:\n<a href=\"https://www.imaios.com/en/e-Anatomy/Head-and-Neck/Brain-MRI-in-axial-slices\">https://www.imaios.com/en/e-Anatomy/Head-and-Neck/Brain-MRI-in-axial-slices</a></p>",
      "rawMarkdown": "good online dicom viewer:\nhttps://www.imaios.com/\n\nexample:\nhttps://www.imaios.com/en/e-Anatomy/Head-and-Neck/Brain-MRI-in-axial-slices"
    },
    {
      "id": 663792,
      "postDate": "2019-11-02T17:08:20.910Z",
      "content": "<p>288x288-resnet34-learnable-sigmoid-window: LB 0.078</p>\n\n<p>local validation LB = 0.078\nlocal train loss = 0.078</p>\n\n<p>estimate kaggle score = 0.072112 </p>\n\n<p>``` <br>\n      #assume probability p is 100% calibrated to true probability, then num_truth/num = p</p>\n\n<pre><code>    eps = 1e-15\n    p = np.clip(  probability, eps, 1-eps)\n    n = np.clip(1-probability, eps, 1-eps)\n    s = - p*np.log(p) - n*np.log(n)\n    estimate kaggle score = s.mean()\n</code></pre>\n\n<p>```</p>",
      "rawMarkdown": "288x288-resnet34-learnable-sigmoid-window: LB 0.078\n\nlocal validation LB = 0.078\nlocal train loss = 0.078\n \n\nestimate kaggle score = 0.072112 \n\n```   \n      #assume probability p is 100% calibrated to true probability, then num_truth/num = p\n \n        eps = 1e-15\n        p = np.clip(  probability, eps, 1-eps)\n        n = np.clip(1-probability, eps, 1-eps)\n        s = - p*np.log(p) - n*np.log(n)\n        estimate kaggle score = s.mean()\n```",
      "replies": [
        {
          "id": 663796,
          "postDate": "2019-11-02T17:13:11.440Z",
          "content": "<p>how the learned window looks like</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fc9055379265adaa3d1bcda8f563ae129%2F26.png?generation=1572714788338465&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "how the learned window looks like\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fc9055379265adaa3d1bcda8f563ae129%2F26.png?generation=1572714788338465&amp;alt=media)\n"
        },
        {
          "id": 719001,
          "postDate": "2020-01-15T02:30:26.433Z",
          "content": "<p><a href=\"/hengck23\">@hengck23</a> , Hello Heng, In your model.py, I don't find the function \"soft_window_to_param\" and \"hard_window_to_param\", so can you provide it?</p>",
          "rawMarkdown": "@hengck23 , Hello Heng, In your model.py, I don't find the function \"soft\\_window\\_to\\_param\" and \"hard\\_window\\_to\\_param\", so can you provide it?"
        }
      ]
    },
    {
      "id": 663751,
      "postDate": "2019-11-02T16:14:58.260Z",
      "content": "<p>additional results:</p>\n\n<p>288x288-desenset121: LB 0.73 (train with 40% of train data for quick experiment)\nlocal cross validation without patient id overlap 0.101</p>",
      "rawMarkdown": "additional results:\n\n288x288-desenset121: LB 0.73 (train with 40% of train data for quick experiment)\nlocal cross validation without patient id overlap 0.101\n\n",
      "replies": [
        {
          "id": 663752,
          "postDate": "2019-11-02T16:19:15.637Z",
          "content": "<p>logfile</p>",
          "rawMarkdown": "logfile"
        }
      ]
    },
    {
      "id": 662479,
      "postDate": "2019-10-31T15:49:08.027Z",
      "content": "<p>lstm improvement:</p>\n\n<p>lb 0.080 resnet34-288 (single, tta= null, flip_lr)</p>\n\n<p>lb 0.077 resnet34-lstm-288 (single, tta= null, flip_lr)</p>\n\n<p>```\nclass Net(nn.Module):</p>\n\n<pre><code>def load_pretrain(self, skip=['logit.', 'block0'], is_print=True):\n    load_pretrain(self, skip, pretrain_file=PRETRAIN_FILE, conversion=CONVERSION, is_print=is_print)\n\n\ndef __init__(self, num_class=6):\n    super(Net, self).__init__()\n\n    e = ResNet34()\n    self.block0 = e.block0 #128\n    self.block1 = e.block1 # 64\n    self.block2 = e.block2 # 32\n    self.block3 = e.block3 # 16\n    self.block4 = e.block4 #  8\n    e = None  #dropped\n\n    self.logit = nn.Linear(512+num_class, num_class)\n\n\ndef forward(self, x):\n    batch_size,C,H,W = x.shape\n    x = F.interpolate(x, size=(288,288),mode='bilinear', align_corners=False)\n\n    x = self.block0(x)\n    x = self.block1(x)\n    x = self.block2(x)\n    x = self.block3(x)\n    x = self.block4(x)\n\n    x = F.dropout(x,0.5,training=self.training)\n    x = F.adaptive_avg_pool2d(x,1).view(batch_size,-1)\n\n    ## unrolled lstm here ---\n\n    probability = torch.zeros(batch_size,6).to(x.device)\n    logit0 = self.logit(torch.cat([x,probability],1))\n    logit1 = self.logit(torch.cat([x,torch.sigmoid(logit0)],1))\n    logit2 = self.logit(torch.cat([x,torch.sigmoid(logit1)],1))\n\n    ## ----------------------\n\n    return logit0,logit1,logit2\n</code></pre>\n\n<p>```</p>\n\n<p>for more advanced lstm, refer to </p>\n\n<p>\"cnn-rnn: a unified framework for multi-label image classification\"\n<a href=\"https://arxiv.org/abs/1604.04573\">https://arxiv.org/abs/1604.04573</a></p>",
      "rawMarkdown": "lstm improvement:\n\nlb 0.080 resnet34-288 (single, tta= null, flip_lr)\n\nlb 0.077 resnet34-lstm-288 (single, tta= null, flip_lr)\n\n```\nclass Net(nn.Module):\n\n    def load_pretrain(self, skip=['logit.', 'block0'], is_print=True):\n        load_pretrain(self, skip, pretrain_file=PRETRAIN_FILE, conversion=CONVERSION, is_print=is_print)\n\n\n    def __init__(self, num_class=6):\n        super(Net, self).__init__()\n\n        e = ResNet34()\n        self.block0 = e.block0 #128\n        self.block1 = e.block1 # 64\n        self.block2 = e.block2 # 32\n        self.block3 = e.block3 # 16\n        self.block4 = e.block4 #  8\n        e = None  #dropped\n\n        self.logit = nn.Linear(512+num_class, num_class)\n\n\n    def forward(self, x):\n        batch_size,C,H,W = x.shape\n        x = F.interpolate(x, size=(288,288),mode='bilinear', align_corners=False)\n\n        x = self.block0(x)\n        x = self.block1(x)\n        x = self.block2(x)\n        x = self.block3(x)\n        x = self.block4(x)\n\n        x = F.dropout(x,0.5,training=self.training)\n        x = F.adaptive_avg_pool2d(x,1).view(batch_size,-1)\n\n        ## unrolled lstm here ---\n\n        probability = torch.zeros(batch_size,6).to(x.device)\n        logit0 = self.logit(torch.cat([x,probability],1))\n        logit1 = self.logit(torch.cat([x,torch.sigmoid(logit0)],1))\n        logit2 = self.logit(torch.cat([x,torch.sigmoid(logit1)],1))\n\n        ## ----------------------\n\n        return logit0,logit1,logit2\n```\n\nfor more advanced lstm, refer to \n\n\"cnn-rnn: a unified framework for multi-label image classification\"\nhttps://arxiv.org/abs/1604.04573\n\n",
      "replies": [
        {
          "id": 662819,
          "postDate": "2019-11-01T03:17:59.087Z",
          "content": "<p><a href=\"/hengck23\">@hengck23</a> Hi, This idea is so COOL! When you calculate the loss, Do you use logit2 or combine these outputs to produce a joint loss?</p>",
          "rawMarkdown": "@hengck23 Hi, This idea is so COOL! When you calculate the loss, Do you use logit2 or combine these outputs to produce a joint loss?"
        },
        {
          "id": 662874,
          "postDate": "2019-11-01T06:00:39.340Z",
          "content": "<p>The loss needed to be computed recursively too for rnn. All three loss need to backprop</p>",
          "rawMarkdown": "The loss needed to be computed recursively too for rnn. All three loss need to backprop"
        }
      ]
    },
    {
      "id": 661587,
      "postDate": "2019-10-30T13:37:04.693Z",
      "content": "<p>i find a useful augmentation:\n```\nnow subdural_image   = window_image(hu, 80, 200), where level,window = 80,200</p>\n\n<p>you can shift the slicing by window_image(hu, 80+noise, 200) and window_image(hu, 80, 200+noise),\n```</p>",
      "rawMarkdown": "i find a useful augmentation:\n```\nnow subdural_image   = window_image(hu, 80, 200), where level,window = 80,200\n\nyou can shift the slicing by window_image(hu, 80+noise, 200) and window_image(hu, 80, 200+noise),\n```",
      "replies": [
        {
          "id": 666366,
          "postDate": "2019-11-06T03:22:42.387Z",
          "content": "<p>Thanks for this advice, but I wonder what is the range of this noise, maybe ± 10% I guess?</p>",
          "rawMarkdown": "Thanks for this advice, but I wonder what is the range of this noise, maybe ± 10% I guess?"
        },
        {
          "id": 666385,
          "postDate": "2019-11-06T03:54:24.210Z",
          "content": "<p>check the window in the dicom files. that will given you a good range</p>",
          "rawMarkdown": "check the window in the dicom files. that will given you a good range"
        }
      ]
    },
    {
      "id": 660802,
      "postDate": "2019-10-29T16:23:39.957Z",
      "content": "<p>more tta for your experimetation .... it haven't submit to LB yet</p>\n\n<p>``` \n            #----\n            #random noise\n            #if 'noise' in augment:\n            if 1:\n                for t in range(4):\n                    logit = data_parallel(net, input + torch.randn(input.shape).cuda()*0.05)\n                    probability  = torch.sigmoid(logit)</p>\n\n<pre><code>                probability_label += 0.25*probability\n                num_augment += 0.25\n        #----\n\n        #----\n        #scale and crop\n        if 'scale_crop' in augment:\n        #if 1:\n            input = F.interpolate(input,(600,600), mode='bilinear', align_corners=True)\n            for x,y in [(0,0),(88,0),(0,88),(88,88)]:\n                logit = data_parallel(net,input[:,:,y:y+512,x:x+512])\n                probability  = torch.sigmoid(logit)\n\n                probability_label += 0.25*probability\n                num_augment += 0.25\n        #----\n\n\n        probability_label = probability_label/num_augment\n</code></pre>\n\n<p>```</p>",
      "rawMarkdown": "more tta for your experimetation .... it haven't submit to LB yet\n\n``` \n            #----\n            #random noise\n            #if 'noise' in augment:\n            if 1:\n                for t in range(4):\n                    logit = data_parallel(net, input + torch.randn(input.shape).cuda()*0.05)\n                    probability  = torch.sigmoid(logit)\n\n                    probability_label += 0.25*probability\n                    num_augment += 0.25\n            #----\n\n            #----\n            #scale and crop\n            if 'scale_crop' in augment:\n            #if 1:\n                input = F.interpolate(input,(600,600), mode='bilinear', align_corners=True)\n                for x,y in [(0,0),(88,0),(0,88),(88,88)]:\n                    logit = data_parallel(net,input[:,:,y:y+512,x:x+512])\n                    probability  = torch.sigmoid(logit)\n\n                    probability_label += 0.25*probability\n                    num_augment += 0.25\n            #----\n\n\n            probability_label = probability_label/num_augment\n\n```"
    },
    {
      "id": 660737,
      "postDate": "2019-10-29T14:28:39.903Z",
      "content": "<p>super link for pretrain model:\n<a href=\"https://github.com/osmr/imgclsmob\">https://github.com/osmr/imgclsmob</a></p>",
      "rawMarkdown": "super link for pretrain model:\nhttps://github.com/osmr/imgclsmob"
    },
    {
      "id": 650328,
      "postDate": "2019-10-16T09:25:21.857Z",
      "content": "<p><a href=\"/hengck23\">@hengck23</a> itll be great if you will comment your code a lil bit. As there are multiple files to read and run..\nand as always thanks for your great starterKits :)</p>",
      "rawMarkdown": "@hengck23 itll be great if you will comment your code a lil bit. As there are multiple files to read and run..\nand as always thanks for your great starterKits :)"
    },
    {
      "id": 643756,
      "postDate": "2019-10-07T21:50:44.440Z",
      "content": "<p><a href=\"/hengck23\">@hengck23</a>  Requesting TTA, <a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/110221#636658\">One cycle learning rate</a>, <a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109995#634580\">binary focal loss</a> and <a href=\"https://arxiv.org/pdf/1710.09412.pdf\">mixup</a> in your starter kit. Thanks :)</p>",
      "rawMarkdown": "@hengck23  Requesting TTA, [One cycle learning rate](https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/110221#636658), [binary focal loss](https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109995#634580) and [mixup](https://arxiv.org/pdf/1710.09412.pdf) in your starter kit. Thanks :)",
      "replies": [
        {
          "id": 644687,
          "postDate": "2019-10-09T07:29:26.803Z",
          "content": "<p><a href=\"/teeyee314\">@teeyee314</a> Here's a present for you :) <a href=\"https://github.com/pudae/kaggle-hpa/blob/master/losses/loss_factory.py\">https://github.com/pudae/kaggle-hpa/blob/master/losses/loss_factory.py</a></p>",
          "rawMarkdown": "@teeyee314 Here's a present for you :) https://github.com/pudae/kaggle-hpa/blob/master/losses/loss_factory.py",
          "votes": 1
        },
        {
          "id": 660570,
          "postDate": "2019-10-29T10:44:23.477Z",
          "content": "<p>\"binary focal loss and mixup \"   these distort the data distribution and may not be useful for the log loss (but they are useful for classification in which we care only about the decision boundary and not the decision values)</p>",
          "rawMarkdown": "\"binary focal loss and mixup \"   these distort the data distribution and may not be useful for the log loss (but they are useful for classification in which we care only about the decision boundary and not the decision values)\n",
          "votes": 2
        }
      ]
    },
    {
      "id": 635993,
      "postDate": "2019-09-28T15:07:55.433Z",
      "content": "<p>I think Heng is busy on another competition ,  waiting for update !</p>",
      "rawMarkdown": "I think Heng is busy on another competition ,  waiting for update !"
    },
    {
      "id": 629086,
      "postDate": "2019-09-18T10:08:38.820Z",
      "rawMarkdown": "",
      "votes": -2,
      "isDeleted": true
    },
    {
      "id": 629081,
      "postDate": "2019-09-18T10:01:23.543Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 661193,
      "author_name": "takuoko",
      "author_url": "",
      "post_date": "2019-10-30T02:07:12.430000",
      "content": "<p>Thank you for your sharing, but I feel it is not allowed to share high score script within one weeks before deadline.  </p>",
      "votes": 9,
      "replies": []
    },
    {
      "id": 661699,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2019-10-30T15:53:09.837000",
      "content": "<p>i wonder if WindowCenter WindowWidth  in dicom is correlated to the target label?</p>\n\n<p>this look fishy ...</p>\n\n<p>```\ntest data:</p>\n\n<h2>only 11 unique combinations ????</h2>\n\n<pre><code>WindowCenter  WindowWidth   Freq\n</code></pre>\n\n<p>0             30           80  69272\n1             35          100   4172\n2             35          135    762\n3             36           80   2157\n4             40           80    985\n5             40           85     22\n6             40          100    136\n7             40          110     68\n8             40          120     16\n9             40          150    921\n10            47           80     34</p>\n\n<p>train data:</p>\n\n<pre><code>WindowCenter  WindowWidth    Freq\n</code></pre>\n\n<p>0             25           60      16\n1             25          100      43\n2             27           90      33\n3             28          335      34\n4             29           78      34\n5             30           80  211860\n6             30           95     344\n7             30          100    1121\n8             30          120      43\n...</p>\n\n<p>93           500         2500    531\n94           600         4000    102\n95           600         4095     32\n96           650         4000    191\n97           800         3000     32\n```</p>",
      "votes": 3,
      "replies": [
        {
          "id": 663234,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2019-11-01T16:43:01.550000",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F29ca2ea15571b5c8333ea68f4c1c6f94%2FSelection_079.png?generation=1572626578238224&amp;alt=media\" alt=\"\"></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 660567,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2019-10-29T10:34:58.413000",
      "content": "<p>updated finally!</p>\n\n<p>the results is close to @Appian at</p>\n\n<p>\"Tips to get 0.066 on LB (with code on github)\"\n<a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/112819#latest-660428\">https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/112819#latest-660428</a></p>",
      "votes": 3,
      "replies": [
        {
          "id": 660608,
          "author_name": "Appian",
          "author_url": "",
          "post_date": "2019-10-29T11:22:31.643000",
          "content": "<p>Thank you for sharing. \nNow I see you use three windows. Have you tried without windowing?</p>\n\n<p>I have a hunch that with windowing\n- It can be trained much faster because cnn can focus on the most important range of the HU values.\n- It loses some information which can be valuable for discriminating normal brains and the others.</p>\n\n<p>but haven't tried without windowing yet.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 630637,
      "author_name": "Adams",
      "author_url": "",
      "post_date": "2019-09-20T14:30:00.447000",
      "content": "<p>[place holder] questions regarding the details of the described steps\n... in peraparation ...</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 628922,
      "author_name": "seefun",
      "author_url": "",
      "post_date": "2019-09-18T05:49:04.840000",
      "content": "<p>[place holder] \n… many words of appreciate in preparation …</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 664376,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2019-11-03T15:33:57.267000",
      "content": "<p>i only know today that the position of the CT slice is given in the dicom file.\nif you arrange the image in order, the label are continuous without break.</p>\n\n<p>this is so especially for the 'any' class.  </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F0c83f69c86b7dedfa7d6d23e726fd130%2FSelection_096.png?generation=1572798390097701&amp;alt=media\" alt=\"\"></p>",
      "votes": 1,
      "replies": [
        {
          "id": 664436,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2019-11-03T16:58:50.960000",
          "content": "<p>it is actually mentioned in the early post here : <a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109953\">https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109953</a></p>\n\n<p>but i missed the post as i enter the competition late</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 663236,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2019-11-01T16:48:27.710000",
      "content": "<p>possible trick?</p>\n\n<p>input = 3 nearby slices of same patient, (if not available, just copy center slice)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 660579,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2019-10-29T11:03:37.373000",
      "content": "<p>other insight:</p>\n\n<p>initially i was using balanced sampler i.e. each train batch consist of equal images randomly draw for each class. But i got very bad results. I was very puzzled, especially after reading the forum where even simple models can give good LB score. This wasted about 2 days. But it make me realized that in order to get good Lb score, it is important:</p>\n\n<ul>\n<li><p>your probability score should be smoothed (hence ensemble is important). It should be thoroughly smoothed at each probability value</p></li>\n<li><p>your probability score must reflect the truth distribution (i.e. calibrated probability). Now the kaggle measures \"performance at each probability value\" for log loss.</p></li>\n</ul>",
      "votes": 1,
      "replies": [
        {
          "id": 660609,
          "author_name": "Appian",
          "author_url": "",
          "post_date": "2019-10-29T11:23:21.127000",
          "content": "<p>I had exact same experience with balanced batch sampler and was not sure why. Thank you for your insight.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 663398,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-02T00:21:37.900000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 660569,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2019-10-29T10:39:39.633000",
      "content": "<p>how a LB 0.067 ensemble results would look like:</p>\n\n<p>```\ntest submission .... @ ensemble\ncsv_file=/root/share/project/kaggle/2019/intracranial_hemorrhage/result-ensemble/xxx3/ensemble-xxx.csv</p>\n\n<h2>                          |  ratio  truth_num | pred_num |  0.0            0.2           0.4           0.6           0.8           1.0</h2>\n\n<pre><code>           any    0   |   0.138   10830   |    9518  |     66095 (0.841)  2189 (0.028)  1507 (0.019)  1623 (0.021)  7131 (0.091)\n      epidural    1   |   0.005     384   |      76  |     78164 (0.995)   252 (0.003)    85 (0.001)    36 (0.000)     8 (0.000)\n</code></pre>\n\n<p>intraparenchymal    2   |   0.045    3554   |    3001  |     74449 (0.948)   833 (0.011)   543 (0.007)   529 (0.007)  2191 (0.028)\n  intraventricular    3   |   0.031    2439   |    2174  |     75712 (0.964)   480 (0.006)   354 (0.005)   429 (0.005)  1570 (0.020)\n      subarachnoid    4   |   0.045    3553   |    2541  |     73954 (0.942)  1567 (0.020)   888 (0.011)   735 (0.009)  1401 (0.018)\n          subdural    5   |   0.059    4670   |    3577  |     72393 (0.922)  1904 (0.024)  1261 (0.016)  1208 (0.015)  1779 (0.023)</p>\n\n<p>```</p>",
      "votes": 1,
      "replies": [
        {
          "id": 661229,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2019-10-30T03:31:21.243000",
          "content": "<p>i just realized that you can use the following to see how well your probability is calibrated:</p>\n\n<p>```\npredicted_score = -sum( p[i]*log(p[i]) + (1-p[i])*log(1-p[i]))</p>\n\n<p>then compare predicted_score with returned lb_score</p>\n\n<p>if your probability is well estimated, then p[i] is close to the actually value :\nnum_positive_with_predict_probe_close_to_p_i / num_all_with_predict_probe_close_to_p_i </p>\n\n<p>```</p>\n\n<p>now assume that you divide the test samples into different groups according to predicted probability, e.g. [0 to 0.2], [0.2 to 0.4] , ....</p>\n\n<p>you can make a partial submission for a single group only. you can then  find out which group is better calibrated then others.</p>\n\n<p>in fact you can repeat this probing and reverse engineer to guess the test labels with high confidence.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 629144,
      "author_name": "Tian",
      "author_url": "",
      "post_date": "2019-09-18T12:24:56.913000",
      "content": "<p>This guy is really nice!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 628928,
      "author_name": "SeuTao",
      "author_url": "",
      "post_date": "2019-09-18T06:02:50.573000",
      "content": "<p>人类starterkit精华</p>",
      "votes": 1,
      "replies": [
        {
          "id": 629263,
          "author_name": "Bibek",
          "author_url": "",
          "post_date": "2019-09-18T15:27:26.227000",
          "content": "<p>[Google Translate] Human starterkit essence\nDon't think itz the right translation....Could u help??</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 629440,
          "author_name": "Yee Ng",
          "author_url": "",
          "post_date": "2019-09-18T19:19:30.653000",
          "content": "<p>close...</p>\n\n<p>The crème de la crème of creating starter kits</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 629652,
          "author_name": "interneuron",
          "author_url": "",
          "post_date": "2019-09-19T03:03:20.713000",
          "content": "<p>Mmmmmm....human essence </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 629831,
          "author_name": "seefun",
          "author_url": "",
          "post_date": "2019-09-19T08:42:28.503000",
          "content": "<p>The best starterkit that humans can make</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 635994,
          "author_name": "liuzhangzhen",
          "author_url": "",
          "post_date": "2019-09-28T15:08:47.037000",
          "content": "<p>Perfect translation!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 664258,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2019-11-03T12:13:15.833000",
      "content": "<p>good online dicom viewer:\n<a href=\"https://www.imaios.com/\">https://www.imaios.com/</a></p>\n\n<p>example:\n<a href=\"https://www.imaios.com/en/e-Anatomy/Head-and-Neck/Brain-MRI-in-axial-slices\">https://www.imaios.com/en/e-Anatomy/Head-and-Neck/Brain-MRI-in-axial-slices</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 663792,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2019-11-02T17:08:20.910000",
      "content": "<p>288x288-resnet34-learnable-sigmoid-window: LB 0.078</p>\n\n<p>local validation LB = 0.078\nlocal train loss = 0.078</p>\n\n<p>estimate kaggle score = 0.072112 </p>\n\n<p>``` <br>\n      #assume probability p is 100% calibrated to true probability, then num_truth/num = p</p>\n\n<pre><code>    eps = 1e-15\n    p = np.clip(  probability, eps, 1-eps)\n    n = np.clip(1-probability, eps, 1-eps)\n    s = - p*np.log(p) - n*np.log(n)\n    estimate kaggle score = s.mean()\n</code></pre>\n\n<p>```</p>",
      "votes": 0,
      "replies": [
        {
          "id": 663796,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2019-11-02T17:13:11.440000",
          "content": "<p>how the learned window looks like</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fc9055379265adaa3d1bcda8f563ae129%2F26.png?generation=1572714788338465&amp;alt=media\" alt=\"\"></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 719001,
          "author_name": "tik_boa",
          "author_url": "",
          "post_date": "2020-01-15T02:30:26.433000",
          "content": "<p><a href=\"/hengck23\">@hengck23</a> , Hello Heng, In your model.py, I don't find the function \"soft_window_to_param\" and \"hard_window_to_param\", so can you provide it?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 663751,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2019-11-02T16:14:58.260000",
      "content": "<p>additional results:</p>\n\n<p>288x288-desenset121: LB 0.73 (train with 40% of train data for quick experiment)\nlocal cross validation without patient id overlap 0.101</p>",
      "votes": 0,
      "replies": [
        {
          "id": 663752,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2019-11-02T16:19:15.637000",
          "content": "<p>logfile</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 662479,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2019-10-31T15:49:08.027000",
      "content": "<p>lstm improvement:</p>\n\n<p>lb 0.080 resnet34-288 (single, tta= null, flip_lr)</p>\n\n<p>lb 0.077 resnet34-lstm-288 (single, tta= null, flip_lr)</p>\n\n<p>```\nclass Net(nn.Module):</p>\n\n<pre><code>def load_pretrain(self, skip=['logit.', 'block0'], is_print=True):\n    load_pretrain(self, skip, pretrain_file=PRETRAIN_FILE, conversion=CONVERSION, is_print=is_print)\n\n\ndef __init__(self, num_class=6):\n    super(Net, self).__init__()\n\n    e = ResNet34()\n    self.block0 = e.block0 #128\n    self.block1 = e.block1 # 64\n    self.block2 = e.block2 # 32\n    self.block3 = e.block3 # 16\n    self.block4 = e.block4 #  8\n    e = None  #dropped\n\n    self.logit = nn.Linear(512+num_class, num_class)\n\n\ndef forward(self, x):\n    batch_size,C,H,W = x.shape\n    x = F.interpolate(x, size=(288,288),mode='bilinear', align_corners=False)\n\n    x = self.block0(x)\n    x = self.block1(x)\n    x = self.block2(x)\n    x = self.block3(x)\n    x = self.block4(x)\n\n    x = F.dropout(x,0.5,training=self.training)\n    x = F.adaptive_avg_pool2d(x,1).view(batch_size,-1)\n\n    ## unrolled lstm here ---\n\n    probability = torch.zeros(batch_size,6).to(x.device)\n    logit0 = self.logit(torch.cat([x,probability],1))\n    logit1 = self.logit(torch.cat([x,torch.sigmoid(logit0)],1))\n    logit2 = self.logit(torch.cat([x,torch.sigmoid(logit1)],1))\n\n    ## ----------------------\n\n    return logit0,logit1,logit2\n</code></pre>\n\n<p>```</p>\n\n<p>for more advanced lstm, refer to </p>\n\n<p>\"cnn-rnn: a unified framework for multi-label image classification\"\n<a href=\"https://arxiv.org/abs/1604.04573\">https://arxiv.org/abs/1604.04573</a></p>",
      "votes": 0,
      "replies": [
        {
          "id": 662819,
          "author_name": "tik_boa",
          "author_url": "",
          "post_date": "2019-11-01T03:17:59.087000",
          "content": "<p><a href=\"/hengck23\">@hengck23</a> Hi, This idea is so COOL! When you calculate the loss, Do you use logit2 or combine these outputs to produce a joint loss?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 662874,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2019-11-01T06:00:39.340000",
          "content": "<p>The loss needed to be computed recursively too for rnn. All three loss need to backprop</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 661587,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2019-10-30T13:37:04.693000",
      "content": "<p>i find a useful augmentation:\n```\nnow subdural_image   = window_image(hu, 80, 200), where level,window = 80,200</p>\n\n<p>you can shift the slicing by window_image(hu, 80+noise, 200) and window_image(hu, 80, 200+noise),\n```</p>",
      "votes": 0,
      "replies": [
        {
          "id": 666366,
          "author_name": "Jiayu Huo",
          "author_url": "",
          "post_date": "2019-11-06T03:22:42.387000",
          "content": "<p>Thanks for this advice, but I wonder what is the range of this noise, maybe ± 10% I guess?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 666385,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2019-11-06T03:54:24.210000",
          "content": "<p>check the window in the dicom files. that will given you a good range</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 660802,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2019-10-29T16:23:39.957000",
      "content": "<p>more tta for your experimetation .... it haven't submit to LB yet</p>\n\n<p>``` \n            #----\n            #random noise\n            #if 'noise' in augment:\n            if 1:\n                for t in range(4):\n                    logit = data_parallel(net, input + torch.randn(input.shape).cuda()*0.05)\n                    probability  = torch.sigmoid(logit)</p>\n\n<pre><code>                probability_label += 0.25*probability\n                num_augment += 0.25\n        #----\n\n        #----\n        #scale and crop\n        if 'scale_crop' in augment:\n        #if 1:\n            input = F.interpolate(input,(600,600), mode='bilinear', align_corners=True)\n            for x,y in [(0,0),(88,0),(0,88),(88,88)]:\n                logit = data_parallel(net,input[:,:,y:y+512,x:x+512])\n                probability  = torch.sigmoid(logit)\n\n                probability_label += 0.25*probability\n                num_augment += 0.25\n        #----\n\n\n        probability_label = probability_label/num_augment\n</code></pre>\n\n<p>```</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 660737,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2019-10-29T14:28:39.903000",
      "content": "<p>super link for pretrain model:\n<a href=\"https://github.com/osmr/imgclsmob\">https://github.com/osmr/imgclsmob</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 650328,
      "author_name": "Kartik Nighania",
      "author_url": "",
      "post_date": "2019-10-16T09:25:21.857000",
      "content": "<p><a href=\"/hengck23\">@hengck23</a> itll be great if you will comment your code a lil bit. As there are multiple files to read and run..\nand as always thanks for your great starterKits :)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 643756,
      "author_name": "Tim Yee",
      "author_url": "",
      "post_date": "2019-10-07T21:50:44.440000",
      "content": "<p><a href=\"/hengck23\">@hengck23</a>  Requesting TTA, <a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/110221#636658\">One cycle learning rate</a>, <a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109995#634580\">binary focal loss</a> and <a href=\"https://arxiv.org/pdf/1710.09412.pdf\">mixup</a> in your starter kit. Thanks :)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 644687,
          "author_name": "datasaurus",
          "author_url": "",
          "post_date": "2019-10-09T07:29:26.803000",
          "content": "<p><a href=\"/teeyee314\">@teeyee314</a> Here's a present for you :) <a href=\"https://github.com/pudae/kaggle-hpa/blob/master/losses/loss_factory.py\">https://github.com/pudae/kaggle-hpa/blob/master/losses/loss_factory.py</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 660570,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2019-10-29T10:44:23.477000",
          "content": "<p>\"binary focal loss and mixup \"   these distort the data distribution and may not be useful for the log loss (but they are useful for classification in which we care only about the decision boundary and not the decision values)</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 635993,
      "author_name": "liuzhangzhen",
      "author_url": "",
      "post_date": "2019-09-28T15:07:55.433000",
      "content": "<p>I think Heng is busy on another competition ,  waiting for update !</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 629086,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-18T10:08:38.820000",
      "content": "",
      "votes": -2,
      "replies": []
    },
    {
      "id": 629081,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-18T10:01:23.543000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "628907": "i started to join in rather late (after https://www.kaggle.com/c/severstal-steel-defect-detection).\nthe dataset was rather large. it took me a few days to setup and train the model.\n\n\nHere is the starter kit.\n\nyou should be about to get:\n\n- lb 0.070 resnet34-512 (single, tta= null, flip_lr)\n- lb 0.080 resnet34-288 (single, tta= null, flip_lr)\n- lb 0.070 se-resnext50-512 (single, tta= null, flip_lr)\n- lb 0.081 efficientb4-224 (single, tta= null, flip_lr)\n\nensemble will give about lb 0.067\n\nxception, efficientb0, inception3 are also implemented but not tested.\n\nyou can use BalanceClassSampler2 in fisrt epoch to speed up training, it is important to use RandomSampler at the final training (so that the distribution of the data is not distorted)",
    "661193": "Thank you for your sharing, but I feel it is not allowed to share high score script within one weeks before deadline.  ",
    "661699": "i wonder if WindowCenter WindowWidth  in dicom is correlated to the target label?\n\nthis look fishy ...\n\n```\ntest data:\n## only 11 unique combinations ????\n    WindowCenter  WindowWidth   Freq\n0             30           80  69272\n1             35          100   4172\n2             35          135    762\n3             36           80   2157\n4             40           80    985\n5             40           85     22\n6             40          100    136\n7             40          110     68\n8             40          120     16\n9             40          150    921\n10            47           80     34\n\n\ntrain data:\n\n    WindowCenter  WindowWidth    Freq\n0             25           60      16\n1             25          100      43\n2             27           90      33\n3             28          335      34\n4             29           78      34\n5             30           80  211860\n6             30           95     344\n7             30          100    1121\n8             30          120      43\n...\n\n93           500         2500    531\n94           600         4000    102\n95           600         4095     32\n96           650         4000    191\n97           800         3000     32\n```",
    "660567": "updated finally!\n\nthe results is close to @Appian at\n\n\"Tips to get 0.066 on LB (with code on github)\"\nhttps://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/112819#latest-660428\n\n\n",
    "630637": "[place holder] questions regarding the details of the described steps\n... in peraparation ...",
    "628922": "[place holder] \n… many words of appreciate in preparation …",
    "664376": "i only know today that the position of the CT slice is given in the dicom file.\nif you arrange the image in order, the label are continuous without break.\n\n\nthis is so especially for the 'any' class.  \n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F0c83f69c86b7dedfa7d6d23e726fd130%2FSelection_096.png?generation=1572798390097701&amp;alt=media)\n",
    "663236": "possible trick?\n\ninput = 3 nearby slices of same patient, (if not available, just copy center slice)",
    "660579": "other insight:\n\ninitially i was using balanced sampler i.e. each train batch consist of equal images randomly draw for each class. But i got very bad results. I was very puzzled, especially after reading the forum where even simple models can give good LB score. This wasted about 2 days. But it make me realized that in order to get good Lb score, it is important:\n\n- your probability score should be smoothed (hence ensemble is important). It should be thoroughly smoothed at each probability value\n\n- your probability score must reflect the truth distribution (i.e. calibrated probability). Now the kaggle measures \"performance at each probability value\" for log loss.\n\n",
    "660569": "how a LB 0.067 ensemble results would look like:\n\n```\ntest submission .... @ ensemble\ncsv_file=/root/share/project/kaggle/2019/intracranial_hemorrhage/result-ensemble/xxx3/ensemble-xxx.csv\n\n                          |  ratio  truth_num | pred_num |  0.0            0.2           0.4           0.6           0.8           1.0\n---------------------------------------------------------------------------------------------------------------------------------------\n               any    0   |   0.138   10830   |    9518  |     66095 (0.841)  2189 (0.028)  1507 (0.019)  1623 (0.021)  7131 (0.091)\n          epidural    1   |   0.005     384   |      76  |     78164 (0.995)   252 (0.003)    85 (0.001)    36 (0.000)     8 (0.000)\n  intraparenchymal    2   |   0.045    3554   |    3001  |     74449 (0.948)   833 (0.011)   543 (0.007)   529 (0.007)  2191 (0.028)\n  intraventricular    3   |   0.031    2439   |    2174  |     75712 (0.964)   480 (0.006)   354 (0.005)   429 (0.005)  1570 (0.020)\n      subarachnoid    4   |   0.045    3553   |    2541  |     73954 (0.942)  1567 (0.020)   888 (0.011)   735 (0.009)  1401 (0.018)\n          subdural    5   |   0.059    4670   |    3577  |     72393 (0.922)  1904 (0.024)  1261 (0.016)  1208 (0.015)  1779 (0.023)\n\n```",
    "629144": "This guy is really nice!",
    "628928": "人类starterkit精华",
    "664258": "good online dicom viewer:\nhttps://www.imaios.com/\n\nexample:\nhttps://www.imaios.com/en/e-Anatomy/Head-and-Neck/Brain-MRI-in-axial-slices",
    "663792": "288x288-resnet34-learnable-sigmoid-window: LB 0.078\n\nlocal validation LB = 0.078\nlocal train loss = 0.078\n \n\nestimate kaggle score = 0.072112 \n\n```   \n      #assume probability p is 100% calibrated to true probability, then num_truth/num = p\n \n        eps = 1e-15\n        p = np.clip(  probability, eps, 1-eps)\n        n = np.clip(1-probability, eps, 1-eps)\n        s = - p*np.log(p) - n*np.log(n)\n        estimate kaggle score = s.mean()\n```",
    "663751": "additional results:\n\n288x288-desenset121: LB 0.73 (train with 40% of train data for quick experiment)\nlocal cross validation without patient id overlap 0.101\n\n",
    "662479": "lstm improvement:\n\nlb 0.080 resnet34-288 (single, tta= null, flip_lr)\n\nlb 0.077 resnet34-lstm-288 (single, tta= null, flip_lr)\n\n```\nclass Net(nn.Module):\n\n    def load_pretrain(self, skip=['logit.', 'block0'], is_print=True):\n        load_pretrain(self, skip, pretrain_file=PRETRAIN_FILE, conversion=CONVERSION, is_print=is_print)\n\n\n    def __init__(self, num_class=6):\n        super(Net, self).__init__()\n\n        e = ResNet34()\n        self.block0 = e.block0 #128\n        self.block1 = e.block1 # 64\n        self.block2 = e.block2 # 32\n        self.block3 = e.block3 # 16\n        self.block4 = e.block4 #  8\n        e = None  #dropped\n\n        self.logit = nn.Linear(512+num_class, num_class)\n\n\n    def forward(self, x):\n        batch_size,C,H,W = x.shape\n        x = F.interpolate(x, size=(288,288),mode='bilinear', align_corners=False)\n\n        x = self.block0(x)\n        x = self.block1(x)\n        x = self.block2(x)\n        x = self.block3(x)\n        x = self.block4(x)\n\n        x = F.dropout(x,0.5,training=self.training)\n        x = F.adaptive_avg_pool2d(x,1).view(batch_size,-1)\n\n        ## unrolled lstm here ---\n\n        probability = torch.zeros(batch_size,6).to(x.device)\n        logit0 = self.logit(torch.cat([x,probability],1))\n        logit1 = self.logit(torch.cat([x,torch.sigmoid(logit0)],1))\n        logit2 = self.logit(torch.cat([x,torch.sigmoid(logit1)],1))\n\n        ## ----------------------\n\n        return logit0,logit1,logit2\n```\n\nfor more advanced lstm, refer to \n\n\"cnn-rnn: a unified framework for multi-label image classification\"\nhttps://arxiv.org/abs/1604.04573\n\n",
    "661587": "i find a useful augmentation:\n```\nnow subdural_image   = window_image(hu, 80, 200), where level,window = 80,200\n\nyou can shift the slicing by window_image(hu, 80+noise, 200) and window_image(hu, 80, 200+noise),\n```",
    "660802": "more tta for your experimetation .... it haven't submit to LB yet\n\n``` \n            #----\n            #random noise\n            #if 'noise' in augment:\n            if 1:\n                for t in range(4):\n                    logit = data_parallel(net, input + torch.randn(input.shape).cuda()*0.05)\n                    probability  = torch.sigmoid(logit)\n\n                    probability_label += 0.25*probability\n                    num_augment += 0.25\n            #----\n\n            #----\n            #scale and crop\n            if 'scale_crop' in augment:\n            #if 1:\n                input = F.interpolate(input,(600,600), mode='bilinear', align_corners=True)\n                for x,y in [(0,0),(88,0),(0,88),(88,88)]:\n                    logit = data_parallel(net,input[:,:,y:y+512,x:x+512])\n                    probability  = torch.sigmoid(logit)\n\n                    probability_label += 0.25*probability\n                    num_augment += 0.25\n            #----\n\n\n            probability_label = probability_label/num_augment\n\n```",
    "660737": "super link for pretrain model:\nhttps://github.com/osmr/imgclsmob",
    "650328": "@hengck23 itll be great if you will comment your code a lil bit. As there are multiple files to read and run..\nand as always thanks for your great starterKits :)",
    "643756": "@hengck23  Requesting TTA, [One cycle learning rate](https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/110221#636658), [binary focal loss](https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109995#634580) and [mixup](https://arxiv.org/pdf/1710.09412.pdf) in your starter kit. Thanks :)",
    "635993": "I think Heng is busy on another competition ,  waiting for update !",
    "629086": "",
    "629081": ""
  }
}