{
  "id": 117417,
  "title": "16th place solution",
  "url": "/competitions/rsna-intracranial-hemorrhage-detection/discussion/117417",
  "author_name": "Beomhee Park",
  "post_date": "2019-11-15T11:24:49.043000",
  "votes": 12,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Congratulations to all winners and thanks to kaggle and organizers for opening this learning space.</p>\n\n<p>We started relatively late, but we made a good starting point with the code that <a href=\"/appian\">@appian</a> shared. Great thanks to <a href=\"/appian\">@appian</a> </p>\n\n<p>Our overall procedure is as follows.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2276152%2F770038c6cb65c7774c82ff8b1e1c5877%2F17th_solution_figure2.png?generation=1573814036534212&amp;alt=media\" alt=\"\"></p>\n\n<h3>In step 1</h3>\n\n<ul>\n<li>Basic training is performed by considering an image as an independent input.\n<ul><li>input shape : (batch_size, 512, 512, 3)\n<ul><li>4th axis (3) means 3 channels with multiple windowing parameters</li></ul></li>\n<li>output shape : (batch_size, 6)</li>\n<li>CNN Architectures : SE-ResNeXt-101 and EfficientNet-B6</li>\n<li>loss : weighted log loss (weights = [2/7, 1/7, 1/7, 1/7, 1/7, 1/7])</li>\n<li>optimizer : Adam (with learning rate from 1e-4 to 1e-5)</li>\n<li>sampling : random sampling or location based sampling (sampling middle slices more from image series in patient-level)</li>\n<li>5 folds or 7 folds training</li></ul></li>\n</ul>\n\n<h3>In step 2</h3>\n\n<ul>\n<li>We wanted to calibrate the output distributions considering the relation of labels or adjacent image slices,  so we recognized the outputs of patient-level images as a signal and trained the model.</li>\n<li>Output distributions are extracted from the validation set. (For example, 5 models from 5 folds can make total training dataset.)</li>\n<li>If about 640,000 images are used in step1, about 19,500 output signals (the number of patients) are used in step 2.\n<ul><li>input shape : (batch_size, None, 6, 1)\n<ul><li>1 axis (None) means the length of signal (the number of slices)</li>\n<li>2 axis (6) means the number of labels</li></ul></li>\n<li>output shape : (batch_size, None, 6, 1)</li>\n<li>CNN Architecture : simple CNN model with 4 convolution layers having 5x6 matrix</li>\n<li>loss : weighted log loss (weights = [2/7, 1/7, 1/7, 1/7, 1/7, 1/7])</li>\n<li>optimizer : Adam (with learning rate 1e-5)</li>\n<li>5 folds training</li></ul></li>\n</ul>\n\n<p>```</p>\n\n<hr>\n\n<h1>Layer (type)                 Output Shape              Param #   </h1>\n\n<p>input_1 (InputLayer)         (None, None, 6, 1)        0         </p>\n\n<hr>\n\n<p>conv2d_1 (Conv2D)            (None, None, 6, 64)       1984      </p>\n\n<hr>\n\n<p>conv2d_2 (Conv2D)            (None, None, 6, 64)       122944    </p>\n\n<hr>\n\n<p>conv2d_3 (Conv2D)            (None, None, 6, 64)       122944    </p>\n\n<hr>\n\n<p>conv2d_4 (Conv2D)            (None, None, 6, 64)       122944    </p>\n\n<hr>\n\n<h1>conv2d_5 (Conv2D)            (None, None, 6, 1)        65        </h1>\n\n<p>```</p>\n\n<p>We also thought to handle sequential information in image-level, but the deadline was short, so the process was split into two steps and output signals with relatively small dimensions were used as the next best thing. </p>\n\n<p><strong>The results are as follows.</strong>\n*<em>step 1 result : 0.05425 (private score)</em>*\n<strong>step 2 result : 0.04793 (private score)</strong>\n*<em>We think that the core processing of our team, like other teams, also was to reflect sequential information.</em>*</p>",
  "messages": [
    {
      "id": 673711,
      "postDate": "2019-11-15T11:24:49.043Z",
      "content": "<p>Congratulations to all winners and thanks to kaggle and organizers for opening this learning space.</p>\n\n<p>We started relatively late, but we made a good starting point with the code that <a href=\"/appian\">@appian</a> shared. Great thanks to <a href=\"/appian\">@appian</a> </p>\n\n<p>Our overall procedure is as follows.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2276152%2F770038c6cb65c7774c82ff8b1e1c5877%2F17th_solution_figure2.png?generation=1573814036534212&amp;alt=media\" alt=\"\"></p>\n\n<h3>In step 1</h3>\n\n<ul>\n<li>Basic training is performed by considering an image as an independent input.\n<ul><li>input shape : (batch_size, 512, 512, 3)\n<ul><li>4th axis (3) means 3 channels with multiple windowing parameters</li></ul></li>\n<li>output shape : (batch_size, 6)</li>\n<li>CNN Architectures : SE-ResNeXt-101 and EfficientNet-B6</li>\n<li>loss : weighted log loss (weights = [2/7, 1/7, 1/7, 1/7, 1/7, 1/7])</li>\n<li>optimizer : Adam (with learning rate from 1e-4 to 1e-5)</li>\n<li>sampling : random sampling or location based sampling (sampling middle slices more from image series in patient-level)</li>\n<li>5 folds or 7 folds training</li></ul></li>\n</ul>\n\n<h3>In step 2</h3>\n\n<ul>\n<li>We wanted to calibrate the output distributions considering the relation of labels or adjacent image slices,  so we recognized the outputs of patient-level images as a signal and trained the model.</li>\n<li>Output distributions are extracted from the validation set. (For example, 5 models from 5 folds can make total training dataset.)</li>\n<li>If about 640,000 images are used in step1, about 19,500 output signals (the number of patients) are used in step 2.\n<ul><li>input shape : (batch_size, None, 6, 1)\n<ul><li>1 axis (None) means the length of signal (the number of slices)</li>\n<li>2 axis (6) means the number of labels</li></ul></li>\n<li>output shape : (batch_size, None, 6, 1)</li>\n<li>CNN Architecture : simple CNN model with 4 convolution layers having 5x6 matrix</li>\n<li>loss : weighted log loss (weights = [2/7, 1/7, 1/7, 1/7, 1/7, 1/7])</li>\n<li>optimizer : Adam (with learning rate 1e-5)</li>\n<li>5 folds training</li></ul></li>\n</ul>\n\n<p>```</p>\n\n<hr>\n\n<h1>Layer (type)                 Output Shape              Param #   </h1>\n\n<p>input_1 (InputLayer)         (None, None, 6, 1)        0         </p>\n\n<hr>\n\n<p>conv2d_1 (Conv2D)            (None, None, 6, 64)       1984      </p>\n\n<hr>\n\n<p>conv2d_2 (Conv2D)            (None, None, 6, 64)       122944    </p>\n\n<hr>\n\n<p>conv2d_3 (Conv2D)            (None, None, 6, 64)       122944    </p>\n\n<hr>\n\n<p>conv2d_4 (Conv2D)            (None, None, 6, 64)       122944    </p>\n\n<hr>\n\n<h1>conv2d_5 (Conv2D)            (None, None, 6, 1)        65        </h1>\n\n<p>```</p>\n\n<p>We also thought to handle sequential information in image-level, but the deadline was short, so the process was split into two steps and output signals with relatively small dimensions were used as the next best thing. </p>\n\n<p><strong>The results are as follows.</strong>\n*<em>step 1 result : 0.05425 (private score)</em>*\n<strong>step 2 result : 0.04793 (private score)</strong>\n*<em>We think that the core processing of our team, like other teams, also was to reflect sequential information.</em>*</p>",
      "rawMarkdown": "Congratulations to all winners and thanks to kaggle and organizers for opening this learning space.\n\nWe started relatively late, but we made a good starting point with the code that @appian shared. Great thanks to @appian \n\nOur overall procedure is as follows.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2276152%2F770038c6cb65c7774c82ff8b1e1c5877%2F17th_solution_figure2.png?generation=1573814036534212&amp;alt=media)\n\n### In step 1\n* Basic training is performed by considering an image as an independent input.\n    * input shape : (batch_size, 512, 512, 3)\n        * 4th axis (3) means 3 channels with multiple windowing parameters\n    * output shape : (batch_size, 6)\n    * CNN Architectures : SE-ResNeXt-101 and EfficientNet-B6\n    * loss : weighted log loss (weights = [2/7, 1/7, 1/7, 1/7, 1/7, 1/7])\n    * optimizer : Adam (with learning rate from 1e-4 to 1e-5)\n    * sampling : random sampling or location based sampling (sampling middle slices more from image series in patient-level)\n    * 5 folds or 7 folds training\n\n### In step 2\n* We wanted to calibrate the output distributions considering the relation of labels or adjacent image slices,  so we recognized the outputs of patient-level images as a signal and trained the model.\n* Output distributions are extracted from the validation set. (For example, 5 models from 5 folds can make total training dataset.)\n* If about 640,000 images are used in step1, about 19,500 output signals (the number of patients) are used in step 2.\n    * input shape : (batch_size, None, 6, 1)\n        * 1 axis (None) means the length of signal (the number of slices)\n        * 2 axis (6) means the number of labels\n    * output shape : (batch_size, None, 6, 1)\n    * CNN Architecture : simple CNN model with 4 convolution layers having 5x6 matrix\n    * loss : weighted log loss (weights = [2/7, 1/7, 1/7, 1/7, 1/7, 1/7])\n    * optimizer : Adam (with learning rate 1e-5)\n    * 5 folds training\n\n```\n_________________________________________________________________\nLayer (type)                 Output Shape              Param #   \n=================================================================\ninput_1 (InputLayer)         (None, None, 6, 1)        0         \n_________________________________________________________________\nconv2d_1 (Conv2D)            (None, None, 6, 64)       1984      \n_________________________________________________________________\nconv2d_2 (Conv2D)            (None, None, 6, 64)       122944    \n_________________________________________________________________\nconv2d_3 (Conv2D)            (None, None, 6, 64)       122944    \n_________________________________________________________________\nconv2d_4 (Conv2D)            (None, None, 6, 64)       122944    \n_________________________________________________________________\nconv2d_5 (Conv2D)            (None, None, 6, 1)        65        \n=================================================================\n```\n\nWe also thought to handle sequential information in image-level, but the deadline was short, so the process was split into two steps and output signals with relatively small dimensions were used as the next best thing. \n\n**The results are as follows.**\n**step 1 result : 0.05425 (private score)**\n**step 2 result : 0.04793 (private score)**\n**We think that the core processing of our team, like other teams, also was to reflect sequential information.**",
      "votes": 12
    },
    {
      "id": 675180,
      "postDate": "2019-11-17T17:57:35.327Z",
      "content": "<p>nice work. congrats. thanks for sharing.</p>",
      "rawMarkdown": "nice work. congrats. thanks for sharing.",
      "votes": 1
    },
    {
      "id": 673718,
      "postDate": "2019-11-15T11:47:05.337Z",
      "content": "<p>Awesome. Congrats. Nice model, thanks for sharing. <a href=\"/beomheep\">@beomheep</a> </p>",
      "rawMarkdown": "Awesome. Congrats. Nice model, thanks for sharing. @beomheep ",
      "votes": 1
    },
    {
      "id": 3422257,
      "postDate": "2026-03-17T10:26:16.513Z",
      "content": "<p>how are you splitting test and train data? Isn't there slices of same patient gets into both? Can it be handled using meta data?? I am trying a subset of this data set for binary classification and test set evaluation gives this!.. I can smell leak😅</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F24327501%2F106e69ea1341c6c503500f3db3732680%2Ffvgzegsd.png?generation=1773743136671651&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "how are you splitting test and train data? Isn't there slices of same patient gets into both? Can it be handled using meta data?? I am trying a subset of this data set for binary classification and test set evaluation gives this!.. I can smell leak😅\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F24327501%2F106e69ea1341c6c503500f3db3732680%2Ffvgzegsd.png?generation=1773743136671651&alt=media)"
    }
  ],
  "comments": [
    {
      "id": 675180,
      "author_name": "M Arrabi",
      "author_url": "",
      "post_date": "2019-11-17T17:57:35.327000",
      "content": "<p>nice work. congrats. thanks for sharing.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 673718,
      "author_name": "Ailurophile",
      "author_url": "",
      "post_date": "2019-11-15T11:47:05.337000",
      "content": "<p>Awesome. Congrats. Nice model, thanks for sharing. <a href=\"/beomheep\">@beomheep</a> </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3422257,
      "author_name": "Aviral Srivastava",
      "author_url": "",
      "post_date": "2026-03-17T10:26:16.513000",
      "content": "<p>how are you splitting test and train data? Isn't there slices of same patient gets into both? Can it be handled using meta data?? I am trying a subset of this data set for binary classification and test set evaluation gives this!.. I can smell leak😅</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F24327501%2F106e69ea1341c6c503500f3db3732680%2Ffvgzegsd.png?generation=1773743136671651&amp;alt=media\" alt=\"\"></p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "673711": "Congratulations to all winners and thanks to kaggle and organizers for opening this learning space.\n\nWe started relatively late, but we made a good starting point with the code that @appian shared. Great thanks to @appian \n\nOur overall procedure is as follows.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2276152%2F770038c6cb65c7774c82ff8b1e1c5877%2F17th_solution_figure2.png?generation=1573814036534212&amp;alt=media)\n\n### In step 1\n* Basic training is performed by considering an image as an independent input.\n    * input shape : (batch_size, 512, 512, 3)\n        * 4th axis (3) means 3 channels with multiple windowing parameters\n    * output shape : (batch_size, 6)\n    * CNN Architectures : SE-ResNeXt-101 and EfficientNet-B6\n    * loss : weighted log loss (weights = [2/7, 1/7, 1/7, 1/7, 1/7, 1/7])\n    * optimizer : Adam (with learning rate from 1e-4 to 1e-5)\n    * sampling : random sampling or location based sampling (sampling middle slices more from image series in patient-level)\n    * 5 folds or 7 folds training\n\n### In step 2\n* We wanted to calibrate the output distributions considering the relation of labels or adjacent image slices,  so we recognized the outputs of patient-level images as a signal and trained the model.\n* Output distributions are extracted from the validation set. (For example, 5 models from 5 folds can make total training dataset.)\n* If about 640,000 images are used in step1, about 19,500 output signals (the number of patients) are used in step 2.\n    * input shape : (batch_size, None, 6, 1)\n        * 1 axis (None) means the length of signal (the number of slices)\n        * 2 axis (6) means the number of labels\n    * output shape : (batch_size, None, 6, 1)\n    * CNN Architecture : simple CNN model with 4 convolution layers having 5x6 matrix\n    * loss : weighted log loss (weights = [2/7, 1/7, 1/7, 1/7, 1/7, 1/7])\n    * optimizer : Adam (with learning rate 1e-5)\n    * 5 folds training\n\n```\n_________________________________________________________________\nLayer (type)                 Output Shape              Param #   \n=================================================================\ninput_1 (InputLayer)         (None, None, 6, 1)        0         \n_________________________________________________________________\nconv2d_1 (Conv2D)            (None, None, 6, 64)       1984      \n_________________________________________________________________\nconv2d_2 (Conv2D)            (None, None, 6, 64)       122944    \n_________________________________________________________________\nconv2d_3 (Conv2D)            (None, None, 6, 64)       122944    \n_________________________________________________________________\nconv2d_4 (Conv2D)            (None, None, 6, 64)       122944    \n_________________________________________________________________\nconv2d_5 (Conv2D)            (None, None, 6, 1)        65        \n=================================================================\n```\n\nWe also thought to handle sequential information in image-level, but the deadline was short, so the process was split into two steps and output signals with relatively small dimensions were used as the next best thing. \n\n**The results are as follows.**\n**step 1 result : 0.05425 (private score)**\n**step 2 result : 0.04793 (private score)**\n**We think that the core processing of our team, like other teams, also was to reflect sequential information.**",
    "675180": "nice work. congrats. thanks for sharing.",
    "673718": "Awesome. Congrats. Nice model, thanks for sharing. @beomheep ",
    "3422257": "how are you splitting test and train data? Isn't there slices of same patient gets into both? Can it be handled using meta data?? I am trying a subset of this data set for binary classification and test set evaluation gives this!.. I can smell leak😅\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F24327501%2F106e69ea1341c6c503500f3db3732680%2Ffvgzegsd.png?generation=1773743136671651&alt=media)"
  }
}