{
  "id": 117493,
  "title": "Academic paper opportunity",
  "url": "/competitions/rsna-intracranial-hemorrhage-detection/discussion/117493",
  "author_name": "Jeremy Howard",
  "post_date": "2019-11-15T22:28:11.978000",
  "votes": 28,
  "comment_count": 21,
  "views": 0,
  "content": "<p>Hi all! Congrats on all the great results in this competition. So many really interesting approaches. :)</p>\n\n<p>Unfortunately, the learnings from these competitions sometimes don't impact the wider research and practitioner community as much as they should. The topic of this competition is too important to allow that to happen here. So let's make sure it doesn't!</p>\n\n<p>As some of you know, I chair an academic initiative called <a href=\"https://wamri.ai/\">WAMRI</a>, which aims to advance the state of the art in medical AI, and also to make it more accessible to both researchers and clinicians. We are working with brilliant collaborators at institutions including UCSF, Stanford, Harvard, Salk Institute, and many more. This competition is a great opportunity to help towards our goal, so I'm hoping to organize a project to collect the best approaches that were developed and tested here, see which ones combine well together, and publish both an academic paper describing the results of that research, as well as publishing one or more pretrained models that anyone can use to help kick start their own models. I'm also hoping to publish research results showing the impact of transfer learning best practices in medical imaging, using these pretrained models as the basis of that analysis.</p>\n\n<p>Of course, anyone who created a novel approach during the competition that is used in this research would be credited in the paper. So, if you're interested in contributing to this important next step in turning this competition into real medical outcomes, and you have a strong single model (or think that you might), could you please do the following:</p>\n\n<ol>\n<li>Create a submission from your best single model, not including any kind of TTA or other ensembling (e.g. also not ensembling across checkpoints), and do a post-competition submission</li>\n<li>Add a reply in this thread stating what private leaderboard result you get with that submission (NB: ignore the public leaderboard result, since that is the 1% subset they used for stage 2; just look at the private leaderboard result, which will be for the stage 2 test set)</li>\n<li>In your reply, summarize the key pieces of this model (or link to an existing forum post you've made that describes it, mentioning which is your best single model)</li>\n<li>Also, let me know in your reply if you're interested in helping with this post-competition research, either in an informal way (i.e. just answering a few questions as we try to replicate your result), or more formally (i.e. assisting in the followup research and writing).</li>\n</ol>\n\n<p>I've started some initial discussions with some folks involved in organizing the competition, and it looks like there might be quite a bit of interest from all sides in contributing to this. However, it's early days, so the exact final nature and makeup of this collaboration isn't yet well defined at all. So it might take a while for this all to come to fruition - which may require some patience on all sides... I'll be sure to keep this thread updated with news as I have it.</p>\n\n<p>cc all the (provisional) gold medalists: @scp173 @shentao @currylc @lanjunyelan @scusywxy @dmitrylarko @darraghdog @takuok @liut0012 @mfang10 @antherxu @yunpchen @anjum48 @tarobxl @mathormad @msl23518 @backaggle @godaibo @thomasal @wowfattie @dmytropoplavskiy @meanshift @maciejbudys @nordberdt @tgilewicz @antorsae @macayaven @cristinagrs @bacterio @shimacos @sugawarya @losveria @tkyyym @appian @yuval6967 @zaharch </p>",
  "messages": [
    {
      "id": 674104,
      "postDate": "2019-11-15T22:28:11.977Z",
      "content": "<p>Hi all! Congrats on all the great results in this competition. So many really interesting approaches. :)</p>\n\n<p>Unfortunately, the learnings from these competitions sometimes don't impact the wider research and practitioner community as much as they should. The topic of this competition is too important to allow that to happen here. So let's make sure it doesn't!</p>\n\n<p>As some of you know, I chair an academic initiative called <a href=\"https://wamri.ai/\">WAMRI</a>, which aims to advance the state of the art in medical AI, and also to make it more accessible to both researchers and clinicians. We are working with brilliant collaborators at institutions including UCSF, Stanford, Harvard, Salk Institute, and many more. This competition is a great opportunity to help towards our goal, so I'm hoping to organize a project to collect the best approaches that were developed and tested here, see which ones combine well together, and publish both an academic paper describing the results of that research, as well as publishing one or more pretrained models that anyone can use to help kick start their own models. I'm also hoping to publish research results showing the impact of transfer learning best practices in medical imaging, using these pretrained models as the basis of that analysis.</p>\n\n<p>Of course, anyone who created a novel approach during the competition that is used in this research would be credited in the paper. So, if you're interested in contributing to this important next step in turning this competition into real medical outcomes, and you have a strong single model (or think that you might), could you please do the following:</p>\n\n<ol>\n<li>Create a submission from your best single model, not including any kind of TTA or other ensembling (e.g. also not ensembling across checkpoints), and do a post-competition submission</li>\n<li>Add a reply in this thread stating what private leaderboard result you get with that submission (NB: ignore the public leaderboard result, since that is the 1% subset they used for stage 2; just look at the private leaderboard result, which will be for the stage 2 test set)</li>\n<li>In your reply, summarize the key pieces of this model (or link to an existing forum post you've made that describes it, mentioning which is your best single model)</li>\n<li>Also, let me know in your reply if you're interested in helping with this post-competition research, either in an informal way (i.e. just answering a few questions as we try to replicate your result), or more formally (i.e. assisting in the followup research and writing).</li>\n</ol>\n\n<p>I've started some initial discussions with some folks involved in organizing the competition, and it looks like there might be quite a bit of interest from all sides in contributing to this. However, it's early days, so the exact final nature and makeup of this collaboration isn't yet well defined at all. So it might take a while for this all to come to fruition - which may require some patience on all sides... I'll be sure to keep this thread updated with news as I have it.</p>\n\n<p>cc all the (provisional) gold medalists: @scp173 @shentao @currylc @lanjunyelan @scusywxy @dmitrylarko @darraghdog @takuok @liut0012 @mfang10 @antherxu @yunpchen @anjum48 @tarobxl @mathormad @msl23518 @backaggle @godaibo @thomasal @wowfattie @dmytropoplavskiy @meanshift @maciejbudys @nordberdt @tgilewicz @antorsae @macayaven @cristinagrs @bacterio @shimacos @sugawarya @losveria @tkyyym @appian @yuval6967 @zaharch </p>",
      "rawMarkdown": "Hi all! Congrats on all the great results in this competition. So many really interesting approaches. :)\n\nUnfortunately, the learnings from these competitions sometimes don't impact the wider research and practitioner community as much as they should. The topic of this competition is too important to allow that to happen here. So let's make sure it doesn't!\n\nAs some of you know, I chair an academic initiative called [WAMRI](https://wamri.ai/), which aims to advance the state of the art in medical AI, and also to make it more accessible to both researchers and clinicians. We are working with brilliant collaborators at institutions including UCSF, Stanford, Harvard, Salk Institute, and many more. This competition is a great opportunity to help towards our goal, so I'm hoping to organize a project to collect the best approaches that were developed and tested here, see which ones combine well together, and publish both an academic paper describing the results of that research, as well as publishing one or more pretrained models that anyone can use to help kick start their own models. I'm also hoping to publish research results showing the impact of transfer learning best practices in medical imaging, using these pretrained models as the basis of that analysis.\n\nOf course, anyone who created a novel approach during the competition that is used in this research would be credited in the paper. So, if you're interested in contributing to this important next step in turning this competition into real medical outcomes, and you have a strong single model (or think that you might), could you please do the following:\n\n1. Create a submission from your best single model, not including any kind of TTA or other ensembling (e.g. also not ensembling across checkpoints), and do a post-competition submission\n1. Add a reply in this thread stating what private leaderboard result you get with that submission (NB: ignore the public leaderboard result, since that is the 1% subset they used for stage 2; just look at the private leaderboard result, which will be for the stage 2 test set)\n1. In your reply, summarize the key pieces of this model (or link to an existing forum post you've made that describes it, mentioning which is your best single model)\n1. Also, let me know in your reply if you're interested in helping with this post-competition research, either in an informal way (i.e. just answering a few questions as we try to replicate your result), or more formally (i.e. assisting in the followup research and writing).\n\nI've started some initial discussions with some folks involved in organizing the competition, and it looks like there might be quite a bit of interest from all sides in contributing to this. However, it's early days, so the exact final nature and makeup of this collaboration isn't yet well defined at all. So it might take a while for this all to come to fruition - which may require some patience on all sides... I'll be sure to keep this thread updated with news as I have it.\n\ncc all the (provisional) gold medalists: @scp173 @shentao @currylc @lanjunyelan @scusywxy @dmitrylarko @darraghdog @takuok @liut0012 @mfang10 @antherxu @yunpchen @anjum48 @tarobxl @mathormad @msl23518 @backaggle @godaibo @thomasal @wowfattie @dmytropoplavskiy @meanshift @maciejbudys @nordberdt @tgilewicz @antorsae @macayaven @cristinagrs @bacterio @shimacos @sugawarya @losveria @tkyyym @appian @yuval6967 @zaharch ",
      "votes": 28
    },
    {
      "id": 681597,
      "postDate": "2019-11-26T09:53:56.480Z",
      "content": "<p>Our best single model was SE-ResNext50 with 7 input slices and output predictions for single slice with score = 0.05061. We also got a simple, plain Resnet18 with single input slice with a nice score of 0.05174. We used our own non-linear, hemorrhage-oriented window. It was a single window only, hence convenient and simple to use.</p>\n\n<p>Single fold results without TTA for all our models are listed in our <a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/118780\">post</a>.</p>\n\n<p>We think that post-competition research is a great idea and we're interested in active participation.</p>",
      "rawMarkdown": "Our best single model was SE-ResNext50 with 7 input slices and output predictions for single slice with score = 0.05061. We also got a simple, plain Resnet18 with single input slice with a nice score of 0.05174. We used our own non-linear, hemorrhage-oriented window. It was a single window only, hence convenient and simple to use.\n\nSingle fold results without TTA for all our models are listed in our [post](https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/118780).\n\nWe think that post-competition research is a great idea and we're interested in active participation.",
      "votes": 5
    },
    {
      "id": 674184,
      "postDate": "2019-11-16T01:18:42.330Z",
      "content": "<p>My best single model is SE-ResNeXt50, input 3 slices and output 3 slices.\nIt used horizontal flip TTA at now, but the private score was 0.04832 without 2nd stage retraining.\nIt used imagenet pretrained and subdural window.\nI will check score without TTA. I feel multi slices input and output improves 2D CNN score.\nCode is <a href=\"https://github.com/okotaku/kaggle_rsna2019_3rd_solution/blob/master/exp/exp34_seres_threetarget.py\">here</a></p>",
      "rawMarkdown": "My best single model is SE-ResNeXt50, input 3 slices and output 3 slices.\nIt used horizontal flip TTA at now, but the private score was 0.04832 without 2nd stage retraining.\nIt used imagenet pretrained and subdural window.\nI will check score without TTA. I feel multi slices input and output improves 2D CNN score.\nCode is [here](https://github.com/okotaku/kaggle_rsna2019_3rd_solution/blob/master/exp/exp34_seres_threetarget.py)",
      "votes": 5,
      "replies": [
        {
          "id": 674186,
          "postDate": "2019-11-16T01:30:14.070Z",
          "content": "<p>This was a really cool idea. Congrats on 3rd place!</p>",
          "rawMarkdown": "This was a really cool idea. Congrats on 3rd place!",
          "votes": 1
        },
        {
          "id": 674200,
          "postDate": "2019-11-16T01:54:44.017Z",
          "content": "<p>Nice! I did something similar (but too late to submit properly to the comp) - but did a whole CT scan of slices as both input and output. I used 3d convs in the head, and 2d in the body. I'll try to train and submit a single model soon(ish) since it would be interesting to compare.</p>",
          "rawMarkdown": "Nice! I did something similar (but too late to submit properly to the comp) - but did a whole CT scan of slices as both input and output. I used 3d convs in the head, and 2d in the body. I'll try to train and submit a single model soon(ish) since it would be interesting to compare.",
          "votes": 1
        }
      ]
    },
    {
      "id": 674978,
      "postDate": "2019-11-17T11:30:45.377Z",
      "content": "<p>Our best single model, single fold, no tta, no retrain on stage1 test data is 0.05087. That is seresnext101_32x4d using pretrained imagenet weight, 384x384 image using 3 windows (brain, subdural and bony). </p>\n\n<p>Our approach also exploit neighbor slices information, but for each study, we take 10 contiguous slices, extract the features, then put those features into a bilstm decoder before  FC layers.</p>\n\n<p>So the flow basically is:\n<code>\nInput (bs,nslices,channels,heigh,width) --Encoder (resnet)--&amp;gt; (bs,nslices,2048) --Decoder (bilstm+fc)--&amp;gt; (bs,nslices,nclasses)\n</code></p>\n\n<p>The decoder code <a href=\"https://github.com/dattran2346/kaggle-rsna-2019/blob/36da5416cf616bbe6e0b1d4ab6dbc874603f46cc/train.py#L92\">link</a></p>\n\n<p>The training loop <a href=\"https://github.com/dattran2346/kaggle-rsna-2019/blob/36da5416cf616bbe6e0b1d4ab6dbc874603f46cc/train.py#L321\">link</a>, we train decoder and encoder end-to-end.</p>\n\n<p>The best  training setting is (batch size here is the number of study, not image):\n<code>\n+-----------------------+------------------+\n| Setting               | Value            |\n+-----------------------+------------------+\n| Backbone              | resnext101_32x4d |\n| Image size            | 384              |\n| Fold                  | 0                |\n| Epochs                | 30               |\n| Warmup                | 3                |\n| Batch size            | 4                |\n| Gradient accumulation | 16               |\n| Lr                    | 0.004            |\n| Optimizer             | adam             |\n| Scheduler             | cos              |\n| Mix window            | 3                |\n| #slices                | 10               |\n+-----------------------+------------------+\n</code>\nThis take about 1h per epochs on gcp v100.</p>\n\n<p>The pretrained model <a href=\"https://www.kaggle.com/dattran2346/rsnamodels\">link</a></p>\n\n<p>We are happy to contribute.</p>",
      "rawMarkdown": "Our best single model, single fold, no tta, no retrain on stage1 test data is 0.05087. That is seresnext101_32x4d using pretrained imagenet weight, 384x384 image using 3 windows (brain, subdural and bony). \n\nOur approach also exploit neighbor slices information, but for each study, we take 10 contiguous slices, extract the features, then put those features into a bilstm decoder before  FC layers.\n\nSo the flow basically is:\n```\nInput (bs,nslices,channels,heigh,width) --Encoder (resnet)--&gt; (bs,nslices,2048) --Decoder (bilstm+fc)--&gt; (bs,nslices,nclasses)\n```\n\nThe decoder code [link](https://github.com/dattran2346/kaggle-rsna-2019/blob/36da5416cf616bbe6e0b1d4ab6dbc874603f46cc/train.py#L92)\n\nThe training loop [link](https://github.com/dattran2346/kaggle-rsna-2019/blob/36da5416cf616bbe6e0b1d4ab6dbc874603f46cc/train.py#L321), we train decoder and encoder end-to-end.\n\nThe best  training setting is (batch size here is the number of study, not image):\n```\n+-----------------------+------------------+\n| Setting               | Value            |\n+-----------------------+------------------+\n| Backbone              | resnext101_32x4d |\n| Image size            | 384              |\n| Fold                  | 0                |\n| Epochs                | 30               |\n| Warmup                | 3                |\n| Batch size            | 4                |\n| Gradient accumulation | 16               |\n| Lr                    | 0.004            |\n| Optimizer             | adam             |\n| Scheduler             | cos              |\n| Mix window            | 3                |\n| #slices                | 10               |\n+-----------------------+------------------+\n```\nThis take about 1h per epochs on gcp v100.\n\nThe pretrained model [link](https://www.kaggle.com/dattran2346/rsnamodels)\n\nWe are happy to contribute.",
      "votes": 4
    },
    {
      "id": 674526,
      "postDate": "2019-11-16T16:39:16.357Z",
      "content": "<p>Two submissions, both same single fold from 5 folds (so 80% of train stg1 used); one with stg1 data only <code>0.04903</code> and another where the additional stg2 training data was included - <code>0.04795</code>. \nSolution write up is <a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/117228#latest-673431\">here</a>\nModel is broken over two parts, image model and sequential model, for the <code>0.04795</code> result, \n- Image model code <a href=\"https://github.com/darraghdog/rsna/blob/master/scripts/resnext101v14/trainorig.py\">link</a> and params (<code>--epochs 3 --fold 0  --lr 0.00002 --batchsize 64 --size 480</code> ); chose epoch to use based on val results. \n- LSTM code <a href=\"https://github.com/darraghdog/rsna/blob/master/scripts/resnext101v14/trainlstmsngl.py\">link</a> and params (<code>--epochs 5 --fold 0  --lr 0.00001 --batchsize 4  --lrgamma 0.95 --nbags 1 --globalepoch 3  lstm_units 2048</code>) - again chose on best val result.</p>\n\n<p>For inference two parts (classifier and sequential) could be easily combined to one end to end model. Let me know if you'd like me to code it. For training, I'm not sure if this would work well; but could work better than training independently. </p>\n\n<p>A few notes on the solution write up. Much of the solution is the same as at least a few of the other solutions which were posted; I think the not so common parts which helped (excl any TTA/bagging etc.) are,\n- on lstm concat on the deltas between current and previous/next embeddings <a href=\"https://github.com/darraghdog/rsna/blob/1d27f95eb5c12a04e2235f17b4798a430c7a3da5/scripts/resnext101v14/trainlstmsngl.py#L150\">link</a> <br>\n- depth of the network on lstm, and summing the embeddings back on to the LSTM outputs (just the original emb, not the delta)  <a href=\"https://github.com/darraghdog/rsna/blob/1d27f95eb5c12a04e2235f17b4798a430c7a3da5/scripts/resnext101v14/trainlstmsngl.py#L317\">link</a>\n- maybe how black space was cut, however in retrospect I see the minimum kernel mentioned in solution write up was not used in the code ( :palmface: ) so this was used - <a href=\"https://github.com/darraghdog/rsna/blob/1d27f95eb5c12a04e2235f17b4798a430c7a3da5/scripts/resnext101v14/trainorig.py#L139\">link</a></p>\n\n<p>Would be very happy to contribute. </p>",
      "rawMarkdown": "Two submissions, both same single fold from 5 folds (so 80% of train stg1 used); one with stg1 data only `0.04903` and another where the additional stg2 training data was included - `0.04795`. \nSolution write up is [here](https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/117228#latest-673431)\nModel is broken over two parts, image model and sequential model, for the `0.04795` result, \n- Image model code [link](https://github.com/darraghdog/rsna/blob/master/scripts/resnext101v14/trainorig.py) and params (`--epochs 3 --fold 0  --lr 0.00002 --batchsize 64 --size 480` ); chose epoch to use based on val results. \n- LSTM code [link](https://github.com/darraghdog/rsna/blob/master/scripts/resnext101v14/trainlstmsngl.py) and params (`--epochs 5 --fold 0  --lr 0.00001 --batchsize 4  --lrgamma 0.95 --nbags 1 --globalepoch 3  lstm_units 2048`) - again chose on best val result.\n\nFor inference two parts (classifier and sequential) could be easily combined to one end to end model. Let me know if you'd like me to code it. For training, I'm not sure if this would work well; but could work better than training independently. \n\nA few notes on the solution write up. Much of the solution is the same as at least a few of the other solutions which were posted; I think the not so common parts which helped (excl any TTA/bagging etc.) are,\n- on lstm concat on the deltas between current and previous/next embeddings [link](https://github.com/darraghdog/rsna/blob/1d27f95eb5c12a04e2235f17b4798a430c7a3da5/scripts/resnext101v14/trainlstmsngl.py#L150)   \n- depth of the network on lstm, and summing the embeddings back on to the LSTM outputs (just the original emb, not the delta)  [link](https://github.com/darraghdog/rsna/blob/1d27f95eb5c12a04e2235f17b4798a430c7a3da5/scripts/resnext101v14/trainlstmsngl.py#L317)\n- maybe how black space was cut, however in retrospect I see the minimum kernel mentioned in solution write up was not used in the code ( :palmface: ) so this was used - [link](https://github.com/darraghdog/rsna/blob/1d27f95eb5c12a04e2235f17b4798a430c7a3da5/scripts/resnext101v14/trainorig.py#L139)\n\nWould be very happy to contribute. ",
      "votes": 1
    },
    {
      "id": 675032,
      "postDate": "2019-11-17T13:27:23.207Z",
      "content": "<p>Thanks <a href=\"/jhoward\">@jhoward</a>, that's a great idea. We would really love to participate.\nWe had an ensemble of two types of models:\nOne type is a simple resnet101 with three windowing values for the channels and resolution of 384. Its score on stage 2 private LB is 0.06262. We then use another 1d CNN over slices belonging to the same series to further post process our predictions. This lowered the LB score to 0.05331.</p>\n\n<p>Second type is a se_resnext50_32x4d, again three channels for different windowing values. Here we predicted on three adjacent slices from the same series which were concatenated together and pass through a dense layer to predict the labels of the middle slice. For this model the LB score is 0.05481. Here we used the same post-processing as before to get LB score of 0.05367.</p>",
      "rawMarkdown": "Thanks @jhoward, that's a great idea. We would really love to participate.\nWe had an ensemble of two types of models:\nOne type is a simple resnet101 with three windowing values for the channels and resolution of 384. Its score on stage 2 private LB is 0.06262. We then use another 1d CNN over slices belonging to the same series to further post process our predictions. This lowered the LB score to 0.05331.\n\nSecond type is a se_resnext50_32x4d, again three channels for different windowing values. Here we predicted on three adjacent slices from the same series which were concatenated together and pass through a dense layer to predict the labels of the middle slice. For this model the LB score is 0.05481. Here we used the same post-processing as before to get LB score of 0.05367.",
      "votes": 2
    },
    {
      "id": 674272,
      "postDate": "2019-11-16T06:19:12.623Z",
      "content": "<p>Hi Jeremy, I'm interested in this and have a few questions regarding <code>ensembling</code>. </p>\n\n<ol>\n<li><p>Does this only apply to the test part? <a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/117330\">My solution</a> involves second level model which requires out-of-fold predictions from 8-folds resnext as input for training. But the testing part can be easily adapted to a single model.</p></li>\n<li><p>Does this only apply to CNN? I use LightGBM, Catboost and XGB for second level model and average their predictions. It can be easily adapted to use one of them if it's not adequate but would like to make sure which is the case.</p></li>\n</ol>",
      "rawMarkdown": "Hi Jeremy, I'm interested in this and have a few questions regarding `ensembling`. \n\n1. Does this only apply to the test part? [My solution](https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/117330) involves second level model which requires out-of-fold predictions from 8-folds resnext as input for training. But the testing part can be easily adapted to a single model.\n\n2. Does this only apply to CNN? I use LightGBM, Catboost and XGB for second level model and average their predictions. It can be easily adapted to use one of them if it's not adequate but would like to make sure which is the case.",
      "votes": 2,
      "replies": [
        {
          "id": 674274,
          "postDate": "2019-11-16T06:32:47.747Z",
          "content": "<p>Yeah your solution is a bit tricky to fit into my simplified request! In your case, it would be most helpful to see the result where you used just one checkpoint, just one fold, and just one slice model.</p>\n\n<p>Basically, I'm looking for something that can be fairly readily picked up and used (and fine-tuned) in a new setting. And I'm also looking to compare as directly as possible the various ways that different parts of the problem can be solved. One can always make things a bit better by doing more folds, more TTA, etc - but it makes assessing the contribution of each piece harder.</p>\n\n<p>I hope this answers your question.</p>\n\n<p>PS: Thanks a lot for the really helpful code release - it was great for saving time in finding good hyper-parameters for fast training. It's amazing you found params that worked so well in just 3 epochs, and without even using weighted sampling!</p>",
          "rawMarkdown": "Yeah your solution is a bit tricky to fit into my simplified request! In your case, it would be most helpful to see the result where you used just one checkpoint, just one fold, and just one slice model.\n\nBasically, I'm looking for something that can be fairly readily picked up and used (and fine-tuned) in a new setting. And I'm also looking to compare as directly as possible the various ways that different parts of the problem can be solved. One can always make things a bit better by doing more folds, more TTA, etc - but it makes assessing the contribution of each piece harder.\n\nI hope this answers your question.\n\nPS: Thanks a lot for the really helpful code release - it was great for saving time in finding good hyper-parameters for fast training. It's amazing you found params that worked so well in just 3 epochs, and without even using weighted sampling!",
          "votes": 2
        },
        {
          "id": 674295,
          "postDate": "2019-11-16T07:27:14.707Z",
          "content": "<p>Thank you. I'm glad to hear it was helpful.\nYes, it seems tricky to fit into the request. Thanks for the answer. </p>",
          "rawMarkdown": "Thank you. I'm glad to hear it was helpful.\nYes, it seems tricky to fit into the request. Thanks for the answer. "
        }
      ]
    },
    {
      "id": 674168,
      "postDate": "2019-11-16T00:35:23.407Z",
      "content": "<p>My best single fold, single model is just an EfficientNet-B5 CNN that takes 5-channel input (slice-1, slice x3, slice+1) to generate a single slice prediction. I also use a trainable windowing module and smooth the probabilities for the series by using a weighted average of the probabilities of a given slice and its adjacent slices. It scores 0.05389. </p>\n\n<p>I would be interested in helping out. At my institution, we have already implemented a very early version of the model and have it running in the background on head CTs as they come in. </p>",
      "rawMarkdown": "My best single fold, single model is just an EfficientNet-B5 CNN that takes 5-channel input (slice-1, slice x3, slice+1) to generate a single slice prediction. I also use a trainable windowing module and smooth the probabilities for the series by using a weighted average of the probabilities of a given slice and its adjacent slices. It scores 0.05389. \n\nI would be interested in helping out. At my institution, we have already implemented a very early version of the model and have it running in the background on head CTs as they come in. ",
      "votes": 2,
      "replies": [
        {
          "id": 674171,
          "postDate": "2019-11-16T00:42:30.650Z",
          "content": "<p>Thanks <a href=\"/vaillant\">@vaillant</a> that's great to know. <a href=\"/alexandrecc\">@alexandrecc</a> has often sung your praises, so I'm particularly looking forward to working with you on this!</p>\n\n<p>I assume slice x3 is normal windowing (brain, subdural, soft tissue)? And was slice -1/+1 just brain windowing? Did you start with a pretrained imagenet model? If so, how did you add the extra channels?</p>\n\n<p>May I ask - what hyperparams (lr, optimizer, # epochs, etc) did you use for training your EfficientNet? Did you get good results for B2 as well, or did you find you needed the bigger models? Did you find it rather slow and memory intensive (and if not, what implementation did you use and what hardware)?</p>\n\n<p>(I'd love to hear more about the early implementation you have, but I'll contact you privately about that.)</p>",
          "rawMarkdown": "Thanks @vaillant that's great to know. @alexandrecc has often sung your praises, so I'm particularly looking forward to working with you on this!\n\nI assume slice x3 is normal windowing (brain, subdural, soft tissue)? And was slice -1/+1 just brain windowing? Did you start with a pretrained imagenet model? If so, how did you add the extra channels?\n\nMay I ask - what hyperparams (lr, optimizer, # epochs, etc) did you use for training your EfficientNet? Did you get good results for B2 as well, or did you find you needed the bigger models? Did you find it rather slow and memory intensive (and if not, what implementation did you use and what hardware)?\n\n(I'd love to hear more about the early implementation you have, but I'll contact you privately about that.)"
        },
        {
          "id": 674182,
          "postDate": "2019-11-16T01:12:54.400Z",
          "content": "<p>Likewise! </p>\n\n<p>Calling it a 5-channel input is a little misleading. The input starts as 5 slices: the target slice times 3, the slice below, and the slice above (padded with zeros if it is at the top/bottom). Each of these slices are in their raw Hounsfield units and is passed through a trainable windowing module (initialized with 3 windows: brain, subdural, blood) which converts each slice into a 3-channel, 8-bit image. Then, I concatenate them all channel wise. So each input now has shape (15, H, W). </p>\n\n<p>This input is passed into an EfficientNet-B5 ImageNet-pretrained CNN. I modified the input layer of the CNN to accept 15-channel input vs. 3-channel. I used the RAdam optimizer, initial learning rate of 1e-4, final learning rate of 1e-6, with cosine annealing and warm restarts for 120 epochs, 6 cycles. For the single fold, single model, I chose the snapshot with the lowest validation loss (epoch 98/cycle 5). I tried smaller EfficientNets, but they always had higher losses (though I imagine in terms of AUC the difference would be negligible). I used PyTorch and this implementation of EfficientNets (<a href=\"https://github.com/lukemelas/EfficientNet-PyTorch\">https://github.com/lukemelas/EfficientNet-PyTorch</a>). I trained at 512 x 512 resolution with a batch size of 16 on a NVIDIA V100 32GB. It took about 1 week to finish training this model. I think the bulk of the time though was related to augmentation and I/O, since the I/O on the computer I was using is pretty slow, and loading 1 batch requires loading and augmenting 48 images. Trained with data augmentation (probability 0.8) which was done on the raw HU array using flips/shift-scale-rotate/blur/uniformly adding a random number of Hounsfield units. </p>",
          "rawMarkdown": "Likewise! \n\nCalling it a 5-channel input is a little misleading. The input starts as 5 slices: the target slice times 3, the slice below, and the slice above (padded with zeros if it is at the top/bottom). Each of these slices are in their raw Hounsfield units and is passed through a trainable windowing module (initialized with 3 windows: brain, subdural, blood) which converts each slice into a 3-channel, 8-bit image. Then, I concatenate them all channel wise. So each input now has shape (15, H, W). \n\nThis input is passed into an EfficientNet-B5 ImageNet-pretrained CNN. I modified the input layer of the CNN to accept 15-channel input vs. 3-channel. I used the RAdam optimizer, initial learning rate of 1e-4, final learning rate of 1e-6, with cosine annealing and warm restarts for 120 epochs, 6 cycles. For the single fold, single model, I chose the snapshot with the lowest validation loss (epoch 98/cycle 5). I tried smaller EfficientNets, but they always had higher losses (though I imagine in terms of AUC the difference would be negligible). I used PyTorch and this implementation of EfficientNets (https://github.com/lukemelas/EfficientNet-PyTorch). I trained at 512 x 512 resolution with a batch size of 16 on a NVIDIA V100 32GB. It took about 1 week to finish training this model. I think the bulk of the time though was related to augmentation and I/O, since the I/O on the computer I was using is pretty slow, and loading 1 batch requires loading and augmenting 48 images. Trained with data augmentation (probability 0.8) which was done on the raw HU array using flips/shift-scale-rotate/blur/uniformly adding a random number of Hounsfield units. ",
          "votes": 2
        },
        {
          "id": 675506,
          "postDate": "2019-11-18T06:27:13.240Z",
          "content": "<p>Hi, <a href=\"/vaillant\">@vaillant</a>, I also use this method(WSO). But is there any reason you to use <strong>the target slice times 3</strong> (I use 3 channel, below, target, above).</p>\n\n<p>P.S. <a href=\"/jhoward\">@jhoward</a>, is there any chance to help this project if my method is almost 90% similar to <a href=\"/vaillant\">@vaillant</a>'s</p>",
          "rawMarkdown": "Hi, @vaillant, I also use this method(WSO). But is there any reason you to use **the target slice times 3** (I use 3 channel, below, target, above).\n\nP.S. @jhoward, is there any chance to help this project if my method is almost 90% similar to @vaillant's"
        }
      ]
    },
    {
      "id": 674123,
      "postDate": "2019-11-15T23:00:25.110Z",
      "content": "<p>Exciting, great idea!</p>",
      "rawMarkdown": "Exciting, great idea!",
      "votes": 2
    },
    {
      "id": 682677,
      "postDate": "2019-11-27T18:34:30.847Z",
      "content": "<p><a href=\"/jhoward\">@jhoward</a>/all I put together an excel here with a rough summary of the models/approaches/posts/code from folks who posted here to get a quick overview of the similarities and differences of each approach. </p>\n\n<p><a href=\"https://1drv.ms/x/s!AmMpcG2zZmnxx1ILR0InSX_Tf0ok\">https://1drv.ms/x/s!AmMpcG2zZmnxx1ILR0InSX_Tf0ok</a></p>\n\n<p>Happy for any feedback, will hope to replicate one of the single model approaches in fastai2 for my own learning too :)  </p>",
      "rawMarkdown": "@jhoward/all I put together an excel here with a rough summary of the models/approaches/posts/code from folks who posted here to get a quick overview of the similarities and differences of each approach. \n\nhttps://1drv.ms/x/s!AmMpcG2zZmnxx1ILR0InSX_Tf0ok\n\nHappy for any feedback, will hope to replicate one of the single model approaches in fastai2 for my own learning too :)  "
    },
    {
      "id": 724811,
      "postDate": "2020-01-21T14:02:03.553Z",
      "content": "<p>Oh wow, I totally missed this. What a great idea. Apologies if this is too late, but my best single model was EfficientNet-B5 which scored 0.04830 on private LB.</p>\n\n<p>The preprocessing was done as per this kernel: <a href=\"https://www.kaggle.com/anjum48/5th-preprocessing-adjacent-images-and-cropping\">https://www.kaggle.com/anjum48/5th-preprocessing-adjacent-images-and-cropping</a></p>\n\n<p>A video of the full solution which contains other single model results can be found here: <a href=\"https://www.youtube.com/watch?v=1zLBxwTAcAs\">https://www.youtube.com/watch?v=1zLBxwTAcAs</a></p>",
      "rawMarkdown": "Oh wow, I totally missed this. What a great idea. Apologies if this is too late, but my best single model was EfficientNet-B5 which scored 0.04830 on private LB.\n\nThe preprocessing was done as per this kernel: https://www.kaggle.com/anjum48/5th-preprocessing-adjacent-images-and-cropping\n\nA video of the full solution which contains other single model results can be found here: https://www.youtube.com/watch?v=1zLBxwTAcAs",
      "replies": [
        {
          "id": 725072,
          "postDate": "2020-01-21T19:01:44.997Z",
          "content": "<p>Thanks! No not too late at all - we're still working on this; it's a lot to get through! We'll check out your kernel and video.</p>",
          "rawMarkdown": "Thanks! No not too late at all - we're still working on this; it's a lot to get through! We'll check out your kernel and video.",
          "votes": 1
        },
        {
          "id": 725081,
          "postDate": "2020-01-21T19:18:21.327Z",
          "content": "<p>I just remembered that score is with TTA. I will try and get a submission done without TTA for comparison</p>",
          "rawMarkdown": "I just remembered that score is with TTA. I will try and get a submission done without TTA for comparison"
        },
        {
          "id": 725175,
          "postDate": "2020-01-21T21:31:59.743Z",
          "content": "<p>Thanks - that would be very helpful.</p>",
          "rawMarkdown": "Thanks - that would be very helpful."
        },
        {
          "id": 728083,
          "postDate": "2020-01-24T11:36:46.967Z",
          "content": "<p>I just re-ran my model and a single fold private LB score without TTA is 0.05553. This is using the EfficientNet-B5 model and the adjacent slice method (subdural window). Here's a scruffy looking ROC-AUC plot (class 0 is \"any\")</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F421965%2Faae646eb9d28994a578014838583e968%2Froc_fold_1.png?generation=1579865677136018&amp;alt=media\" alt=\"\"></p>\n\n<p>If you need any other help, please let me know. I'm also happy to share model weights etc.</p>",
          "rawMarkdown": "I just re-ran my model and a single fold private LB score without TTA is 0.05553. This is using the EfficientNet-B5 model and the adjacent slice method (subdural window). Here's a scruffy looking ROC-AUC plot (class 0 is \"any\")\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F421965%2Faae646eb9d28994a578014838583e968%2Froc_fold_1.png?generation=1579865677136018&amp;alt=media)\n\nIf you need any other help, please let me know. I'm also happy to share model weights etc.\n"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 681597,
      "author_name": "Adam Brzeski",
      "author_url": "",
      "post_date": "2019-11-26T09:53:56.480000",
      "content": "<p>Our best single model was SE-ResNext50 with 7 input slices and output predictions for single slice with score = 0.05061. We also got a simple, plain Resnet18 with single input slice with a nice score of 0.05174. We used our own non-linear, hemorrhage-oriented window. It was a single window only, hence convenient and simple to use.</p>\n\n<p>Single fold results without TTA for all our models are listed in our <a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/118780\">post</a>.</p>\n\n<p>We think that post-competition research is a great idea and we're interested in active participation.</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 674184,
      "author_name": "takuoko",
      "author_url": "",
      "post_date": "2019-11-16T01:18:42.330000",
      "content": "<p>My best single model is SE-ResNeXt50, input 3 slices and output 3 slices.\nIt used horizontal flip TTA at now, but the private score was 0.04832 without 2nd stage retraining.\nIt used imagenet pretrained and subdural window.\nI will check score without TTA. I feel multi slices input and output improves 2D CNN score.\nCode is <a href=\"https://github.com/okotaku/kaggle_rsna2019_3rd_solution/blob/master/exp/exp34_seres_threetarget.py\">here</a></p>",
      "votes": 5,
      "replies": [
        {
          "id": 674186,
          "author_name": "Ian Pan",
          "author_url": "",
          "post_date": "2019-11-16T01:30:14.070000",
          "content": "<p>This was a really cool idea. Congrats on 3rd place!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 674200,
          "author_name": "Jeremy Howard",
          "author_url": "",
          "post_date": "2019-11-16T01:54:44.017000",
          "content": "<p>Nice! I did something similar (but too late to submit properly to the comp) - but did a whole CT scan of slices as both input and output. I used 3d convs in the head, and 2d in the body. I'll try to train and submit a single model soon(ish) since it would be interesting to compare.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 674978,
      "author_name": "nan",
      "author_url": "",
      "post_date": "2019-11-17T11:30:45.377000",
      "content": "<p>Our best single model, single fold, no tta, no retrain on stage1 test data is 0.05087. That is seresnext101_32x4d using pretrained imagenet weight, 384x384 image using 3 windows (brain, subdural and bony). </p>\n\n<p>Our approach also exploit neighbor slices information, but for each study, we take 10 contiguous slices, extract the features, then put those features into a bilstm decoder before  FC layers.</p>\n\n<p>So the flow basically is:\n<code>\nInput (bs,nslices,channels,heigh,width) --Encoder (resnet)--&amp;gt; (bs,nslices,2048) --Decoder (bilstm+fc)--&amp;gt; (bs,nslices,nclasses)\n</code></p>\n\n<p>The decoder code <a href=\"https://github.com/dattran2346/kaggle-rsna-2019/blob/36da5416cf616bbe6e0b1d4ab6dbc874603f46cc/train.py#L92\">link</a></p>\n\n<p>The training loop <a href=\"https://github.com/dattran2346/kaggle-rsna-2019/blob/36da5416cf616bbe6e0b1d4ab6dbc874603f46cc/train.py#L321\">link</a>, we train decoder and encoder end-to-end.</p>\n\n<p>The best  training setting is (batch size here is the number of study, not image):\n<code>\n+-----------------------+------------------+\n| Setting               | Value            |\n+-----------------------+------------------+\n| Backbone              | resnext101_32x4d |\n| Image size            | 384              |\n| Fold                  | 0                |\n| Epochs                | 30               |\n| Warmup                | 3                |\n| Batch size            | 4                |\n| Gradient accumulation | 16               |\n| Lr                    | 0.004            |\n| Optimizer             | adam             |\n| Scheduler             | cos              |\n| Mix window            | 3                |\n| #slices                | 10               |\n+-----------------------+------------------+\n</code>\nThis take about 1h per epochs on gcp v100.</p>\n\n<p>The pretrained model <a href=\"https://www.kaggle.com/dattran2346/rsnamodels\">link</a></p>\n\n<p>We are happy to contribute.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 674526,
      "author_name": "Darragh",
      "author_url": "",
      "post_date": "2019-11-16T16:39:16.357000",
      "content": "<p>Two submissions, both same single fold from 5 folds (so 80% of train stg1 used); one with stg1 data only <code>0.04903</code> and another where the additional stg2 training data was included - <code>0.04795</code>. \nSolution write up is <a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/117228#latest-673431\">here</a>\nModel is broken over two parts, image model and sequential model, for the <code>0.04795</code> result, \n- Image model code <a href=\"https://github.com/darraghdog/rsna/blob/master/scripts/resnext101v14/trainorig.py\">link</a> and params (<code>--epochs 3 --fold 0  --lr 0.00002 --batchsize 64 --size 480</code> ); chose epoch to use based on val results. \n- LSTM code <a href=\"https://github.com/darraghdog/rsna/blob/master/scripts/resnext101v14/trainlstmsngl.py\">link</a> and params (<code>--epochs 5 --fold 0  --lr 0.00001 --batchsize 4  --lrgamma 0.95 --nbags 1 --globalepoch 3  lstm_units 2048</code>) - again chose on best val result.</p>\n\n<p>For inference two parts (classifier and sequential) could be easily combined to one end to end model. Let me know if you'd like me to code it. For training, I'm not sure if this would work well; but could work better than training independently. </p>\n\n<p>A few notes on the solution write up. Much of the solution is the same as at least a few of the other solutions which were posted; I think the not so common parts which helped (excl any TTA/bagging etc.) are,\n- on lstm concat on the deltas between current and previous/next embeddings <a href=\"https://github.com/darraghdog/rsna/blob/1d27f95eb5c12a04e2235f17b4798a430c7a3da5/scripts/resnext101v14/trainlstmsngl.py#L150\">link</a> <br>\n- depth of the network on lstm, and summing the embeddings back on to the LSTM outputs (just the original emb, not the delta)  <a href=\"https://github.com/darraghdog/rsna/blob/1d27f95eb5c12a04e2235f17b4798a430c7a3da5/scripts/resnext101v14/trainlstmsngl.py#L317\">link</a>\n- maybe how black space was cut, however in retrospect I see the minimum kernel mentioned in solution write up was not used in the code ( :palmface: ) so this was used - <a href=\"https://github.com/darraghdog/rsna/blob/1d27f95eb5c12a04e2235f17b4798a430c7a3da5/scripts/resnext101v14/trainorig.py#L139\">link</a></p>\n\n<p>Would be very happy to contribute. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 675032,
      "author_name": "Ohad Silbert",
      "author_url": "",
      "post_date": "2019-11-17T13:27:23.207000",
      "content": "<p>Thanks <a href=\"/jhoward\">@jhoward</a>, that's a great idea. We would really love to participate.\nWe had an ensemble of two types of models:\nOne type is a simple resnet101 with three windowing values for the channels and resolution of 384. Its score on stage 2 private LB is 0.06262. We then use another 1d CNN over slices belonging to the same series to further post process our predictions. This lowered the LB score to 0.05331.</p>\n\n<p>Second type is a se_resnext50_32x4d, again three channels for different windowing values. Here we predicted on three adjacent slices from the same series which were concatenated together and pass through a dense layer to predict the labels of the middle slice. For this model the LB score is 0.05481. Here we used the same post-processing as before to get LB score of 0.05367.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 674272,
      "author_name": "Appian",
      "author_url": "",
      "post_date": "2019-11-16T06:19:12.623000",
      "content": "<p>Hi Jeremy, I'm interested in this and have a few questions regarding <code>ensembling</code>. </p>\n\n<ol>\n<li><p>Does this only apply to the test part? <a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/117330\">My solution</a> involves second level model which requires out-of-fold predictions from 8-folds resnext as input for training. But the testing part can be easily adapted to a single model.</p></li>\n<li><p>Does this only apply to CNN? I use LightGBM, Catboost and XGB for second level model and average their predictions. It can be easily adapted to use one of them if it's not adequate but would like to make sure which is the case.</p></li>\n</ol>",
      "votes": 2,
      "replies": [
        {
          "id": 674274,
          "author_name": "Jeremy Howard",
          "author_url": "",
          "post_date": "2019-11-16T06:32:47.747000",
          "content": "<p>Yeah your solution is a bit tricky to fit into my simplified request! In your case, it would be most helpful to see the result where you used just one checkpoint, just one fold, and just one slice model.</p>\n\n<p>Basically, I'm looking for something that can be fairly readily picked up and used (and fine-tuned) in a new setting. And I'm also looking to compare as directly as possible the various ways that different parts of the problem can be solved. One can always make things a bit better by doing more folds, more TTA, etc - but it makes assessing the contribution of each piece harder.</p>\n\n<p>I hope this answers your question.</p>\n\n<p>PS: Thanks a lot for the really helpful code release - it was great for saving time in finding good hyper-parameters for fast training. It's amazing you found params that worked so well in just 3 epochs, and without even using weighted sampling!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 674295,
          "author_name": "Appian",
          "author_url": "",
          "post_date": "2019-11-16T07:27:14.707000",
          "content": "<p>Thank you. I'm glad to hear it was helpful.\nYes, it seems tricky to fit into the request. Thanks for the answer. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 674168,
      "author_name": "Ian Pan",
      "author_url": "",
      "post_date": "2019-11-16T00:35:23.407000",
      "content": "<p>My best single fold, single model is just an EfficientNet-B5 CNN that takes 5-channel input (slice-1, slice x3, slice+1) to generate a single slice prediction. I also use a trainable windowing module and smooth the probabilities for the series by using a weighted average of the probabilities of a given slice and its adjacent slices. It scores 0.05389. </p>\n\n<p>I would be interested in helping out. At my institution, we have already implemented a very early version of the model and have it running in the background on head CTs as they come in. </p>",
      "votes": 2,
      "replies": [
        {
          "id": 674171,
          "author_name": "Jeremy Howard",
          "author_url": "",
          "post_date": "2019-11-16T00:42:30.650000",
          "content": "<p>Thanks <a href=\"/vaillant\">@vaillant</a> that's great to know. <a href=\"/alexandrecc\">@alexandrecc</a> has often sung your praises, so I'm particularly looking forward to working with you on this!</p>\n\n<p>I assume slice x3 is normal windowing (brain, subdural, soft tissue)? And was slice -1/+1 just brain windowing? Did you start with a pretrained imagenet model? If so, how did you add the extra channels?</p>\n\n<p>May I ask - what hyperparams (lr, optimizer, # epochs, etc) did you use for training your EfficientNet? Did you get good results for B2 as well, or did you find you needed the bigger models? Did you find it rather slow and memory intensive (and if not, what implementation did you use and what hardware)?</p>\n\n<p>(I'd love to hear more about the early implementation you have, but I'll contact you privately about that.)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 674182,
          "author_name": "Ian Pan",
          "author_url": "",
          "post_date": "2019-11-16T01:12:54.400000",
          "content": "<p>Likewise! </p>\n\n<p>Calling it a 5-channel input is a little misleading. The input starts as 5 slices: the target slice times 3, the slice below, and the slice above (padded with zeros if it is at the top/bottom). Each of these slices are in their raw Hounsfield units and is passed through a trainable windowing module (initialized with 3 windows: brain, subdural, blood) which converts each slice into a 3-channel, 8-bit image. Then, I concatenate them all channel wise. So each input now has shape (15, H, W). </p>\n\n<p>This input is passed into an EfficientNet-B5 ImageNet-pretrained CNN. I modified the input layer of the CNN to accept 15-channel input vs. 3-channel. I used the RAdam optimizer, initial learning rate of 1e-4, final learning rate of 1e-6, with cosine annealing and warm restarts for 120 epochs, 6 cycles. For the single fold, single model, I chose the snapshot with the lowest validation loss (epoch 98/cycle 5). I tried smaller EfficientNets, but they always had higher losses (though I imagine in terms of AUC the difference would be negligible). I used PyTorch and this implementation of EfficientNets (<a href=\"https://github.com/lukemelas/EfficientNet-PyTorch\">https://github.com/lukemelas/EfficientNet-PyTorch</a>). I trained at 512 x 512 resolution with a batch size of 16 on a NVIDIA V100 32GB. It took about 1 week to finish training this model. I think the bulk of the time though was related to augmentation and I/O, since the I/O on the computer I was using is pretty slow, and loading 1 batch requires loading and augmenting 48 images. Trained with data augmentation (probability 0.8) which was done on the raw HU array using flips/shift-scale-rotate/blur/uniformly adding a random number of Hounsfield units. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 675506,
          "author_name": "Kirayue",
          "author_url": "",
          "post_date": "2019-11-18T06:27:13.240000",
          "content": "<p>Hi, <a href=\"/vaillant\">@vaillant</a>, I also use this method(WSO). But is there any reason you to use <strong>the target slice times 3</strong> (I use 3 channel, below, target, above).</p>\n\n<p>P.S. <a href=\"/jhoward\">@jhoward</a>, is there any chance to help this project if my method is almost 90% similar to <a href=\"/vaillant\">@vaillant</a>'s</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 674123,
      "author_name": "Darragh",
      "author_url": "",
      "post_date": "2019-11-15T23:00:25.110000",
      "content": "<p>Exciting, great idea!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 682677,
      "author_name": "morg",
      "author_url": "",
      "post_date": "2019-11-27T18:34:30.847000",
      "content": "<p><a href=\"/jhoward\">@jhoward</a>/all I put together an excel here with a rough summary of the models/approaches/posts/code from folks who posted here to get a quick overview of the similarities and differences of each approach. </p>\n\n<p><a href=\"https://1drv.ms/x/s!AmMpcG2zZmnxx1ILR0InSX_Tf0ok\">https://1drv.ms/x/s!AmMpcG2zZmnxx1ILR0InSX_Tf0ok</a></p>\n\n<p>Happy for any feedback, will hope to replicate one of the single model approaches in fastai2 for my own learning too :)  </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 724811,
      "author_name": "datasaurus",
      "author_url": "",
      "post_date": "2020-01-21T14:02:03.553000",
      "content": "<p>Oh wow, I totally missed this. What a great idea. Apologies if this is too late, but my best single model was EfficientNet-B5 which scored 0.04830 on private LB.</p>\n\n<p>The preprocessing was done as per this kernel: <a href=\"https://www.kaggle.com/anjum48/5th-preprocessing-adjacent-images-and-cropping\">https://www.kaggle.com/anjum48/5th-preprocessing-adjacent-images-and-cropping</a></p>\n\n<p>A video of the full solution which contains other single model results can be found here: <a href=\"https://www.youtube.com/watch?v=1zLBxwTAcAs\">https://www.youtube.com/watch?v=1zLBxwTAcAs</a></p>",
      "votes": 0,
      "replies": [
        {
          "id": 725072,
          "author_name": "Jeremy Howard",
          "author_url": "",
          "post_date": "2020-01-21T19:01:44.997000",
          "content": "<p>Thanks! No not too late at all - we're still working on this; it's a lot to get through! We'll check out your kernel and video.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 725081,
          "author_name": "datasaurus",
          "author_url": "",
          "post_date": "2020-01-21T19:18:21.327000",
          "content": "<p>I just remembered that score is with TTA. I will try and get a submission done without TTA for comparison</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 725175,
          "author_name": "Jeremy Howard",
          "author_url": "",
          "post_date": "2020-01-21T21:31:59.743000",
          "content": "<p>Thanks - that would be very helpful.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 728083,
          "author_name": "datasaurus",
          "author_url": "",
          "post_date": "2020-01-24T11:36:46.967000",
          "content": "<p>I just re-ran my model and a single fold private LB score without TTA is 0.05553. This is using the EfficientNet-B5 model and the adjacent slice method (subdural window). Here's a scruffy looking ROC-AUC plot (class 0 is \"any\")</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F421965%2Faae646eb9d28994a578014838583e968%2Froc_fold_1.png?generation=1579865677136018&amp;alt=media\" alt=\"\"></p>\n\n<p>If you need any other help, please let me know. I'm also happy to share model weights etc.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "674104": "Hi all! Congrats on all the great results in this competition. So many really interesting approaches. :)\n\nUnfortunately, the learnings from these competitions sometimes don't impact the wider research and practitioner community as much as they should. The topic of this competition is too important to allow that to happen here. So let's make sure it doesn't!\n\nAs some of you know, I chair an academic initiative called [WAMRI](https://wamri.ai/), which aims to advance the state of the art in medical AI, and also to make it more accessible to both researchers and clinicians. We are working with brilliant collaborators at institutions including UCSF, Stanford, Harvard, Salk Institute, and many more. This competition is a great opportunity to help towards our goal, so I'm hoping to organize a project to collect the best approaches that were developed and tested here, see which ones combine well together, and publish both an academic paper describing the results of that research, as well as publishing one or more pretrained models that anyone can use to help kick start their own models. I'm also hoping to publish research results showing the impact of transfer learning best practices in medical imaging, using these pretrained models as the basis of that analysis.\n\nOf course, anyone who created a novel approach during the competition that is used in this research would be credited in the paper. So, if you're interested in contributing to this important next step in turning this competition into real medical outcomes, and you have a strong single model (or think that you might), could you please do the following:\n\n1. Create a submission from your best single model, not including any kind of TTA or other ensembling (e.g. also not ensembling across checkpoints), and do a post-competition submission\n1. Add a reply in this thread stating what private leaderboard result you get with that submission (NB: ignore the public leaderboard result, since that is the 1% subset they used for stage 2; just look at the private leaderboard result, which will be for the stage 2 test set)\n1. In your reply, summarize the key pieces of this model (or link to an existing forum post you've made that describes it, mentioning which is your best single model)\n1. Also, let me know in your reply if you're interested in helping with this post-competition research, either in an informal way (i.e. just answering a few questions as we try to replicate your result), or more formally (i.e. assisting in the followup research and writing).\n\nI've started some initial discussions with some folks involved in organizing the competition, and it looks like there might be quite a bit of interest from all sides in contributing to this. However, it's early days, so the exact final nature and makeup of this collaboration isn't yet well defined at all. So it might take a while for this all to come to fruition - which may require some patience on all sides... I'll be sure to keep this thread updated with news as I have it.\n\ncc all the (provisional) gold medalists: @scp173 @shentao @currylc @lanjunyelan @scusywxy @dmitrylarko @darraghdog @takuok @liut0012 @mfang10 @antherxu @yunpchen @anjum48 @tarobxl @mathormad @msl23518 @backaggle @godaibo @thomasal @wowfattie @dmytropoplavskiy @meanshift @maciejbudys @nordberdt @tgilewicz @antorsae @macayaven @cristinagrs @bacterio @shimacos @sugawarya @losveria @tkyyym @appian @yuval6967 @zaharch ",
    "681597": "Our best single model was SE-ResNext50 with 7 input slices and output predictions for single slice with score = 0.05061. We also got a simple, plain Resnet18 with single input slice with a nice score of 0.05174. We used our own non-linear, hemorrhage-oriented window. It was a single window only, hence convenient and simple to use.\n\nSingle fold results without TTA for all our models are listed in our [post](https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/118780).\n\nWe think that post-competition research is a great idea and we're interested in active participation.",
    "674184": "My best single model is SE-ResNeXt50, input 3 slices and output 3 slices.\nIt used horizontal flip TTA at now, but the private score was 0.04832 without 2nd stage retraining.\nIt used imagenet pretrained and subdural window.\nI will check score without TTA. I feel multi slices input and output improves 2D CNN score.\nCode is [here](https://github.com/okotaku/kaggle_rsna2019_3rd_solution/blob/master/exp/exp34_seres_threetarget.py)",
    "674978": "Our best single model, single fold, no tta, no retrain on stage1 test data is 0.05087. That is seresnext101_32x4d using pretrained imagenet weight, 384x384 image using 3 windows (brain, subdural and bony). \n\nOur approach also exploit neighbor slices information, but for each study, we take 10 contiguous slices, extract the features, then put those features into a bilstm decoder before  FC layers.\n\nSo the flow basically is:\n```\nInput (bs,nslices,channels,heigh,width) --Encoder (resnet)--&gt; (bs,nslices,2048) --Decoder (bilstm+fc)--&gt; (bs,nslices,nclasses)\n```\n\nThe decoder code [link](https://github.com/dattran2346/kaggle-rsna-2019/blob/36da5416cf616bbe6e0b1d4ab6dbc874603f46cc/train.py#L92)\n\nThe training loop [link](https://github.com/dattran2346/kaggle-rsna-2019/blob/36da5416cf616bbe6e0b1d4ab6dbc874603f46cc/train.py#L321), we train decoder and encoder end-to-end.\n\nThe best  training setting is (batch size here is the number of study, not image):\n```\n+-----------------------+------------------+\n| Setting               | Value            |\n+-----------------------+------------------+\n| Backbone              | resnext101_32x4d |\n| Image size            | 384              |\n| Fold                  | 0                |\n| Epochs                | 30               |\n| Warmup                | 3                |\n| Batch size            | 4                |\n| Gradient accumulation | 16               |\n| Lr                    | 0.004            |\n| Optimizer             | adam             |\n| Scheduler             | cos              |\n| Mix window            | 3                |\n| #slices                | 10               |\n+-----------------------+------------------+\n```\nThis take about 1h per epochs on gcp v100.\n\nThe pretrained model [link](https://www.kaggle.com/dattran2346/rsnamodels)\n\nWe are happy to contribute.",
    "674526": "Two submissions, both same single fold from 5 folds (so 80% of train stg1 used); one with stg1 data only `0.04903` and another where the additional stg2 training data was included - `0.04795`. \nSolution write up is [here](https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/117228#latest-673431)\nModel is broken over two parts, image model and sequential model, for the `0.04795` result, \n- Image model code [link](https://github.com/darraghdog/rsna/blob/master/scripts/resnext101v14/trainorig.py) and params (`--epochs 3 --fold 0  --lr 0.00002 --batchsize 64 --size 480` ); chose epoch to use based on val results. \n- LSTM code [link](https://github.com/darraghdog/rsna/blob/master/scripts/resnext101v14/trainlstmsngl.py) and params (`--epochs 5 --fold 0  --lr 0.00001 --batchsize 4  --lrgamma 0.95 --nbags 1 --globalepoch 3  lstm_units 2048`) - again chose on best val result.\n\nFor inference two parts (classifier and sequential) could be easily combined to one end to end model. Let me know if you'd like me to code it. For training, I'm not sure if this would work well; but could work better than training independently. \n\nA few notes on the solution write up. Much of the solution is the same as at least a few of the other solutions which were posted; I think the not so common parts which helped (excl any TTA/bagging etc.) are,\n- on lstm concat on the deltas between current and previous/next embeddings [link](https://github.com/darraghdog/rsna/blob/1d27f95eb5c12a04e2235f17b4798a430c7a3da5/scripts/resnext101v14/trainlstmsngl.py#L150)   \n- depth of the network on lstm, and summing the embeddings back on to the LSTM outputs (just the original emb, not the delta)  [link](https://github.com/darraghdog/rsna/blob/1d27f95eb5c12a04e2235f17b4798a430c7a3da5/scripts/resnext101v14/trainlstmsngl.py#L317)\n- maybe how black space was cut, however in retrospect I see the minimum kernel mentioned in solution write up was not used in the code ( :palmface: ) so this was used - [link](https://github.com/darraghdog/rsna/blob/1d27f95eb5c12a04e2235f17b4798a430c7a3da5/scripts/resnext101v14/trainorig.py#L139)\n\nWould be very happy to contribute. ",
    "675032": "Thanks @jhoward, that's a great idea. We would really love to participate.\nWe had an ensemble of two types of models:\nOne type is a simple resnet101 with three windowing values for the channels and resolution of 384. Its score on stage 2 private LB is 0.06262. We then use another 1d CNN over slices belonging to the same series to further post process our predictions. This lowered the LB score to 0.05331.\n\nSecond type is a se_resnext50_32x4d, again three channels for different windowing values. Here we predicted on three adjacent slices from the same series which were concatenated together and pass through a dense layer to predict the labels of the middle slice. For this model the LB score is 0.05481. Here we used the same post-processing as before to get LB score of 0.05367.",
    "674272": "Hi Jeremy, I'm interested in this and have a few questions regarding `ensembling`. \n\n1. Does this only apply to the test part? [My solution](https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/117330) involves second level model which requires out-of-fold predictions from 8-folds resnext as input for training. But the testing part can be easily adapted to a single model.\n\n2. Does this only apply to CNN? I use LightGBM, Catboost and XGB for second level model and average their predictions. It can be easily adapted to use one of them if it's not adequate but would like to make sure which is the case.",
    "674168": "My best single fold, single model is just an EfficientNet-B5 CNN that takes 5-channel input (slice-1, slice x3, slice+1) to generate a single slice prediction. I also use a trainable windowing module and smooth the probabilities for the series by using a weighted average of the probabilities of a given slice and its adjacent slices. It scores 0.05389. \n\nI would be interested in helping out. At my institution, we have already implemented a very early version of the model and have it running in the background on head CTs as they come in. ",
    "674123": "Exciting, great idea!",
    "682677": "@jhoward/all I put together an excel here with a rough summary of the models/approaches/posts/code from folks who posted here to get a quick overview of the similarities and differences of each approach. \n\nhttps://1drv.ms/x/s!AmMpcG2zZmnxx1ILR0InSX_Tf0ok\n\nHappy for any feedback, will hope to replicate one of the single model approaches in fastai2 for my own learning too :)  ",
    "724811": "Oh wow, I totally missed this. What a great idea. Apologies if this is too late, but my best single model was EfficientNet-B5 which scored 0.04830 on private LB.\n\nThe preprocessing was done as per this kernel: https://www.kaggle.com/anjum48/5th-preprocessing-adjacent-images-and-cropping\n\nA video of the full solution which contains other single model results can be found here: https://www.youtube.com/watch?v=1zLBxwTAcAs"
  }
}