{
  "id": 193400,
  "title": "[54th solution] Baseline inference optimization with TTA (No additional training)",
  "url": "/competitions/rsna-str-pulmonary-embolism-detection/discussion/193400",
  "author_name": "Stanley Zheng",
  "post_date": "2020-10-26T23:59:58.297000",
  "votes": 12,
  "comment_count": 14,
  "views": 0,
  "content": "<p>This is the solution for my team, \"Speed is all you need\". We were about 40th before the public baseline. TLDR at bottom.</p>\n<p>First of all, thanks to my team, especially <a href=\"https://www.kaggle.com/Neurallmonk\" target=\"_blank\">@Neurallmonk</a>, who has been exceedingly helpful, especially considering it is his first competition. Thanks to Ian Pan for the dataset, upon which our first month of experimenting was based on, and thank you to Kun for the baseline, though I was frustrated when it was initially released, I have learned so much from the code, and a medal is only a medal, but learning the code lasts forever.</p>\n<p>Our solution is nothing special - just hours and hours of optimizing Kun's baseline inference. We did not have the compute resources in order to retrain the first stage models, and a few of our teammates were busy during the last week of the competition, myself included as I had exams throughout. Otherwise, we would have liked to replace GRU with LSTM and ensemble.</p>\n<p>I optimized Kun's inference by moving model loading outside of for loops, replacing zoom with albumentations' resize, and hundreds of other small changes which combined together to make the inference over 3x faster. This results in a score of 0.229 public, 0.230 private LB in 2.5 hours (original notebook was 0.233 public, 0.232 private in 9 hours). Admittedly, I have no clue why this notebook has a slightly better public score, but the main goal was to be faster. This 2.5 hour long notebook is <a href=\"https://www.kaggle.com/stanleyjzheng/fast-baseline-with-tta?scriptVersionId=45225265\" target=\"_blank\">Here</a></p>\n<p>Then, we added TTA. We used TTAx3 with flips, CLAHE, brightness/contrast/hue. TTAx3 was the maximum we could fit in 9 hours, and these augmentations were found with trial and error on the public LB. Our final notebook with TTA is <a href=\"https://www.kaggle.com/stanleyjzheng/fast-baseline-with-tta-after-gru-8e13cc?scriptVersionId=45424927\" target=\"_blank\">Here</a></p>\n<p>TLDR: We used the original models trained in the public baseline, and changed purely inference. After optimization, our notebook runs in 2.5 hours (0.229 public, 0.230 private), and after TTAx3, we achieve a score of 0.221 public, 0.216 private.</p>",
  "messages": [
    {
      "id": 1061318,
      "postDate": "2020-10-26T23:59:58.297Z",
      "content": "<p>This is the solution for my team, \"Speed is all you need\". We were about 40th before the public baseline. TLDR at bottom.</p>\n<p>First of all, thanks to my team, especially <a href=\"https://www.kaggle.com/Neurallmonk\" target=\"_blank\">@Neurallmonk</a>, who has been exceedingly helpful, especially considering it is his first competition. Thanks to Ian Pan for the dataset, upon which our first month of experimenting was based on, and thank you to Kun for the baseline, though I was frustrated when it was initially released, I have learned so much from the code, and a medal is only a medal, but learning the code lasts forever.</p>\n<p>Our solution is nothing special - just hours and hours of optimizing Kun's baseline inference. We did not have the compute resources in order to retrain the first stage models, and a few of our teammates were busy during the last week of the competition, myself included as I had exams throughout. Otherwise, we would have liked to replace GRU with LSTM and ensemble.</p>\n<p>I optimized Kun's inference by moving model loading outside of for loops, replacing zoom with albumentations' resize, and hundreds of other small changes which combined together to make the inference over 3x faster. This results in a score of 0.229 public, 0.230 private LB in 2.5 hours (original notebook was 0.233 public, 0.232 private in 9 hours). Admittedly, I have no clue why this notebook has a slightly better public score, but the main goal was to be faster. This 2.5 hour long notebook is <a href=\"https://www.kaggle.com/stanleyjzheng/fast-baseline-with-tta?scriptVersionId=45225265\" target=\"_blank\">Here</a></p>\n<p>Then, we added TTA. We used TTAx3 with flips, CLAHE, brightness/contrast/hue. TTAx3 was the maximum we could fit in 9 hours, and these augmentations were found with trial and error on the public LB. Our final notebook with TTA is <a href=\"https://www.kaggle.com/stanleyjzheng/fast-baseline-with-tta-after-gru-8e13cc?scriptVersionId=45424927\" target=\"_blank\">Here</a></p>\n<p>TLDR: We used the original models trained in the public baseline, and changed purely inference. After optimization, our notebook runs in 2.5 hours (0.229 public, 0.230 private), and after TTAx3, we achieve a score of 0.221 public, 0.216 private.</p>",
      "rawMarkdown": "This is the solution for my team, \"Speed is all you need\". We were about 40th before the public baseline. TLDR at bottom.\n\nFirst of all, thanks to my team, especially @Neurallmonk, who has been exceedingly helpful, especially considering it is his first competition. Thanks to Ian Pan for the dataset, upon which our first month of experimenting was based on, and thank you to Kun for the baseline, though I was frustrated when it was initially released, I have learned so much from the code, and a medal is only a medal, but learning the code lasts forever.\n\nOur solution is nothing special - just hours and hours of optimizing Kun's baseline inference. We did not have the compute resources in order to retrain the first stage models, and a few of our teammates were busy during the last week of the competition, myself included as I had exams throughout. Otherwise, we would have liked to replace GRU with LSTM and ensemble.\n\nI optimized Kun's inference by moving model loading outside of for loops, replacing zoom with albumentations' resize, and hundreds of other small changes which combined together to make the inference over 3x faster. This results in a score of 0.229 public, 0.230 private LB in 2.5 hours (original notebook was 0.233 public, 0.232 private in 9 hours). Admittedly, I have no clue why this notebook has a slightly better public score, but the main goal was to be faster. This 2.5 hour long notebook is [Here](https://www.kaggle.com/stanleyjzheng/fast-baseline-with-tta?scriptVersionId=45225265)\n\nThen, we added TTA. We used TTAx3 with flips, CLAHE, brightness/contrast/hue. TTAx3 was the maximum we could fit in 9 hours, and these augmentations were found with trial and error on the public LB. Our final notebook with TTA is [Here](https://www.kaggle.com/stanleyjzheng/fast-baseline-with-tta-after-gru-8e13cc?scriptVersionId=45424927)\n\nTLDR: We used the original models trained in the public baseline, and changed purely inference. After optimization, our notebook runs in 2.5 hours (0.229 public, 0.230 private), and after TTAx3, we achieve a score of 0.221 public, 0.216 private.",
      "votes": 12
    },
    {
      "id": 1061464,
      "postDate": "2020-10-27T03:00:31.183Z",
      "content": "<p>Congratulations Stanley and team. You did many things to speed up public notebook. That's great. Increasing speed always helps. Using TTA was smart.</p>",
      "rawMarkdown": "Congratulations Stanley and team. You did many things to speed up public notebook. That's great. Increasing speed always helps. Using TTA was smart.",
      "votes": 3,
      "replies": [
        {
          "id": 1061471,
          "postDate": "2020-10-27T03:05:23.260Z",
          "content": "<p>Thank you so much Chris!</p>",
          "rawMarkdown": "Thank you so much Chris!",
          "votes": 3
        }
      ]
    },
    {
      "id": 1061460,
      "postDate": "2020-10-27T02:55:48.910Z",
      "content": "<p>Nice job <a href=\"https://www.kaggle.com/stanleyjzheng\" target=\"_blank\">@stanleyjzheng</a> </p>",
      "rawMarkdown": "Nice job @stanleyjzheng ",
      "votes": 1,
      "replies": [
        {
          "id": 1061472,
          "postDate": "2020-10-27T03:05:33.930Z",
          "content": "<p>Thanks Parker! </p>",
          "rawMarkdown": "Thanks Parker! "
        }
      ]
    },
    {
      "id": 1061366,
      "postDate": "2020-10-27T00:49:57.260Z",
      "content": "<p>Congrats on results and thanks for sharing solution <a href=\"https://www.kaggle.com/stanleyjzheng\" target=\"_blank\">@stanleyjzheng</a> </p>",
      "rawMarkdown": "Congrats on results and thanks for sharing solution @stanleyjzheng ",
      "votes": 1,
      "replies": [
        {
          "id": 1061368,
          "postDate": "2020-10-27T00:52:15.890Z",
          "content": "<p>Thank you!</p>",
          "rawMarkdown": "Thank you!"
        }
      ]
    },
    {
      "id": 1061354,
      "postDate": "2020-10-27T00:39:56.677Z",
      "content": "<p>3x faster, great work! I try to optimize but failed😆</p>",
      "rawMarkdown": "3x faster, great work! I try to optimize but failed😆\n",
      "votes": 1,
      "replies": [
        {
          "id": 1061357,
          "postDate": "2020-10-27T00:41:49.690Z",
          "content": "<p>Thank you! Congratulations on your medal too.</p>",
          "rawMarkdown": "Thank you! Congratulations on your medal too."
        }
      ]
    },
    {
      "id": 1061323,
      "postDate": "2020-10-27T00:06:27.223Z",
      "content": "<p><a href=\"https://www.kaggle.com/stanleyjzheng\" target=\"_blank\">@stanleyjzheng</a> Optimization OP</p>",
      "rawMarkdown": "@stanleyjzheng Optimization OP",
      "votes": 1
    },
    {
      "id": 1061444,
      "postDate": "2020-10-27T02:34:41.483Z",
      "content": "<p>Nice work! \"Speed is all you need\" is a perfect team name for this solution. Very impressive how fast you got it running. Did you predict only private test? If not you probably could've made it even faster by doing so.</p>",
      "rawMarkdown": "Nice work! \"Speed is all you need\" is a perfect team name for this solution. Very impressive how fast you got it running. Did you predict only private test? If not you probably could've made it even faster by doing so.",
      "votes": 2,
      "replies": [
        {
          "id": 1061449,
          "postDate": "2020-10-27T02:37:44.603Z",
          "content": "<p>Thanks so much, never thought I'd get complemented by a GM when I started! We predicted on both public and private test on all submissions - we considered private only, but we only got TTA working 2 days before the end, so with only 10 submissions to figure out which augmentations were most helpful, we needed the public LB scores.</p>",
          "rawMarkdown": "Thanks so much, never thought I'd get complemented by a GM when I started! We predicted on both public and private test on all submissions - we considered private only, but we only got TTA working 2 days before the end, so with only 10 submissions to figure out which augmentations were most helpful, we needed the public LB scores.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1061330,
      "postDate": "2020-10-27T00:13:16.427Z",
      "content": "<p>So you improved by 0.014. Is that a lot.</p>\n<p>I'm genuinely curious, I'm not familiar with the competition's metric. Please don't downvote me 😇</p>",
      "rawMarkdown": "So you improved by 0.014. Is that a lot.\n\nI'm genuinely curious, I'm not familiar with the competition's metric. Please don't downvote me 😇",
      "votes": -2,
      "replies": [
        {
          "id": 1061333,
          "postDate": "2020-10-27T00:16:28.950Z",
          "content": "<p>Don't worry about downvotes, it's a great question. I would not quantify 0.014 log loss as \"a lot\", but in the scheme of this competition, it's the difference between 150th and 56th, so to me, it matters a lot :)</p>",
          "rawMarkdown": "Don't worry about downvotes, it's a great question. I would not quantify 0.014 log loss as \"a lot\", but in the scheme of this competition, it's the difference between 150th and 56th, so to me, it matters a lot :)\n"
        },
        {
          "id": 1061343,
          "postDate": "2020-10-27T00:25:37.510Z",
          "content": "<p>So I was right! Besides, it's not log loss it's weighted log loss. I didn't have time to experiment with the metric so .013  can be game changing for all I know.</p>",
          "rawMarkdown": "So I was right! Besides, it's not log loss it's weighted log loss. I didn't have time to experiment with the metric so .013  can be game changing for all I know.",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1061464,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-10-27T03:00:31.183000",
      "content": "<p>Congratulations Stanley and team. You did many things to speed up public notebook. That's great. Increasing speed always helps. Using TTA was smart.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1061471,
          "author_name": "Stanley Zheng",
          "author_url": "",
          "post_date": "2020-10-27T03:05:23.260000",
          "content": "<p>Thank you so much Chris!</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1061460,
      "author_name": "cswwp",
      "author_url": "",
      "post_date": "2020-10-27T02:55:48.910000",
      "content": "<p>Nice job <a href=\"https://www.kaggle.com/stanleyjzheng\" target=\"_blank\">@stanleyjzheng</a> </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1061472,
          "author_name": "Stanley Zheng",
          "author_url": "",
          "post_date": "2020-10-27T03:05:33.930000",
          "content": "<p>Thanks Parker! </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1061366,
      "author_name": "KhanhVD",
      "author_url": "",
      "post_date": "2020-10-27T00:49:57.260000",
      "content": "<p>Congrats on results and thanks for sharing solution <a href=\"https://www.kaggle.com/stanleyjzheng\" target=\"_blank\">@stanleyjzheng</a> </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1061368,
          "author_name": "Stanley Zheng",
          "author_url": "",
          "post_date": "2020-10-27T00:52:15.890000",
          "content": "<p>Thank you!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1061354,
      "author_name": "Salaryman",
      "author_url": "",
      "post_date": "2020-10-27T00:39:56.677000",
      "content": "<p>3x faster, great work! I try to optimize but failed😆</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1061357,
          "author_name": "Stanley Zheng",
          "author_url": "",
          "post_date": "2020-10-27T00:41:49.690000",
          "content": "<p>Thank you! Congratulations on your medal too.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1061323,
      "author_name": "Saurabh dubey",
      "author_url": "",
      "post_date": "2020-10-27T00:06:27.223000",
      "content": "<p><a href=\"https://www.kaggle.com/stanleyjzheng\" target=\"_blank\">@stanleyjzheng</a> Optimization OP</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1061444,
      "author_name": "Rob Mulla",
      "author_url": "",
      "post_date": "2020-10-27T02:34:41.483000",
      "content": "<p>Nice work! \"Speed is all you need\" is a perfect team name for this solution. Very impressive how fast you got it running. Did you predict only private test? If not you probably could've made it even faster by doing so.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1061449,
          "author_name": "Stanley Zheng",
          "author_url": "",
          "post_date": "2020-10-27T02:37:44.603000",
          "content": "<p>Thanks so much, never thought I'd get complemented by a GM when I started! We predicted on both public and private test on all submissions - we considered private only, but we only got TTA working 2 days before the end, so with only 10 submissions to figure out which augmentations were most helpful, we needed the public LB scores.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1061330,
      "author_name": "DarkCube",
      "author_url": "",
      "post_date": "2020-10-27T00:13:16.427000",
      "content": "<p>So you improved by 0.014. Is that a lot.</p>\n<p>I'm genuinely curious, I'm not familiar with the competition's metric. Please don't downvote me 😇</p>",
      "votes": -2,
      "replies": [
        {
          "id": 1061333,
          "author_name": "Stanley Zheng",
          "author_url": "",
          "post_date": "2020-10-27T00:16:28.950000",
          "content": "<p>Don't worry about downvotes, it's a great question. I would not quantify 0.014 log loss as \"a lot\", but in the scheme of this competition, it's the difference between 150th and 56th, so to me, it matters a lot :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1061343,
          "author_name": "DarkCube",
          "author_url": "",
          "post_date": "2020-10-27T00:25:37.510000",
          "content": "<p>So I was right! Besides, it's not log loss it's weighted log loss. I didn't have time to experiment with the metric so .013  can be game changing for all I know.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1061318": "This is the solution for my team, \"Speed is all you need\". We were about 40th before the public baseline. TLDR at bottom.\n\nFirst of all, thanks to my team, especially @Neurallmonk, who has been exceedingly helpful, especially considering it is his first competition. Thanks to Ian Pan for the dataset, upon which our first month of experimenting was based on, and thank you to Kun for the baseline, though I was frustrated when it was initially released, I have learned so much from the code, and a medal is only a medal, but learning the code lasts forever.\n\nOur solution is nothing special - just hours and hours of optimizing Kun's baseline inference. We did not have the compute resources in order to retrain the first stage models, and a few of our teammates were busy during the last week of the competition, myself included as I had exams throughout. Otherwise, we would have liked to replace GRU with LSTM and ensemble.\n\nI optimized Kun's inference by moving model loading outside of for loops, replacing zoom with albumentations' resize, and hundreds of other small changes which combined together to make the inference over 3x faster. This results in a score of 0.229 public, 0.230 private LB in 2.5 hours (original notebook was 0.233 public, 0.232 private in 9 hours). Admittedly, I have no clue why this notebook has a slightly better public score, but the main goal was to be faster. This 2.5 hour long notebook is [Here](https://www.kaggle.com/stanleyjzheng/fast-baseline-with-tta?scriptVersionId=45225265)\n\nThen, we added TTA. We used TTAx3 with flips, CLAHE, brightness/contrast/hue. TTAx3 was the maximum we could fit in 9 hours, and these augmentations were found with trial and error on the public LB. Our final notebook with TTA is [Here](https://www.kaggle.com/stanleyjzheng/fast-baseline-with-tta-after-gru-8e13cc?scriptVersionId=45424927)\n\nTLDR: We used the original models trained in the public baseline, and changed purely inference. After optimization, our notebook runs in 2.5 hours (0.229 public, 0.230 private), and after TTAx3, we achieve a score of 0.221 public, 0.216 private.",
    "1061464": "Congratulations Stanley and team. You did many things to speed up public notebook. That's great. Increasing speed always helps. Using TTA was smart.",
    "1061460": "Nice job @stanleyjzheng ",
    "1061366": "Congrats on results and thanks for sharing solution @stanleyjzheng ",
    "1061354": "3x faster, great work! I try to optimize but failed😆\n",
    "1061323": "@stanleyjzheng Optimization OP",
    "1061444": "Nice work! \"Speed is all you need\" is a perfect team name for this solution. Very impressive how fast you got it running. Did you predict only private test? If not you probably could've made it even faster by doing so.",
    "1061330": "So you improved by 0.014. Is that a lot.\n\nI'm genuinely curious, I'm not familiar with the competition's metric. Please don't downvote me 😇"
  }
}