{
  "id": 441557,
  "title": "Tips on Beating the Baseline",
  "url": "/competitions/rsna-2023-abdominal-trauma-detection/discussion/441557",
  "author_name": "Theo Viel",
  "post_date": "2023-09-19T09:13:48.177000",
  "votes": 74,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Because of several factors, this competition is extremely challenging to tackle, to name a few :</p>\n<ul>\n<li>The log loss heavily penalizes mistakes hence a not so strong model will not beat the weighted mean baseline</li>\n<li>3D data is tough to handle, much more than 2D</li>\n</ul>\n<p>Here are a few tips to get started, that can also be relevant for other competitions.</p>\n<h4>1. Data</h4>\n<p>Work with preprocessed pngs, the datasets I've shared here are strong enough to achieve LB 0.4. <a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/427427\" target=\"_blank\">Discussion link</a>, please upvote the datasets !<br>\nPNG is lossless, and usually a safer choice than jpg. Images are not resized so you're not losing any information, and you can downsize them later in your pipeline if needed. It's usually a good idea to start experimenting with 256x256 and then move to bigger image sizes. Overall, the only information loss is windowing which if done correctly has proven to be efficient in previous competitions.</p>\n<h4>2. Previous competitions</h4>\n<p>I highly encourage everyone to study (by study I mean make sure you understand everything) top solutions from the previous <a href=\"https://www.kaggle.com/competitions/rsna-2022-cervical-spine-fracture-detection\" target=\"_blank\">RSNA competition</a>. Some things do not apply here but the data and problem setting are very similar.</p>\n<p>Furthermore, make sure you understand the challenges specific to our data when training models.</p>\n<h4>3. Experimenting</h4>\n<p>Although the competition metric is what we aim to optimize, I heavily recommend <strong>tracking the AUC</strong> when experimenting. The reason is two-fold :</p>\n<ul>\n<li>A model scoring &lt;0.6 AUC is weak and not learning, whereas it's quite hard to infer a similar relationship with the weighted loss</li>\n<li>The AUC cannot be gamed with scaling. Scaling your predictions will decrease the log loss without increasing the quality of your models. First make sure your models are good, then scale your predictions.</li>\n</ul>\n<p>Alongside with that, I like to train models on the validation data to make sure that there is signal to learn and no bug in my pipeline. A small bug or wrong choice in the pipeline design can be the difference between a model that learns and one that does not.</p>\n<h4>4. Submitting</h4>\n<p>Improving your LB score is tempting, but not useful in the process of beating the baseline. A good CV scheme and metric implementation will be more reliable than public LB on that matter (make sure you have both !). Submitting a model takes time, time that could instead be allocated to improving your pipeline. In the early stage of the competition your pipeline can change a lot and you want to avoid spending 4 hours writing inference code every time. </p>\n<p>My advice is to first build a strong model, then work on the inference code. When working on your inference code, make sure you get the <strong>same result</strong> on the Kaggle environment as locally, to avoid any bug. </p>\n<p><em>That's about it, I hope this is helpful to anyone getting started.</em></p>\n<p>Also, do not ask for hints regarding my solution. Refer to <strong>2.</strong> and start coding 🙃</p>",
  "messages": [
    {
      "id": 2446128,
      "postDate": "2023-09-19T09:13:48.177Z",
      "content": "<p>Because of several factors, this competition is extremely challenging to tackle, to name a few :</p>\n<ul>\n<li>The log loss heavily penalizes mistakes hence a not so strong model will not beat the weighted mean baseline</li>\n<li>3D data is tough to handle, much more than 2D</li>\n</ul>\n<p>Here are a few tips to get started, that can also be relevant for other competitions.</p>\n<h4>1. Data</h4>\n<p>Work with preprocessed pngs, the datasets I've shared here are strong enough to achieve LB 0.4. <a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/427427\" target=\"_blank\">Discussion link</a>, please upvote the datasets !<br>\nPNG is lossless, and usually a safer choice than jpg. Images are not resized so you're not losing any information, and you can downsize them later in your pipeline if needed. It's usually a good idea to start experimenting with 256x256 and then move to bigger image sizes. Overall, the only information loss is windowing which if done correctly has proven to be efficient in previous competitions.</p>\n<h4>2. Previous competitions</h4>\n<p>I highly encourage everyone to study (by study I mean make sure you understand everything) top solutions from the previous <a href=\"https://www.kaggle.com/competitions/rsna-2022-cervical-spine-fracture-detection\" target=\"_blank\">RSNA competition</a>. Some things do not apply here but the data and problem setting are very similar.</p>\n<p>Furthermore, make sure you understand the challenges specific to our data when training models.</p>\n<h4>3. Experimenting</h4>\n<p>Although the competition metric is what we aim to optimize, I heavily recommend <strong>tracking the AUC</strong> when experimenting. The reason is two-fold :</p>\n<ul>\n<li>A model scoring &lt;0.6 AUC is weak and not learning, whereas it's quite hard to infer a similar relationship with the weighted loss</li>\n<li>The AUC cannot be gamed with scaling. Scaling your predictions will decrease the log loss without increasing the quality of your models. First make sure your models are good, then scale your predictions.</li>\n</ul>\n<p>Alongside with that, I like to train models on the validation data to make sure that there is signal to learn and no bug in my pipeline. A small bug or wrong choice in the pipeline design can be the difference between a model that learns and one that does not.</p>\n<h4>4. Submitting</h4>\n<p>Improving your LB score is tempting, but not useful in the process of beating the baseline. A good CV scheme and metric implementation will be more reliable than public LB on that matter (make sure you have both !). Submitting a model takes time, time that could instead be allocated to improving your pipeline. In the early stage of the competition your pipeline can change a lot and you want to avoid spending 4 hours writing inference code every time. </p>\n<p>My advice is to first build a strong model, then work on the inference code. When working on your inference code, make sure you get the <strong>same result</strong> on the Kaggle environment as locally, to avoid any bug. </p>\n<p><em>That's about it, I hope this is helpful to anyone getting started.</em></p>\n<p>Also, do not ask for hints regarding my solution. Refer to <strong>2.</strong> and start coding 🙃</p>",
      "rawMarkdown": "Because of several factors, this competition is extremely challenging to tackle, to name a few :\n- The log loss heavily penalizes mistakes hence a not so strong model will not beat the weighted mean baseline\n- 3D data is tough to handle, much more than 2D\n\nHere are a few tips to get started, that can also be relevant for other competitions.\n\n#### 1. Data\n\nWork with preprocessed pngs, the datasets I've shared here are strong enough to achieve LB 0.4. [Discussion link](https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/427427), please upvote the datasets !\nPNG is lossless, and usually a safer choice than jpg. Images are not resized so you're not losing any information, and you can downsize them later in your pipeline if needed. It's usually a good idea to start experimenting with 256x256 and then move to bigger image sizes. Overall, the only information loss is windowing which if done correctly has proven to be efficient in previous competitions.\n\n#### 2. Previous competitions\n\nI highly encourage everyone to study (by study I mean make sure you understand everything) top solutions from the previous [RSNA competition](https://www.kaggle.com/competitions/rsna-2022-cervical-spine-fracture-detection). Some things do not apply here but the data and problem setting are very similar.\n\nFurthermore, make sure you understand the challenges specific to our data when training models.\n\n#### 3. Experimenting\n\nAlthough the competition metric is what we aim to optimize, I heavily recommend **tracking the AUC** when experimenting. The reason is two-fold :\n- A model scoring <0.6 AUC is weak and not learning, whereas it's quite hard to infer a similar relationship with the weighted loss\n- The AUC cannot be gamed with scaling. Scaling your predictions will decrease the log loss without increasing the quality of your models. First make sure your models are good, then scale your predictions.\n\nAlongside with that, I like to train models on the validation data to make sure that there is signal to learn and no bug in my pipeline. A small bug or wrong choice in the pipeline design can be the difference between a model that learns and one that does not.\n\n#### 4. Submitting\n\nImproving your LB score is tempting, but not useful in the process of beating the baseline. A good CV scheme and metric implementation will be more reliable than public LB on that matter (make sure you have both !). Submitting a model takes time, time that could instead be allocated to improving your pipeline. In the early stage of the competition your pipeline can change a lot and you want to avoid spending 4 hours writing inference code every time. \n\nMy advice is to first build a strong model, then work on the inference code. When working on your inference code, make sure you get the **same result** on the Kaggle environment as locally, to avoid any bug. \n\n\n\n*That's about it, I hope this is helpful to anyone getting started.*\n\nAlso, do not ask for hints regarding my solution. Refer to **2.** and start coding 🙃",
      "votes": 74
    },
    {
      "id": 2460349,
      "postDate": "2023-09-28T18:02:28.487Z",
      "content": "<p>This video was recorded by someone who did well in the previous RSNA competition: <a href=\"https://youtu.be/x3BLoHkzUYc?si=Tgcv9tRNdxjYJRe8\" target=\"_blank\">https://youtu.be/x3BLoHkzUYc?si=Tgcv9tRNdxjYJRe8</a></p>\n<p>“highly encourage everyone to study (by study I mean make sure you understand everything) top solutions from the previous RSNA competition”</p>\n<p>Thank you for these tips!</p>",
      "rawMarkdown": "This video was recorded by someone who did well in the previous RSNA competition: https://youtu.be/x3BLoHkzUYc?si=Tgcv9tRNdxjYJRe8\n\n“highly encourage everyone to study (by study I mean make sure you understand everything) top solutions from the previous RSNA competition”\n\nThank you for these tips!",
      "votes": 3
    },
    {
      "id": 2473738,
      "postDate": "2023-10-08T14:35:58.347Z",
      "content": "<p>Just curious: how much time did your submission take? :) I mean, how much time did your submission take to calculate the score?</p>",
      "rawMarkdown": "Just curious: how much time did your submission take? :) I mean, how much time did your submission take to calculate the score?",
      "votes": 1
    },
    {
      "id": 2457410,
      "postDate": "2023-09-26T21:19:01.820Z",
      "content": "<p>Thanks for the tips - especially the AUC tracking.</p>",
      "rawMarkdown": "Thanks for the tips - especially the AUC tracking.",
      "votes": 1
    },
    {
      "id": 2446501,
      "postDate": "2023-09-19T13:12:25.067Z",
      "content": "<p>Very cool, literally a legend hahah</p>",
      "rawMarkdown": "Very cool, literally a legend hahah",
      "votes": 1
    },
    {
      "id": 2473406,
      "postDate": "2023-10-08T08:43:00.290Z",
      "content": "<blockquote>\n  <p>I like to train models on the validation data to make sure that there is signal to learn and no bug in my pipeline. </p>\n</blockquote>\n<p>Thank you for the tips! Just wanting to check what you mean by you train on the validation data? Do you mean for the final submission model or just in general in your experimentation process you don't use train/validation split/you assume your models are generalising well?</p>",
      "rawMarkdown": ">I like to train models on the validation data to make sure that there is signal to learn and no bug in my pipeline. \n\nThank you for the tips! Just wanting to check what you mean by you train on the validation data? Do you mean for the final submission model or just in general in your experimentation process you don't use train/validation split/you assume your models are generalising well?",
      "replies": [
        {
          "id": 2473407,
          "postDate": "2023-10-08T08:45:41.537Z",
          "content": "<p>I keep my train/val split, but instead of training on train and validating on val, I train on (train + val) and validate on val. <br>\nIf I don't get (very) high scores on val it means my model is not learning. This is kind of like tracking the evaluation metric on train data.</p>",
          "rawMarkdown": "I keep my train/val split, but instead of training on train and validating on val, I train on (train + val) and validate on val. \nIf I don't get (very) high scores on val it means my model is not learning. This is kind of like tracking the evaluation metric on train data.",
          "votes": 5,
          "replies": [
            {
              "id": 2473430,
              "postDate": "2023-10-08T09:01:48.800Z",
              "content": "<p>Ah okay, interesting. Does this not run the risk of overfitting since you won't know if the model generalises to the test data? Or you just submit to see if that's the case when the score is high enough?</p>",
              "rawMarkdown": "Ah okay, interesting. Does this not run the risk of overfitting since you won't know if the model generalises to the test data? Or you just submit to see if that's the case when the score is high enough?",
              "votes": 1
            },
            {
              "id": 2473452,
              "postDate": "2023-10-08T09:27:10.197Z",
              "content": "<p>I don't submit it nor use it to assess generalization. It's just a quick check.</p>",
              "rawMarkdown": "I don't submit it nor use it to assess generalization. It's just a quick check."
            }
          ]
        }
      ]
    },
    {
      "id": 2483972,
      "postDate": "2023-10-16T06:21:40.817Z",
      "content": "<p>Thank you for your guidance😀</p>",
      "rawMarkdown": "Thank you for your guidance😀"
    },
    {
      "id": 2482494,
      "postDate": "2023-10-15T04:52:49.567Z",
      "content": "<p>Appreciate it thanks a lot</p>",
      "rawMarkdown": "Appreciate it thanks a lot"
    },
    {
      "id": 2446378,
      "postDate": "2023-09-19T11:57:36.193Z",
      "content": "<p><a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> Thanks for your post</p>",
      "rawMarkdown": "@theoviel Thanks for your post"
    }
  ],
  "comments": [
    {
      "id": 2460349,
      "author_name": "RickPack",
      "author_url": "",
      "post_date": "2023-09-28T18:02:28.487000",
      "content": "<p>This video was recorded by someone who did well in the previous RSNA competition: <a href=\"https://youtu.be/x3BLoHkzUYc?si=Tgcv9tRNdxjYJRe8\" target=\"_blank\">https://youtu.be/x3BLoHkzUYc?si=Tgcv9tRNdxjYJRe8</a></p>\n<p>“highly encourage everyone to study (by study I mean make sure you understand everything) top solutions from the previous RSNA competition”</p>\n<p>Thank you for these tips!</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2473738,
      "author_name": "Michalina Hulak",
      "author_url": "",
      "post_date": "2023-10-08T14:35:58.347000",
      "content": "<p>Just curious: how much time did your submission take? :) I mean, how much time did your submission take to calculate the score?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2457410,
      "author_name": "Joshua Tice",
      "author_url": "",
      "post_date": "2023-09-26T21:19:01.820000",
      "content": "<p>Thanks for the tips - especially the AUC tracking.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2446501,
      "author_name": "Man of the year",
      "author_url": "",
      "post_date": "2023-09-19T13:12:25.067000",
      "content": "<p>Very cool, literally a legend hahah</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2473406,
      "author_name": "Mark",
      "author_url": "",
      "post_date": "2023-10-08T08:43:00.290000",
      "content": "<blockquote>\n  <p>I like to train models on the validation data to make sure that there is signal to learn and no bug in my pipeline. </p>\n</blockquote>\n<p>Thank you for the tips! Just wanting to check what you mean by you train on the validation data? Do you mean for the final submission model or just in general in your experimentation process you don't use train/validation split/you assume your models are generalising well?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2473407,
          "author_name": "Theo Viel",
          "author_url": "",
          "post_date": "2023-10-08T08:45:41.537000",
          "content": "<p>I keep my train/val split, but instead of training on train and validating on val, I train on (train + val) and validate on val. <br>\nIf I don't get (very) high scores on val it means my model is not learning. This is kind of like tracking the evaluation metric on train data.</p>",
          "votes": 5,
          "replies": [
            {
              "id": 2473430,
              "author_name": "Mark",
              "author_url": "",
              "post_date": "2023-10-08T09:01:48.800000",
              "content": "<p>Ah okay, interesting. Does this not run the risk of overfitting since you won't know if the model generalises to the test data? Or you just submit to see if that's the case when the score is high enough?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2473452,
              "author_name": "Theo Viel",
              "author_url": "",
              "post_date": "2023-10-08T09:27:10.197000",
              "content": "<p>I don't submit it nor use it to assess generalization. It's just a quick check.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2483972,
      "author_name": "fcc_personal",
      "author_url": "",
      "post_date": "2023-10-16T06:21:40.817000",
      "content": "<p>Thank you for your guidance😀</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2482494,
      "author_name": "kerry sun",
      "author_url": "",
      "post_date": "2023-10-15T04:52:49.567000",
      "content": "<p>Appreciate it thanks a lot</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2446378,
      "author_name": "𝔄ℌ𝔐𝔈𝔇 𝔄𝔖ℌℜ𝔄𝔉",
      "author_url": "",
      "post_date": "2023-09-19T11:57:36.193000",
      "content": "<p><a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> Thanks for your post</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2446128": "Because of several factors, this competition is extremely challenging to tackle, to name a few :\n- The log loss heavily penalizes mistakes hence a not so strong model will not beat the weighted mean baseline\n- 3D data is tough to handle, much more than 2D\n\nHere are a few tips to get started, that can also be relevant for other competitions.\n\n#### 1. Data\n\nWork with preprocessed pngs, the datasets I've shared here are strong enough to achieve LB 0.4. [Discussion link](https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/427427), please upvote the datasets !\nPNG is lossless, and usually a safer choice than jpg. Images are not resized so you're not losing any information, and you can downsize them later in your pipeline if needed. It's usually a good idea to start experimenting with 256x256 and then move to bigger image sizes. Overall, the only information loss is windowing which if done correctly has proven to be efficient in previous competitions.\n\n#### 2. Previous competitions\n\nI highly encourage everyone to study (by study I mean make sure you understand everything) top solutions from the previous [RSNA competition](https://www.kaggle.com/competitions/rsna-2022-cervical-spine-fracture-detection). Some things do not apply here but the data and problem setting are very similar.\n\nFurthermore, make sure you understand the challenges specific to our data when training models.\n\n#### 3. Experimenting\n\nAlthough the competition metric is what we aim to optimize, I heavily recommend **tracking the AUC** when experimenting. The reason is two-fold :\n- A model scoring <0.6 AUC is weak and not learning, whereas it's quite hard to infer a similar relationship with the weighted loss\n- The AUC cannot be gamed with scaling. Scaling your predictions will decrease the log loss without increasing the quality of your models. First make sure your models are good, then scale your predictions.\n\nAlongside with that, I like to train models on the validation data to make sure that there is signal to learn and no bug in my pipeline. A small bug or wrong choice in the pipeline design can be the difference between a model that learns and one that does not.\n\n#### 4. Submitting\n\nImproving your LB score is tempting, but not useful in the process of beating the baseline. A good CV scheme and metric implementation will be more reliable than public LB on that matter (make sure you have both !). Submitting a model takes time, time that could instead be allocated to improving your pipeline. In the early stage of the competition your pipeline can change a lot and you want to avoid spending 4 hours writing inference code every time. \n\nMy advice is to first build a strong model, then work on the inference code. When working on your inference code, make sure you get the **same result** on the Kaggle environment as locally, to avoid any bug. \n\n\n\n*That's about it, I hope this is helpful to anyone getting started.*\n\nAlso, do not ask for hints regarding my solution. Refer to **2.** and start coding 🙃",
    "2460349": "This video was recorded by someone who did well in the previous RSNA competition: https://youtu.be/x3BLoHkzUYc?si=Tgcv9tRNdxjYJRe8\n\n“highly encourage everyone to study (by study I mean make sure you understand everything) top solutions from the previous RSNA competition”\n\nThank you for these tips!",
    "2473738": "Just curious: how much time did your submission take? :) I mean, how much time did your submission take to calculate the score?",
    "2457410": "Thanks for the tips - especially the AUC tracking.",
    "2446501": "Very cool, literally a legend hahah",
    "2473406": ">I like to train models on the validation data to make sure that there is signal to learn and no bug in my pipeline. \n\nThank you for the tips! Just wanting to check what you mean by you train on the validation data? Do you mean for the final submission model or just in general in your experimentation process you don't use train/validation split/you assume your models are generalising well?",
    "2483972": "Thank you for your guidance😀",
    "2482494": "Appreciate it thanks a lot",
    "2446378": "@theoviel Thanks for your post"
  }
}