{
  "id": 357877,
  "title": "It was not all randomness ! (maybe) - 5th Place Secret Sauce",
  "url": "/competitions/mayo-clinic-strip-ai/discussion/357877",
  "author_name": "Theo Viel",
  "post_date": "2022-10-06T00:21:53.350000",
  "votes": 29,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Although public LB was random (20 samples is not enough), you could actually get fairly consistent results.</p>\n<p>Here are a few ideas, I will work on a more detailed write-up tomorrow :</p>\n<p>1) Optimize the AUC, then you will know whether your model discriminates or not<br>\nMy models achieve AUCs around 0.67, which is quite low but not random.</p>\n<p>2) Understand the metric : the log loss heavily penalizes mistakes when the model is confident and the reward for getting a correct guess in comparison is much lower. <br>\nSince our models are not so good, you want to stay on the sweet spot where your loss does not get heavily increased because of the mistakes your model makes. </p>\n<p>I did so by scaling predictions to the [0.15, 0.85] range, and then clipping them to [0.25, 0.75]. This was tweaked on CV, my best private achieved a 0.64 CV.</p>\n<p>Sure it's only 0.05 lower than random predictions but that's probably close to the lowest you could get with the provided data :)</p>",
  "messages": [
    {
      "id": 1973920,
      "postDate": "2022-10-06T00:21:53.350Z",
      "content": "<p>Although public LB was random (20 samples is not enough), you could actually get fairly consistent results.</p>\n<p>Here are a few ideas, I will work on a more detailed write-up tomorrow :</p>\n<p>1) Optimize the AUC, then you will know whether your model discriminates or not<br>\nMy models achieve AUCs around 0.67, which is quite low but not random.</p>\n<p>2) Understand the metric : the log loss heavily penalizes mistakes when the model is confident and the reward for getting a correct guess in comparison is much lower. <br>\nSince our models are not so good, you want to stay on the sweet spot where your loss does not get heavily increased because of the mistakes your model makes. </p>\n<p>I did so by scaling predictions to the [0.15, 0.85] range, and then clipping them to [0.25, 0.75]. This was tweaked on CV, my best private achieved a 0.64 CV.</p>\n<p>Sure it's only 0.05 lower than random predictions but that's probably close to the lowest you could get with the provided data :)</p>",
      "rawMarkdown": "Although public LB was random (20 samples is not enough), you could actually get fairly consistent results.\n\nHere are a few ideas, I will work on a more detailed write-up tomorrow :\n\n1) Optimize the AUC, then you will know whether your model discriminates or not\nMy models achieve AUCs around 0.67, which is quite low but not random.\n\n2) Understand the metric : the log loss heavily penalizes mistakes when the model is confident and the reward for getting a correct guess in comparison is much lower. \nSince our models are not so good, you want to stay on the sweet spot where your loss does not get heavily increased because of the mistakes your model makes. \n\nI did so by scaling predictions to the [0.15, 0.85] range, and then clipping them to [0.25, 0.75]. This was tweaked on CV, my best private achieved a 0.64 CV.\n\nSure it's only 0.05 lower than random predictions but that's probably close to the lowest you could get with the provided data :)",
      "votes": 28
    },
    {
      "id": 1974769,
      "postDate": "2022-10-06T12:52:23.810Z",
      "content": "<p>Congrats on your finish! <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a><br>\nYour insights and discussions on the metric and its behavior helped me greatly! 😄</p>",
      "rawMarkdown": "Congrats on your finish! @theoviel\nYour insights and discussions on the metric and its behavior helped me greatly! 😄",
      "votes": 1
    },
    {
      "id": 1974568,
      "postDate": "2022-10-06T10:16:37.997Z",
      "content": "<p>Congratulations. Could you elaborate on how you scaled predictions to the [0.15, 0.85] range? That seems to potentially be useful for other future competitions. I tried the following non-linear stretching idea, which based on the private leaderboard, didn't help:</p>\n<pre><code>def stretch_prediction(x):\n    \"\"\" Based on input, stretch prediction to upper or lower\n    \"\"\"\n    low_bound = 0.49\n    high_bound = 0.51\n    if (x &gt; low_bound) &amp; (x &lt; high_bound):\n        return 0.50\n    elif x &lt; low_bound:\n        diff_low = low_bound - x\n        scaled_diff = diff_low ** 0.75\n        final_x = low_bound - scaled_diff\n        return np.max([0, final_x])\n    else:\n        diff_high = x - high_bound\n        scaled_diff = diff_high ** 0.75\n        final_x = high_bound + scaled_diff\n        return np.min([1, final_x])\n</code></pre>",
      "rawMarkdown": "Congratulations. Could you elaborate on how you scaled predictions to the [0.15, 0.85] range? That seems to potentially be useful for other future competitions. I tried the following non-linear stretching idea, which based on the private leaderboard, didn't help:\n\n```\ndef stretch_prediction(x):\n    \"\"\" Based on input, stretch prediction to upper or lower\n    \"\"\"\n    low_bound = 0.49\n    high_bound = 0.51\n    if (x > low_bound) & (x < high_bound):\n        return 0.50\n    elif x < low_bound:\n        diff_low = low_bound - x\n        scaled_diff = diff_low ** 0.75\n        final_x = low_bound - scaled_diff\n        return np.max([0, final_x])\n    else:\n        diff_high = x - high_bound\n        scaled_diff = diff_high ** 0.75\n        final_x = high_bound + scaled_diff\n        return np.min([1, final_x])\n```",
      "votes": 1,
      "replies": [
        {
          "id": 1974590,
          "postDate": "2022-10-06T10:34:23.257Z",
          "content": "<pre><code>def scale(preds, min_=0.2, max_=0.8):\n    preds = (preds - preds.min()) / (preds.max() - preds.min())\n    preds = preds * (max_ - min_) + min_\n    return preds\n</code></pre>",
          "rawMarkdown": "```\ndef scale(preds, min_=0.2, max_=0.8):\n    preds = (preds - preds.min()) / (preds.max() - preds.min())\n    preds = preds * (max_ - min_) + min_\n    return preds\n```",
          "votes": 2
        },
        {
          "id": 1974601,
          "postDate": "2022-10-06T10:53:40.070Z",
          "content": "<p>Neat, thanks for the secret sauce. Clipping like this seems worth trying for future competitions with log loss.</p>",
          "rawMarkdown": "Neat, thanks for the secret sauce. Clipping like this seems worth trying for future competitions with log loss.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1973952,
      "postDate": "2022-10-06T01:25:40.060Z",
      "content": "<p><a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> point 2) was a good call, and now makes a lot of sense… I did the same after some of us had discussions on the <a href=\"https://www.kaggle.com/competitions/mayo-clinic-strip-ai/discussion/354911\" target=\"_blank\">other thread</a> about scores being random… I have to thank you guys, that gave me confidence to select the notebook that scored best for me  :)</p>",
      "rawMarkdown": "@theoviel point 2) was a good call, and now makes a lot of sense... I did the same after some of us had discussions on the [other thread](https://www.kaggle.com/competitions/mayo-clinic-strip-ai/discussion/354911) about scores being random... I have to thank you guys, that gave me confidence to select the notebook that scored best for me  :)",
      "votes": 1,
      "replies": [
        {
          "id": 1974440,
          "postDate": "2022-10-06T08:50:53.043Z",
          "content": "<p>There was indeed some hints in the forum :)<br>\nCongratz on 7th place !</p>",
          "rawMarkdown": "There was indeed some hints in the forum :)\nCongratz on 7th place !",
          "votes": 1
        },
        {
          "id": 1975007,
          "postDate": "2022-10-06T15:18:42.593Z",
          "content": "<p>Thanks and congrats to you too :) </p>",
          "rawMarkdown": "Thanks and congrats to you too :) ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1973944,
      "postDate": "2022-10-06T01:07:26.967Z",
      "content": "<p>Did you use non resized tiles to train and predict?</p>",
      "rawMarkdown": "Did you use non resized tiles to train and predict?",
      "votes": 1,
      "replies": [
        {
          "id": 1974437,
          "postDate": "2022-10-06T08:49:53.777Z",
          "content": "<p>I'll post a more detailed write-up soon :)</p>",
          "rawMarkdown": "I'll post a more detailed write-up soon :)"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1974769,
      "author_name": "Yerram Varun",
      "author_url": "",
      "post_date": "2022-10-06T12:52:23.810000",
      "content": "<p>Congrats on your finish! <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a><br>\nYour insights and discussions on the metric and its behavior helped me greatly! 😄</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1974568,
      "author_name": "Joe Marturano",
      "author_url": "",
      "post_date": "2022-10-06T10:16:37.997000",
      "content": "<p>Congratulations. Could you elaborate on how you scaled predictions to the [0.15, 0.85] range? That seems to potentially be useful for other future competitions. I tried the following non-linear stretching idea, which based on the private leaderboard, didn't help:</p>\n<pre><code>def stretch_prediction(x):\n    \"\"\" Based on input, stretch prediction to upper or lower\n    \"\"\"\n    low_bound = 0.49\n    high_bound = 0.51\n    if (x &gt; low_bound) &amp; (x &lt; high_bound):\n        return 0.50\n    elif x &lt; low_bound:\n        diff_low = low_bound - x\n        scaled_diff = diff_low ** 0.75\n        final_x = low_bound - scaled_diff\n        return np.max([0, final_x])\n    else:\n        diff_high = x - high_bound\n        scaled_diff = diff_high ** 0.75\n        final_x = high_bound + scaled_diff\n        return np.min([1, final_x])\n</code></pre>",
      "votes": 1,
      "replies": [
        {
          "id": 1974590,
          "author_name": "Theo Viel",
          "author_url": "",
          "post_date": "2022-10-06T10:34:23.257000",
          "content": "<pre><code>def scale(preds, min_=0.2, max_=0.8):\n    preds = (preds - preds.min()) / (preds.max() - preds.min())\n    preds = preds * (max_ - min_) + min_\n    return preds\n</code></pre>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1974601,
          "author_name": "Joe Marturano",
          "author_url": "",
          "post_date": "2022-10-06T10:53:40.070000",
          "content": "<p>Neat, thanks for the secret sauce. Clipping like this seems worth trying for future competitions with log loss.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1973952,
      "author_name": "tdiceman",
      "author_url": "",
      "post_date": "2022-10-06T01:25:40.060000",
      "content": "<p><a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> point 2) was a good call, and now makes a lot of sense… I did the same after some of us had discussions on the <a href=\"https://www.kaggle.com/competitions/mayo-clinic-strip-ai/discussion/354911\" target=\"_blank\">other thread</a> about scores being random… I have to thank you guys, that gave me confidence to select the notebook that scored best for me  :)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1974440,
          "author_name": "Theo Viel",
          "author_url": "",
          "post_date": "2022-10-06T08:50:53.043000",
          "content": "<p>There was indeed some hints in the forum :)<br>\nCongratz on 7th place !</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1975007,
          "author_name": "tdiceman",
          "author_url": "",
          "post_date": "2022-10-06T15:18:42.593000",
          "content": "<p>Thanks and congrats to you too :) </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1973944,
      "author_name": "Pierre Tisseur",
      "author_url": "",
      "post_date": "2022-10-06T01:07:26.967000",
      "content": "<p>Did you use non resized tiles to train and predict?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1974437,
          "author_name": "Theo Viel",
          "author_url": "",
          "post_date": "2022-10-06T08:49:53.777000",
          "content": "<p>I'll post a more detailed write-up soon :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1973920": "Although public LB was random (20 samples is not enough), you could actually get fairly consistent results.\n\nHere are a few ideas, I will work on a more detailed write-up tomorrow :\n\n1) Optimize the AUC, then you will know whether your model discriminates or not\nMy models achieve AUCs around 0.67, which is quite low but not random.\n\n2) Understand the metric : the log loss heavily penalizes mistakes when the model is confident and the reward for getting a correct guess in comparison is much lower. \nSince our models are not so good, you want to stay on the sweet spot where your loss does not get heavily increased because of the mistakes your model makes. \n\nI did so by scaling predictions to the [0.15, 0.85] range, and then clipping them to [0.25, 0.75]. This was tweaked on CV, my best private achieved a 0.64 CV.\n\nSure it's only 0.05 lower than random predictions but that's probably close to the lowest you could get with the provided data :)",
    "1974769": "Congrats on your finish! @theoviel\nYour insights and discussions on the metric and its behavior helped me greatly! 😄",
    "1974568": "Congratulations. Could you elaborate on how you scaled predictions to the [0.15, 0.85] range? That seems to potentially be useful for other future competitions. I tried the following non-linear stretching idea, which based on the private leaderboard, didn't help:\n\n```\ndef stretch_prediction(x):\n    \"\"\" Based on input, stretch prediction to upper or lower\n    \"\"\"\n    low_bound = 0.49\n    high_bound = 0.51\n    if (x > low_bound) & (x < high_bound):\n        return 0.50\n    elif x < low_bound:\n        diff_low = low_bound - x\n        scaled_diff = diff_low ** 0.75\n        final_x = low_bound - scaled_diff\n        return np.max([0, final_x])\n    else:\n        diff_high = x - high_bound\n        scaled_diff = diff_high ** 0.75\n        final_x = high_bound + scaled_diff\n        return np.min([1, final_x])\n```",
    "1973952": "@theoviel point 2) was a good call, and now makes a lot of sense... I did the same after some of us had discussions on the [other thread](https://www.kaggle.com/competitions/mayo-clinic-strip-ai/discussion/354911) about scores being random... I have to thank you guys, that gave me confidence to select the notebook that scored best for me  :)",
    "1973944": "Did you use non resized tiles to train and predict?"
  }
}