{
  "id": 354911,
  "title": "Why do my models have lower performances than sample submission? it's a weird competition",
  "url": "/competitions/mayo-clinic-strip-ai/discussion/354911",
  "author_name": "kaggler",
  "post_date": "2022-09-24T13:33:33.223000",
  "votes": 8,
  "comment_count": 45,
  "views": 0,
  "content": "<p>Why do my models have lower performances than sample submission? it's a weird competition  <br>\nin local environment, anyway it learns the pattern.  <br>\nbut when i submit it i just score 0.6 or 0.7  </p>\n<h2>it's weird!!  </h2>\n<p>Thanks for the all the discussions! Good Luck For Everyone Here!  </p>",
  "messages": [
    {
      "id": 1953430,
      "postDate": "2022-09-24T13:33:33.223Z",
      "content": "<p>Why do my models have lower performances than sample submission? it's a weird competition  <br>\nin local environment, anyway it learns the pattern.  <br>\nbut when i submit it i just score 0.6 or 0.7  </p>\n<h2>it's weird!!  </h2>\n<p>Thanks for the all the discussions! Good Luck For Everyone Here!  </p>",
      "rawMarkdown": "Why do my models have lower performances than sample submission? it's a weird competition  \nin local environment, anyway it learns the pattern.  \nbut when i submit it i just score 0.6 or 0.7  \nit's weird!!  \n-----------------------------------------------------------------------------\n  Thanks for the all the discussions! Good Luck For Everyone Here!  \n\n",
      "votes": 8
    },
    {
      "id": 1970621,
      "postDate": "2022-10-04T07:02:38.270Z",
      "content": "<p>I am really not sure what is going on in this competition, I just started three days ago !  I did a basic training and got 0.4 at first shot. When I trained with \"focus\" and deeply I got a score of 1.0 and another try score of 0.6!  I didn't really do much and will most likely go down sliding in the private board ! but I don't get my results ! … my 0.4 is a lucky shot maybe ! </p>",
      "rawMarkdown": "I am really not sure what is going on in this competition, I just started three days ago !  I did a basic training and got 0.4 at first shot. When I trained with \"focus\" and deeply I got a score of 1.0 and another try score of 0.6!  I didn't really do much and will most likely go down sliding in the private board ! but I don't get my results ! ... my 0.4 is a lucky shot maybe ! ",
      "votes": 4,
      "replies": [
        {
          "id": 1970627,
          "postDate": "2022-10-04T07:13:03.680Z",
          "content": "<p>Can you share some details like what are your oof scores, what model are you using, tiles or whole image etc.? </p>",
          "rawMarkdown": "Can you share some details like what are your oof scores, what model are you using, tiles or whole image etc.? "
        },
        {
          "id": 1970747,
          "postDate": "2022-10-04T09:08:25.013Z",
          "content": "<p>There isn't much signal, as I originally thought. The 0.4 is likely a lucky shot, unfortunately :(</p>",
          "rawMarkdown": "There isn't much signal, as I originally thought. The 0.4 is likely a lucky shot, unfortunately :(",
          "votes": 2
        },
        {
          "id": 1971406,
          "postDate": "2022-10-04T15:28:52.507Z",
          "content": "<p><a href=\"https://www.kaggle.com/asalhi\" target=\"_blank\">@asalhi</a> thanks for sharing this. All the best.</p>",
          "rawMarkdown": "@asalhi thanks for sharing this. All the best."
        },
        {
          "id": 1971450,
          "postDate": "2022-10-04T16:09:08.207Z",
          "content": "<p><a href=\"https://www.kaggle.com/asalhi\" target=\"_blank\">@asalhi</a> please tell us more details about lucky shot)</p>",
          "rawMarkdown": "@asalhi please tell us more details about lucky shot)"
        },
        {
          "id": 1971471,
          "postDate": "2022-10-04T16:23:17.820Z",
          "content": "<p>I prefer not to go into details now, till the competition ends. But I am pretty sure it's a really super lucky shot that I didn't see coming, it will surely not replace hard efforts done by you and others who surely spent hours on this competition, I just passed by a few days ago and had some free time and was struggling in another competition so I said let's try this!  </p>\n<p>I'm pretty sure my lucky shot will shake badly! since my new tests are far from lucky. I would say. However I will choose it as one of the final submissions, I will share how to regenerate this lucky shot once the competition ends, untill then best of luck to you all! I am just “a passing by guy” in this competition … </p>",
          "rawMarkdown": "I prefer not to go into details now, till the competition ends. But I am pretty sure it's a really super lucky shot that I didn't see coming, it will surely not replace hard efforts done by you and others who surely spent hours on this competition, I just passed by a few days ago and had some free time and was struggling in another competition so I said let's try this!  \n\nI'm pretty sure my lucky shot will shake badly! since my new tests are far from lucky. I would say. However I will choose it as one of the final submissions, I will share how to regenerate this lucky shot once the competition ends, untill then best of luck to you all! I am just “a passing by guy” in this competition ... \n",
          "votes": 4
        },
        {
          "id": 1971487,
          "postDate": "2022-10-04T16:34:04.507Z",
          "content": "<p><a href=\"https://www.kaggle.com/asalhi\" target=\"_blank\">@asalhi</a> I'm sure we'll all learn a lot by the time we are done analyzing  solutions for this competition. Looking forward to understand what worked, and also how can we go wrong in such cases. Hope you do not slide in the rankings.</p>",
          "rawMarkdown": "@asalhi I'm sure we'll all learn a lot by the time we are done analyzing  solutions for this competition. Looking forward to understand what worked, and also how can we go wrong in such cases. Hope you do not slide in the rankings.",
          "votes": 1
        },
        {
          "id": 1971498,
          "postDate": "2022-10-04T16:40:24.823Z",
          "content": "<p>I hope not to go sliding, However  I have a feeling that it will be a rollercoaster hhhh</p>",
          "rawMarkdown": "I hope not to go sliding, However  I have a feeling that it will be a rollercoaster hhhh",
          "votes": 1
        }
      ]
    },
    {
      "id": 1972166,
      "postDate": "2022-10-05T03:05:15.953Z",
      "content": "<p>Well, I still think that anything that produces score less than 0.6-0.5 is a product of a mere luck .</p>",
      "rawMarkdown": "Well, I still think that anything that produces score less than 0.6-0.5 is a product of a mere luck ~~or overfitting that happens in such specific conditions as a lack of signal in the data~~.",
      "votes": 1
    },
    {
      "id": 1971445,
      "postDate": "2022-10-04T16:04:57.503Z",
      "content": "<p>Any last minute guesses what the medal range will be for private LB?</p>\n<p>I think low but possible chance sample submission may make it into bronze lol</p>\n<p>Maybe 0.5 Gold to 0.7 Bronze?</p>",
      "rawMarkdown": "Any last minute guesses what the medal range will be for private LB?\n\nI think low but possible chance sample submission may make it into bronze lol\n\nMaybe 0.5 Gold to 0.7 Bronze?",
      "votes": 1,
      "replies": [
        {
          "id": 1971628,
          "postDate": "2022-10-04T17:44:11.163Z",
          "content": "<p>we can not know about it as discussed above</p>",
          "rawMarkdown": "we can not know about it as discussed above"
        },
        {
          "id": 1971713,
          "postDate": "2022-10-04T18:21:29.543Z",
          "content": "<p>I think private leaderboard will be between 0.63 and 0.75. 0.63 is the gold zone, 0.64 is silver.</p>",
          "rawMarkdown": "I think private leaderboard will be between 0.63 and 0.75. 0.63 is the gold zone, 0.64 is silver.",
          "votes": 4
        }
      ]
    },
    {
      "id": 1965228,
      "postDate": "2022-10-01T07:44:58.250Z",
      "content": "<pre><code>def binary_weighted_log_loss_forcewithme(y_true, y_pred):\n\n    log_loss_positive = log_loss(y_true, y_pred)\n    log_loss_negative = log_loss(y_true, 1 - y_pred)\n    log_loss_weighted = 0.5 * log_loss_positive + 0.5 * log_loss_negative\n\n    return log_loss_positive, log_loss_negative, log_loss_weighted\n\n\ndef binary_weighted_log_loss_theo(y_true, y_pred):\n\n    y_true_positive = y_true == 1\n    y_true_negative = y_true == 0\n\n    log_loss_positive = log_loss(y_true[y_true_positive], y_pred[y_true_positive], labels=[0, 1])\n    log_loss_negative = log_loss(y_true[y_true_negative], y_pred[y_true_negative], labels=[0, 1])\n    log_loss_weighted = 0.5 * log_loss_positive + 0.5 * log_loss_negative\n\n    return log_loss_positive, log_loss_negative, log_loss_weighted\n</code></pre>\n<p>These two functions yield different results. This is my final blend score. I have no idea which one is the correct implementation so I use both of them for now.</p>\n<p><code>Blend - {'accuracy': 0.6286472148541115, 'roc_auc': 0.66280723136299, 'log_loss_positive_forcewithme': 0.6593462375631115, 'log_loss_negative_forcewithme': 0.7662855417748091, 'log_loss_weighted_forcewithme': 0.7128158896689603, 'log_loss_positive_theo': 0.6395257262674887, 'log_loss_negative_theo': 0.6668468698084385, 'log_loss_weighted_theo': 0.6531862980379637} - Predictions Mean: 0.4952 Std: 0.0957 Min: 0.2987 Max: 0.8416</code></p>",
      "rawMarkdown": "```\ndef binary_weighted_log_loss_forcewithme(y_true, y_pred):\n\n    log_loss_positive = log_loss(y_true, y_pred)\n    log_loss_negative = log_loss(y_true, 1 - y_pred)\n    log_loss_weighted = 0.5 * log_loss_positive + 0.5 * log_loss_negative\n\n    return log_loss_positive, log_loss_negative, log_loss_weighted\n\n\ndef binary_weighted_log_loss_theo(y_true, y_pred):\n\n    y_true_positive = y_true == 1\n    y_true_negative = y_true == 0\n\n    log_loss_positive = log_loss(y_true[y_true_positive], y_pred[y_true_positive], labels=[0, 1])\n    log_loss_negative = log_loss(y_true[y_true_negative], y_pred[y_true_negative], labels=[0, 1])\n    log_loss_weighted = 0.5 * log_loss_positive + 0.5 * log_loss_negative\n\n    return log_loss_positive, log_loss_negative, log_loss_weighted\n```\n\nThese two functions yield different results. This is my final blend score. I have no idea which one is the correct implementation so I use both of them for now.\n\n` Blend - {'accuracy': 0.6286472148541115, 'roc_auc': 0.66280723136299, 'log_loss_positive_forcewithme': 0.6593462375631115, 'log_loss_negative_forcewithme': 0.7662855417748091, 'log_loss_weighted_forcewithme': 0.7128158896689603, 'log_loss_positive_theo': 0.6395257262674887, 'log_loss_negative_theo': 0.6668468698084385, 'log_loss_weighted_theo': 0.6531862980379637} - Predictions Mean: 0.4952 Std: 0.0957 Min: 0.2987 Max: 0.8416`",
      "votes": 1,
      "replies": [
        {
          "id": 1965330,
          "postDate": "2022-10-01T09:08:16.457Z",
          "content": "<p>My final blend scores AUC 0.68 / log_loss_weighted_theo 0.64  but I'm afraid I am overfitting my validation scores and this wont translate to private :/</p>",
          "rawMarkdown": "My final blend scores AUC 0.68 / log_loss_weighted_theo 0.64  but I'm afraid I am overfitting my validation scores and this wont translate to private :/",
          "votes": 1
        },
        {
          "id": 1965800,
          "postDate": "2022-10-01T14:37:01.587Z",
          "content": "<p>0.68 AUC is really good. My AUC becomes 0.6725 when I aggregate predictions on patient_id groups. I guess my submissions were kinda consistent with my oof scores. I have no idea how people are reaching below 0.5.</p>",
          "rawMarkdown": "0.68 AUC is really good. My AUC becomes 0.6725 when I aggregate predictions on patient_id groups. I guess my submissions were kinda consistent with my oof scores. I have no idea how people are reaching below 0.5.",
          "votes": 1
        },
        {
          "id": 1965875,
          "postDate": "2022-10-01T15:29:23.657Z",
          "content": "<p>Scores below 0.5 might very well be models overfit for the test set, created by tuning the model on the test set score until it went down. No clue if that's what really happened, but so few people got a score below 0.5 that it makes it seem lowkey overfit on that 7% of data points…</p>",
          "rawMarkdown": "Scores below 0.5 might very well be models overfit for the test set, created by tuning the model on the test set score until it went down. No clue if that's what really happened, but so few people got a score below 0.5 that it makes it seem lowkey overfit on that 7% of data points...",
          "votes": 1
        },
        {
          "id": 1966991,
          "postDate": "2022-10-02T09:21:26.690Z",
          "content": "<p>If I compute my logloss on 20 randomly sampled images from my validation set 1000 times, I get 10% of \"lucky\" runs with a score of 0.4 or lower. </p>\n<p>If you take into account that people also use LB feedback to tweak their models then I'm not really surprised. I will be however if private LBs of 0.5 (or lower) are reached, it's much harder to get a lucky draw with 260 samples.</p>",
          "rawMarkdown": "If I compute my logloss on 20 randomly sampled images from my validation set 1000 times, I get 10% of \"lucky\" runs with a score of 0.4 or lower. \n\nIf you take into account that people also use LB feedback to tweak their models then I'm not really surprised. I will be however if private LBs of 0.5 (or lower) are reached, it's much harder to get a lucky draw with 260 samples.",
          "votes": 3
        },
        {
          "id": 1967022,
          "postDate": "2022-10-02T09:40:26.783Z",
          "content": "<p>Exactly. Not discounting the work from the current top of the LB, but I think it's more likely random chance than actually capturing signal. As a last attempt, I tried doing weakly supervised multiple instance learning with attention to which image in the bag made the bag positive (LAA in my case) - and there was basically nothing by the end.</p>\n<p>I wonder how pathologists figure this out…</p>",
          "rawMarkdown": "Exactly. Not discounting the work from the current top of the LB, but I think it's more likely random chance than actually capturing signal. As a last attempt, I tried doing weakly supervised multiple instance learning with attention to which image in the bag made the bag positive (LAA in my case) - and there was basically nothing by the end.\n\nI wonder how pathologists figure this out..."
        },
        {
          "id": 1967686,
          "postDate": "2022-10-02T16:30:03.797Z",
          "content": "<p><a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> I have observed similar spread of log loss over prediction scores, it could very much be the case here.</p>",
          "rawMarkdown": "@theoviel I have observed similar spread of log loss over prediction scores, it could very much be the case here."
        },
        {
          "id": 1970384,
          "postDate": "2022-10-04T04:53:47.863Z",
          "content": "<p>if my process doesn't deviate from the right, weighted CV scores lower than 0.5 are a very rare case</p>",
          "rawMarkdown": "if my process doesn't deviate from the right, weighted CV scores lower than 0.5 are a very rare case",
          "votes": 2
        }
      ]
    },
    {
      "id": 1953465,
      "postDate": "2022-09-24T13:56:28.853Z",
      "content": "<p>Same here. It's because the public test set is quite small. My cv score is 0.564 but my lb score is 0.8.</p>",
      "rawMarkdown": "Same here. It's because the public test set is quite small. My cv score is 0.564 but my lb score is 0.8.",
      "votes": 1,
      "replies": [
        {
          "id": 1953501,
          "postDate": "2022-09-24T14:24:13.330Z",
          "content": "<p>Nice to see you here again after the hubmap finishes!<br>\nMy 5fold cv scores are around 0.65 anyway it's lower than the sample submission. Thanks for replying!</p>",
          "rawMarkdown": "Nice to see you here again after the hubmap finishes!\nMy 5fold cv scores are around 0.65 anyway it's lower than the sample submission. Thanks for replying!",
          "votes": 1
        },
        {
          "id": 1955494,
          "postDate": "2022-09-26T03:30:03.643Z",
          "content": "<p>Are you sure you are using the right metrics? A 0.564 cv seems very high in this competition.</p>",
          "rawMarkdown": "Are you sure you are using the right metrics? A 0.564 cv seems very high in this competition.",
          "votes": 1
        },
        {
          "id": 1955541,
          "postDate": "2022-09-26T04:00:06.670Z",
          "content": "<p><a href=\"https://www.kaggle.com/forcewithme\" target=\"_blank\">@forcewithme</a> may i ask you what your cv scores are around?</p>",
          "rawMarkdown": "@forcewithme may i ask you what your cv scores are around?"
        },
        {
          "id": 1955556,
          "postDate": "2022-09-26T04:05:53.180Z",
          "content": "<p>I use <code>sklearn.metrics.log_loss</code>. I tried it with both binary and multiclass predictions and the result was same. My current cv score is 0.5598 right now. Am I missing something and overfitting like crazy?</p>",
          "rawMarkdown": "I use `sklearn.metrics.log_loss`. I tried it with both binary and multiclass predictions and the result was same. My current cv score is 0.5598 right now. Am I missing something and overfitting like crazy?",
          "votes": 1
        },
        {
          "id": 1955629,
          "postDate": "2022-09-26T04:40:17.657Z",
          "content": "<p><a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> don't you consider applying weighted log loss?</p>",
          "rawMarkdown": "@gunesevitan don't you consider applying weighted log loss?"
        },
        {
          "id": 1955654,
          "postDate": "2022-09-26T04:54:14.097Z",
          "content": "<p>I didn't use weighted log loss since organizers said that each class is roughly equally important. I expect 0.02 deviation from the regular log loss.</p>",
          "rawMarkdown": "I didn't use weighted log loss since organizers said that each class is roughly equally important. I expect 0.02 deviation from the regular log loss.",
          "votes": 2
        },
        {
          "id": 1955765,
          "postDate": "2022-09-26T06:13:05.793Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a>  <a href=\"https://www.kaggle.com/deepkim\" target=\"_blank\">@deepkim</a> </p>\n<ol>\n<li>I implemented the metric of this competition myself. The code in this <a href=\"https://www.kaggle.com/competitions/mayo-clinic-strip-ai/discussion/353305\" target=\"_blank\">post</a> can get the same result as me.</li>\n<li>Use <code>sklearn.metrics.log_loss</code> to get the score of each class, and average them. For example, with  <code>sklearn.metrics.log_loss</code>, you get 0.5 for CE, get 0.8 for LAA, then your score should be 0.65.</li>\n</ol>\n<p>===================================================<br>\nAbout the weight of each class. I am pretty sure that the weight of each class should be 0.5:0.5. Because the score should be 17.2 when all the predicted ce/laa probability are 0/1(1/0), and 0.6+(actually 0.69) when all the predicted ce/laa probability are 0.5 /0.5. Both scores(17.2 0.6) are verified. </p>",
          "rawMarkdown": "Hi @gunesevitan  @deepkim \n1. I implemented the metric of this competition myself. The code in this [post](https://www.kaggle.com/competitions/mayo-clinic-strip-ai/discussion/353305) can get the same result as me.\n2. Use `sklearn.metrics.log_loss` to get the score of each class, and average them. For example, with  `sklearn.metrics.log_loss`, you get 0.5 for CE, get 0.8 for LAA, then your score should be 0.65.\n\n===================================================\nAbout the weight of each class. I am pretty sure that the weight of each class should be 0.5:0.5. Because the score should be 17.2 when all the predicted ce/laa probability are 0/1(1/0), and 0.6+(actually 0.69) when all the predicted ce/laa probability are 0.5 /0.5. Both scores(17.2 0.6) are verified. ",
          "votes": 3
        },
        {
          "id": 1955781,
          "postDate": "2022-09-26T06:21:07.330Z",
          "content": "<p>So is it like macro log loss? I'm currently using log loss on binary sigmoided logits. I guess I have to create another dimension with <code>1 - y</code> and do <code>mean(log_loss(y_true, y_pred), log_loss(1 - y_true, 1 - y_pred))</code>. It doesn't make much sense though. I think it would have the same score.</p>",
          "rawMarkdown": "So is it like macro log loss? I'm currently using log loss on binary sigmoided logits. I guess I have to create another dimension with `1 - y` and do `mean(log_loss(y_true, y_pred), log_loss(1 - y_true, 1 - y_pred))`. It doesn't make much sense though. I think it would have the same score.",
          "votes": 1
        },
        {
          "id": 1955787,
          "postDate": "2022-09-26T06:23:23.867Z",
          "content": "<p>Yeah, I always created <code>1-y</code> for <code>class 0</code></p>",
          "rawMarkdown": "Yeah, I always created `1-y` for `class 0`",
          "votes": 1
        },
        {
          "id": 1955793,
          "postDate": "2022-09-26T06:25:15.297Z",
          "content": "<p><a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> In my opinion, <br>\n<code>mean(log_loss(y_true, y_pred), log_loss(1 - y_true, 1 - y_pred))</code>  is wrong.<br>\n<code>mean(log_loss(y_true, y_pred), log_loss(y_true, 1 - y_pred))</code> is right.<br>\nPlease correct me if you find something wrong. </p>",
          "rawMarkdown": "@gunesevitan In my opinion, \n`mean(log_loss(y_true, y_pred), log_loss(1 - y_true, 1 - y_pred))`  is wrong.\n`mean(log_loss(y_true, y_pred), log_loss(y_true, 1 - y_pred))` is right.\nPlease correct me if you find something wrong. ",
          "votes": 2
        },
        {
          "id": 1955811,
          "postDate": "2022-09-26T06:32:34.227Z",
          "content": "<p>I created two vectors and their names are self-explanatory…</p>\n<pre><code>y_true = np.random.randint(0, 2, 100)\ny_pred = np.random.rand(100)\n</code></pre>\n<p>As I thought, these two yield same scores</p>\n<pre><code>&gt;&gt;&gt; log_loss(y_true, y_pred)\n0.9869673142342817\n&gt;&gt;&gt; log_loss(1 - y_true, 1 - y_pred)\n0.9869673142342817\n</code></pre>\n<p>but this one yields a slightly different score</p>\n<pre><code>&gt;&gt;&gt; log_loss(y_true, 1 - y_pred)\n0.9635681791900268\n</code></pre>\n<p>and the average of them is</p>\n<pre><code>&gt;&gt;&gt; np.mean([log_loss(y_true, y_pred), log_loss(y_true, 1 - y_pred)])\n0.9752677467121542\n</code></pre>\n<p>so the deviation is even less than 0.02. However that is totally dependent to predictions. I don't think macro and micro log loss scores wouldn't be too much different from each other since both class predictions are dependent to each other.</p>",
          "rawMarkdown": "I created two vectors and their names are self-explanatory...\n\n```\ny_true = np.random.randint(0, 2, 100)\ny_pred = np.random.rand(100)\n```\n\nAs I thought, these two yield same scores\n\n```\n>>> log_loss(y_true, y_pred)\n0.9869673142342817\n>>> log_loss(1 - y_true, 1 - y_pred)\n0.9869673142342817\n```\n\nbut this one yields a slightly different score\n\n```\n>>> log_loss(y_true, 1 - y_pred)\n0.9635681791900268\n```\n\nand the average of them is\n```\n\n>>> np.mean([log_loss(y_true, y_pred), log_loss(y_true, 1 - y_pred)])\n0.9752677467121542\n```\n\nso the deviation is even less than 0.02. However that is totally dependent to predictions. I don't think macro and micro log loss scores wouldn't be too much different from each other since both class predictions are dependent to each other.",
          "votes": 2
        },
        {
          "id": 1956090,
          "postDate": "2022-09-26T09:20:40.017Z",
          "content": "<p>Here's the code I use</p>\n<pre><code># Group by patient\ndfg = df[[\"patient_id\", \"pred\", \"target\"]].groupby(\"patient_id\").mean()\n\n# Separate the two classes\ndf1 = dfg[dfg['target'] == 1]\ndf0 = dfg[dfg['target'] == 0]\n\n# Compute the losses and average\nl0 = log_loss(df0[\"target\"], df0[\"pred\"], labels=[0, 1])\nl1 = log_loss(df1[\"target\"], df1[\"pred\"], labels=[0, 1])\nloss = 0.5 * (l0 + l1)\n\nprint(f\"\\nCV loss : {loss:.3f}  ({l0:.3f}, {l1:.3f})\")\n\n# For comparison, this is the binary logloss\nloss = log_loss(dfg[\"target\"], dfg[\"pred\"])\nprint(f\"\\nlogloss : {loss:.3f}\")\n</code></pre>\n<p>I can get 0.55 logloss but my weighted logloss scores are around 0.65.<br>\nThe binary logloss benefits from the class imbalance hence if I have a model that predicts more CE the loss will be lower.</p>\n<p><a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> What is your CV AUC though ? That's what I'm optimizing on my side =)</p>",
          "rawMarkdown": "Here's the code I use\n``` \n# Group by patient\ndfg = df[[\"patient_id\", \"pred\", \"target\"]].groupby(\"patient_id\").mean()\n\n# Separate the two classes\ndf1 = dfg[dfg['target'] == 1]\ndf0 = dfg[dfg['target'] == 0]\n\n# Compute the losses and average\nl0 = log_loss(df0[\"target\"], df0[\"pred\"], labels=[0, 1])\nl1 = log_loss(df1[\"target\"], df1[\"pred\"], labels=[0, 1])\nloss = 0.5 * (l0 + l1)\n\nprint(f\"\\nCV loss : {loss:.3f}  ({l0:.3f}, {l1:.3f})\")\n\n# For comparison, this is the binary logloss\nloss = log_loss(dfg[\"target\"], dfg[\"pred\"])\nprint(f\"\\nlogloss : {loss:.3f}\")\n``` \n\nI can get 0.55 logloss but my weighted logloss scores are around 0.65.\nThe binary logloss benefits from the class imbalance hence if I have a model that predicts more CE the loss will be lower.\n\n@gunesevitan What is your CV AUC though ? That's what I'm optimizing on my side =)",
          "votes": 2
        },
        {
          "id": 1956117,
          "postDate": "2022-09-26T09:39:06.200Z",
          "content": "<p>That makes sense… My results were close to each other because dummy vectors had normal distribution. AUC seems to be a good auxiliary metric. These are my fold scores. I use stratified group kfold, splits are stratified on label and groups are patient_id. </p>\n<pre><code>{\n  \"fold_scores\": {\n    \"fold1\": {\n      \"accuracy\": 0.7450980392156863,\n      \"roc_auc\": 0.6655420602789024,\n      \"log_loss\": 0.5360595129819867\n    },\n    \"fold2\": {\n      \"accuracy\": 0.6948051948051948,\n      \"roc_auc\": 0.5798369457148538,\n      \"log_loss\": 0.6151186826747733\n    },\n    \"fold3\": {\n      \"accuracy\": 0.7635135135135135,\n      \"roc_auc\": 0.6552579365079365,\n      \"log_loss\": 0.5286322120778464\n    },\n    \"fold4\": {\n      \"accuracy\": 0.74,\n      \"roc_auc\": 0.6325799955247259,\n      \"log_loss\": 0.5594933565209309\n    },\n    \"fold5\": {\n      \"accuracy\": 0.7181208053691275,\n      \"roc_auc\": 0.7004329004329004,\n      \"log_loss\": 0.5591903770679996\n    }\n  },\n  \"oof_scores\": {\n    \"accuracy\": 0.7320954907161804,\n    \"roc_auc\": 0.6483851310176723,\n    \"log_loss\": 0.5599818563222174\n  }\n}\n</code></pre>",
          "rawMarkdown": "That makes sense... My results were close to each other because dummy vectors had normal distribution. AUC seems to be a good auxiliary metric. These are my fold scores. I use stratified group kfold, splits are stratified on label and groups are patient_id. \n\n```\n{\n  \"fold_scores\": {\n    \"fold1\": {\n      \"accuracy\": 0.7450980392156863,\n      \"roc_auc\": 0.6655420602789024,\n      \"log_loss\": 0.5360595129819867\n    },\n    \"fold2\": {\n      \"accuracy\": 0.6948051948051948,\n      \"roc_auc\": 0.5798369457148538,\n      \"log_loss\": 0.6151186826747733\n    },\n    \"fold3\": {\n      \"accuracy\": 0.7635135135135135,\n      \"roc_auc\": 0.6552579365079365,\n      \"log_loss\": 0.5286322120778464\n    },\n    \"fold4\": {\n      \"accuracy\": 0.74,\n      \"roc_auc\": 0.6325799955247259,\n      \"log_loss\": 0.5594933565209309\n    },\n    \"fold5\": {\n      \"accuracy\": 0.7181208053691275,\n      \"roc_auc\": 0.7004329004329004,\n      \"log_loss\": 0.5591903770679996\n    }\n  },\n  \"oof_scores\": {\n    \"accuracy\": 0.7320954907161804,\n    \"roc_auc\": 0.6483851310176723,\n    \"log_loss\": 0.5599818563222174\n  }\n}\n```",
          "votes": 3
        },
        {
          "id": 1956187,
          "postDate": "2022-09-26T10:32:30.100Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a>, I can confirm I have similary AUCs (~0.65) using the same splitting strategy</p>",
          "rawMarkdown": "Thanks @gunesevitan, I can confirm I have similary AUCs (~0.65) using the same splitting strategy",
          "votes": 2
        },
        {
          "id": 1956204,
          "postDate": "2022-09-26T10:43:04.443Z",
          "content": "<p>my weighted log loss converges to 0.65 but in terms of accuracy, lower than <a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> </p>",
          "rawMarkdown": "my weighted log loss converges to 0.65 but in terms of accuracy, lower than @gunesevitan "
        },
        {
          "id": 1956211,
          "postDate": "2022-09-26T10:50:29.147Z",
          "content": "<p>After implementing <a href=\"https://www.kaggle.com/forcewithme\" target=\"_blank\">@forcewithme</a>'s metric as a loss function, my model didn't learn. Weighted log loss and regular log loss starts at 0.69 and finishes like that. Accuracy and ROC AUC randomly oscillates. This is what I'm using right now.</p>\n<pre><code>class MacroBCEWithLogitsLoss(_WeightedLoss):\n\n    def __init__(self, weight=None, reduction='mean'):\n\n        super(MacroBCEWithLogitsLoss, self).__init__(weight=weight, reduction=reduction)\n\n        self.weight = weight\n        self.reduction = reduction\n\n    def forward(self, inputs, targets):\n\n        inputs = torch.sigmoid(inputs)\n        loss_positive = F.binary_cross_entropy(inputs, targets, self.weight)\n        loss_negative = F.binary_cross_entropy(1 - inputs, targets, self.weight)\n        loss = (0.5 * loss_positive) + (0.5 * loss_negative)\n\n        if self.reduction == 'mean':\n            loss = loss.mean()\n        elif self.reduction == 'sum':\n            loss = loss.sum()\n\n        return loss\n</code></pre>\n<p>Edit: Optimizing this loss function didn't work for me. After one epoch, model learns to make safe predictions between 0.48-0.52, and loss doesn't improve after 0.69. This behavior continues when I tweak the weights.</p>",
          "rawMarkdown": "After implementing @forcewithme's metric as a loss function, my model didn't learn. Weighted log loss and regular log loss starts at 0.69 and finishes like that. Accuracy and ROC AUC randomly oscillates. This is what I'm using right now.\n\n```\nclass MacroBCEWithLogitsLoss(_WeightedLoss):\n\n    def __init__(self, weight=None, reduction='mean'):\n\n        super(MacroBCEWithLogitsLoss, self).__init__(weight=weight, reduction=reduction)\n\n        self.weight = weight\n        self.reduction = reduction\n\n    def forward(self, inputs, targets):\n\n        inputs = torch.sigmoid(inputs)\n        loss_positive = F.binary_cross_entropy(inputs, targets, self.weight)\n        loss_negative = F.binary_cross_entropy(1 - inputs, targets, self.weight)\n        loss = (0.5 * loss_positive) + (0.5 * loss_negative)\n\n        if self.reduction == 'mean':\n            loss = loss.mean()\n        elif self.reduction == 'sum':\n            loss = loss.sum()\n\n        return loss\n\n```\n\nEdit: Optimizing this loss function didn't work for me. After one epoch, model learns to make safe predictions between 0.48-0.52, and loss doesn't improve after 0.69. This behavior continues when I tweak the weights.",
          "votes": 2
        },
        {
          "id": 1956287,
          "postDate": "2022-09-26T11:24:12.887Z",
          "content": "<p>I only used the competition metric for validation and save weights. And I can't get a significantly better score than <code>sample_submission.csv</code> 🤕</p>",
          "rawMarkdown": "I only used the competition metric for validation and save weights. And I can't get a significantly better score than `sample_submission.csv` 🤕",
          "votes": 1
        },
        {
          "id": 1956774,
          "postDate": "2022-09-26T15:38:35.550Z",
          "content": "<p>I believe public LB score is only on 20 images out of 280, there is good chance that private LB will look very different. Plus do also check the spread of log loss over the predictions, mean values only tell part of the story. Even if I get a mean log loss of under 0.6, the spread is from 0.2 to 0.8, which hints at some overfitting for sure - I guess this is where this dataset is hard. Also waiting to see top private LB solutions :)</p>",
          "rawMarkdown": "I believe public LB score is only on 20 images out of 280, there is good chance that private LB will look very different. Plus do also check the spread of log loss over the predictions, mean values only tell part of the story. Even if I get a mean log loss of under 0.6, the spread is from 0.2 to 0.8, which hints at some overfitting for sure - I guess this is where this dataset is hard. Also waiting to see top private LB solutions :)",
          "votes": 1
        },
        {
          "id": 1956964,
          "postDate": "2022-09-26T17:36:43.627Z",
          "content": "<p><a href=\"https://www.kaggle.com/icemantd\" target=\"_blank\">@icemantd</a> checking the spread of log loss over the predictions seems really good strategy!!</p>",
          "rawMarkdown": "@icemantd checking the spread of log loss over the predictions seems really good strategy!!"
        }
      ]
    },
    {
      "id": 1970760,
      "postDate": "2022-10-04T09:19:08.317Z",
      "content": "<p>Turns out submitting <code>1 - preds</code> scores better on LB for me.</p>\n<p>This should happen in ~6% of the cases if my LB predictions follow the same distribution as my oof ones, but still I am a bit worried.</p>",
      "rawMarkdown": "Turns out submitting `1 - preds` scores better on LB for me.\n\nThis should happen in ~6% of the cases if my LB predictions follow the same distribution as my oof ones, but still I am a bit worried.\n\n",
      "votes": 2,
      "replies": [
        {
          "id": 1971341,
          "postDate": "2022-10-04T14:39:37.203Z",
          "content": "<p>Same here. I tried also reversing CE and LAA, it gave me better score in LB (0.7 -&gt; 0.6)  <br>\ndoes it imply there is no signal?</p>",
          "rawMarkdown": "Same here. I tried also reversing CE and LAA, it gave me better score in LB (0.7 -> 0.6)  \ndoes it imply there is no signal?",
          "votes": 1
        },
        {
          "id": 1971410,
          "postDate": "2022-10-04T15:30:45.380Z",
          "content": "<p>We'll find out soon enough. Hold on to your two submissions with some popcorn :D</p>",
          "rawMarkdown": "We'll find out soon enough. Hold on to your two submissions with some popcorn :D",
          "votes": 2
        },
        {
          "id": 1972627,
          "postDate": "2022-10-05T08:53:51.617Z",
          "content": "<p>My best submission's AUC is 0.6842 and weighted log loss is 0.6387. It's weird that we are so close to each other in public leaderboard <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a>, but I haven't inverted my predictions.</p>",
          "rawMarkdown": "My best submission's AUC is 0.6842 and weighted log loss is 0.6387. It's weird that we are so close to each other in public leaderboard @theoviel, but I haven't inverted my predictions.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1955382,
      "postDate": "2022-09-26T01:15:27.510Z",
      "content": "<p>I'm looking forward to the \"1st place solution\" post by sample_submission.csv.</p>",
      "rawMarkdown": "I'm looking forward to the \"1st place solution\" post by sample_submission.csv.",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 1970621,
      "author_name": "Ali",
      "author_url": "",
      "post_date": "2022-10-04T07:02:38.270000",
      "content": "<p>I am really not sure what is going on in this competition, I just started three days ago !  I did a basic training and got 0.4 at first shot. When I trained with \"focus\" and deeply I got a score of 1.0 and another try score of 0.6!  I didn't really do much and will most likely go down sliding in the private board ! but I don't get my results ! … my 0.4 is a lucky shot maybe ! </p>",
      "votes": 4,
      "replies": [
        {
          "id": 1970627,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2022-10-04T07:13:03.680000",
          "content": "<p>Can you share some details like what are your oof scores, what model are you using, tiles or whole image etc.? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1970747,
          "author_name": "David Landup",
          "author_url": "",
          "post_date": "2022-10-04T09:08:25.013000",
          "content": "<p>There isn't much signal, as I originally thought. The 0.4 is likely a lucky shot, unfortunately :(</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1971406,
          "author_name": "tdiceman",
          "author_url": "",
          "post_date": "2022-10-04T15:28:52.507000",
          "content": "<p><a href=\"https://www.kaggle.com/asalhi\" target=\"_blank\">@asalhi</a> thanks for sharing this. All the best.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1971450,
          "author_name": "anthony",
          "author_url": "",
          "post_date": "2022-10-04T16:09:08.207000",
          "content": "<p><a href=\"https://www.kaggle.com/asalhi\" target=\"_blank\">@asalhi</a> please tell us more details about lucky shot)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1971471,
          "author_name": "Ali",
          "author_url": "",
          "post_date": "2022-10-04T16:23:17.820000",
          "content": "<p>I prefer not to go into details now, till the competition ends. But I am pretty sure it's a really super lucky shot that I didn't see coming, it will surely not replace hard efforts done by you and others who surely spent hours on this competition, I just passed by a few days ago and had some free time and was struggling in another competition so I said let's try this!  </p>\n<p>I'm pretty sure my lucky shot will shake badly! since my new tests are far from lucky. I would say. However I will choose it as one of the final submissions, I will share how to regenerate this lucky shot once the competition ends, untill then best of luck to you all! I am just “a passing by guy” in this competition … </p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1971487,
          "author_name": "tdiceman",
          "author_url": "",
          "post_date": "2022-10-04T16:34:04.507000",
          "content": "<p><a href=\"https://www.kaggle.com/asalhi\" target=\"_blank\">@asalhi</a> I'm sure we'll all learn a lot by the time we are done analyzing  solutions for this competition. Looking forward to understand what worked, and also how can we go wrong in such cases. Hope you do not slide in the rankings.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1971498,
          "author_name": "Ali",
          "author_url": "",
          "post_date": "2022-10-04T16:40:24.823000",
          "content": "<p>I hope not to go sliding, However  I have a feeling that it will be a rollercoaster hhhh</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1972166,
      "author_name": "majoraregalia",
      "author_url": "",
      "post_date": "2022-10-05T03:05:15.953000",
      "content": "<p>Well, I still think that anything that produces score less than 0.6-0.5 is a product of a mere luck .</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1971445,
      "author_name": "JM",
      "author_url": "",
      "post_date": "2022-10-04T16:04:57.503000",
      "content": "<p>Any last minute guesses what the medal range will be for private LB?</p>\n<p>I think low but possible chance sample submission may make it into bronze lol</p>\n<p>Maybe 0.5 Gold to 0.7 Bronze?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1971628,
          "author_name": "kaggler",
          "author_url": "",
          "post_date": "2022-10-04T17:44:11.163000",
          "content": "<p>we can not know about it as discussed above</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1971713,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2022-10-04T18:21:29.543000",
          "content": "<p>I think private leaderboard will be between 0.63 and 0.75. 0.63 is the gold zone, 0.64 is silver.</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 1965228,
      "author_name": "Gunes Evitan",
      "author_url": "",
      "post_date": "2022-10-01T07:44:58.250000",
      "content": "<pre><code>def binary_weighted_log_loss_forcewithme(y_true, y_pred):\n\n    log_loss_positive = log_loss(y_true, y_pred)\n    log_loss_negative = log_loss(y_true, 1 - y_pred)\n    log_loss_weighted = 0.5 * log_loss_positive + 0.5 * log_loss_negative\n\n    return log_loss_positive, log_loss_negative, log_loss_weighted\n\n\ndef binary_weighted_log_loss_theo(y_true, y_pred):\n\n    y_true_positive = y_true == 1\n    y_true_negative = y_true == 0\n\n    log_loss_positive = log_loss(y_true[y_true_positive], y_pred[y_true_positive], labels=[0, 1])\n    log_loss_negative = log_loss(y_true[y_true_negative], y_pred[y_true_negative], labels=[0, 1])\n    log_loss_weighted = 0.5 * log_loss_positive + 0.5 * log_loss_negative\n\n    return log_loss_positive, log_loss_negative, log_loss_weighted\n</code></pre>\n<p>These two functions yield different results. This is my final blend score. I have no idea which one is the correct implementation so I use both of them for now.</p>\n<p><code>Blend - {'accuracy': 0.6286472148541115, 'roc_auc': 0.66280723136299, 'log_loss_positive_forcewithme': 0.6593462375631115, 'log_loss_negative_forcewithme': 0.7662855417748091, 'log_loss_weighted_forcewithme': 0.7128158896689603, 'log_loss_positive_theo': 0.6395257262674887, 'log_loss_negative_theo': 0.6668468698084385, 'log_loss_weighted_theo': 0.6531862980379637} - Predictions Mean: 0.4952 Std: 0.0957 Min: 0.2987 Max: 0.8416</code></p>",
      "votes": 1,
      "replies": [
        {
          "id": 1965330,
          "author_name": "Theo Viel",
          "author_url": "",
          "post_date": "2022-10-01T09:08:16.457000",
          "content": "<p>My final blend scores AUC 0.68 / log_loss_weighted_theo 0.64  but I'm afraid I am overfitting my validation scores and this wont translate to private :/</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1965800,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2022-10-01T14:37:01.587000",
          "content": "<p>0.68 AUC is really good. My AUC becomes 0.6725 when I aggregate predictions on patient_id groups. I guess my submissions were kinda consistent with my oof scores. I have no idea how people are reaching below 0.5.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1965875,
          "author_name": "David Landup",
          "author_url": "",
          "post_date": "2022-10-01T15:29:23.657000",
          "content": "<p>Scores below 0.5 might very well be models overfit for the test set, created by tuning the model on the test set score until it went down. No clue if that's what really happened, but so few people got a score below 0.5 that it makes it seem lowkey overfit on that 7% of data points…</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1966991,
          "author_name": "Theo Viel",
          "author_url": "",
          "post_date": "2022-10-02T09:21:26.690000",
          "content": "<p>If I compute my logloss on 20 randomly sampled images from my validation set 1000 times, I get 10% of \"lucky\" runs with a score of 0.4 or lower. </p>\n<p>If you take into account that people also use LB feedback to tweak their models then I'm not really surprised. I will be however if private LBs of 0.5 (or lower) are reached, it's much harder to get a lucky draw with 260 samples.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1967022,
          "author_name": "David Landup",
          "author_url": "",
          "post_date": "2022-10-02T09:40:26.783000",
          "content": "<p>Exactly. Not discounting the work from the current top of the LB, but I think it's more likely random chance than actually capturing signal. As a last attempt, I tried doing weakly supervised multiple instance learning with attention to which image in the bag made the bag positive (LAA in my case) - and there was basically nothing by the end.</p>\n<p>I wonder how pathologists figure this out…</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1967686,
          "author_name": "tdiceman",
          "author_url": "",
          "post_date": "2022-10-02T16:30:03.797000",
          "content": "<p><a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> I have observed similar spread of log loss over prediction scores, it could very much be the case here.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1970384,
          "author_name": "kaggler",
          "author_url": "",
          "post_date": "2022-10-04T04:53:47.863000",
          "content": "<p>if my process doesn't deviate from the right, weighted CV scores lower than 0.5 are a very rare case</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1953465,
      "author_name": "Gunes Evitan",
      "author_url": "",
      "post_date": "2022-09-24T13:56:28.853000",
      "content": "<p>Same here. It's because the public test set is quite small. My cv score is 0.564 but my lb score is 0.8.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1953501,
          "author_name": "kaggler",
          "author_url": "",
          "post_date": "2022-09-24T14:24:13.330000",
          "content": "<p>Nice to see you here again after the hubmap finishes!<br>\nMy 5fold cv scores are around 0.65 anyway it's lower than the sample submission. Thanks for replying!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1955494,
          "author_name": "ForcewithMe",
          "author_url": "",
          "post_date": "2022-09-26T03:30:03.643000",
          "content": "<p>Are you sure you are using the right metrics? A 0.564 cv seems very high in this competition.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1955541,
          "author_name": "kaggler",
          "author_url": "",
          "post_date": "2022-09-26T04:00:06.670000",
          "content": "<p><a href=\"https://www.kaggle.com/forcewithme\" target=\"_blank\">@forcewithme</a> may i ask you what your cv scores are around?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1955556,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2022-09-26T04:05:53.180000",
          "content": "<p>I use <code>sklearn.metrics.log_loss</code>. I tried it with both binary and multiclass predictions and the result was same. My current cv score is 0.5598 right now. Am I missing something and overfitting like crazy?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1955629,
          "author_name": "kaggler",
          "author_url": "",
          "post_date": "2022-09-26T04:40:17.657000",
          "content": "<p><a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> don't you consider applying weighted log loss?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1955654,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2022-09-26T04:54:14.097000",
          "content": "<p>I didn't use weighted log loss since organizers said that each class is roughly equally important. I expect 0.02 deviation from the regular log loss.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1955765,
          "author_name": "ForcewithMe",
          "author_url": "",
          "post_date": "2022-09-26T06:13:05.793000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a>  <a href=\"https://www.kaggle.com/deepkim\" target=\"_blank\">@deepkim</a> </p>\n<ol>\n<li>I implemented the metric of this competition myself. The code in this <a href=\"https://www.kaggle.com/competitions/mayo-clinic-strip-ai/discussion/353305\" target=\"_blank\">post</a> can get the same result as me.</li>\n<li>Use <code>sklearn.metrics.log_loss</code> to get the score of each class, and average them. For example, with  <code>sklearn.metrics.log_loss</code>, you get 0.5 for CE, get 0.8 for LAA, then your score should be 0.65.</li>\n</ol>\n<p>===================================================<br>\nAbout the weight of each class. I am pretty sure that the weight of each class should be 0.5:0.5. Because the score should be 17.2 when all the predicted ce/laa probability are 0/1(1/0), and 0.6+(actually 0.69) when all the predicted ce/laa probability are 0.5 /0.5. Both scores(17.2 0.6) are verified. </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1955781,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2022-09-26T06:21:07.330000",
          "content": "<p>So is it like macro log loss? I'm currently using log loss on binary sigmoided logits. I guess I have to create another dimension with <code>1 - y</code> and do <code>mean(log_loss(y_true, y_pred), log_loss(1 - y_true, 1 - y_pred))</code>. It doesn't make much sense though. I think it would have the same score.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1955787,
          "author_name": "ForcewithMe",
          "author_url": "",
          "post_date": "2022-09-26T06:23:23.867000",
          "content": "<p>Yeah, I always created <code>1-y</code> for <code>class 0</code></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1955793,
          "author_name": "ForcewithMe",
          "author_url": "",
          "post_date": "2022-09-26T06:25:15.297000",
          "content": "<p><a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> In my opinion, <br>\n<code>mean(log_loss(y_true, y_pred), log_loss(1 - y_true, 1 - y_pred))</code>  is wrong.<br>\n<code>mean(log_loss(y_true, y_pred), log_loss(y_true, 1 - y_pred))</code> is right.<br>\nPlease correct me if you find something wrong. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1955811,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2022-09-26T06:32:34.227000",
          "content": "<p>I created two vectors and their names are self-explanatory…</p>\n<pre><code>y_true = np.random.randint(0, 2, 100)\ny_pred = np.random.rand(100)\n</code></pre>\n<p>As I thought, these two yield same scores</p>\n<pre><code>&gt;&gt;&gt; log_loss(y_true, y_pred)\n0.9869673142342817\n&gt;&gt;&gt; log_loss(1 - y_true, 1 - y_pred)\n0.9869673142342817\n</code></pre>\n<p>but this one yields a slightly different score</p>\n<pre><code>&gt;&gt;&gt; log_loss(y_true, 1 - y_pred)\n0.9635681791900268\n</code></pre>\n<p>and the average of them is</p>\n<pre><code>&gt;&gt;&gt; np.mean([log_loss(y_true, y_pred), log_loss(y_true, 1 - y_pred)])\n0.9752677467121542\n</code></pre>\n<p>so the deviation is even less than 0.02. However that is totally dependent to predictions. I don't think macro and micro log loss scores wouldn't be too much different from each other since both class predictions are dependent to each other.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1956090,
          "author_name": "Theo Viel",
          "author_url": "",
          "post_date": "2022-09-26T09:20:40.017000",
          "content": "<p>Here's the code I use</p>\n<pre><code># Group by patient\ndfg = df[[\"patient_id\", \"pred\", \"target\"]].groupby(\"patient_id\").mean()\n\n# Separate the two classes\ndf1 = dfg[dfg['target'] == 1]\ndf0 = dfg[dfg['target'] == 0]\n\n# Compute the losses and average\nl0 = log_loss(df0[\"target\"], df0[\"pred\"], labels=[0, 1])\nl1 = log_loss(df1[\"target\"], df1[\"pred\"], labels=[0, 1])\nloss = 0.5 * (l0 + l1)\n\nprint(f\"\\nCV loss : {loss:.3f}  ({l0:.3f}, {l1:.3f})\")\n\n# For comparison, this is the binary logloss\nloss = log_loss(dfg[\"target\"], dfg[\"pred\"])\nprint(f\"\\nlogloss : {loss:.3f}\")\n</code></pre>\n<p>I can get 0.55 logloss but my weighted logloss scores are around 0.65.<br>\nThe binary logloss benefits from the class imbalance hence if I have a model that predicts more CE the loss will be lower.</p>\n<p><a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> What is your CV AUC though ? That's what I'm optimizing on my side =)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1956117,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2022-09-26T09:39:06.200000",
          "content": "<p>That makes sense… My results were close to each other because dummy vectors had normal distribution. AUC seems to be a good auxiliary metric. These are my fold scores. I use stratified group kfold, splits are stratified on label and groups are patient_id. </p>\n<pre><code>{\n  \"fold_scores\": {\n    \"fold1\": {\n      \"accuracy\": 0.7450980392156863,\n      \"roc_auc\": 0.6655420602789024,\n      \"log_loss\": 0.5360595129819867\n    },\n    \"fold2\": {\n      \"accuracy\": 0.6948051948051948,\n      \"roc_auc\": 0.5798369457148538,\n      \"log_loss\": 0.6151186826747733\n    },\n    \"fold3\": {\n      \"accuracy\": 0.7635135135135135,\n      \"roc_auc\": 0.6552579365079365,\n      \"log_loss\": 0.5286322120778464\n    },\n    \"fold4\": {\n      \"accuracy\": 0.74,\n      \"roc_auc\": 0.6325799955247259,\n      \"log_loss\": 0.5594933565209309\n    },\n    \"fold5\": {\n      \"accuracy\": 0.7181208053691275,\n      \"roc_auc\": 0.7004329004329004,\n      \"log_loss\": 0.5591903770679996\n    }\n  },\n  \"oof_scores\": {\n    \"accuracy\": 0.7320954907161804,\n    \"roc_auc\": 0.6483851310176723,\n    \"log_loss\": 0.5599818563222174\n  }\n}\n</code></pre>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1956187,
          "author_name": "Theo Viel",
          "author_url": "",
          "post_date": "2022-09-26T10:32:30.100000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a>, I can confirm I have similary AUCs (~0.65) using the same splitting strategy</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1956204,
          "author_name": "kaggler",
          "author_url": "",
          "post_date": "2022-09-26T10:43:04.443000",
          "content": "<p>my weighted log loss converges to 0.65 but in terms of accuracy, lower than <a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1956211,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2022-09-26T10:50:29.147000",
          "content": "<p>After implementing <a href=\"https://www.kaggle.com/forcewithme\" target=\"_blank\">@forcewithme</a>'s metric as a loss function, my model didn't learn. Weighted log loss and regular log loss starts at 0.69 and finishes like that. Accuracy and ROC AUC randomly oscillates. This is what I'm using right now.</p>\n<pre><code>class MacroBCEWithLogitsLoss(_WeightedLoss):\n\n    def __init__(self, weight=None, reduction='mean'):\n\n        super(MacroBCEWithLogitsLoss, self).__init__(weight=weight, reduction=reduction)\n\n        self.weight = weight\n        self.reduction = reduction\n\n    def forward(self, inputs, targets):\n\n        inputs = torch.sigmoid(inputs)\n        loss_positive = F.binary_cross_entropy(inputs, targets, self.weight)\n        loss_negative = F.binary_cross_entropy(1 - inputs, targets, self.weight)\n        loss = (0.5 * loss_positive) + (0.5 * loss_negative)\n\n        if self.reduction == 'mean':\n            loss = loss.mean()\n        elif self.reduction == 'sum':\n            loss = loss.sum()\n\n        return loss\n</code></pre>\n<p>Edit: Optimizing this loss function didn't work for me. After one epoch, model learns to make safe predictions between 0.48-0.52, and loss doesn't improve after 0.69. This behavior continues when I tweak the weights.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1956287,
          "author_name": "ForcewithMe",
          "author_url": "",
          "post_date": "2022-09-26T11:24:12.887000",
          "content": "<p>I only used the competition metric for validation and save weights. And I can't get a significantly better score than <code>sample_submission.csv</code> 🤕</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1956774,
          "author_name": "tdiceman",
          "author_url": "",
          "post_date": "2022-09-26T15:38:35.550000",
          "content": "<p>I believe public LB score is only on 20 images out of 280, there is good chance that private LB will look very different. Plus do also check the spread of log loss over the predictions, mean values only tell part of the story. Even if I get a mean log loss of under 0.6, the spread is from 0.2 to 0.8, which hints at some overfitting for sure - I guess this is where this dataset is hard. Also waiting to see top private LB solutions :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1956964,
          "author_name": "kaggler",
          "author_url": "",
          "post_date": "2022-09-26T17:36:43.627000",
          "content": "<p><a href=\"https://www.kaggle.com/icemantd\" target=\"_blank\">@icemantd</a> checking the spread of log loss over the predictions seems really good strategy!!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1970760,
      "author_name": "Theo Viel",
      "author_url": "",
      "post_date": "2022-10-04T09:19:08.317000",
      "content": "<p>Turns out submitting <code>1 - preds</code> scores better on LB for me.</p>\n<p>This should happen in ~6% of the cases if my LB predictions follow the same distribution as my oof ones, but still I am a bit worried.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1971341,
          "author_name": "kaggler",
          "author_url": "",
          "post_date": "2022-10-04T14:39:37.203000",
          "content": "<p>Same here. I tried also reversing CE and LAA, it gave me better score in LB (0.7 -&gt; 0.6)  <br>\ndoes it imply there is no signal?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1971410,
          "author_name": "tdiceman",
          "author_url": "",
          "post_date": "2022-10-04T15:30:45.380000",
          "content": "<p>We'll find out soon enough. Hold on to your two submissions with some popcorn :D</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1972627,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2022-10-05T08:53:51.617000",
          "content": "<p>My best submission's AUC is 0.6842 and weighted log loss is 0.6387. It's weird that we are so close to each other in public leaderboard <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a>, but I haven't inverted my predictions.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1955382,
      "author_name": "Joe Marturano",
      "author_url": "",
      "post_date": "2022-09-26T01:15:27.510000",
      "content": "<p>I'm looking forward to the \"1st place solution\" post by sample_submission.csv.</p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1953430": "Why do my models have lower performances than sample submission? it's a weird competition  \nin local environment, anyway it learns the pattern.  \nbut when i submit it i just score 0.6 or 0.7  \nit's weird!!  \n-----------------------------------------------------------------------------\n  Thanks for the all the discussions! Good Luck For Everyone Here!  \n\n",
    "1970621": "I am really not sure what is going on in this competition, I just started three days ago !  I did a basic training and got 0.4 at first shot. When I trained with \"focus\" and deeply I got a score of 1.0 and another try score of 0.6!  I didn't really do much and will most likely go down sliding in the private board ! but I don't get my results ! ... my 0.4 is a lucky shot maybe ! ",
    "1972166": "Well, I still think that anything that produces score less than 0.6-0.5 is a product of a mere luck ~~or overfitting that happens in such specific conditions as a lack of signal in the data~~.",
    "1971445": "Any last minute guesses what the medal range will be for private LB?\n\nI think low but possible chance sample submission may make it into bronze lol\n\nMaybe 0.5 Gold to 0.7 Bronze?",
    "1965228": "```\ndef binary_weighted_log_loss_forcewithme(y_true, y_pred):\n\n    log_loss_positive = log_loss(y_true, y_pred)\n    log_loss_negative = log_loss(y_true, 1 - y_pred)\n    log_loss_weighted = 0.5 * log_loss_positive + 0.5 * log_loss_negative\n\n    return log_loss_positive, log_loss_negative, log_loss_weighted\n\n\ndef binary_weighted_log_loss_theo(y_true, y_pred):\n\n    y_true_positive = y_true == 1\n    y_true_negative = y_true == 0\n\n    log_loss_positive = log_loss(y_true[y_true_positive], y_pred[y_true_positive], labels=[0, 1])\n    log_loss_negative = log_loss(y_true[y_true_negative], y_pred[y_true_negative], labels=[0, 1])\n    log_loss_weighted = 0.5 * log_loss_positive + 0.5 * log_loss_negative\n\n    return log_loss_positive, log_loss_negative, log_loss_weighted\n```\n\nThese two functions yield different results. This is my final blend score. I have no idea which one is the correct implementation so I use both of them for now.\n\n` Blend - {'accuracy': 0.6286472148541115, 'roc_auc': 0.66280723136299, 'log_loss_positive_forcewithme': 0.6593462375631115, 'log_loss_negative_forcewithme': 0.7662855417748091, 'log_loss_weighted_forcewithme': 0.7128158896689603, 'log_loss_positive_theo': 0.6395257262674887, 'log_loss_negative_theo': 0.6668468698084385, 'log_loss_weighted_theo': 0.6531862980379637} - Predictions Mean: 0.4952 Std: 0.0957 Min: 0.2987 Max: 0.8416`",
    "1953465": "Same here. It's because the public test set is quite small. My cv score is 0.564 but my lb score is 0.8.",
    "1970760": "Turns out submitting `1 - preds` scores better on LB for me.\n\nThis should happen in ~6% of the cases if my LB predictions follow the same distribution as my oof ones, but still I am a bit worried.\n\n",
    "1955382": "I'm looking forward to the \"1st place solution\" post by sample_submission.csv."
  }
}