{
  "id": 355596,
  "title": "Optimize weighted log loss by post processing",
  "url": "/competitions/mayo-clinic-strip-ai/discussion/355596",
  "author_name": "Gunes Evitan",
  "post_date": "2022-09-27T11:09:12.317000",
  "votes": 3,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I've been thinking about this since yesterday morning and I did some experiments. For some background information, I believe best way to model this data is MIL and I'm using a slightly different variation of it. I use 5 stratified (label) group (patient_id) folds with a good seed. Yes, I tune split seed by minimizing target mean deviation between folds.</p>\n<p>I initially started using <code>BCEWithLogitsLoss</code> for binary classification. Even though I can reach 0.56x log loss consistently,<br>\nmy weighted log loss is always between 0.8 and 0.9. These are my best single model scores.</p>\n<pre><code>{\n  \"fold_scores\": {\n    \"fold1\": {\n      \"accuracy\": 0.7450980392156863,\n      \"roc_auc\": 0.6704903283850652,\n      \"log_loss_positive\": 0.531411202928794,\n      \"log_loss_negative\": 1.2409183252968041,\n      \"log_loss_weighted\": 0.886164764112799\n    },\n    \"fold2\": {\n      \"accuracy\": 0.6948051948051948,\n      \"roc_auc\": 0.5806323324716643,\n      \"log_loss_positive\": 0.6140930646812761,\n      \"log_loss_negative\": 1.0609355782727143,\n      \"log_loss_weighted\": 0.8375143214769952\n    },\n    \"fold3\": {\n      \"accuracy\": 0.7635135135135135,\n      \"roc_auc\": 0.6584821428571429,\n      \"log_loss_positive\": 0.5283679999914523,\n      \"log_loss_negative\": 1.2545960068299964,\n      \"log_loss_weighted\": 0.8914820034107244\n    },\n    \"fold4\": {\n      \"accuracy\": 0.74,\n      \"roc_auc\": 0.6220631013649586,\n      \"log_loss_positive\": 0.5628854739914337,\n      \"log_loss_negative\": 1.132615669965744,\n      \"log_loss_weighted\": 0.8477505719785888\n    },\n    \"fold5\": {\n      \"accuracy\": 0.7181208053691275,\n      \"roc_auc\": 0.6978354978354979,\n      \"log_loss_positive\": 0.5618902654855843,\n      \"log_loss_negative\": 1.2231793243613018,\n      \"log_loss_weighted\": 0.892534794923443\n    }\n  },\n  \"oof_scores\": {\n    \"accuracy\": 0.7320954907161804,\n    \"roc_auc\": 0.6467866006058519,\n    \"log_loss_positive\": 0.5599856507477773,\n    \"log_loss_negative\": 1.1817915937134535,\n    \"log_loss_weighted\": 0.8708886222306154\n  }\n</code></pre>\n<p>I tried pos_weight, label smoothing, focal loss, multiclass classification using BCELoss, but the result is always same. Weighted log loss is always between 0.8 and 0.9. After that, I thought maybe I should optimize weighted log loss directly.</p>\n<pre><code>class WeightedBCEWithLogitsLoss(_WeightedLoss):\n\n    def __init__(self, weight=None, reduction='mean'):\n\n        super(MacroBCEWithLogitsLoss, self).__init__(weight=weight, reduction=reduction)\n\n        self.weight = weight\n        self.reduction = reduction\n\n    def forward(self, inputs, targets):\n\n        inputs = torch.sigmoid(inputs)\n        loss_positive = F.binary_cross_entropy(inputs, targets, self.weight)\n        loss_negative = F.binary_cross_entropy(1 - inputs, targets, self.weight)\n        loss = (0.5 * loss_positive) + (0.5 * loss_negative)\n\n        if self.reduction == 'mean':\n            loss = loss.mean()\n        elif self.reduction == 'sum':\n            loss = loss.sum()\n\n        return loss\n</code></pre>\n<p>When I train my models using this loss function, they don't learn anything and weighted log loss stuck at 0.69x. 0.69x seems better than my previous experiments, right, but it doesn't look good because results were actually random. When I printed my model's predictions, I found that the predictions are always close 0.5. They are randomly separated around 0.5 which means that the model can only make safe predictions. I also printed my previous model's predictions and their mean was 0.2x, so that model was also making safe predictions in its own way. I think I would prefer my previous model's because they captured some signal from the data and their oof AUC score was 0.65.</p>\n<p>I also found that I can improve my oof AUC and log loss by ensembling. This is my final score after blending some similar models.</p>\n<p><code>Blend - {'accuracy': 0.7387267904509284, 'roc_auc': 0.6609172561799539, 'log_loss_positive': 0.5547231636743363, 'log_loss_negative': 1.1302577089303685, 'log_loss_weighted': 0.8424904363023524} - Predictions Mean: 0.2769 Std: 0.0981 Min: 0.0688 Max: 0.6277</code></p>\n<p>Weighted log loss isn't decent yet but I can make it little bit better by shifting the distribution closer to 0.5. When I add 0.2 to my predictions, my weighted log loss improves to 0.71.</p>\n<p><code>Blend - {'accuracy': 0.6803713527851459, 'roc_auc': 0.6609172561799539, 'log_loss_positive': 0.6359318632390547, 'log_loss_negative': 0.7958761961902944, 'log_loss_weighted': 0.7159040297146746} - Predictions Mean: 0.4669 Std: 0.0981 Min: 0.2588 Max: 0.8177</code></p>\n<p>So the point is we can always manipulate our model outputs for a particular metric. Next step would be stretching the distribution tails for making your confident predictions even more confident. This approach is probably quite risky but at least it was intuitive to me.</p>",
  "messages": [
    {
      "id": 1958277,
      "postDate": "2022-09-27T11:09:12.317Z",
      "content": "<p>I've been thinking about this since yesterday morning and I did some experiments. For some background information, I believe best way to model this data is MIL and I'm using a slightly different variation of it. I use 5 stratified (label) group (patient_id) folds with a good seed. Yes, I tune split seed by minimizing target mean deviation between folds.</p>\n<p>I initially started using <code>BCEWithLogitsLoss</code> for binary classification. Even though I can reach 0.56x log loss consistently,<br>\nmy weighted log loss is always between 0.8 and 0.9. These are my best single model scores.</p>\n<pre><code>{\n  \"fold_scores\": {\n    \"fold1\": {\n      \"accuracy\": 0.7450980392156863,\n      \"roc_auc\": 0.6704903283850652,\n      \"log_loss_positive\": 0.531411202928794,\n      \"log_loss_negative\": 1.2409183252968041,\n      \"log_loss_weighted\": 0.886164764112799\n    },\n    \"fold2\": {\n      \"accuracy\": 0.6948051948051948,\n      \"roc_auc\": 0.5806323324716643,\n      \"log_loss_positive\": 0.6140930646812761,\n      \"log_loss_negative\": 1.0609355782727143,\n      \"log_loss_weighted\": 0.8375143214769952\n    },\n    \"fold3\": {\n      \"accuracy\": 0.7635135135135135,\n      \"roc_auc\": 0.6584821428571429,\n      \"log_loss_positive\": 0.5283679999914523,\n      \"log_loss_negative\": 1.2545960068299964,\n      \"log_loss_weighted\": 0.8914820034107244\n    },\n    \"fold4\": {\n      \"accuracy\": 0.74,\n      \"roc_auc\": 0.6220631013649586,\n      \"log_loss_positive\": 0.5628854739914337,\n      \"log_loss_negative\": 1.132615669965744,\n      \"log_loss_weighted\": 0.8477505719785888\n    },\n    \"fold5\": {\n      \"accuracy\": 0.7181208053691275,\n      \"roc_auc\": 0.6978354978354979,\n      \"log_loss_positive\": 0.5618902654855843,\n      \"log_loss_negative\": 1.2231793243613018,\n      \"log_loss_weighted\": 0.892534794923443\n    }\n  },\n  \"oof_scores\": {\n    \"accuracy\": 0.7320954907161804,\n    \"roc_auc\": 0.6467866006058519,\n    \"log_loss_positive\": 0.5599856507477773,\n    \"log_loss_negative\": 1.1817915937134535,\n    \"log_loss_weighted\": 0.8708886222306154\n  }\n</code></pre>\n<p>I tried pos_weight, label smoothing, focal loss, multiclass classification using BCELoss, but the result is always same. Weighted log loss is always between 0.8 and 0.9. After that, I thought maybe I should optimize weighted log loss directly.</p>\n<pre><code>class WeightedBCEWithLogitsLoss(_WeightedLoss):\n\n    def __init__(self, weight=None, reduction='mean'):\n\n        super(MacroBCEWithLogitsLoss, self).__init__(weight=weight, reduction=reduction)\n\n        self.weight = weight\n        self.reduction = reduction\n\n    def forward(self, inputs, targets):\n\n        inputs = torch.sigmoid(inputs)\n        loss_positive = F.binary_cross_entropy(inputs, targets, self.weight)\n        loss_negative = F.binary_cross_entropy(1 - inputs, targets, self.weight)\n        loss = (0.5 * loss_positive) + (0.5 * loss_negative)\n\n        if self.reduction == 'mean':\n            loss = loss.mean()\n        elif self.reduction == 'sum':\n            loss = loss.sum()\n\n        return loss\n</code></pre>\n<p>When I train my models using this loss function, they don't learn anything and weighted log loss stuck at 0.69x. 0.69x seems better than my previous experiments, right, but it doesn't look good because results were actually random. When I printed my model's predictions, I found that the predictions are always close 0.5. They are randomly separated around 0.5 which means that the model can only make safe predictions. I also printed my previous model's predictions and their mean was 0.2x, so that model was also making safe predictions in its own way. I think I would prefer my previous model's because they captured some signal from the data and their oof AUC score was 0.65.</p>\n<p>I also found that I can improve my oof AUC and log loss by ensembling. This is my final score after blending some similar models.</p>\n<p><code>Blend - {'accuracy': 0.7387267904509284, 'roc_auc': 0.6609172561799539, 'log_loss_positive': 0.5547231636743363, 'log_loss_negative': 1.1302577089303685, 'log_loss_weighted': 0.8424904363023524} - Predictions Mean: 0.2769 Std: 0.0981 Min: 0.0688 Max: 0.6277</code></p>\n<p>Weighted log loss isn't decent yet but I can make it little bit better by shifting the distribution closer to 0.5. When I add 0.2 to my predictions, my weighted log loss improves to 0.71.</p>\n<p><code>Blend - {'accuracy': 0.6803713527851459, 'roc_auc': 0.6609172561799539, 'log_loss_positive': 0.6359318632390547, 'log_loss_negative': 0.7958761961902944, 'log_loss_weighted': 0.7159040297146746} - Predictions Mean: 0.4669 Std: 0.0981 Min: 0.2588 Max: 0.8177</code></p>\n<p>So the point is we can always manipulate our model outputs for a particular metric. Next step would be stretching the distribution tails for making your confident predictions even more confident. This approach is probably quite risky but at least it was intuitive to me.</p>",
      "rawMarkdown": "I've been thinking about this since yesterday morning and I did some experiments. For some background information, I believe best way to model this data is MIL and I'm using a slightly different variation of it. I use 5 stratified (label) group (patient_id) folds with a good seed. Yes, I tune split seed by minimizing target mean deviation between folds.\n\nI initially started using `BCEWithLogitsLoss` for binary classification. Even though I can reach 0.56x log loss consistently,\nmy weighted log loss is always between 0.8 and 0.9. These are my best single model scores.\n\n```\n{\n  \"fold_scores\": {\n    \"fold1\": {\n      \"accuracy\": 0.7450980392156863,\n      \"roc_auc\": 0.6704903283850652,\n      \"log_loss_positive\": 0.531411202928794,\n      \"log_loss_negative\": 1.2409183252968041,\n      \"log_loss_weighted\": 0.886164764112799\n    },\n    \"fold2\": {\n      \"accuracy\": 0.6948051948051948,\n      \"roc_auc\": 0.5806323324716643,\n      \"log_loss_positive\": 0.6140930646812761,\n      \"log_loss_negative\": 1.0609355782727143,\n      \"log_loss_weighted\": 0.8375143214769952\n    },\n    \"fold3\": {\n      \"accuracy\": 0.7635135135135135,\n      \"roc_auc\": 0.6584821428571429,\n      \"log_loss_positive\": 0.5283679999914523,\n      \"log_loss_negative\": 1.2545960068299964,\n      \"log_loss_weighted\": 0.8914820034107244\n    },\n    \"fold4\": {\n      \"accuracy\": 0.74,\n      \"roc_auc\": 0.6220631013649586,\n      \"log_loss_positive\": 0.5628854739914337,\n      \"log_loss_negative\": 1.132615669965744,\n      \"log_loss_weighted\": 0.8477505719785888\n    },\n    \"fold5\": {\n      \"accuracy\": 0.7181208053691275,\n      \"roc_auc\": 0.6978354978354979,\n      \"log_loss_positive\": 0.5618902654855843,\n      \"log_loss_negative\": 1.2231793243613018,\n      \"log_loss_weighted\": 0.892534794923443\n    }\n  },\n  \"oof_scores\": {\n    \"accuracy\": 0.7320954907161804,\n    \"roc_auc\": 0.6467866006058519,\n    \"log_loss_positive\": 0.5599856507477773,\n    \"log_loss_negative\": 1.1817915937134535,\n    \"log_loss_weighted\": 0.8708886222306154\n  }\n```\n\nI tried pos_weight, label smoothing, focal loss, multiclass classification using BCELoss, but the result is always same. Weighted log loss is always between 0.8 and 0.9. After that, I thought maybe I should optimize weighted log loss directly.\n```\n\nclass WeightedBCEWithLogitsLoss(_WeightedLoss):\n\n    def __init__(self, weight=None, reduction='mean'):\n\n        super(MacroBCEWithLogitsLoss, self).__init__(weight=weight, reduction=reduction)\n\n        self.weight = weight\n        self.reduction = reduction\n\n    def forward(self, inputs, targets):\n\n        inputs = torch.sigmoid(inputs)\n        loss_positive = F.binary_cross_entropy(inputs, targets, self.weight)\n        loss_negative = F.binary_cross_entropy(1 - inputs, targets, self.weight)\n        loss = (0.5 * loss_positive) + (0.5 * loss_negative)\n\n        if self.reduction == 'mean':\n            loss = loss.mean()\n        elif self.reduction == 'sum':\n            loss = loss.sum()\n\n        return loss\n```\n\nWhen I train my models using this loss function, they don't learn anything and weighted log loss stuck at 0.69x. 0.69x seems better than my previous experiments, right, but it doesn't look good because results were actually random. When I printed my model's predictions, I found that the predictions are always close 0.5. They are randomly separated around 0.5 which means that the model can only make safe predictions. I also printed my previous model's predictions and their mean was 0.2x, so that model was also making safe predictions in its own way. I think I would prefer my previous model's because they captured some signal from the data and their oof AUC score was 0.65.\n\nI also found that I can improve my oof AUC and log loss by ensembling. This is my final score after blending some similar models.\n\n`Blend - {'accuracy': 0.7387267904509284, 'roc_auc': 0.6609172561799539, 'log_loss_positive': 0.5547231636743363, 'log_loss_negative': 1.1302577089303685, 'log_loss_weighted': 0.8424904363023524} - Predictions Mean: 0.2769 Std: 0.0981 Min: 0.0688 Max: 0.6277`\n\nWeighted log loss isn't decent yet but I can make it little bit better by shifting the distribution closer to 0.5. When I add 0.2 to my predictions, my weighted log loss improves to 0.71.\n\n`Blend - {'accuracy': 0.6803713527851459, 'roc_auc': 0.6609172561799539, 'log_loss_positive': 0.6359318632390547, 'log_loss_negative': 0.7958761961902944, 'log_loss_weighted': 0.7159040297146746} - Predictions Mean: 0.4669 Std: 0.0981 Min: 0.2588 Max: 0.8177`\n\nSo the point is we can always manipulate our model outputs for a particular metric. Next step would be stretching the distribution tails for making your confident predictions even more confident. This approach is probably quite risky but at least it was intuitive to me.",
      "votes": 3
    },
    {
      "id": 1965084,
      "postDate": "2022-10-01T06:22:09.763Z",
      "content": "<p>I did lots of experiments on using different loss functions but they all ended up in failure. I think ROC AUC score is the proper auxiliary metric that we should monitor. The problem can be simplified to ranking of positive and negative examples regardless of the scale. After making your predictions 0.5 centered, you get your potential highest weighted log loss.</p>",
      "rawMarkdown": "I did lots of experiments on using different loss functions but they all ended up in failure. I think ROC AUC score is the proper auxiliary metric that we should monitor. The problem can be simplified to ranking of positive and negative examples regardless of the scale. After making your predictions 0.5 centered, you get your potential highest weighted log loss.",
      "votes": 1
    },
    {
      "id": 1959811,
      "postDate": "2022-09-28T10:35:49.940Z",
      "content": "<p>Thanks for sharing your experience. Fascinating that shifting the predictions seems to help. Have you tried randomly shuffling the label and re-running your CV? Something like this:</p>\n<p><code>y_shuffle = df[\"label\"].sample(frac=1, replace=False)</code></p>\n<p>That helped me understand what was a real signal or not.</p>",
      "rawMarkdown": "Thanks for sharing your experience. Fascinating that shifting the predictions seems to help. Have you tried randomly shuffling the label and re-running your CV? Something like this:\n\n`y_shuffle = df[\"label\"].sample(frac=1, replace=False)`\n\nThat helped me understand what was a real signal or not.",
      "votes": 1,
      "replies": [
        {
          "id": 1959818,
          "postDate": "2022-09-28T10:43:08.037Z",
          "content": "<p>I haven't tried that yet. What was the result after you did that? Did you get significantly worse scores?</p>",
          "rawMarkdown": "I haven't tried that yet. What was the result after you did that? Did you get significantly worse scores?"
        },
        {
          "id": 1960150,
          "postDate": "2022-09-28T14:02:03.210Z",
          "content": "<p>Yes, I'm running a 5-fold CV where all images from a single patient are constrained to the same fold. When I randomly shuffle the label, I made two observations:</p>\n<ul>\n<li>The mean ROCAUC on the validation folds decreased from 0.60 (true labels) to 0.48 (shuffled labels)</li>\n<li>The prediction probabilities on the shuffled labels were 0.5 +/- 0.005 (mean +/- StD)</li>\n</ul>\n<p>So my conclusion was that there is likely some small amount of 'real' signal in the dataset, and that predictions from 0.49 to 0.51 are not different than random guesses.</p>",
          "rawMarkdown": "Yes, I'm running a 5-fold CV where all images from a single patient are constrained to the same fold. When I randomly shuffle the label, I made two observations:\n\n- The mean ROCAUC on the validation folds decreased from 0.60 (true labels) to 0.48 (shuffled labels)\n- The prediction probabilities on the shuffled labels were 0.5 +/- 0.005 (mean +/- StD)\n\nSo my conclusion was that there is likely some small amount of 'real' signal in the dataset, and that predictions from 0.49 to 0.51 are not different than random guesses.",
          "votes": 1
        },
        {
          "id": 1960185,
          "postDate": "2022-09-28T14:15:05.713Z",
          "content": "<p>I agree with your conclusion. I also think there is small amount of signal in images. There was a similar competition with very weak signal last year, RSNA-MICCAI Brain Tumor Radiogenomic Classification. If I remember correctly, my models were more unstable and my AUC was bouncing around 0.55 and 0.60 when I was training a classifier for that competition. I can consistently reach 0.64 AUC in this dataset and my ensembles AUC is 0.66. However, I couldn't find a decent way to boost weighted log loss.</p>",
          "rawMarkdown": "I agree with your conclusion. I also think there is small amount of signal in images. There was a similar competition with very weak signal last year, RSNA-MICCAI Brain Tumor Radiogenomic Classification. If I remember correctly, my models were more unstable and my AUC was bouncing around 0.55 and 0.60 when I was training a classifier for that competition. I can consistently reach 0.64 AUC in this dataset and my ensembles AUC is 0.66. However, I couldn't find a decent way to boost weighted log loss.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1958486,
      "postDate": "2022-09-27T13:15:14.330Z",
      "content": "<p>I went down this path yesterday too, I noticed shifting the predictions also takes my weighted log loss from 0.74 to 0.601, my assumption is this is just overfitting the oof predictions and on public lb doubt it will have same effect.. Would test it out but with 1 submission a day not sure its worth it lol</p>\n<p>Have you tried your 0.2 shift on public LB?</p>",
      "rawMarkdown": "I went down this path yesterday too, I noticed shifting the predictions also takes my weighted log loss from 0.74 to 0.601, my assumption is this is just overfitting the oof predictions and on public lb doubt it will have same effect.. Would test it out but with 1 submission a day not sure its worth it lol\n\nHave you tried your 0.2 shift on public LB?",
      "replies": [
        {
          "id": 1958523,
          "postDate": "2022-09-27T13:36:56.997Z",
          "content": "<p>I don't think training and private test sets are different so it wouldn't be overfitting.   </p>",
          "rawMarkdown": "I don't think training and private test sets are different so it wouldn't be overfitting.   "
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1965084,
      "author_name": "Gunes Evitan",
      "author_url": "",
      "post_date": "2022-10-01T06:22:09.763000",
      "content": "<p>I did lots of experiments on using different loss functions but they all ended up in failure. I think ROC AUC score is the proper auxiliary metric that we should monitor. The problem can be simplified to ranking of positive and negative examples regardless of the scale. After making your predictions 0.5 centered, you get your potential highest weighted log loss.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1959811,
      "author_name": "Joe Marturano",
      "author_url": "",
      "post_date": "2022-09-28T10:35:49.940000",
      "content": "<p>Thanks for sharing your experience. Fascinating that shifting the predictions seems to help. Have you tried randomly shuffling the label and re-running your CV? Something like this:</p>\n<p><code>y_shuffle = df[\"label\"].sample(frac=1, replace=False)</code></p>\n<p>That helped me understand what was a real signal or not.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1959818,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2022-09-28T10:43:08.037000",
          "content": "<p>I haven't tried that yet. What was the result after you did that? Did you get significantly worse scores?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1960150,
          "author_name": "Joe Marturano",
          "author_url": "",
          "post_date": "2022-09-28T14:02:03.210000",
          "content": "<p>Yes, I'm running a 5-fold CV where all images from a single patient are constrained to the same fold. When I randomly shuffle the label, I made two observations:</p>\n<ul>\n<li>The mean ROCAUC on the validation folds decreased from 0.60 (true labels) to 0.48 (shuffled labels)</li>\n<li>The prediction probabilities on the shuffled labels were 0.5 +/- 0.005 (mean +/- StD)</li>\n</ul>\n<p>So my conclusion was that there is likely some small amount of 'real' signal in the dataset, and that predictions from 0.49 to 0.51 are not different than random guesses.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1960185,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2022-09-28T14:15:05.713000",
          "content": "<p>I agree with your conclusion. I also think there is small amount of signal in images. There was a similar competition with very weak signal last year, RSNA-MICCAI Brain Tumor Radiogenomic Classification. If I remember correctly, my models were more unstable and my AUC was bouncing around 0.55 and 0.60 when I was training a classifier for that competition. I can consistently reach 0.64 AUC in this dataset and my ensembles AUC is 0.66. However, I couldn't find a decent way to boost weighted log loss.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1958486,
      "author_name": "JM",
      "author_url": "",
      "post_date": "2022-09-27T13:15:14.330000",
      "content": "<p>I went down this path yesterday too, I noticed shifting the predictions also takes my weighted log loss from 0.74 to 0.601, my assumption is this is just overfitting the oof predictions and on public lb doubt it will have same effect.. Would test it out but with 1 submission a day not sure its worth it lol</p>\n<p>Have you tried your 0.2 shift on public LB?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1958523,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2022-09-27T13:36:56.997000",
          "content": "<p>I don't think training and private test sets are different so it wouldn't be overfitting.   </p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1958277": "I've been thinking about this since yesterday morning and I did some experiments. For some background information, I believe best way to model this data is MIL and I'm using a slightly different variation of it. I use 5 stratified (label) group (patient_id) folds with a good seed. Yes, I tune split seed by minimizing target mean deviation between folds.\n\nI initially started using `BCEWithLogitsLoss` for binary classification. Even though I can reach 0.56x log loss consistently,\nmy weighted log loss is always between 0.8 and 0.9. These are my best single model scores.\n\n```\n{\n  \"fold_scores\": {\n    \"fold1\": {\n      \"accuracy\": 0.7450980392156863,\n      \"roc_auc\": 0.6704903283850652,\n      \"log_loss_positive\": 0.531411202928794,\n      \"log_loss_negative\": 1.2409183252968041,\n      \"log_loss_weighted\": 0.886164764112799\n    },\n    \"fold2\": {\n      \"accuracy\": 0.6948051948051948,\n      \"roc_auc\": 0.5806323324716643,\n      \"log_loss_positive\": 0.6140930646812761,\n      \"log_loss_negative\": 1.0609355782727143,\n      \"log_loss_weighted\": 0.8375143214769952\n    },\n    \"fold3\": {\n      \"accuracy\": 0.7635135135135135,\n      \"roc_auc\": 0.6584821428571429,\n      \"log_loss_positive\": 0.5283679999914523,\n      \"log_loss_negative\": 1.2545960068299964,\n      \"log_loss_weighted\": 0.8914820034107244\n    },\n    \"fold4\": {\n      \"accuracy\": 0.74,\n      \"roc_auc\": 0.6220631013649586,\n      \"log_loss_positive\": 0.5628854739914337,\n      \"log_loss_negative\": 1.132615669965744,\n      \"log_loss_weighted\": 0.8477505719785888\n    },\n    \"fold5\": {\n      \"accuracy\": 0.7181208053691275,\n      \"roc_auc\": 0.6978354978354979,\n      \"log_loss_positive\": 0.5618902654855843,\n      \"log_loss_negative\": 1.2231793243613018,\n      \"log_loss_weighted\": 0.892534794923443\n    }\n  },\n  \"oof_scores\": {\n    \"accuracy\": 0.7320954907161804,\n    \"roc_auc\": 0.6467866006058519,\n    \"log_loss_positive\": 0.5599856507477773,\n    \"log_loss_negative\": 1.1817915937134535,\n    \"log_loss_weighted\": 0.8708886222306154\n  }\n```\n\nI tried pos_weight, label smoothing, focal loss, multiclass classification using BCELoss, but the result is always same. Weighted log loss is always between 0.8 and 0.9. After that, I thought maybe I should optimize weighted log loss directly.\n```\n\nclass WeightedBCEWithLogitsLoss(_WeightedLoss):\n\n    def __init__(self, weight=None, reduction='mean'):\n\n        super(MacroBCEWithLogitsLoss, self).__init__(weight=weight, reduction=reduction)\n\n        self.weight = weight\n        self.reduction = reduction\n\n    def forward(self, inputs, targets):\n\n        inputs = torch.sigmoid(inputs)\n        loss_positive = F.binary_cross_entropy(inputs, targets, self.weight)\n        loss_negative = F.binary_cross_entropy(1 - inputs, targets, self.weight)\n        loss = (0.5 * loss_positive) + (0.5 * loss_negative)\n\n        if self.reduction == 'mean':\n            loss = loss.mean()\n        elif self.reduction == 'sum':\n            loss = loss.sum()\n\n        return loss\n```\n\nWhen I train my models using this loss function, they don't learn anything and weighted log loss stuck at 0.69x. 0.69x seems better than my previous experiments, right, but it doesn't look good because results were actually random. When I printed my model's predictions, I found that the predictions are always close 0.5. They are randomly separated around 0.5 which means that the model can only make safe predictions. I also printed my previous model's predictions and their mean was 0.2x, so that model was also making safe predictions in its own way. I think I would prefer my previous model's because they captured some signal from the data and their oof AUC score was 0.65.\n\nI also found that I can improve my oof AUC and log loss by ensembling. This is my final score after blending some similar models.\n\n`Blend - {'accuracy': 0.7387267904509284, 'roc_auc': 0.6609172561799539, 'log_loss_positive': 0.5547231636743363, 'log_loss_negative': 1.1302577089303685, 'log_loss_weighted': 0.8424904363023524} - Predictions Mean: 0.2769 Std: 0.0981 Min: 0.0688 Max: 0.6277`\n\nWeighted log loss isn't decent yet but I can make it little bit better by shifting the distribution closer to 0.5. When I add 0.2 to my predictions, my weighted log loss improves to 0.71.\n\n`Blend - {'accuracy': 0.6803713527851459, 'roc_auc': 0.6609172561799539, 'log_loss_positive': 0.6359318632390547, 'log_loss_negative': 0.7958761961902944, 'log_loss_weighted': 0.7159040297146746} - Predictions Mean: 0.4669 Std: 0.0981 Min: 0.2588 Max: 0.8177`\n\nSo the point is we can always manipulate our model outputs for a particular metric. Next step would be stretching the distribution tails for making your confident predictions even more confident. This approach is probably quite risky but at least it was intuitive to me.",
    "1965084": "I did lots of experiments on using different loss functions but they all ended up in failure. I think ROC AUC score is the proper auxiliary metric that we should monitor. The problem can be simplified to ranking of positive and negative examples regardless of the scale. After making your predictions 0.5 centered, you get your potential highest weighted log loss.",
    "1959811": "Thanks for sharing your experience. Fascinating that shifting the predictions seems to help. Have you tried randomly shuffling the label and re-running your CV? Something like this:\n\n`y_shuffle = df[\"label\"].sample(frac=1, replace=False)`\n\nThat helped me understand what was a real signal or not.",
    "1958486": "I went down this path yesterday too, I noticed shifting the predictions also takes my weighted log loss from 0.74 to 0.601, my assumption is this is just overfitting the oof predictions and on public lb doubt it will have same effect.. Would test it out but with 1 submission a day not sure its worth it lol\n\nHave you tried your 0.2 shift on public LB?"
  }
}