{
  "id": 370276,
  "title": "What's your local CV without binarizing predictions?",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/370276",
  "author_name": "moth",
  "post_date": "2022-12-03T21:41:42.408000",
  "votes": 11,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Clearly the Public LB has shown us that binarizing predictions over a threshold can lead to better scores. </p>\n<p>I've trained a baseline model and my local CV was around +0.10, but still haven't tried binarizing.</p>\n<p>What is your local CV before binarizing? </p>",
  "messages": [
    {
      "id": 2054121,
      "postDate": "2022-12-03T21:41:42.407Z",
      "content": "<p>Clearly the Public LB has shown us that binarizing predictions over a threshold can lead to better scores. </p>\n<p>I've trained a baseline model and my local CV was around +0.10, but still haven't tried binarizing.</p>\n<p>What is your local CV before binarizing? </p>",
      "rawMarkdown": "Clearly the Public LB has shown us that binarizing predictions over a threshold can lead to better scores. \n\nI've trained a baseline model and my local CV was around +0.10, but still haven't tried binarizing.\n\nWhat is your local CV before binarizing? ",
      "votes": 11
    },
    {
      "id": 2054706,
      "postDate": "2022-12-04T11:23:30.820Z",
      "content": "<p>I am getting to around <code>0.08</code> on a good run on 256px with resnet18 🙂 </p>",
      "rawMarkdown": "I am getting to around `0.08` on a good run on 256px with resnet18 🙂 ",
      "votes": 3,
      "replies": [
        {
          "id": 2059296,
          "postDate": "2022-12-08T17:25:02.133Z",
          "content": "<p>I'm trying resnet18 too, but a have a bad and weird result 😔</p>\n<p>dataset: <a href=\"https://www.kaggle.com/datasets/theoviel/rsna-breast-cancer-512-pngs\" target=\"_blank\">https://www.kaggle.com/datasets/theoviel/rsna-breast-cancer-512-pngs</a><br>\ndata split: <a href=\"https://www.kaggle.com/code/awsaf49/rsna-bcd-efficientnet-tf-tpu-1vm-train\" target=\"_blank\">https://www.kaggle.com/code/awsaf49/rsna-bcd-efficientnet-tf-tpu-1vm-train</a></p>\n<p>BCELoss <br>\ncv 0.033 </p>\n<p>FocalLoss: <a href=\"https://amaarora.github.io/2020/06/29/FocalLoss.html\" target=\"_blank\">https://amaarora.github.io/2020/06/29/FocalLoss.html</a><br>\ncv 0.03</p>\n<p>BCEloss + class_weight<br>\ncv 0.041</p>\n<p>All train:</p>\n<p>epochs : 10<br>\nlr: 1e-4<br>\nconsine-scheduler + warmup<br>\nwarmup: 0.12<br>\nmin_lr: 5e-6<br>\nmodel: resnet18</p>\n<p>I think my model is not learning cancer samples  … my loss decreases on each epoch, but Pf1 decreases or increases a lit bit , but always around of Pf1 first epoch</p>",
          "rawMarkdown": "I'm trying resnet18 too, but a have a bad and weird result 😔\n\ndataset: https://www.kaggle.com/datasets/theoviel/rsna-breast-cancer-512-pngs\ndata split: https://www.kaggle.com/code/awsaf49/rsna-bcd-efficientnet-tf-tpu-1vm-train\n\nBCELoss \ncv 0.033 \n\nFocalLoss: https://amaarora.github.io/2020/06/29/FocalLoss.html\ncv 0.03\n\nBCEloss + class_weight\ncv 0.041\n\nAll train:\n\nepochs : 10\nlr: 1e-4\nconsine-scheduler + warmup\nwarmup: 0.12\nmin_lr: 5e-6\nmodel: resnet18\n\nI think my model is not learning cancer samples  ... my loss decreases on each epoch, but Pf1 decreases or increases a lit bit , but always around of Pf1 first epoch"
        }
      ]
    },
    {
      "id": 2059081,
      "postDate": "2022-12-08T13:20:33.363Z",
      "content": "<p>Mine is <code>0.125</code> with roi  with <code>768x384 ~ 512x512</code> img-size using this <a href=\"https://www.kaggle.com/code/awsaf49/rsna-bcd-efficientnet-tf-tpu-1vm-train\" target=\"_blank\">notebook</a> . Sadly, due to timeout haven't been able to submit yet =(</p>",
      "rawMarkdown": "Mine is `0.125` with roi  with `768x384 ~ 512x512` img-size using this [notebook](https://www.kaggle.com/code/awsaf49/rsna-bcd-efficientnet-tf-tpu-1vm-train) . Sadly, due to timeout haven't been able to submit yet =(",
      "votes": 2,
      "replies": [
        {
          "id": 2059544,
          "postDate": "2022-12-09T02:10:26.640Z",
          "content": "<p>It is sad to see that about ~6 hrs of inference is just reading dicom files and transforming to pngs. So in reality you need to make a prediction for the whole hidden set in just 3 hours. Goodbye big ensembles (?)</p>",
          "rawMarkdown": "It is sad to see that about ~6 hrs of inference is just reading dicom files and transforming to pngs. So in reality you need to make a prediction for the whole hidden set in just 3 hours. Goodbye big ensembles (?)"
        }
      ]
    },
    {
      "id": 2102934,
      "postDate": "2023-01-16T22:28:58.367Z",
      "content": "<p>Could you explain to me the meaning of binaries ?</p>\n<p>Thank in Advance </p>",
      "rawMarkdown": "Could you explain to me the meaning of binaries ?\n\nThank in Advance ",
      "replies": [
        {
          "id": 2103058,
          "postDate": "2023-01-17T01:16:42.973Z",
          "content": "<p>Models output continuous predictions on the [0,1] range. You can group by predictions on <code>prediction_id</code>, perform an aggregation (like max or mean) and then binarize the results using a threshold. For example, if the prediction is less than <code>threshold=0.4</code> then its 0 and if it's equal or higher its 1.</p>",
          "rawMarkdown": "Models output continuous predictions on the [0,1] range. You can group by predictions on `prediction_id`, perform an aggregation (like max or mean) and then binarize the results using a threshold. For example, if the prediction is less than `threshold=0.4` then its 0 and if it's equal or higher its 1.",
          "replies": [
            {
              "id": 2103107,
              "postDate": "2023-01-17T02:18:08.803Z",
              "content": "<p>Hi, would you be so kind as to show me any example code please or notebook where I can see what you told me ?</p>\n<p>Thank you in advance</p>",
              "rawMarkdown": "Hi, would you be so kind as to show me any example code please or notebook where I can see what you told me ?\n\nThank you in advance"
            },
            {
              "id": 2103898,
              "postDate": "2023-01-17T12:53:04.673Z",
              "content": "<pre><code>def binarize(x,threshold):\n    if x&gt;=threshold:\n        x=1\n    else:\n        x=0\n    return x\n\ndf = df.groupby([\"prediction_id\"]).mean() # could be .max() etc\ndf[\"cancer\"] = df[\"cancer\"].apply(lambda x: binarize(x, threshold))\n</code></pre>",
              "rawMarkdown": "```\ndef binarize(x,threshold):\n    if x>=threshold:\n        x=1\n    else:\n        x=0\n    return x\n\ndf = df.groupby([\"prediction_id\"]).mean() # could be .max() etc\ndf[\"cancer\"] = df[\"cancer\"].apply(lambda x: binarize(x, threshold))\n```"
            }
          ]
        }
      ]
    },
    {
      "id": 2059303,
      "postDate": "2022-12-08T17:33:52.870Z",
      "content": "<p>Mine is OOF CV: ~<code>0.09</code> after binarizing <code>0.13</code>. Public LB <code>0.11</code></p>",
      "rawMarkdown": "Mine is OOF CV: ~`0.09` after binarizing `0.13`. Public LB `0.11`"
    }
  ],
  "comments": [
    {
      "id": 2054706,
      "author_name": "Radek Osmulski",
      "author_url": "",
      "post_date": "2022-12-04T11:23:30.820000",
      "content": "<p>I am getting to around <code>0.08</code> on a good run on 256px with resnet18 🙂 </p>",
      "votes": 3,
      "replies": [
        {
          "id": 2059296,
          "author_name": "adriel cabral",
          "author_url": "",
          "post_date": "2022-12-08T17:25:02.133000",
          "content": "<p>I'm trying resnet18 too, but a have a bad and weird result 😔</p>\n<p>dataset: <a href=\"https://www.kaggle.com/datasets/theoviel/rsna-breast-cancer-512-pngs\" target=\"_blank\">https://www.kaggle.com/datasets/theoviel/rsna-breast-cancer-512-pngs</a><br>\ndata split: <a href=\"https://www.kaggle.com/code/awsaf49/rsna-bcd-efficientnet-tf-tpu-1vm-train\" target=\"_blank\">https://www.kaggle.com/code/awsaf49/rsna-bcd-efficientnet-tf-tpu-1vm-train</a></p>\n<p>BCELoss <br>\ncv 0.033 </p>\n<p>FocalLoss: <a href=\"https://amaarora.github.io/2020/06/29/FocalLoss.html\" target=\"_blank\">https://amaarora.github.io/2020/06/29/FocalLoss.html</a><br>\ncv 0.03</p>\n<p>BCEloss + class_weight<br>\ncv 0.041</p>\n<p>All train:</p>\n<p>epochs : 10<br>\nlr: 1e-4<br>\nconsine-scheduler + warmup<br>\nwarmup: 0.12<br>\nmin_lr: 5e-6<br>\nmodel: resnet18</p>\n<p>I think my model is not learning cancer samples  … my loss decreases on each epoch, but Pf1 decreases or increases a lit bit , but always around of Pf1 first epoch</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2059081,
      "author_name": "Awsaf",
      "author_url": "",
      "post_date": "2022-12-08T13:20:33.363000",
      "content": "<p>Mine is <code>0.125</code> with roi  with <code>768x384 ~ 512x512</code> img-size using this <a href=\"https://www.kaggle.com/code/awsaf49/rsna-bcd-efficientnet-tf-tpu-1vm-train\" target=\"_blank\">notebook</a> . Sadly, due to timeout haven't been able to submit yet =(</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2059544,
          "author_name": "moth",
          "author_url": "",
          "post_date": "2022-12-09T02:10:26.640000",
          "content": "<p>It is sad to see that about ~6 hrs of inference is just reading dicom files and transforming to pngs. So in reality you need to make a prediction for the whole hidden set in just 3 hours. Goodbye big ensembles (?)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2102934,
      "author_name": "Pablo Larrosa",
      "author_url": "",
      "post_date": "2023-01-16T22:28:58.367000",
      "content": "<p>Could you explain to me the meaning of binaries ?</p>\n<p>Thank in Advance </p>",
      "votes": 0,
      "replies": [
        {
          "id": 2103058,
          "author_name": "moth",
          "author_url": "",
          "post_date": "2023-01-17T01:16:42.973000",
          "content": "<p>Models output continuous predictions on the [0,1] range. You can group by predictions on <code>prediction_id</code>, perform an aggregation (like max or mean) and then binarize the results using a threshold. For example, if the prediction is less than <code>threshold=0.4</code> then its 0 and if it's equal or higher its 1.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2103107,
              "author_name": "Pablo Larrosa",
              "author_url": "",
              "post_date": "2023-01-17T02:18:08.803000",
              "content": "<p>Hi, would you be so kind as to show me any example code please or notebook where I can see what you told me ?</p>\n<p>Thank you in advance</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2103898,
              "author_name": "moth",
              "author_url": "",
              "post_date": "2023-01-17T12:53:04.673000",
              "content": "<pre><code>def binarize(x,threshold):\n    if x&gt;=threshold:\n        x=1\n    else:\n        x=0\n    return x\n\ndf = df.groupby([\"prediction_id\"]).mean() # could be .max() etc\ndf[\"cancer\"] = df[\"cancer\"].apply(lambda x: binarize(x, threshold))\n</code></pre>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2059303,
      "author_name": "moth",
      "author_url": "",
      "post_date": "2022-12-08T17:33:52.870000",
      "content": "<p>Mine is OOF CV: ~<code>0.09</code> after binarizing <code>0.13</code>. Public LB <code>0.11</code></p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2054121": "Clearly the Public LB has shown us that binarizing predictions over a threshold can lead to better scores. \n\nI've trained a baseline model and my local CV was around +0.10, but still haven't tried binarizing.\n\nWhat is your local CV before binarizing? ",
    "2054706": "I am getting to around `0.08` on a good run on 256px with resnet18 🙂 ",
    "2059081": "Mine is `0.125` with roi  with `768x384 ~ 512x512` img-size using this [notebook](https://www.kaggle.com/code/awsaf49/rsna-bcd-efficientnet-tf-tpu-1vm-train) . Sadly, due to timeout haven't been able to submit yet =(",
    "2102934": "Could you explain to me the meaning of binaries ?\n\nThank in Advance ",
    "2059303": "Mine is OOF CV: ~`0.09` after binarizing `0.13`. Public LB `0.11`"
  }
}