{
  "id": 371805,
  "title": "thoughts on reaching LB 0.42",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/371805",
  "author_name": "Radek Osmulski",
  "post_date": "2022-12-12T13:44:05.285000",
  "votes": 65,
  "comment_count": 25,
  "views": 0,
  "content": "<p>Hey,</p>\n<p>First and foremost, data processed by <code>dicomsdl</code> seems perfectly fine to train on. Processing data like this gives us a lot of time in the pipeline for inference, ensembling, etc.</p>\n<p>(sharing this up front as I believe this has been one of the unknowns at this point in the competition)</p>\n<p>I share the data I trained on here:</p>\n<p><a href=\"https://www.kaggle.com/datasets/radek1/rsna-mammo-dicomsdl-1024\" target=\"_blank\">RSNA Mammo dicomsdl 1024</a></p>\n<p>I shared how I processed the data in the following notebook:</p>\n<p><a href=\"https://www.kaggle.com/code/radek1/how-to-process-dicom-images-to-pngs\" target=\"_blank\">💡 how to process DICOM images to PNGs</a></p>\n<p>I didn't upload it to the dataset with all the other resolutions as at 13GB of size my Internet connection gave up, so I created this dataset directly from my GCP vm using the <code>kaggle cli</code>.</p>\n<h3>What are my observations from reaching LB 0.42?</h3>\n<p>It is best to approach this competition like any other computer vision project but with two important characteristics.</p>\n<ul>\n<li>the dataset is imbalanced which makes training hard (plus some important features become discernable only at higher resolutions)</li>\n<li>because of the class imbalance (there being very few positive examples) the metric is very noisy</li>\n</ul>\n<p>The first point dictates that a major challenge of this competition will be helping our model to train. There are several ways one can go about this.</p>\n<p>I share the code I used for training and inference here:</p>\n<p><a href=\"https://www.kaggle.com/code/radek1/fast-ai-starter-pack-train-inference\" target=\"_blank\">🤖 [fast.ai starter pack] train + inference 🚀</a></p>\n<p>But I haven't updated the notebook with the hyperparameters I used for the last couple of runs. Here they are:</p>\n<ul>\n<li>I trained with a weighted cross entropy loss (weight 20 for the positive class), this goes back to the idea of making training for our model easier</li>\n<li>I upsampled the positive class by 3x</li>\n<li>I trained on a resolution of 1024x1024 across 4 folds, using one cycle policy with a batch size of 24 and a lr of 2.25e-4 (this way I could train on a single GPU with 48GB of RAM, I trained for 7 or 8 epochs, training took around 30 minutes per epoch, you can achieve similar results with gradient accumulation -- natively supported by fast.ai -- on smaller GPUs)</li>\n</ul>\n<p>It is very important though to take into account the metric being very noisy. You need to look for hyperparameters to maximize your performance across all your folds.</p>\n<p>My LB score is an ensemble of the models across the 4 folds. Essentially, addressing the metric being so noisy (how to train, how to figure out which are genuine improvements, how not to overfit to a specific fold, etc) will be a big part of this competition.</p>\n<p>I just hope there is enough data in test and that it is similar to the data in train so that this doesn't turn into a luckfest when it comes to the private LB!</p>\n<p>Anyhow, hope this can be of help 🙏</p>\n<p>Wishing you the best of luck in the competition! Happy Kaggling! 🙂</p>",
  "messages": [
    {
      "id": 2062859,
      "postDate": "2022-12-12T13:44:05.287Z",
      "content": "<p>Hey,</p>\n<p>First and foremost, data processed by <code>dicomsdl</code> seems perfectly fine to train on. Processing data like this gives us a lot of time in the pipeline for inference, ensembling, etc.</p>\n<p>(sharing this up front as I believe this has been one of the unknowns at this point in the competition)</p>\n<p>I share the data I trained on here:</p>\n<p><a href=\"https://www.kaggle.com/datasets/radek1/rsna-mammo-dicomsdl-1024\" target=\"_blank\">RSNA Mammo dicomsdl 1024</a></p>\n<p>I shared how I processed the data in the following notebook:</p>\n<p><a href=\"https://www.kaggle.com/code/radek1/how-to-process-dicom-images-to-pngs\" target=\"_blank\">💡 how to process DICOM images to PNGs</a></p>\n<p>I didn't upload it to the dataset with all the other resolutions as at 13GB of size my Internet connection gave up, so I created this dataset directly from my GCP vm using the <code>kaggle cli</code>.</p>\n<h3>What are my observations from reaching LB 0.42?</h3>\n<p>It is best to approach this competition like any other computer vision project but with two important characteristics.</p>\n<ul>\n<li>the dataset is imbalanced which makes training hard (plus some important features become discernable only at higher resolutions)</li>\n<li>because of the class imbalance (there being very few positive examples) the metric is very noisy</li>\n</ul>\n<p>The first point dictates that a major challenge of this competition will be helping our model to train. There are several ways one can go about this.</p>\n<p>I share the code I used for training and inference here:</p>\n<p><a href=\"https://www.kaggle.com/code/radek1/fast-ai-starter-pack-train-inference\" target=\"_blank\">🤖 [fast.ai starter pack] train + inference 🚀</a></p>\n<p>But I haven't updated the notebook with the hyperparameters I used for the last couple of runs. Here they are:</p>\n<ul>\n<li>I trained with a weighted cross entropy loss (weight 20 for the positive class), this goes back to the idea of making training for our model easier</li>\n<li>I upsampled the positive class by 3x</li>\n<li>I trained on a resolution of 1024x1024 across 4 folds, using one cycle policy with a batch size of 24 and a lr of 2.25e-4 (this way I could train on a single GPU with 48GB of RAM, I trained for 7 or 8 epochs, training took around 30 minutes per epoch, you can achieve similar results with gradient accumulation -- natively supported by fast.ai -- on smaller GPUs)</li>\n</ul>\n<p>It is very important though to take into account the metric being very noisy. You need to look for hyperparameters to maximize your performance across all your folds.</p>\n<p>My LB score is an ensemble of the models across the 4 folds. Essentially, addressing the metric being so noisy (how to train, how to figure out which are genuine improvements, how not to overfit to a specific fold, etc) will be a big part of this competition.</p>\n<p>I just hope there is enough data in test and that it is similar to the data in train so that this doesn't turn into a luckfest when it comes to the private LB!</p>\n<p>Anyhow, hope this can be of help 🙏</p>\n<p>Wishing you the best of luck in the competition! Happy Kaggling! 🙂</p>",
      "rawMarkdown": "Hey,\n\nFirst and foremost, data processed by `dicomsdl` seems perfectly fine to train on. Processing data like this gives us a lot of time in the pipeline for inference, ensembling, etc.\n\n(sharing this up front as I believe this has been one of the unknowns at this point in the competition)\n\nI share the data I trained on here:\n\n[RSNA Mammo dicomsdl 1024](https://www.kaggle.com/datasets/radek1/rsna-mammo-dicomsdl-1024)\n\nI shared how I processed the data in the following notebook:\n\n[💡 how to process DICOM images to PNGs](https://www.kaggle.com/code/radek1/how-to-process-dicom-images-to-pngs)\n\nI didn't upload it to the dataset with all the other resolutions as at 13GB of size my Internet connection gave up, so I created this dataset directly from my GCP vm using the `kaggle cli`.\n\n### What are my observations from reaching LB 0.42?\n\nIt is best to approach this competition like any other computer vision project but with two important characteristics.\n\n* the dataset is imbalanced which makes training hard (plus some important features become discernable only at higher resolutions)\n* because of the class imbalance (there being very few positive examples) the metric is very noisy\n\nThe first point dictates that a major challenge of this competition will be helping our model to train. There are several ways one can go about this.\n\nI share the code I used for training and inference here:\n\n[🤖 [fast.ai starter pack] train + inference 🚀](https://www.kaggle.com/code/radek1/fast-ai-starter-pack-train-inference)\n\nBut I haven't updated the notebook with the hyperparameters I used for the last couple of runs. Here they are:\n\n* I trained with a weighted cross entropy loss (weight 20 for the positive class), this goes back to the idea of making training for our model easier\n* I upsampled the positive class by 3x\n* I trained on a resolution of 1024x1024 across 4 folds, using one cycle policy with a batch size of 24 and a lr of 2.25e-4 (this way I could train on a single GPU with 48GB of RAM, I trained for 7 or 8 epochs, training took around 30 minutes per epoch, you can achieve similar results with gradient accumulation -- natively supported by fast.ai -- on smaller GPUs)\n\nIt is very important though to take into account the metric being very noisy. You need to look for hyperparameters to maximize your performance across all your folds.\n\nMy LB score is an ensemble of the models across the 4 folds. Essentially, addressing the metric being so noisy (how to train, how to figure out which are genuine improvements, how not to overfit to a specific fold, etc) will be a big part of this competition.\n\nI just hope there is enough data in test and that it is similar to the data in train so that this doesn't turn into a luckfest when it comes to the private LB!\n\nAnyhow, hope this can be of help 🙏\n\nWishing you the best of luck in the competition! Happy Kaggling! 🙂",
      "votes": 62
    },
    {
      "id": 2062897,
      "postDate": "2022-12-12T14:05:18.793Z",
      "content": "<p>Nice, thanks for sharing. In my case I am using a standard pytorch pipeline, no tricks just some moderate augmentations and image size 1024 to achieve lb 0.51. I wonder if dicomsdl api gives the same image quality compared to pydicom.</p>\n<p>Maybe we can train models with image size 1024 using the 2 T4 GPU kaggle offer, that is 30 vram, I believe that is enough to train a small efficientnet with image size 1024</p>",
      "rawMarkdown": "Nice, thanks for sharing. In my case I am using a standard pytorch pipeline, no tricks just some moderate augmentations and image size 1024 to achieve lb 0.51. I wonder if dicomsdl api gives the same image quality compared to pydicom.\n\nMaybe we can train models with image size 1024 using the 2 T4 GPU kaggle offer, that is 30 vram, I believe that is enough to train a small efficientnet with image size 1024",
      "votes": 9,
      "replies": [
        {
          "id": 2062923,
          "postDate": "2022-12-12T14:20:01.737Z",
          "content": "<p>I tried many times, but the score did not reach 0.4, are you currently using ensemble models？</p>",
          "rawMarkdown": "I tried many times, but the score did not reach 0.4, are you currently using ensemble models？",
          "votes": 1
        },
        {
          "id": 2062955,
          "postDate": "2022-12-12T14:41:25.887Z",
          "content": "<p>Nop, 5 KFold single model</p>",
          "rawMarkdown": "Nop, 5 KFold single model"
        },
        {
          "id": 2062965,
          "postDate": "2022-12-12T14:49:28.887Z",
          "content": "<p>Received, thank you for your reply, I still need to continue to improve the local cv score, which is currently around 0.3 (binarized preds)</p>",
          "rawMarkdown": "Received, thank you for your reply, I still need to continue to improve the local cv score, which is currently around 0.3 (binarized preds)",
          "votes": 1
        },
        {
          "id": 2063379,
          "postDate": "2022-12-12T20:29:49.910Z",
          "content": "<p><a href=\"https://www.kaggle.com/ragnar123\" target=\"_blank\">@ragnar123</a> - efficientnet or something else?</p>",
          "rawMarkdown": "@ragnar123 - efficientnet or something else?",
          "votes": 1
        },
        {
          "id": 2064110,
          "postDate": "2022-12-13T14:24:21.750Z",
          "content": "<p>Yes, efficientnet, tried other backbones and they also perform good. </p>",
          "rawMarkdown": "Yes, efficientnet, tried other backbones and they also perform good. ",
          "votes": 4
        },
        {
          "id": 2064693,
          "postDate": "2022-12-14T03:46:54.387Z",
          "content": "<p>Did you use pretrained weight?</p>",
          "rawMarkdown": "Did you use pretrained weight?"
        },
        {
          "id": 2069555,
          "postDate": "2022-12-19T05:32:09.637Z",
          "content": "<p>Great result! What is your local cv tho?</p>",
          "rawMarkdown": "Great result! What is your local cv tho?"
        },
        {
          "id": 2080239,
          "postDate": "2022-12-30T01:44:37.970Z",
          "content": "<p>How do you sample the data?</p>",
          "rawMarkdown": "How do you sample the data?"
        }
      ]
    },
    {
      "id": 2065068,
      "postDate": "2022-12-14T09:31:59.270Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> , thank you for all your posts ! If I could I ask a beginner question… I understood you trained your model with image resized to 1024 with dicomsdl, but I'm wondering how you manage this for the submission (same size for test images transformation ?), because I faced issues \"out of time\" ? </p>",
      "rawMarkdown": "Hi @radek1 , thank you for all your posts ! If I could I ask a beginner question... I understood you trained your model with image resized to 1024 with dicomsdl, but I'm wondering how you manage this for the submission (same size for test images transformation ?), because I faced issues \"out of time\" ? ",
      "votes": 1,
      "replies": [
        {
          "id": 2065112,
          "postDate": "2022-12-14T10:12:44.340Z",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/pourchot\" target=\"_blank\">@pourchot</a>! </p>\n<p><code>dicomsdl</code> is much faster than <code>pydicom</code>! I was able to resize the images to 1024x1024 and infer using 4 models.</p>\n<p>For the code I used for processing of the DICOM files, please see here (I ran the same operation on the test files):</p>\n<p><a href=\"https://www.kaggle.com/code/radek1/how-to-process-dicom-images-to-pngs\" target=\"_blank\">💡 how to process DICOM images to PNGs</a></p>",
          "rawMarkdown": "Hey @pourchot! \n\n`dicomsdl` is much faster than `pydicom`! I was able to resize the images to 1024x1024 and infer using 4 models.\n\nFor the code I used for processing of the DICOM files, please see here (I ran the same operation on the test files):\n\n[💡 how to process DICOM images to PNGs](https://www.kaggle.com/code/radek1/how-to-process-dicom-images-to-pngs)",
          "votes": 1
        }
      ]
    },
    {
      "id": 2065909,
      "postDate": "2022-12-15T07:29:42.003Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/redek1\" target=\"_blank\">@redek1</a>, I am beginner and thank you for very helpful post<br>\nCan I ask you one question?<br>\nI understand that pf1 is a kind of the F1 score for probability<br>\nTherefore, it seems to reasonable that the input of pf1 should be probability number(output of sigmod or softmax) <br>\nbut in the code shared by you, a binary value(thresholded number)<br>\nI wonder why binary value is passed not probability</p>",
      "rawMarkdown": "Hi @redek1, I am beginner and thank you for very helpful post\nCan I ask you one question?\nI understand that pf1 is a kind of the F1 score for probability\nTherefore, it seems to reasonable that the input of pf1 should be probability number(output of sigmod or softmax) \nbut in the code shared by you, a binary value(thresholded number)\nI wonder why binary value is passed not probability",
      "votes": 2,
      "replies": [
        {
          "id": 2065918,
          "postDate": "2022-12-15T07:34:32.787Z",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/hyunseoki1\" target=\"_blank\">@hyunseoki1</a>! Please see <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369886\" target=\"_blank\">here</a> for a discussion of this</p>\n<p>Essentially, this is a trick that improves the score given class imbalance!</p>",
          "rawMarkdown": "Hey @hyunseoki1! Please see [here](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369886) for a discussion of this\n\nEssentially, this is a trick that improves the score given class imbalance!",
          "votes": 2
        }
      ]
    },
    {
      "id": 2064003,
      "postDate": "2022-12-13T13:03:41.097Z",
      "content": "<p>Interesting <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> - I am not able to setup bs more then 32 images on A100 (80GB vRAM) … (A6000 - bs = 16) for 1024x1024.</p>",
      "rawMarkdown": "Interesting @radek1 - I am not able to setup bs more then 32 images on A100 (80GB vRAM) ... (A6000 - bs = 16) for 1024x1024.",
      "votes": 2,
      "replies": [
        {
          "id": 2064208,
          "postDate": "2022-12-13T15:20:49.867Z",
          "content": "<p>Mhmm not sure what might be going on there, something must be off. I was able to train with a BS of 24 with 48GB of RAM (using fp16, maybe that is the missing part?)</p>",
          "rawMarkdown": "Mhmm not sure what might be going on there, something must be off. I was able to train with a BS of 24 with 48GB of RAM (using fp16, maybe that is the missing part?)",
          "votes": 1
        },
        {
          "id": 2064212,
          "postDate": "2022-12-13T15:23:40.307Z",
          "content": "<p>I use pure Pytorch code - implemented Automatic Mixed Precision inside train loop. Interesting. As I can see most of people has bigger bs than I have. </p>",
          "rawMarkdown": "I use pure Pytorch code - implemented Automatic Mixed Precision inside train loop. Interesting. As I can see most of people has bigger bs than I have. ",
          "votes": 1
        },
        {
          "id": 2064242,
          "postDate": "2022-12-13T16:05:26.123Z",
          "content": "<p>Maybe they are using p.p. img backgr. reduc. or accum. bs.</p>",
          "rawMarkdown": "Maybe they are using p.p. img backgr. reduc. or accum. bs.",
          "votes": 1
        },
        {
          "id": 2065341,
          "postDate": "2022-12-14T14:54:17.587Z",
          "content": "<p>I fixed some lines of code and now bs for 1024 is 72 :) </p>",
          "rawMarkdown": "I fixed some lines of code and now bs for 1024 is 72 :) ",
          "votes": 3
        }
      ]
    },
    {
      "id": 2073704,
      "postDate": "2022-12-23T09:29:35.580Z",
      "content": "<p>Thank you for sharing. I am new to this competition. May I ask why you call metric being very noisy? </p>",
      "rawMarkdown": "Thank you for sharing. I am new to this competition. May I ask why you call metric being very noisy? ",
      "replies": [
        {
          "id": 2131758,
          "postDate": "2023-02-06T10:26:35.433Z",
          "content": "<p>Hi there, another beginner here, so don't take my word for it. I think he means by the metric being noisy that it's difficult to trust the metric. E.g. if you have one fold with a better metric value, it doesn't necessarily mean that the model is also better (generalizes to the test set), it might be just that the specific fold is easier to predict </p>",
          "rawMarkdown": "Hi there, another beginner here, so don't take my word for it. I think he means by the metric being noisy that it's difficult to trust the metric. E.g. if you have one fold with a better metric value, it doesn't necessarily mean that the model is also better (generalizes to the test set), it might be just that the specific fold is easier to predict "
        }
      ]
    },
    {
      "id": 2072644,
      "postDate": "2022-12-22T09:52:41.610Z",
      "content": "<p>Did you try Focal Loss? In my experiment, it increases my pfbeta score.</p>",
      "rawMarkdown": "Did you try Focal Loss? In my experiment, it increases my pfbeta score."
    },
    {
      "id": 2069556,
      "postDate": "2022-12-19T05:34:26.470Z",
      "content": "<p>Hi! Do you mind sharing your local CV and the code for upsampling?</p>",
      "rawMarkdown": "Hi! Do you mind sharing your local CV and the code for upsampling?"
    },
    {
      "id": 2069468,
      "postDate": "2022-12-19T02:43:04.530Z",
      "content": "<p><a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> I am a beginner, Thank you for all your posts, learning a lot.</p>",
      "rawMarkdown": "@radek1 I am a beginner, Thank you for all your posts, learning a lot."
    },
    {
      "id": 2072180,
      "postDate": "2022-12-21T19:51:10.777Z",
      "content": "<p>Thanks for this.</p>",
      "rawMarkdown": "Thanks for this."
    },
    {
      "id": 2066748,
      "postDate": "2022-12-16T02:53:15.337Z",
      "content": "<p>i am a biginner thank you</p>",
      "rawMarkdown": "i am a biginner thank you"
    }
  ],
  "comments": [
    {
      "id": 2062897,
      "author_name": "Martin Kovacevic Buvinic",
      "author_url": "",
      "post_date": "2022-12-12T14:05:18.793000",
      "content": "<p>Nice, thanks for sharing. In my case I am using a standard pytorch pipeline, no tricks just some moderate augmentations and image size 1024 to achieve lb 0.51. I wonder if dicomsdl api gives the same image quality compared to pydicom.</p>\n<p>Maybe we can train models with image size 1024 using the 2 T4 GPU kaggle offer, that is 30 vram, I believe that is enough to train a small efficientnet with image size 1024</p>",
      "votes": 9,
      "replies": [
        {
          "id": 2062923,
          "author_name": "yanqiangmiffy",
          "author_url": "",
          "post_date": "2022-12-12T14:20:01.737000",
          "content": "<p>I tried many times, but the score did not reach 0.4, are you currently using ensemble models？</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2062955,
          "author_name": "Martin Kovacevic Buvinic",
          "author_url": "",
          "post_date": "2022-12-12T14:41:25.887000",
          "content": "<p>Nop, 5 KFold single model</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2062965,
          "author_name": "yanqiangmiffy",
          "author_url": "",
          "post_date": "2022-12-12T14:49:28.887000",
          "content": "<p>Received, thank you for your reply, I still need to continue to improve the local cv score, which is currently around 0.3 (binarized preds)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2063379,
          "author_name": "Remek Kinas",
          "author_url": "",
          "post_date": "2022-12-12T20:29:49.910000",
          "content": "<p><a href=\"https://www.kaggle.com/ragnar123\" target=\"_blank\">@ragnar123</a> - efficientnet or something else?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2064110,
          "author_name": "Martin Kovacevic Buvinic",
          "author_url": "",
          "post_date": "2022-12-13T14:24:21.750000",
          "content": "<p>Yes, efficientnet, tried other backbones and they also perform good. </p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 2064693,
          "author_name": "yrnmmh",
          "author_url": "",
          "post_date": "2022-12-14T03:46:54.387000",
          "content": "<p>Did you use pretrained weight?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2069555,
          "author_name": "Feng Qilong",
          "author_url": "",
          "post_date": "2022-12-19T05:32:09.637000",
          "content": "<p>Great result! What is your local cv tho?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2080239,
          "author_name": "kk",
          "author_url": "",
          "post_date": "2022-12-30T01:44:37.970000",
          "content": "<p>How do you sample the data?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2065068,
      "author_name": "Laurent Pourchot",
      "author_url": "",
      "post_date": "2022-12-14T09:31:59.270000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> , thank you for all your posts ! If I could I ask a beginner question… I understood you trained your model with image resized to 1024 with dicomsdl, but I'm wondering how you manage this for the submission (same size for test images transformation ?), because I faced issues \"out of time\" ? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2065112,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2022-12-14T10:12:44.340000",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/pourchot\" target=\"_blank\">@pourchot</a>! </p>\n<p><code>dicomsdl</code> is much faster than <code>pydicom</code>! I was able to resize the images to 1024x1024 and infer using 4 models.</p>\n<p>For the code I used for processing of the DICOM files, please see here (I ran the same operation on the test files):</p>\n<p><a href=\"https://www.kaggle.com/code/radek1/how-to-process-dicom-images-to-pngs\" target=\"_blank\">💡 how to process DICOM images to PNGs</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2065909,
      "author_name": "hyunseoki",
      "author_url": "",
      "post_date": "2022-12-15T07:29:42.003000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/redek1\" target=\"_blank\">@redek1</a>, I am beginner and thank you for very helpful post<br>\nCan I ask you one question?<br>\nI understand that pf1 is a kind of the F1 score for probability<br>\nTherefore, it seems to reasonable that the input of pf1 should be probability number(output of sigmod or softmax) <br>\nbut in the code shared by you, a binary value(thresholded number)<br>\nI wonder why binary value is passed not probability</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2065918,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2022-12-15T07:34:32.787000",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/hyunseoki1\" target=\"_blank\">@hyunseoki1</a>! Please see <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369886\" target=\"_blank\">here</a> for a discussion of this</p>\n<p>Essentially, this is a trick that improves the score given class imbalance!</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2064003,
      "author_name": "Remek Kinas",
      "author_url": "",
      "post_date": "2022-12-13T13:03:41.097000",
      "content": "<p>Interesting <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> - I am not able to setup bs more then 32 images on A100 (80GB vRAM) … (A6000 - bs = 16) for 1024x1024.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2064208,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2022-12-13T15:20:49.867000",
          "content": "<p>Mhmm not sure what might be going on there, something must be off. I was able to train with a BS of 24 with 48GB of RAM (using fp16, maybe that is the missing part?)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2064212,
          "author_name": "Remek Kinas",
          "author_url": "",
          "post_date": "2022-12-13T15:23:40.307000",
          "content": "<p>I use pure Pytorch code - implemented Automatic Mixed Precision inside train loop. Interesting. As I can see most of people has bigger bs than I have. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2064242,
          "author_name": "Kirderf",
          "author_url": "",
          "post_date": "2022-12-13T16:05:26.123000",
          "content": "<p>Maybe they are using p.p. img backgr. reduc. or accum. bs.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2065341,
          "author_name": "Remek Kinas",
          "author_url": "",
          "post_date": "2022-12-14T14:54:17.587000",
          "content": "<p>I fixed some lines of code and now bs for 1024 is 72 :) </p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 2073704,
      "author_name": "Chenjie",
      "author_url": "",
      "post_date": "2022-12-23T09:29:35.580000",
      "content": "<p>Thank you for sharing. I am new to this competition. May I ask why you call metric being very noisy? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 2131758,
          "author_name": "Lucas",
          "author_url": "",
          "post_date": "2023-02-06T10:26:35.433000",
          "content": "<p>Hi there, another beginner here, so don't take my word for it. I think he means by the metric being noisy that it's difficult to trust the metric. E.g. if you have one fold with a better metric value, it doesn't necessarily mean that the model is also better (generalizes to the test set), it might be just that the specific fold is easier to predict </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2072644,
      "author_name": "Khang Duong",
      "author_url": "",
      "post_date": "2022-12-22T09:52:41.610000",
      "content": "<p>Did you try Focal Loss? In my experiment, it increases my pfbeta score.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2069556,
      "author_name": "Feng Qilong",
      "author_url": "",
      "post_date": "2022-12-19T05:34:26.470000",
      "content": "<p>Hi! Do you mind sharing your local CV and the code for upsampling?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2069468,
      "author_name": "Tajinder",
      "author_url": "",
      "post_date": "2022-12-19T02:43:04.530000",
      "content": "<p><a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> I am a beginner, Thank you for all your posts, learning a lot.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2072180,
      "author_name": "Marcus Gray",
      "author_url": "",
      "post_date": "2022-12-21T19:51:10.777000",
      "content": "<p>Thanks for this.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2066748,
      "author_name": "Sudarshan Sahane",
      "author_url": "",
      "post_date": "2022-12-16T02:53:15.337000",
      "content": "<p>i am a biginner thank you</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2062859": "Hey,\n\nFirst and foremost, data processed by `dicomsdl` seems perfectly fine to train on. Processing data like this gives us a lot of time in the pipeline for inference, ensembling, etc.\n\n(sharing this up front as I believe this has been one of the unknowns at this point in the competition)\n\nI share the data I trained on here:\n\n[RSNA Mammo dicomsdl 1024](https://www.kaggle.com/datasets/radek1/rsna-mammo-dicomsdl-1024)\n\nI shared how I processed the data in the following notebook:\n\n[💡 how to process DICOM images to PNGs](https://www.kaggle.com/code/radek1/how-to-process-dicom-images-to-pngs)\n\nI didn't upload it to the dataset with all the other resolutions as at 13GB of size my Internet connection gave up, so I created this dataset directly from my GCP vm using the `kaggle cli`.\n\n### What are my observations from reaching LB 0.42?\n\nIt is best to approach this competition like any other computer vision project but with two important characteristics.\n\n* the dataset is imbalanced which makes training hard (plus some important features become discernable only at higher resolutions)\n* because of the class imbalance (there being very few positive examples) the metric is very noisy\n\nThe first point dictates that a major challenge of this competition will be helping our model to train. There are several ways one can go about this.\n\nI share the code I used for training and inference here:\n\n[🤖 [fast.ai starter pack] train + inference 🚀](https://www.kaggle.com/code/radek1/fast-ai-starter-pack-train-inference)\n\nBut I haven't updated the notebook with the hyperparameters I used for the last couple of runs. Here they are:\n\n* I trained with a weighted cross entropy loss (weight 20 for the positive class), this goes back to the idea of making training for our model easier\n* I upsampled the positive class by 3x\n* I trained on a resolution of 1024x1024 across 4 folds, using one cycle policy with a batch size of 24 and a lr of 2.25e-4 (this way I could train on a single GPU with 48GB of RAM, I trained for 7 or 8 epochs, training took around 30 minutes per epoch, you can achieve similar results with gradient accumulation -- natively supported by fast.ai -- on smaller GPUs)\n\nIt is very important though to take into account the metric being very noisy. You need to look for hyperparameters to maximize your performance across all your folds.\n\nMy LB score is an ensemble of the models across the 4 folds. Essentially, addressing the metric being so noisy (how to train, how to figure out which are genuine improvements, how not to overfit to a specific fold, etc) will be a big part of this competition.\n\nI just hope there is enough data in test and that it is similar to the data in train so that this doesn't turn into a luckfest when it comes to the private LB!\n\nAnyhow, hope this can be of help 🙏\n\nWishing you the best of luck in the competition! Happy Kaggling! 🙂",
    "2062897": "Nice, thanks for sharing. In my case I am using a standard pytorch pipeline, no tricks just some moderate augmentations and image size 1024 to achieve lb 0.51. I wonder if dicomsdl api gives the same image quality compared to pydicom.\n\nMaybe we can train models with image size 1024 using the 2 T4 GPU kaggle offer, that is 30 vram, I believe that is enough to train a small efficientnet with image size 1024",
    "2065068": "Hi @radek1 , thank you for all your posts ! If I could I ask a beginner question... I understood you trained your model with image resized to 1024 with dicomsdl, but I'm wondering how you manage this for the submission (same size for test images transformation ?), because I faced issues \"out of time\" ? ",
    "2065909": "Hi @redek1, I am beginner and thank you for very helpful post\nCan I ask you one question?\nI understand that pf1 is a kind of the F1 score for probability\nTherefore, it seems to reasonable that the input of pf1 should be probability number(output of sigmod or softmax) \nbut in the code shared by you, a binary value(thresholded number)\nI wonder why binary value is passed not probability",
    "2064003": "Interesting @radek1 - I am not able to setup bs more then 32 images on A100 (80GB vRAM) ... (A6000 - bs = 16) for 1024x1024.",
    "2073704": "Thank you for sharing. I am new to this competition. May I ask why you call metric being very noisy? ",
    "2072644": "Did you try Focal Loss? In my experiment, it increases my pfbeta score.",
    "2069556": "Hi! Do you mind sharing your local CV and the code for upsampling?",
    "2069468": "@radek1 I am a beginner, Thank you for all your posts, learning a lot.",
    "2072180": "Thanks for this.",
    "2066748": "i am a biginner thank you"
  }
}