{
  "id": 378521,
  "title": "share you top models validation metrics performance",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/378521",
  "author_name": "hengck23",
  "post_date": "2023-01-16T05:26:59.654000",
  "votes": 69,
  "comment_count": 57,
  "views": 0,
  "content": "<p>code to generate the results is in the attachment.<br>\nclick on the image for a bigger and clear one.</p>\n<p>I showed results for single fold model of different input image sizes, different model complexity.<br>\nAll use sampling of 1 pos in batch=8, except for effB2-2048x2848 that uses sampling of 1 pos in batch=16.<br>\nAll are trained on the same fold .</p>\n<p>observations:  </p>\n<ol>\n<li>prefer threshold that gives high percision  for approximate the same local validation f1.  </li>\n<li>site 1 and 2 have different bias. sometimes you will get high validation because you are overfitting stite 2 (at the expense of underfitting site 1). Check your site 1 validation score separately and make sure tat it is not to bad.  </li>\n</ol>\n<p>But we cannot confirm the correlations of these obersvations with the public LB unless we have more observations. So if you have top formance models, you can use my code to generate results and post it here!</p>\n<p><a href=\"https://ibb.co/QKs71mr\"><img src=\"https://i.ibb.co/CHxpT0h/Selection-559.png\" alt=\"Selection-559\"></a><br>\n<a href=\"https://ibb.co/F7GtH4b\"><img src=\"https://i.ibb.co/zfKw6Vh/Selection-558.png\" alt=\"Selection-558\"></a><br>\n<a href=\"https://ibb.co/hXxZq61\"><img src=\"https://i.ibb.co/gvbPn2F/Selection-557.png\" alt=\"Selection-557\"></a><br>\n<a href=\"https://ibb.co/NC8nZzV\"><img src=\"https://i.ibb.co/dMdJGyW/Selection-556.png\" alt=\"Selection-556\"></a><br>\n<a href=\"https://ibb.co/kGj80mS\"><img src=\"https://i.ibb.co/Fh2wJHD/Selection-555.png\" alt=\"Selection-555\"></a> </p>",
  "messages": [
    {
      "id": 2101688,
      "postDate": "2023-01-16T05:26:59.653Z",
      "content": "<p>code to generate the results is in the attachment.<br>\nclick on the image for a bigger and clear one.</p>\n<p>I showed results for single fold model of different input image sizes, different model complexity.<br>\nAll use sampling of 1 pos in batch=8, except for effB2-2048x2848 that uses sampling of 1 pos in batch=16.<br>\nAll are trained on the same fold .</p>\n<p>observations:  </p>\n<ol>\n<li>prefer threshold that gives high percision  for approximate the same local validation f1.  </li>\n<li>site 1 and 2 have different bias. sometimes you will get high validation because you are overfitting stite 2 (at the expense of underfitting site 1). Check your site 1 validation score separately and make sure tat it is not to bad.  </li>\n</ol>\n<p>But we cannot confirm the correlations of these obersvations with the public LB unless we have more observations. So if you have top formance models, you can use my code to generate results and post it here!</p>\n<p><a href=\"https://ibb.co/QKs71mr\"><img src=\"https://i.ibb.co/CHxpT0h/Selection-559.png\" alt=\"Selection-559\"></a><br>\n<a href=\"https://ibb.co/F7GtH4b\"><img src=\"https://i.ibb.co/zfKw6Vh/Selection-558.png\" alt=\"Selection-558\"></a><br>\n<a href=\"https://ibb.co/hXxZq61\"><img src=\"https://i.ibb.co/gvbPn2F/Selection-557.png\" alt=\"Selection-557\"></a><br>\n<a href=\"https://ibb.co/NC8nZzV\"><img src=\"https://i.ibb.co/dMdJGyW/Selection-556.png\" alt=\"Selection-556\"></a><br>\n<a href=\"https://ibb.co/kGj80mS\"><img src=\"https://i.ibb.co/Fh2wJHD/Selection-555.png\" alt=\"Selection-555\"></a> </p>",
      "rawMarkdown": "code to generate the results is in the attachment.\nclick on the image for a bigger and clear one.\n\nI showed results for single fold model of different input image sizes, different model complexity.\nAll use sampling of 1 pos in batch=8, except for effB2-2048x2848 that uses sampling of 1 pos in batch=16.\nAll are trained on the same fold .\n\n\nobservations:  \n1. prefer threshold that gives high percision  for approximate the same local validation f1.  \n2. site 1 and 2 have different bias. sometimes you will get high validation because you are overfitting stite 2 (at the expense of underfitting site 1). Check your site 1 validation score separately and make sure tat it is not to bad.  \n\nBut we cannot confirm the correlations of these obersvations with the public LB unless we have more observations. So if you have top formance models, you can use my code to generate results and post it here!\n\n\n<a href=\"https://ibb.co/QKs71mr\"><img src=\"https://i.ibb.co/CHxpT0h/Selection-559.png\" alt=\"Selection-559\" border=\"0\"></a>\n<a href=\"https://ibb.co/F7GtH4b\"><img src=\"https://i.ibb.co/zfKw6Vh/Selection-558.png\" alt=\"Selection-558\" border=\"0\"></a>\n<a href=\"https://ibb.co/hXxZq61\"><img src=\"https://i.ibb.co/gvbPn2F/Selection-557.png\" alt=\"Selection-557\" border=\"0\"></a>\n<a href=\"https://ibb.co/NC8nZzV\"><img src=\"https://i.ibb.co/dMdJGyW/Selection-556.png\" alt=\"Selection-556\" border=\"0\"></a>\n<a href=\"https://ibb.co/kGj80mS\"><img src=\"https://i.ibb.co/Fh2wJHD/Selection-555.png\" alt=\"Selection-555\" border=\"0\"></a> ",
      "votes": 69
    },
    {
      "id": 2107587,
      "postDate": "2023-01-19T21:57:45.773Z",
      "content": "<p>I think the precision part being more important than recall for the LB has to do with the false positive and false negative counts being present in the precision and recall and denominators respectively. <br>\nAs an example we might have 10000 negative cases and 250 positive cases in a fold, the maximum amount of false positives we can have is 10000, whereas we can only have 250 false negatives.<br>\nSo in general i think its way more punishing for the LB score if the model is predicting too many positive cases because it reduces the precision significantly, whereas with recall that is not the case. This is completely visible with any model when playing with different thresholds trying to increase recall.<br>\nIn my experiments this leads to the negative class having more impact on the LB score. surprisingly i have a better LB score when i give a bigger class weight to negative cases rather than positive ones. </p>\n<p>Or i might be completely wrong because my model is not complex enough to learn positive cases effectively. ( it's effnetB2)</p>\n<p>Edit: my setup is efficientnetB2, 2048x1024<br>\nmetrics: val_auc: 0.8427 - val_f1_score: 0.3931 - val_presicion: 0.6939 - val_recall: 0.2742 - val_acc: 0.9806 - LB=0.50@th:0.25, one fold only</p>",
      "rawMarkdown": "I think the precision part being more important than recall for the LB has to do with the false positive and false negative counts being present in the precision and recall and denominators respectively. \nAs an example we might have 10000 negative cases and 250 positive cases in a fold, the maximum amount of false positives we can have is 10000, whereas we can only have 250 false negatives.\nSo in general i think its way more punishing for the LB score if the model is predicting too many positive cases because it reduces the precision significantly, whereas with recall that is not the case. This is completely visible with any model when playing with different thresholds trying to increase recall.\nIn my experiments this leads to the negative class having more impact on the LB score. surprisingly i have a better LB score when i give a bigger class weight to negative cases rather than positive ones. \n\nOr i might be completely wrong because my model is not complex enough to learn positive cases effectively. ( it's effnetB2)\n\nEdit: my setup is efficientnetB2, 2048x1024\nmetrics: val_auc: 0.8427 - val_f1_score: 0.3931 - val_presicion: 0.6939 - val_recall: 0.2742 - val_acc: 0.9806 - LB=0.50@th:0.25, one fold only",
      "votes": 13
    },
    {
      "id": 2124002,
      "postDate": "2023-01-31T19:39:57.640Z",
      "content": "<p>Today a little bit better model (I will check it tomorrow on LB). I managed to get val probf1 &gt; 0.52 (single model, single fold, no tta). But still:</p>\n<ul>\n<li>site 1 only 0.35 (low prec/recall)</li>\n<li>site 2 high 0.71</li>\n</ul>\n<p><img src=\"https://i.ibb.co/txJ4H4k/002.jpg\" alt=\"\"></p>",
      "rawMarkdown": "Today a little bit better model (I will check it tomorrow on LB). I managed to get val probf1 > 0.52 (single model, single fold, no tta). But still:\n- site 1 only 0.35 (low prec/recall)\n- site 2 high 0.71\n\n![](https://i.ibb.co/txJ4H4k/002.jpg)",
      "votes": 7,
      "replies": [
        {
          "id": 2124044,
          "postDate": "2023-01-31T20:04:41.210Z",
          "content": "<p>If you don't me asking, what type of model is this?</p>\n<p>p.s. switched to paperspace growth (from your comments) and I've been getting a better CV. I'm porting a lot of code so I can take advantage of TPU here. Thanks for your tips!</p>",
          "rawMarkdown": "If you don't me asking, what type of model is this?\n\np.s. switched to paperspace growth (from your comments) and I've been getting a better CV. I'm porting a lot of code so I can take advantage of TPU here. Thanks for your tips!",
          "replies": [
            {
              "id": 2124101,
              "postDate": "2023-01-31T20:37:49.643Z",
              "content": "<p>Paperspace Growth works for me in this competition as well. </p>\n<p>Maybe I'll surprise you, but for me the model doesn't matter that much at the moment. I get very similar results (0.54-0.57) on any (effnet_v2_b2, nextvit proposed by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>, maxxvit, convnext or even … inception-v4). Probably the most important thing I introduced is:</p>\n<ul>\n<li>a different way of processing dataset - I think that most of the presented methods are not optimal (are not consistent)</li>\n<li>augmentation - some transformations work and some cause the result to drop a lot (although intuitively it would seem that they should help)</li>\n<li>control of the number of positive images in the batch (I wrote my own sampler)</li>\n<li>appropriate weight in BCEloss</li>\n<li>properly set dropouts in the network</li>\n</ul>\n<p>Now my single model (I do not blend models - maybe soon) is 0.57 LB. It is OK but … it is not able to see \"cancer\" in many positive images. For sure some post-processing is needed as well. So far I have not touched this part. </p>",
              "rawMarkdown": "Paperspace Growth works for me in this competition as well. \n\nMaybe I'll surprise you, but for me the model doesn't matter that much at the moment. I get very similar results (0.54-0.57) on any (effnet_v2_b2, nextvit proposed by @hengck23, maxxvit, convnext or even ... inception-v4). Probably the most important thing I introduced is:\n- a different way of processing dataset - I think that most of the presented methods are not optimal (are not consistent)\n- augmentation - some transformations work and some cause the result to drop a lot (although intuitively it would seem that they should help)\n- control of the number of positive images in the batch (I wrote my own sampler)\n- appropriate weight in BCEloss\n- properly set dropouts in the network\n\nNow my single model (I do not blend models - maybe soon) is 0.57 LB. It is OK but ... it is not able to see \"cancer\" in many positive images. For sure some post-processing is needed as well. So far I have not touched this part. ",
              "votes": 9
            },
            {
              "id": 2124152,
              "postDate": "2023-01-31T21:02:42.973Z",
              "content": "<p>Thanks for your comment. I have a few ideas for postprocessing but the dataset is killing me and I will do more testing and I think a different sampler is a good idea</p>",
              "rawMarkdown": "Thanks for your comment. I have a few ideas for postprocessing but the dataset is killing me and I will do more testing and I think a different sampler is a good idea"
            },
            {
              "id": 2127435,
              "postDate": "2023-02-02T22:07:18.083Z",
              "content": "<p><a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> </p>\n<p>I tested the data sampling strategy and I am confused about why is there such a big difference in F1 scores. Is a model able to generalize better in these noisy (or occulted) data points with a sampling strategy that ensures positive labels in each batch?</p>\n<p>Compared to another strategy that is random with the same positive labels per epoch my CV jumped by 0.05 F1 mean! (1 pos per 8 batch).</p>\n<p>Sorry if it is a simple question but I am just trying to understand the training better. It would be very helpful if you can point me to different resources to explore.</p>",
              "rawMarkdown": "@remekkinas @hengck23 \n\nI tested the data sampling strategy and I am confused about why is there such a big difference in F1 scores. Is a model able to generalize better in these noisy (or occulted) data points with a sampling strategy that ensures positive labels in each batch?\n\nCompared to another strategy that is random with the same positive labels per epoch my CV jumped by 0.05 F1 mean! (1 pos per 8 batch).\n\nSorry if it is a simple question but I am just trying to understand the training better. It would be very helpful if you can point me to different resources to explore.",
              "votes": 1
            },
            {
              "id": 2136293,
              "postDate": "2023-02-09T08:14:55.420Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2136294,
              "postDate": "2023-02-09T08:15:56.820Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2139601,
              "postDate": "2023-02-10T22:44:16.693Z",
              "content": "<p>I am finally able to achieve similar results with eff-b0 groupby mean 0.52F1 val. Thanks for your help and motivation!</p>",
              "rawMarkdown": "I am finally able to achieve similar results with eff-b0 groupby mean 0.52F1 val. Thanks for your help and motivation!"
            }
          ]
        },
        {
          "id": 2124072,
          "postDate": "2023-01-31T20:18:28.493Z",
          "content": "<p>Awesome.<br>\ndid you use full competition dataset for getting such awesome CV or used downsampled dataset?<br>\nwith full dataset i can't achieve such high CV. i don't know why but i have a feeling that there exist many label noise in the dataset,specially in negative samples,sorry if i am wrong</p>",
          "rawMarkdown": "Awesome.\ndid you use full competition dataset for getting such awesome CV or used downsampled dataset?\nwith full dataset i can't achieve such high CV. i don't know why but i have a feeling that there exist many label noise in the dataset,specially in negative samples,sorry if i am wrong",
          "replies": [
            {
              "id": 2124112,
              "postDate": "2023-01-31T20:43:21.473Z",
              "content": "<p><a href=\"https://www.kaggle.com/mobassir\" target=\"_blank\">@mobassir</a> I use full dataset (no external dataset). Made 4 folds (by patient_id). Result presented is from one fold only (validated on one fold, trained on 3 folds). I have not downsampled. I upsampled positive cancer by using custom sampler (I have not seen such implementation here).</p>\n<p>During my experiments I noticed that downsampling even by factor 0.4 did not influence final score a lot but allow to train faster. I use this during my experimentation sessions.   </p>",
              "rawMarkdown": "@mobassir I use full dataset (no external dataset). Made 4 folds (by patient_id). Result presented is from one fold only (validated on one fold, trained on 3 folds). I have not downsampled. I upsampled positive cancer by using custom sampler (I have not seen such implementation here).\n\nDuring my experiments I noticed that downsampling even by factor 0.4 did not influence final score a lot but allow to train faster. I use this during my experimentation sessions.   ",
              "votes": 4
            },
            {
              "id": 2124120,
              "postDate": "2023-01-31T20:49:01.850Z",
              "content": "<p>nice,so you think it's the custom sampler which is giving high CV boost for you?<br>\ndid you use weighted loss?</p>",
              "rawMarkdown": "nice,so you think it's the custom sampler which is giving high CV boost for you?\ndid you use weighted loss?"
            },
            {
              "id": 2124137,
              "postDate": "2023-01-31T20:56:37.120Z",
              "content": "<p>This is not sampler alone. In my case this is:</p>\n<ul>\n<li>dataset </li>\n<li>augumentation</li>\n<li>sampling ratio</li>\n<li>I use weighted loss (tested weighted bce, ce, focal loss, ldam and combination focall loss + ldam, + label smoothing) -&gt; the best for me is weighted bce  </li>\n</ul>\n<p>I tested many network architecture + combinations and it appeare that basic things works the best for me. But as you can see still not over 0.6 so … I missed something.</p>",
              "rawMarkdown": "This is not sampler alone. In my case this is:\n- dataset \n- augumentation\n- sampling ratio\n- I use weighted loss (tested weighted bce, ce, focal loss, ldam and combination focall loss + ldam, + label smoothing) -> the best for me is weighted bce  \n\nI tested many network architecture + combinations and it appeare that basic things works the best for me. But as you can see still not over 0.6 so ... I missed something.",
              "votes": 3
            },
            {
              "id": 2124160,
              "postDate": "2023-01-31T21:11:58.953Z",
              "content": "<p>\" I missed something.\"</p>\n<p>TTA will add 0.01<br>\ni haven't submit, but i think Multiple Image Prediction should add  &gt;0.01?</p>\n<p>\"effnet_v2_b2, nextvit proposed by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>, maxxvit, convnext or even … inception-v4\"  <br>\nthe surprise is inception-v4. but with many models tested, probabling the missing part is not the model.</p>\n<p>time to explore other option …</p>\n<p>explore to exploit rule is 30%  to 70% … don't stuck in your own local minimium<br>\n(Multi-Armed Bandits: Exploration versus Exploitation)</p>",
              "rawMarkdown": "\" I missed something.\"\n\nTTA will add 0.01\ni haven't submit, but i think Multiple Image Prediction should add  >0.01?\n\n\n\"effnet_v2_b2, nextvit proposed by @hengck23, maxxvit, convnext or even … inception-v4\"  \nthe surprise is inception-v4. but with many models tested, probabling the missing part is not the model.\n\ntime to explore other option ...\n\nexplore to exploit rule is 30%  to 70% ... don't stuck in your own local minimium\n(Multi-Armed Bandits: Exploration versus Exploitation)",
              "votes": 1
            },
            {
              "id": 2124167,
              "postDate": "2023-01-31T21:17:51.030Z",
              "content": "<p>Thank you! <br>\ninception-v4 gave me good result but … it has tendency to \"overfit\" very fast. In my tests they were very similar to effnet_b2. I agree with you - I have tested a lot of different models. Model was not game changer for me.</p>",
              "rawMarkdown": "Thank you! \ninception-v4 gave me good result but ... it has tendency to \"overfit\" very fast. In my tests they were very similar to effnet_b2. I agree with you - I have tested a lot of different models. Model was not game changer for me."
            },
            {
              "id": 2125253,
              "postDate": "2023-02-01T14:49:18.513Z",
              "content": "<p>Thanks for sharing, nice jump <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> </p>\n<p>Also to add, make sure to tune the threshold based on TTA predictions, I have been using threshold generated using the same model w/o TTA assuming there won't be much change in predictions, and consistently got bad results on lb.</p>",
              "rawMarkdown": "Thanks for sharing, nice jump @remekkinas \n\nAlso to add, make sure to tune the threshold based on TTA predictions, I have been using threshold generated using the same model w/o TTA assuming there won't be much change in predictions, and consistently got bad results on lb.",
              "votes": 1
            },
            {
              "id": 2125271,
              "postDate": "2023-02-01T15:04:53.213Z",
              "content": "<p>I am devastated by this competition 😂😨😂😂 <br>\nNow I have (in theory) better model (local validation) but it performs on LB worse then my previous one (from the same fold and model architecture). </p>\n<p><a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> I agree with you. Now I can see … but … I am afraid about this:</p>\n<ul>\n<li>TH=0.5 -&gt; LB = 0.49</li>\n<li>TH=0.6 -&gt; LB = 0.56</li>\n</ul>\n<p>If we have only part of test set … I even do not want to think about full test score 😂😂😂</p>",
              "rawMarkdown": "I am devastated by this competition 😂😨😂😂 \nNow I have (in theory) better model (local validation) but it performs on LB worse then my previous one (from the same fold and model architecture). \n\n@nischaydnk I agree with you. Now I can see ... but ... I am afraid about this:\n- TH=0.5 -> LB = 0.49\n- TH=0.6 -> LB = 0.56\n\nIf we have only part of test set ... I even do not want to think about full test score 😂😂😂",
              "votes": 4
            },
            {
              "id": 2126660,
              "postDate": "2023-02-02T11:23:04.830Z",
              "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> - do you remember my local val score from first post (probf1: 0.52) ? max LB score I was able to get is 0.47 😂😨😂😂</p>",
              "rawMarkdown": "@hengck23 - do you remember my local val score from first post (probf1: 0.52) ? max LB score I was able to get is 0.47 😂😨😂😂\n",
              "votes": 1
            }
          ]
        },
        {
          "id": 2129001,
          "postDate": "2023-02-04T08:45:51.020Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a>, thank for your contribution. <br>\nCan you explain a litttle bit about your result feature:<br>\nWhat is \"Single image\", \"group mean()\", \"group max()\"?</p>",
          "rawMarkdown": "Hi @remekkinas, thank for your contribution. \nCan you explain a litttle bit about your result feature:\nWhat is \"Single image\", \"group mean()\", \"group max()\"?",
          "replies": [
            {
              "id": 2129120,
              "postDate": "2023-02-04T10:52:28.483Z",
              "content": "<p>This is validation script provided by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>. </p>\n<p>Single image is validation result (probf1 score binarized) based on single images (without grouping).<br>\nmax - aggregation by prediction_id and max proba<br>\nmean - aggregation by prediction_id and mean proba </p>\n<p>If you have any question let me know. I will help and answer.</p>",
              "rawMarkdown": "This is validation script provided by @hengck23. \n\nSingle image is validation result (probf1 score binarized) based on single images (without grouping).\nmax - aggregation by prediction_id and max proba\nmean - aggregation by prediction_id and mean proba \n\nIf you have any question let me know. I will help and answer.",
              "votes": 2
            },
            {
              "id": 2129788,
              "postDate": "2023-02-04T21:37:08.047Z",
              "content": "<p>Thank you for your answering 🔥💯<br>\nI have 2 more questions: <br>\n1, What does each of this block refer to? <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6230460%2F9d34d94823239475a94eb883e73f1d58%2FAnnotation%202023-02-05%20043215.png?generation=1675546361728079&amp;alt=media\" alt=\"\"></p>\n<p>2, I still don't get the idea of single image: Let's say I have 1000 images on validation set, so each image will have 1 result (probf1 score binarized). =&gt; 1000 results =&gt; which results is logged? </p>",
              "rawMarkdown": "Thank you for your answering 🔥💯\nI have 2 more questions: \n1, What does each of this block refer to? \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6230460%2F9d34d94823239475a94eb883e73f1d58%2FAnnotation%202023-02-05%20043215.png?generation=1675546361728079&alt=media)\n\n2, I still don't get the idea of single image: Let's say I have 1000 images on validation set, so each image will have 1 result (probf1 score binarized). => 1000 results => which results is logged? "
            },
            {
              "id": 2129833,
              "postDate": "2023-02-04T22:52:37.273Z",
              "content": "<p>Site:<br>\n0 - global (site_id_1+site_id_2)<br>\n1 - site_id = 1<br>\n2 - site_id = 2</p>",
              "rawMarkdown": "Site:\n0 - global (site_id_1+site_id_2)\n1 - site_id = 1\n2 - site_id = 2"
            },
            {
              "id": 2141106,
              "postDate": "2023-02-12T13:06:42.090Z",
              "content": "<p><a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> I'm confused about the \"group by mean\", \"group by max\" part too, could you please tell me more details or make an example about it:) Thank you!</p>",
              "rawMarkdown": "@remekkinas I'm confused about the \"group by mean\", \"group by max\" part too, could you please tell me more details or make an example about it:) Thank you!"
            }
          ]
        }
      ]
    },
    {
      "id": 2106196,
      "postDate": "2023-01-19T02:37:53.293Z",
      "content": "<p>Are the other fold scores about the same?<br>\nJust as we should not put too much trust in PublicLB, we should not put too much trust in 1fold results. The f1 score in particular seems to have a lot of variance.</p>",
      "rawMarkdown": "Are the other fold scores about the same?\nJust as we should not put too much trust in PublicLB, we should not put too much trust in 1fold results. The f1 score in particular seems to have a lot of variance.",
      "votes": 1,
      "replies": [
        {
          "id": 2106268,
          "postDate": "2023-01-19T04:05:10.640Z",
          "content": "<p><img src=\"https://i.ibb.co/FssPwXN/Selection-588.png\" alt=\"https://i.ibb.co/FssPwXN/Selection-588.png\"></p>\n<p>fold0 and fold1 for next-vit-B.</p>\n<p>i will says that :</p>\n<ol>\n<li>the predicted negative distribution is consistent</li>\n<li>at high precision, results are more stable</li>\n</ol>\n<p>we can only know the answer if more kagglers post their results.</p>",
          "rawMarkdown": "![https://i.ibb.co/FssPwXN/Selection-588.png](https://i.ibb.co/FssPwXN/Selection-588.png)\n\nfold0 and fold1 for next-vit-B.\n\ni will says that :\n1. the predicted negative distribution is consistent\n2. at high precision, results are more stable\n\nwe can only know the answer if more kagglers post their results.",
          "votes": 3,
          "replies": [
            {
              "id": 2106346,
              "postDate": "2023-01-19T05:16:20.677Z",
              "content": "<p>So fold1 is a better score. Thanks!</p>",
              "rawMarkdown": "So fold1 is a better score. Thanks!"
            }
          ]
        }
      ]
    },
    {
      "id": 2102632,
      "postDate": "2023-01-16T19:02:00.527Z",
      "content": "<p>while it is training in progress, here is the interesting results of convnextv2 tiny<br>\n<img src=\"https://i.ibb.co/LCs2LFS/Selection-577.png\" alt=\"https://i.ibb.co/LCs2LFS/Selection-577.png\"></p>",
      "rawMarkdown": "while it is training in progress, here is the interesting results of convnextv2 tiny\n![https://i.ibb.co/LCs2LFS/Selection-577.png](https://i.ibb.co/LCs2LFS/Selection-577.png)",
      "votes": 1
    },
    {
      "id": 2102265,
      "postDate": "2023-01-16T13:48:40.770Z",
      "content": "<p><a href=\"https://imgbb.com/\"><img src=\"https://i.ibb.co/Tb5p4mw/Selection-564.png\" alt=\"Selection-564\"></a></p>\n<p>i do not know which of the above should be optimum validation distribution (i.e. not overfitted).<br>\nbut you can get different distribution by:</p>\n<ol>\n<li>control the number of training epoches</li>\n<li>increase of decrease the aggressive of loss: e.g. BCE is essentially log loss of log(1+exp(-yf)). the more aggresive version is exp loss exp(-yf)</li>\n<li>margin or class, sample weighning</li>\n<li>post-mapping, e.g. apply sqrt to predicted probability</li>\n</ol>",
      "rawMarkdown": "<a href=\"https://imgbb.com/\"><img src=\"https://i.ibb.co/Tb5p4mw/Selection-564.png\" alt=\"Selection-564\" border=\"0\"></a>\n\ni do not know which of the above should be optimum validation distribution (i.e. not overfitted).\nbut you can get different distribution by:\n1. control the number of training epoches\n2. increase of decrease the aggressive of loss: e.g. BCE is essentially log loss of log(1+exp(-yf)). the more aggresive version is exp loss exp(-yf)\n3. margin or class, sample weighning\n4. post-mapping, e.g. apply sqrt to predicted probability\n",
      "votes": 1,
      "replies": [
        {
          "id": 2102943,
          "postDate": "2023-01-16T22:42:38.693Z",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> you could try training on predicted oof data as pseudo labels  and see how well it scores when trying to predict in fold data as a way of measuring over fitting.</p>\n<p>See my post here - <a href=\"https://www.kaggle.com/competitions/novozymes-enzyme-stability-prediction/discussion/378527\" target=\"_blank\">https://www.kaggle.com/competitions/novozymes-enzyme-stability-prediction/discussion/378527</a></p>\n<p>Theory is that there should be greater coherence / less overfitting when pseudo labels score best when used for training </p>",
          "rawMarkdown": "@hengck23 you could try training on predicted oof data as pseudo labels  and see how well it scores when trying to predict in fold data as a way of measuring over fitting.\n\nSee my post here - https://www.kaggle.com/competitions/novozymes-enzyme-stability-prediction/discussion/378527\n\nTheory is that there should be greater coherence / less overfitting when pseudo labels score best when used for training ",
          "replies": [
            {
              "id": 2106529,
              "postDate": "2023-01-19T08:00:27.810Z",
              "content": "<p>fwiw, this technique also worked with the latest playground here - <a href=\"https://www.kaggle.com/competitions/playground-series-s3e3/discussion/379347\" target=\"_blank\">https://www.kaggle.com/competitions/playground-series-s3e3/discussion/379347</a></p>",
              "rawMarkdown": "fwiw, this technique also worked with the latest playground here - https://www.kaggle.com/competitions/playground-series-s3e3/discussion/379347"
            },
            {
              "id": 2107836,
              "postDate": "2023-01-20T05:34:34.220Z",
              "content": "<p>it is better to use e.g. convnext, efficientl2(finetuned on some dataset or other correlated label #)  to annotate token label (pseudo label) and then use this as aux loss to train vision trasnformer</p>\n<p>[1]All Tokens Matter: Token Labeling for Training Better Vision Transformers<br>\n<a href=\"https://proceedings.neurips.cc/paper/2021/file/9a49a25d845a483fae4be7e341368e36-Paper.pdf\" target=\"_blank\">https://proceedings.neurips.cc/paper/2021/file/9a49a25d845a483fae4be7e341368e36-Paper.pdf</a></p>\n<p>e.g. use can train a classifier to label biopsy and non biopsy with class activation map an use it as token label</p>",
              "rawMarkdown": "it is better to use e.g. convnext, efficientl2(finetuned on some dataset or other correlated label #)  to annotate token label (pseudo label) and then use this as aux loss to train vision trasnformer\n\n[1]All Tokens Matter: Token Labeling for Training Better Vision Transformers\nhttps://proceedings.neurips.cc/paper/2021/file/9a49a25d845a483fae4be7e341368e36-Paper.pdf\n\ne.g. use can train a classifier to label biopsy and non biopsy with class activation map an use it as token label",
              "votes": 1
            },
            {
              "id": 2108529,
              "postDate": "2023-01-20T15:43:48.717Z",
              "content": "<p>Hmm, not sure you got my intention.  It isn't really for pseudo labeling / supplementing training data but rather measuring fitness of the fold.</p>",
              "rawMarkdown": "Hmm, not sure you got my intention.  It isn't really for pseudo labeling / supplementing training data but rather measuring fitness of the fold."
            }
          ]
        }
      ]
    },
    {
      "id": 2152074,
      "postDate": "2023-02-20T14:49:23.160Z",
      "content": "<p>Hi! <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>, I got a image-level pF1 score of 0.41, after aggregated by mean, I got a pF1 score of 0.52, But my LB only got 0.5, I wonder what happened? I was expecting to get a 0.6 LB score… </p>",
      "rawMarkdown": "Hi! @hengck23, I got a image-level pF1 score of 0.41, after aggregated by mean, I got a pF1 score of 0.52, But my LB only got 0.5, I wonder what happened? I was expecting to get a 0.6 LB score... "
    },
    {
      "id": 2123196,
      "postDate": "2023-01-31T11:24:04.680Z",
      "content": "<p>thanx a lot for your sharing, it seems useful)</p>",
      "rawMarkdown": "thanx a lot for your sharing, it seems useful)"
    },
    {
      "id": 2121433,
      "postDate": "2023-01-30T10:06:25.823Z",
      "content": "<p>it's my best LB result (LB=0.56)<br>\nThanks.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3830852%2F7a0b307594493d43b6ec481570003dfd%2FRNSA_results.png?generation=1675073213216103&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3830852%2F99977bd5960aedaf3d9aff8ef07440e9%2FRNSA_results2.png?generation=1675073223983235&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "it's my best LB result (LB=0.56)\nThanks.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3830852%2F7a0b307594493d43b6ec481570003dfd%2FRNSA_results.png?generation=1675073213216103&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3830852%2F99977bd5960aedaf3d9aff8ef07440e9%2FRNSA_results2.png?generation=1675073223983235&alt=media)",
      "replies": [
        {
          "id": 2121486,
          "postDate": "2023-01-30T10:54:44.897Z",
          "content": "<p><a href=\"https://www.kaggle.com/taruto1215\" target=\"_blank\">@taruto1215</a> <br>\ncheck this as well: <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/370333#2121485\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/370333#2121485</a></p>",
          "rawMarkdown": "@taruto1215 \ncheck this as well: https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/370333#2121485",
          "votes": 1,
          "replies": [
            {
              "id": 2121499,
              "postDate": "2023-01-30T11:13:37.273Z",
              "content": "<p>A reasonable idea! Thanks!</p>",
              "rawMarkdown": "A reasonable idea! Thanks!"
            }
          ]
        }
      ]
    },
    {
      "id": 2116712,
      "postDate": "2023-01-26T16:56:51.847Z",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> when you report your metrics are they with respect to the validation set or an OOF? Are you using the same augmentations on the validation set that you used in the training set?</p>",
      "rawMarkdown": "@hengck23 when you report your metrics are they with respect to the validation set or an OOF? Are you using the same augmentations on the validation set that you used in the training set?",
      "replies": [
        {
          "id": 2116801,
          "postDate": "2023-01-26T18:28:29.937Z",
          "content": "<p>only one fold is shown (so there is no OOF)</p>",
          "rawMarkdown": "only one fold is shown (so there is no OOF)"
        }
      ]
    },
    {
      "id": 2112478,
      "postDate": "2023-01-23T16:45:34.733Z",
      "content": "<p>Thank you so much for sharing, I learned a lot and I have a question: Are the above calculations based on the pretrained model or calculated from scratch?</p>",
      "rawMarkdown": "Thank you so much for sharing, I learned a lot and I have a question: Are the above calculations based on the pretrained model or calculated from scratch?"
    },
    {
      "id": 2108883,
      "postDate": "2023-01-20T22:49:45Z",
      "content": "<p>Can I ask how do you ensure at least one positive sample in each batch? are there any tutorials or tools for that?</p>",
      "rawMarkdown": "Can I ask how do you ensure at least one positive sample in each batch? are there any tutorials or tools for that?",
      "replies": [
        {
          "id": 2109148,
          "postDate": "2023-01-21T07:02:12.003Z",
          "content": "<p>You can use balancesampler in dataloader. Code Frog Guy (hengck23) has provided.</p>",
          "rawMarkdown": "You can use balancesampler in dataloader. Code Frog Guy (hengck23) has provided."
        }
      ]
    },
    {
      "id": 2108383,
      "postDate": "2023-01-20T13:46:07.993Z",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> , Great post! How long does NextVit converge in your experiment? In my experiment, nextvit converge much slower than efficientnet.</p>",
      "rawMarkdown": "@hengck23 , Great post! How long does NextVit converge in your experiment? In my experiment, nextvit converge much slower than efficientnet."
    },
    {
      "id": 2108212,
      "postDate": "2023-01-20T10:53:28.607Z",
      "content": "<p>Thanks for all your contributions and posts! 👍 Is it 3 folds we see in the picture, results seperate per fold or together? E.g. [3] is 3 folds average or fold #3.</p>\n<p>Is it the same configuration in the benchmark besides the model specific setup, e.g. same batch size etc per dim. training?</p>",
      "rawMarkdown": "Thanks for all your contributions and posts! 👍 Is it 3 folds we see in the picture, results seperate per fold or together? E.g. [3] is 3 folds average or fold #3.\n\nIs it the same configuration in the benchmark besides the model specific setup, e.g. same batch size etc per dim. training?",
      "replies": [
        {
          "id": 2108238,
          "postDate": "2023-01-20T11:12:37.333Z",
          "content": "<p>[0] all images<br>\n[1] images from site id 1<br>\n[2] images from site id 2</p>\n<p>same fold means same train and validation images.<br>\nother setting are mainly the same, but there are some small differences<br>\n(e.g. difference batch size due to memory constraint image resolution, different learning rate for CNN, VIT)</p>",
          "rawMarkdown": "[0] all images\n[1] images from site id 1\n[2] images from site id 2\n\nsame fold means same train and validation images.\nother setting are mainly the same, but there are some small differences\n(e.g. difference batch size due to memory constraint image resolution, different learning rate for CNN, VIT)"
        }
      ]
    },
    {
      "id": 2102609,
      "postDate": "2023-01-16T18:38:23.073Z",
      "content": "<p>thank you for sharing. [0] [1] [2] are different folds same setup? they seem to have huge differences.</p>",
      "rawMarkdown": "thank you for sharing. [0] [1] [2] are different folds same setup? they seem to have huge differences.",
      "replies": [
        {
          "id": 2102626,
          "postDate": "2023-01-16T18:57:07.830Z",
          "content": "<p>[1] is site id 1, [2] is site id 2, [0] is both site id (i.e. one fold)  </p>\n<blockquote>\n  <p>I showed results for single fold model of different input image sizes, different model complexity. </p>\n</blockquote>",
          "rawMarkdown": "[1] is site id 1, [2] is site id 2, [0] is both site id (i.e. one fold)  \n\n>I showed results for single fold model of different input image sizes, different model complexity. ",
          "votes": 1,
          "replies": [
            {
              "id": 2102644,
              "postDate": "2023-01-16T19:12:24.603Z",
              "content": "<p>ah makes sense, thank you</p>",
              "rawMarkdown": "ah makes sense, thank you",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2102601,
      "postDate": "2023-01-16T18:23:03.330Z",
      "content": "<p>Thank you for sharing.</p>\n<p>BTW Are you using excel for experiment tracking? If yes why do you not use a tailored tool e.g. Weights&amp;Biases, Neptune.ai, …?</p>",
      "rawMarkdown": "Thank you for sharing.\n\nBTW Are you using excel for experiment tracking? If yes why do you not use a tailored tool e.g. Weights&Biases, Neptune.ai, ...?"
    },
    {
      "id": 2102111,
      "postDate": "2023-01-16T12:00:57.157Z",
      "content": "<p>May I ask how you set the validation dataset? I use one fold as the validation but the auc is like 0.97, close to the neg/pos proportion.<br>\nThanks!</p>",
      "rawMarkdown": "May I ask how you set the validation dataset? I use one fold as the validation but the auc is like 0.97, close to the neg/pos proportion.\nThanks!"
    },
    {
      "id": 2102097,
      "postDate": "2023-01-16T11:56:20.707Z",
      "content": "<p>Thanks! This is a great work! May I ask what is site1 and 2 ? Sorry about my rookie question…</p>",
      "rawMarkdown": "Thanks! This is a great work! May I ask what is site1 and 2 ? Sorry about my rookie question...",
      "replies": [
        {
          "id": 2102129,
          "postDate": "2023-01-16T12:08:18.667Z",
          "content": "<p>Look into dataset - location where images were taken.</p>",
          "rawMarkdown": "Look into dataset - location where images were taken."
        }
      ]
    },
    {
      "id": 2101933,
      "postDate": "2023-01-16T10:00:32.037Z",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> Hi, Can I ask you that you crop breast during training or crop in advance then save then training?</p>",
      "rawMarkdown": "@hengck23 Hi, Can I ask you that you crop breast during training or crop in advance then save then training?",
      "replies": [
        {
          "id": 2108524,
          "postDate": "2023-01-20T15:41:12.423Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 2102317,
      "postDate": "2023-01-16T14:27:46.690Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 2102347,
          "postDate": "2023-01-16T14:44:59.050Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 2102291,
      "postDate": "2023-01-16T14:04:43.773Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2107587,
      "author_name": "Abdolkarim Saeedi",
      "author_url": "",
      "post_date": "2023-01-19T21:57:45.773000",
      "content": "<p>I think the precision part being more important than recall for the LB has to do with the false positive and false negative counts being present in the precision and recall and denominators respectively. <br>\nAs an example we might have 10000 negative cases and 250 positive cases in a fold, the maximum amount of false positives we can have is 10000, whereas we can only have 250 false negatives.<br>\nSo in general i think its way more punishing for the LB score if the model is predicting too many positive cases because it reduces the precision significantly, whereas with recall that is not the case. This is completely visible with any model when playing with different thresholds trying to increase recall.<br>\nIn my experiments this leads to the negative class having more impact on the LB score. surprisingly i have a better LB score when i give a bigger class weight to negative cases rather than positive ones. </p>\n<p>Or i might be completely wrong because my model is not complex enough to learn positive cases effectively. ( it's effnetB2)</p>\n<p>Edit: my setup is efficientnetB2, 2048x1024<br>\nmetrics: val_auc: 0.8427 - val_f1_score: 0.3931 - val_presicion: 0.6939 - val_recall: 0.2742 - val_acc: 0.9806 - LB=0.50@th:0.25, one fold only</p>",
      "votes": 13,
      "replies": []
    },
    {
      "id": 2124002,
      "author_name": "Remek Kinas",
      "author_url": "",
      "post_date": "2023-01-31T19:39:57.640000",
      "content": "<p>Today a little bit better model (I will check it tomorrow on LB). I managed to get val probf1 &gt; 0.52 (single model, single fold, no tta). But still:</p>\n<ul>\n<li>site 1 only 0.35 (low prec/recall)</li>\n<li>site 2 high 0.71</li>\n</ul>\n<p><img src=\"https://i.ibb.co/txJ4H4k/002.jpg\" alt=\"\"></p>",
      "votes": 7,
      "replies": [
        {
          "id": 2124044,
          "author_name": "outwrest",
          "author_url": "",
          "post_date": "2023-01-31T20:04:41.210000",
          "content": "<p>If you don't me asking, what type of model is this?</p>\n<p>p.s. switched to paperspace growth (from your comments) and I've been getting a better CV. I'm porting a lot of code so I can take advantage of TPU here. Thanks for your tips!</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2124101,
              "author_name": "Remek Kinas",
              "author_url": "",
              "post_date": "2023-01-31T20:37:49.643000",
              "content": "<p>Paperspace Growth works for me in this competition as well. </p>\n<p>Maybe I'll surprise you, but for me the model doesn't matter that much at the moment. I get very similar results (0.54-0.57) on any (effnet_v2_b2, nextvit proposed by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>, maxxvit, convnext or even … inception-v4). Probably the most important thing I introduced is:</p>\n<ul>\n<li>a different way of processing dataset - I think that most of the presented methods are not optimal (are not consistent)</li>\n<li>augmentation - some transformations work and some cause the result to drop a lot (although intuitively it would seem that they should help)</li>\n<li>control of the number of positive images in the batch (I wrote my own sampler)</li>\n<li>appropriate weight in BCEloss</li>\n<li>properly set dropouts in the network</li>\n</ul>\n<p>Now my single model (I do not blend models - maybe soon) is 0.57 LB. It is OK but … it is not able to see \"cancer\" in many positive images. For sure some post-processing is needed as well. So far I have not touched this part. </p>",
              "votes": 9,
              "replies": []
            },
            {
              "id": 2124152,
              "author_name": "outwrest",
              "author_url": "",
              "post_date": "2023-01-31T21:02:42.973000",
              "content": "<p>Thanks for your comment. I have a few ideas for postprocessing but the dataset is killing me and I will do more testing and I think a different sampler is a good idea</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2127435,
              "author_name": "outwrest",
              "author_url": "",
              "post_date": "2023-02-02T22:07:18.083000",
              "content": "<p><a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> </p>\n<p>I tested the data sampling strategy and I am confused about why is there such a big difference in F1 scores. Is a model able to generalize better in these noisy (or occulted) data points with a sampling strategy that ensures positive labels in each batch?</p>\n<p>Compared to another strategy that is random with the same positive labels per epoch my CV jumped by 0.05 F1 mean! (1 pos per 8 batch).</p>\n<p>Sorry if it is a simple question but I am just trying to understand the training better. It would be very helpful if you can point me to different resources to explore.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2136293,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-02-09T08:14:55.420000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2136294,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-02-09T08:15:56.820000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2139601,
              "author_name": "outwrest",
              "author_url": "",
              "post_date": "2023-02-10T22:44:16.693000",
              "content": "<p>I am finally able to achieve similar results with eff-b0 groupby mean 0.52F1 val. Thanks for your help and motivation!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2124072,
          "author_name": "Mobassir",
          "author_url": "",
          "post_date": "2023-01-31T20:18:28.493000",
          "content": "<p>Awesome.<br>\ndid you use full competition dataset for getting such awesome CV or used downsampled dataset?<br>\nwith full dataset i can't achieve such high CV. i don't know why but i have a feeling that there exist many label noise in the dataset,specially in negative samples,sorry if i am wrong</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2124112,
              "author_name": "Remek Kinas",
              "author_url": "",
              "post_date": "2023-01-31T20:43:21.473000",
              "content": "<p><a href=\"https://www.kaggle.com/mobassir\" target=\"_blank\">@mobassir</a> I use full dataset (no external dataset). Made 4 folds (by patient_id). Result presented is from one fold only (validated on one fold, trained on 3 folds). I have not downsampled. I upsampled positive cancer by using custom sampler (I have not seen such implementation here).</p>\n<p>During my experiments I noticed that downsampling even by factor 0.4 did not influence final score a lot but allow to train faster. I use this during my experimentation sessions.   </p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 2124120,
              "author_name": "Mobassir",
              "author_url": "",
              "post_date": "2023-01-31T20:49:01.850000",
              "content": "<p>nice,so you think it's the custom sampler which is giving high CV boost for you?<br>\ndid you use weighted loss?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2124137,
              "author_name": "Remek Kinas",
              "author_url": "",
              "post_date": "2023-01-31T20:56:37.120000",
              "content": "<p>This is not sampler alone. In my case this is:</p>\n<ul>\n<li>dataset </li>\n<li>augumentation</li>\n<li>sampling ratio</li>\n<li>I use weighted loss (tested weighted bce, ce, focal loss, ldam and combination focall loss + ldam, + label smoothing) -&gt; the best for me is weighted bce  </li>\n</ul>\n<p>I tested many network architecture + combinations and it appeare that basic things works the best for me. But as you can see still not over 0.6 so … I missed something.</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2124160,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-01-31T21:11:58.953000",
              "content": "<p>\" I missed something.\"</p>\n<p>TTA will add 0.01<br>\ni haven't submit, but i think Multiple Image Prediction should add  &gt;0.01?</p>\n<p>\"effnet_v2_b2, nextvit proposed by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>, maxxvit, convnext or even … inception-v4\"  <br>\nthe surprise is inception-v4. but with many models tested, probabling the missing part is not the model.</p>\n<p>time to explore other option …</p>\n<p>explore to exploit rule is 30%  to 70% … don't stuck in your own local minimium<br>\n(Multi-Armed Bandits: Exploration versus Exploitation)</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2124167,
              "author_name": "Remek Kinas",
              "author_url": "",
              "post_date": "2023-01-31T21:17:51.030000",
              "content": "<p>Thank you! <br>\ninception-v4 gave me good result but … it has tendency to \"overfit\" very fast. In my tests they were very similar to effnet_b2. I agree with you - I have tested a lot of different models. Model was not game changer for me.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2125253,
              "author_name": "Nischay Dhankhar",
              "author_url": "",
              "post_date": "2023-02-01T14:49:18.513000",
              "content": "<p>Thanks for sharing, nice jump <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> </p>\n<p>Also to add, make sure to tune the threshold based on TTA predictions, I have been using threshold generated using the same model w/o TTA assuming there won't be much change in predictions, and consistently got bad results on lb.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2125271,
              "author_name": "Remek Kinas",
              "author_url": "",
              "post_date": "2023-02-01T15:04:53.213000",
              "content": "<p>I am devastated by this competition 😂😨😂😂 <br>\nNow I have (in theory) better model (local validation) but it performs on LB worse then my previous one (from the same fold and model architecture). </p>\n<p><a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> I agree with you. Now I can see … but … I am afraid about this:</p>\n<ul>\n<li>TH=0.5 -&gt; LB = 0.49</li>\n<li>TH=0.6 -&gt; LB = 0.56</li>\n</ul>\n<p>If we have only part of test set … I even do not want to think about full test score 😂😂😂</p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 2126660,
              "author_name": "Remek Kinas",
              "author_url": "",
              "post_date": "2023-02-02T11:23:04.830000",
              "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> - do you remember my local val score from first post (probf1: 0.52) ? max LB score I was able to get is 0.47 😂😨😂😂</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 2129001,
          "author_name": "happymentee",
          "author_url": "",
          "post_date": "2023-02-04T08:45:51.020000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a>, thank for your contribution. <br>\nCan you explain a litttle bit about your result feature:<br>\nWhat is \"Single image\", \"group mean()\", \"group max()\"?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2129120,
              "author_name": "Remek Kinas",
              "author_url": "",
              "post_date": "2023-02-04T10:52:28.483000",
              "content": "<p>This is validation script provided by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>. </p>\n<p>Single image is validation result (probf1 score binarized) based on single images (without grouping).<br>\nmax - aggregation by prediction_id and max proba<br>\nmean - aggregation by prediction_id and mean proba </p>\n<p>If you have any question let me know. I will help and answer.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2129788,
              "author_name": "happymentee",
              "author_url": "",
              "post_date": "2023-02-04T21:37:08.047000",
              "content": "<p>Thank you for your answering 🔥💯<br>\nI have 2 more questions: <br>\n1, What does each of this block refer to? <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6230460%2F9d34d94823239475a94eb883e73f1d58%2FAnnotation%202023-02-05%20043215.png?generation=1675546361728079&amp;alt=media\" alt=\"\"></p>\n<p>2, I still don't get the idea of single image: Let's say I have 1000 images on validation set, so each image will have 1 result (probf1 score binarized). =&gt; 1000 results =&gt; which results is logged? </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2129833,
              "author_name": "Remek Kinas",
              "author_url": "",
              "post_date": "2023-02-04T22:52:37.273000",
              "content": "<p>Site:<br>\n0 - global (site_id_1+site_id_2)<br>\n1 - site_id = 1<br>\n2 - site_id = 2</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2141106,
              "author_name": "Lau2664",
              "author_url": "",
              "post_date": "2023-02-12T13:06:42.090000",
              "content": "<p><a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> I'm confused about the \"group by mean\", \"group by max\" part too, could you please tell me more details or make an example about it:) Thank you!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2106196,
      "author_name": "YujiAriyasu",
      "author_url": "",
      "post_date": "2023-01-19T02:37:53.293000",
      "content": "<p>Are the other fold scores about the same?<br>\nJust as we should not put too much trust in PublicLB, we should not put too much trust in 1fold results. The f1 score in particular seems to have a lot of variance.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2106268,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2023-01-19T04:05:10.640000",
          "content": "<p><img src=\"https://i.ibb.co/FssPwXN/Selection-588.png\" alt=\"https://i.ibb.co/FssPwXN/Selection-588.png\"></p>\n<p>fold0 and fold1 for next-vit-B.</p>\n<p>i will says that :</p>\n<ol>\n<li>the predicted negative distribution is consistent</li>\n<li>at high precision, results are more stable</li>\n</ol>\n<p>we can only know the answer if more kagglers post their results.</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2106346,
              "author_name": "YujiAriyasu",
              "author_url": "",
              "post_date": "2023-01-19T05:16:20.677000",
              "content": "<p>So fold1 is a better score. Thanks!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2102632,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-01-16T19:02:00.527000",
      "content": "<p>while it is training in progress, here is the interesting results of convnextv2 tiny<br>\n<img src=\"https://i.ibb.co/LCs2LFS/Selection-577.png\" alt=\"https://i.ibb.co/LCs2LFS/Selection-577.png\"></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2102265,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-01-16T13:48:40.770000",
      "content": "<p><a href=\"https://imgbb.com/\"><img src=\"https://i.ibb.co/Tb5p4mw/Selection-564.png\" alt=\"Selection-564\"></a></p>\n<p>i do not know which of the above should be optimum validation distribution (i.e. not overfitted).<br>\nbut you can get different distribution by:</p>\n<ol>\n<li>control the number of training epoches</li>\n<li>increase of decrease the aggressive of loss: e.g. BCE is essentially log loss of log(1+exp(-yf)). the more aggresive version is exp loss exp(-yf)</li>\n<li>margin or class, sample weighning</li>\n<li>post-mapping, e.g. apply sqrt to predicted probability</li>\n</ol>",
      "votes": 1,
      "replies": [
        {
          "id": 2102943,
          "author_name": "@kaggleqrdl",
          "author_url": "",
          "post_date": "2023-01-16T22:42:38.693000",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> you could try training on predicted oof data as pseudo labels  and see how well it scores when trying to predict in fold data as a way of measuring over fitting.</p>\n<p>See my post here - <a href=\"https://www.kaggle.com/competitions/novozymes-enzyme-stability-prediction/discussion/378527\" target=\"_blank\">https://www.kaggle.com/competitions/novozymes-enzyme-stability-prediction/discussion/378527</a></p>\n<p>Theory is that there should be greater coherence / less overfitting when pseudo labels score best when used for training </p>",
          "votes": 0,
          "replies": [
            {
              "id": 2106529,
              "author_name": "@kaggleqrdl",
              "author_url": "",
              "post_date": "2023-01-19T08:00:27.810000",
              "content": "<p>fwiw, this technique also worked with the latest playground here - <a href=\"https://www.kaggle.com/competitions/playground-series-s3e3/discussion/379347\" target=\"_blank\">https://www.kaggle.com/competitions/playground-series-s3e3/discussion/379347</a></p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2107836,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-01-20T05:34:34.220000",
              "content": "<p>it is better to use e.g. convnext, efficientl2(finetuned on some dataset or other correlated label #)  to annotate token label (pseudo label) and then use this as aux loss to train vision trasnformer</p>\n<p>[1]All Tokens Matter: Token Labeling for Training Better Vision Transformers<br>\n<a href=\"https://proceedings.neurips.cc/paper/2021/file/9a49a25d845a483fae4be7e341368e36-Paper.pdf\" target=\"_blank\">https://proceedings.neurips.cc/paper/2021/file/9a49a25d845a483fae4be7e341368e36-Paper.pdf</a></p>\n<p>e.g. use can train a classifier to label biopsy and non biopsy with class activation map an use it as token label</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2108529,
              "author_name": "@kaggleqrdl",
              "author_url": "",
              "post_date": "2023-01-20T15:43:48.717000",
              "content": "<p>Hmm, not sure you got my intention.  It isn't really for pseudo labeling / supplementing training data but rather measuring fitness of the fold.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2152074,
      "author_name": "Lau2664",
      "author_url": "",
      "post_date": "2023-02-20T14:49:23.160000",
      "content": "<p>Hi! <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>, I got a image-level pF1 score of 0.41, after aggregated by mean, I got a pF1 score of 0.52, But my LB only got 0.5, I wonder what happened? I was expecting to get a 0.6 LB score… </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2123196,
      "author_name": "nikiduki",
      "author_url": "",
      "post_date": "2023-01-31T11:24:04.680000",
      "content": "<p>thanx a lot for your sharing, it seems useful)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2121433,
      "author_name": "taruto",
      "author_url": "",
      "post_date": "2023-01-30T10:06:25.823000",
      "content": "<p>it's my best LB result (LB=0.56)<br>\nThanks.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3830852%2F7a0b307594493d43b6ec481570003dfd%2FRNSA_results.png?generation=1675073213216103&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3830852%2F99977bd5960aedaf3d9aff8ef07440e9%2FRNSA_results2.png?generation=1675073223983235&amp;alt=media\" alt=\"\"></p>",
      "votes": 0,
      "replies": [
        {
          "id": 2121486,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2023-01-30T10:54:44.897000",
          "content": "<p><a href=\"https://www.kaggle.com/taruto1215\" target=\"_blank\">@taruto1215</a> <br>\ncheck this as well: <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/370333#2121485\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/370333#2121485</a></p>",
          "votes": 1,
          "replies": [
            {
              "id": 2121499,
              "author_name": "taruto",
              "author_url": "",
              "post_date": "2023-01-30T11:13:37.273000",
              "content": "<p>A reasonable idea! Thanks!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2116712,
      "author_name": "moth",
      "author_url": "",
      "post_date": "2023-01-26T16:56:51.847000",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> when you report your metrics are they with respect to the validation set or an OOF? Are you using the same augmentations on the validation set that you used in the training set?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2116801,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2023-01-26T18:28:29.937000",
          "content": "<p>only one fold is shown (so there is no OOF)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2112478,
      "author_name": "Mr.G",
      "author_url": "",
      "post_date": "2023-01-23T16:45:34.733000",
      "content": "<p>Thank you so much for sharing, I learned a lot and I have a question: Are the above calculations based on the pretrained model or calculated from scratch?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2108883,
      "author_name": "Sihead",
      "author_url": "",
      "post_date": "2023-01-20T22:49:45",
      "content": "<p>Can I ask how do you ensure at least one positive sample in each batch? are there any tutorials or tools for that?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2109148,
          "author_name": "huaqiliang",
          "author_url": "",
          "post_date": "2023-01-21T07:02:12.003000",
          "content": "<p>You can use balancesampler in dataloader. Code Frog Guy (hengck23) has provided.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2108383,
      "author_name": "ForcewithMe",
      "author_url": "",
      "post_date": "2023-01-20T13:46:07.993000",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> , Great post! How long does NextVit converge in your experiment? In my experiment, nextvit converge much slower than efficientnet.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2108212,
      "author_name": "Kirderf",
      "author_url": "",
      "post_date": "2023-01-20T10:53:28.607000",
      "content": "<p>Thanks for all your contributions and posts! 👍 Is it 3 folds we see in the picture, results seperate per fold or together? E.g. [3] is 3 folds average or fold #3.</p>\n<p>Is it the same configuration in the benchmark besides the model specific setup, e.g. same batch size etc per dim. training?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2108238,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2023-01-20T11:12:37.333000",
          "content": "<p>[0] all images<br>\n[1] images from site id 1<br>\n[2] images from site id 2</p>\n<p>same fold means same train and validation images.<br>\nother setting are mainly the same, but there are some small differences<br>\n(e.g. difference batch size due to memory constraint image resolution, different learning rate for CNN, VIT)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2102609,
      "author_name": "Eleftherios Fanioudakis",
      "author_url": "",
      "post_date": "2023-01-16T18:38:23.073000",
      "content": "<p>thank you for sharing. [0] [1] [2] are different folds same setup? they seem to have huge differences.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2102626,
          "author_name": "RB",
          "author_url": "",
          "post_date": "2023-01-16T18:57:07.830000",
          "content": "<p>[1] is site id 1, [2] is site id 2, [0] is both site id (i.e. one fold)  </p>\n<blockquote>\n  <p>I showed results for single fold model of different input image sizes, different model complexity. </p>\n</blockquote>",
          "votes": 1,
          "replies": [
            {
              "id": 2102644,
              "author_name": "Eleftherios Fanioudakis",
              "author_url": "",
              "post_date": "2023-01-16T19:12:24.603000",
              "content": "<p>ah makes sense, thank you</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2102601,
      "author_name": "Ali Abdin",
      "author_url": "",
      "post_date": "2023-01-16T18:23:03.330000",
      "content": "<p>Thank you for sharing.</p>\n<p>BTW Are you using excel for experiment tracking? If yes why do you not use a tailored tool e.g. Weights&amp;Biases, Neptune.ai, …?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2102111,
      "author_name": "nicehzj",
      "author_url": "",
      "post_date": "2023-01-16T12:00:57.157000",
      "content": "<p>May I ask how you set the validation dataset? I use one fold as the validation but the auc is like 0.97, close to the neg/pos proportion.<br>\nThanks!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2102097,
      "author_name": "nicehzj",
      "author_url": "",
      "post_date": "2023-01-16T11:56:20.707000",
      "content": "<p>Thanks! This is a great work! May I ask what is site1 and 2 ? Sorry about my rookie question…</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2102129,
          "author_name": "Remek Kinas",
          "author_url": "",
          "post_date": "2023-01-16T12:08:18.667000",
          "content": "<p>Look into dataset - location where images were taken.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2101933,
      "author_name": "Mad_Neil",
      "author_url": "",
      "post_date": "2023-01-16T10:00:32.037000",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> Hi, Can I ask you that you crop breast during training or crop in advance then save then training?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2108524,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-01-20T15:41:12.423000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2102317,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-01-16T14:27:46.690000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 2102347,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-01-16T14:44:59.050000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2102291,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-01-16T14:04:43.773000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2101688": "code to generate the results is in the attachment.\nclick on the image for a bigger and clear one.\n\nI showed results for single fold model of different input image sizes, different model complexity.\nAll use sampling of 1 pos in batch=8, except for effB2-2048x2848 that uses sampling of 1 pos in batch=16.\nAll are trained on the same fold .\n\n\nobservations:  \n1. prefer threshold that gives high percision  for approximate the same local validation f1.  \n2. site 1 and 2 have different bias. sometimes you will get high validation because you are overfitting stite 2 (at the expense of underfitting site 1). Check your site 1 validation score separately and make sure tat it is not to bad.  \n\nBut we cannot confirm the correlations of these obersvations with the public LB unless we have more observations. So if you have top formance models, you can use my code to generate results and post it here!\n\n\n<a href=\"https://ibb.co/QKs71mr\"><img src=\"https://i.ibb.co/CHxpT0h/Selection-559.png\" alt=\"Selection-559\" border=\"0\"></a>\n<a href=\"https://ibb.co/F7GtH4b\"><img src=\"https://i.ibb.co/zfKw6Vh/Selection-558.png\" alt=\"Selection-558\" border=\"0\"></a>\n<a href=\"https://ibb.co/hXxZq61\"><img src=\"https://i.ibb.co/gvbPn2F/Selection-557.png\" alt=\"Selection-557\" border=\"0\"></a>\n<a href=\"https://ibb.co/NC8nZzV\"><img src=\"https://i.ibb.co/dMdJGyW/Selection-556.png\" alt=\"Selection-556\" border=\"0\"></a>\n<a href=\"https://ibb.co/kGj80mS\"><img src=\"https://i.ibb.co/Fh2wJHD/Selection-555.png\" alt=\"Selection-555\" border=\"0\"></a> ",
    "2107587": "I think the precision part being more important than recall for the LB has to do with the false positive and false negative counts being present in the precision and recall and denominators respectively. \nAs an example we might have 10000 negative cases and 250 positive cases in a fold, the maximum amount of false positives we can have is 10000, whereas we can only have 250 false negatives.\nSo in general i think its way more punishing for the LB score if the model is predicting too many positive cases because it reduces the precision significantly, whereas with recall that is not the case. This is completely visible with any model when playing with different thresholds trying to increase recall.\nIn my experiments this leads to the negative class having more impact on the LB score. surprisingly i have a better LB score when i give a bigger class weight to negative cases rather than positive ones. \n\nOr i might be completely wrong because my model is not complex enough to learn positive cases effectively. ( it's effnetB2)\n\nEdit: my setup is efficientnetB2, 2048x1024\nmetrics: val_auc: 0.8427 - val_f1_score: 0.3931 - val_presicion: 0.6939 - val_recall: 0.2742 - val_acc: 0.9806 - LB=0.50@th:0.25, one fold only",
    "2124002": "Today a little bit better model (I will check it tomorrow on LB). I managed to get val probf1 > 0.52 (single model, single fold, no tta). But still:\n- site 1 only 0.35 (low prec/recall)\n- site 2 high 0.71\n\n![](https://i.ibb.co/txJ4H4k/002.jpg)",
    "2106196": "Are the other fold scores about the same?\nJust as we should not put too much trust in PublicLB, we should not put too much trust in 1fold results. The f1 score in particular seems to have a lot of variance.",
    "2102632": "while it is training in progress, here is the interesting results of convnextv2 tiny\n![https://i.ibb.co/LCs2LFS/Selection-577.png](https://i.ibb.co/LCs2LFS/Selection-577.png)",
    "2102265": "<a href=\"https://imgbb.com/\"><img src=\"https://i.ibb.co/Tb5p4mw/Selection-564.png\" alt=\"Selection-564\" border=\"0\"></a>\n\ni do not know which of the above should be optimum validation distribution (i.e. not overfitted).\nbut you can get different distribution by:\n1. control the number of training epoches\n2. increase of decrease the aggressive of loss: e.g. BCE is essentially log loss of log(1+exp(-yf)). the more aggresive version is exp loss exp(-yf)\n3. margin or class, sample weighning\n4. post-mapping, e.g. apply sqrt to predicted probability\n",
    "2152074": "Hi! @hengck23, I got a image-level pF1 score of 0.41, after aggregated by mean, I got a pF1 score of 0.52, But my LB only got 0.5, I wonder what happened? I was expecting to get a 0.6 LB score... ",
    "2123196": "thanx a lot for your sharing, it seems useful)",
    "2121433": "it's my best LB result (LB=0.56)\nThanks.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3830852%2F7a0b307594493d43b6ec481570003dfd%2FRNSA_results.png?generation=1675073213216103&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3830852%2F99977bd5960aedaf3d9aff8ef07440e9%2FRNSA_results2.png?generation=1675073223983235&alt=media)",
    "2116712": "@hengck23 when you report your metrics are they with respect to the validation set or an OOF? Are you using the same augmentations on the validation set that you used in the training set?",
    "2112478": "Thank you so much for sharing, I learned a lot and I have a question: Are the above calculations based on the pretrained model or calculated from scratch?",
    "2108883": "Can I ask how do you ensure at least one positive sample in each batch? are there any tutorials or tools for that?",
    "2108383": "@hengck23 , Great post! How long does NextVit converge in your experiment? In my experiment, nextvit converge much slower than efficientnet.",
    "2108212": "Thanks for all your contributions and posts! 👍 Is it 3 folds we see in the picture, results seperate per fold or together? E.g. [3] is 3 folds average or fold #3.\n\nIs it the same configuration in the benchmark besides the model specific setup, e.g. same batch size etc per dim. training?",
    "2102609": "thank you for sharing. [0] [1] [2] are different folds same setup? they seem to have huge differences.",
    "2102601": "Thank you for sharing.\n\nBTW Are you using excel for experiment tracking? If yes why do you not use a tailored tool e.g. Weights&Biases, Neptune.ai, ...?",
    "2102111": "May I ask how you set the validation dataset? I use one fold as the validation but the auc is like 0.97, close to the neg/pos proportion.\nThanks!",
    "2102097": "Thanks! This is a great work! May I ask what is site1 and 2 ? Sorry about my rookie question...",
    "2101933": "@hengck23 Hi, Can I ask you that you crop breast during training or crop in advance then save then training?",
    "2102317": "",
    "2102291": ""
  }
}