{
  "id": 382198,
  "title": "The score gap is huge ! (CV: 0.395 LB: 0.56)",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/382198",
  "author_name": "taruto",
  "post_date": "2023-01-30T01:04:17.820000",
  "votes": 20,
  "comment_count": 78,
  "views": 0,
  "content": "<p>there is a huge score gap between my cv and my lb CV: 0.395 (after mean aggregation ) LB: 0.56 (thres=0.20, no strict tunning).<br>\nI am very concern about shake down.<br>\nHow is your score gap??<br>\nthanks.</p>\n<p>my cv strategy is below. (n_splits=3)<br>\nI refer <a href=\"https://www.kaggle.com/code/awsaf49/rsna-bcd-efficientnet-tf-tpu-1vm-train\" target=\"_blank\">this notebook</a></p>\n<pre><code>strat_cols = [\n    'laterality', 'view', 'biopsy','invasive', 'BIRADS', 'age_bin',\n    'implant', 'density','machine_id', 'difficult_negative_case',\n    'cancer',\n]\n\ndf['stratify'] = ''\nfor col in strat_cols:\n    df['stratify'] += df[col].astype(str)\n\nkfold = StratifiedGroupKFold(n_splits=Config['SPLITS'],random_state=Config[\"seed\"],shuffle=True)\nfor fold_, (train_idx, valid_idx) in enumerate(kfold.split(df, df['stratify'].values, df['patient_id'].values)):`\n</code></pre>\n<p><strong>update2 : CV score details (n_splits=3)</strong><br>\n<strong>all data：precision=0.485 recall=0.333 pf1=0.395 thres=0.5000 roc_auc=0.873</strong><br>\n<strong>site_id=1 : precision=0.329 recall=0.287 f1=0.307 thres=0.4500 roc_auc=0.826</strong><br>\n<strong>site_id=2 : precision=0.639 recall=0.416 f1=0.504 thres=0.5000 roc_auc=0.906</strong></p>",
  "messages": [
    {
      "id": 2121006,
      "postDate": "2023-01-30T01:04:17.820Z",
      "content": "<p>there is a huge score gap between my cv and my lb CV: 0.395 (after mean aggregation ) LB: 0.56 (thres=0.20, no strict tunning).<br>\nI am very concern about shake down.<br>\nHow is your score gap??<br>\nthanks.</p>\n<p>my cv strategy is below. (n_splits=3)<br>\nI refer <a href=\"https://www.kaggle.com/code/awsaf49/rsna-bcd-efficientnet-tf-tpu-1vm-train\" target=\"_blank\">this notebook</a></p>\n<pre><code>strat_cols = [\n    'laterality', 'view', 'biopsy','invasive', 'BIRADS', 'age_bin',\n    'implant', 'density','machine_id', 'difficult_negative_case',\n    'cancer',\n]\n\ndf['stratify'] = ''\nfor col in strat_cols:\n    df['stratify'] += df[col].astype(str)\n\nkfold = StratifiedGroupKFold(n_splits=Config['SPLITS'],random_state=Config[\"seed\"],shuffle=True)\nfor fold_, (train_idx, valid_idx) in enumerate(kfold.split(df, df['stratify'].values, df['patient_id'].values)):`\n</code></pre>\n<p><strong>update2 : CV score details (n_splits=3)</strong><br>\n<strong>all data：precision=0.485 recall=0.333 pf1=0.395 thres=0.5000 roc_auc=0.873</strong><br>\n<strong>site_id=1 : precision=0.329 recall=0.287 f1=0.307 thres=0.4500 roc_auc=0.826</strong><br>\n<strong>site_id=2 : precision=0.639 recall=0.416 f1=0.504 thres=0.5000 roc_auc=0.906</strong></p>",
      "rawMarkdown": "there is a huge score gap between my cv and my lb CV: 0.395 (after mean aggregation ) LB: 0.56 (thres=0.20, no strict tunning).\nI am very concern about shake down.\nHow is your score gap??\nthanks.\n\nmy cv strategy is below. (n_splits=3)\nI refer [this notebook](https://www.kaggle.com/code/awsaf49/rsna-bcd-efficientnet-tf-tpu-1vm-train)\n\n \n    strat_cols = [\n        'laterality', 'view', 'biopsy','invasive', 'BIRADS', 'age_bin',\n        'implant', 'density','machine_id', 'difficult_negative_case',\n        'cancer',\n    ]\n\n    df['stratify'] = ''\n    for col in strat_cols:\n        df['stratify'] += df[col].astype(str)\n    \n    kfold = StratifiedGroupKFold(n_splits=Config['SPLITS'],random_state=Config[\"seed\"],shuffle=True)\n    for fold_, (train_idx, valid_idx) in enumerate(kfold.split(df, df['stratify'].values, df['patient_id'].values)):`\n\n**update2 : CV score details (n_splits=3)**\n**all data：precision=0.485 recall=0.333 pf1=0.395 thres=0.5000 roc_auc=0.873**\n**site_id=1 : precision=0.329 recall=0.287 f1=0.307 thres=0.4500 roc_auc=0.826**\n**site_id=2 : precision=0.639 recall=0.416 f1=0.504 thres=0.5000 roc_auc=0.906**",
      "votes": 20
    },
    {
      "id": 2131916,
      "postDate": "2023-02-06T13:37:45.833Z",
      "content": "<p>Finally my LB is more than 0.60.<br>\nBut my cv score is 0.42.<br>\nThe gap is too big !!!!!!</p>",
      "rawMarkdown": "Finally my LB is more than 0.60.\nBut my cv score is 0.42.\nThe gap is too big !!!!!!",
      "votes": 7,
      "replies": [
        {
          "id": 2131925,
          "postDate": "2023-02-06T13:42:11.230Z",
          "content": "<p>any post-processing?</p>",
          "rawMarkdown": "any post-processing?",
          "replies": [
            {
              "id": 2131955,
              "postDate": "2023-02-06T14:02:18.667Z",
              "content": "<p>No post-processing</p>",
              "rawMarkdown": "No post-processing"
            }
          ]
        },
        {
          "id": 2132155,
          "postDate": "2023-02-06T16:23:15.843Z",
          "content": "<p>Your LB score is very impressive. Could you tell me your image size did you use?</p>",
          "rawMarkdown": "Your LB score is very impressive. Could you tell me your image size did you use?",
          "replies": [
            {
              "id": 2133133,
              "postDate": "2023-02-07T09:04:35.323Z",
              "content": "<p>Thanks.<br>\ncurrent my image size is 1536x960</p>",
              "rawMarkdown": "Thanks.\ncurrent my image size is 1536x960",
              "votes": 3
            }
          ]
        },
        {
          "id": 2132435,
          "postDate": "2023-02-06T20:31:11.657Z",
          "content": "<p>Amazing result! Congratulations crossing 0.6x.<br>\nDo you still use effnet? </p>\n<p>BTW:<br>\nI have local CV: 0.54 now … and this model is only 0.51 on LB 😂</p>",
          "rawMarkdown": "Amazing result! Congratulations crossing 0.6x.\nDo you still use effnet? \n\nBTW:\nI have local CV: 0.54 now ... and this model is only 0.51 on LB 😂",
          "votes": 1,
          "replies": [
            {
              "id": 2133141,
              "postDate": "2023-02-07T09:11:54.400Z",
              "content": "<p>Thank you!<br>\nI still use the effnet.</p>\n<p>You have a good local CV!<br>\nI don't still know why my CV is too low… (it bothers me)</p>\n<p>How is the image level CV score for your model?</p>",
              "rawMarkdown": "Thank you!\nI still use the effnet.\n\nYou have a good local CV!\nI don't still know why my CV is too low... (it bothers me)\n\nHow is the image level CV score for your model?"
            },
            {
              "id": 2133283,
              "postDate": "2023-02-07T10:44:32.670Z",
              "content": "<p>For single image 0.4.</p>",
              "rawMarkdown": "For single image 0.4.",
              "votes": 1
            },
            {
              "id": 2133298,
              "postDate": "2023-02-07T10:51:47.230Z",
              "content": "<p>it's a also higher score than my image level cv score 0.356.<br>\nthanks for your sharing.</p>",
              "rawMarkdown": "it's a also higher score than my image level cv score 0.356.\nthanks for your sharing."
            },
            {
              "id": 2133300,
              "postDate": "2023-02-07T10:54:18.857Z",
              "content": "<p>If it is possible please share result (validator result). It would be great. Especially I am looking now on predicition dynamics (chart with probability distribution).</p>\n<p>As far as I understand you blend 3 models (3 folds) for final solution?</p>",
              "rawMarkdown": "If it is possible please share result (validator result). It would be great. Especially I am looking now on predicition dynamics (chart with probability distribution).\n\nAs far as I understand you blend 3 models (3 folds) for final solution?"
            },
            {
              "id": 2133432,
              "postDate": "2023-02-07T12:11:27.647Z",
              "content": "<p>I share you my best cv result.<br>\nYes, I just blended 3 models (3folds).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3830852%2F92b091a4351296e585ce247fbd6008c3%2FRSNA_LB063_results2.png?generation=1675771723070708&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3830852%2Fdc4e15caabae71248a9097dfeaebcc08%2FRSNA_LB063_results.png?generation=1675771667026332&amp;alt=media\" alt=\"\"></p>",
              "rawMarkdown": "I share you my best cv result.\nYes, I just blended 3 models (3folds).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3830852%2F92b091a4351296e585ce247fbd6008c3%2FRSNA_LB063_results2.png?generation=1675771723070708&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3830852%2Fdc4e15caabae71248a9097dfeaebcc08%2FRSNA_LB063_results.png?generation=1675771667026332&alt=media)",
              "votes": 3
            },
            {
              "id": 2133443,
              "postDate": "2023-02-07T12:18:58.940Z",
              "content": "<p>Thank you  a lot.👍<br>\nStats you provided are from one model or blended?</p>",
              "rawMarkdown": "Thank you  a lot.👍\nStats you provided are from one model or blended?"
            },
            {
              "id": 2133445,
              "postDate": "2023-02-07T12:22:53.867Z",
              "content": "<p>It's blended from 3 models (3folds)</p>",
              "rawMarkdown": "It's blended from 3 models (3folds)",
              "votes": 1
            },
            {
              "id": 2134843,
              "postDate": "2023-02-08T09:32:32.167Z",
              "content": "<p><a href=\"https://www.kaggle.com/taruto1215\" target=\"_blank\">@taruto1215</a> how many epochs do you train model? Or from what epoch do you choose models from? </p>",
              "rawMarkdown": "@taruto1215 how many epochs do you train model? Or from what epoch do you choose models from? "
            },
            {
              "id": 2134858,
              "postDate": "2023-02-08T09:39:15.480Z",
              "content": "<p><a href=\"https://www.kaggle.com/taruto1215\" target=\"_blank\">@taruto1215</a> may i know what validation pf1,auc and thresholded pf1 you get for each fold? thanks</p>",
              "rawMarkdown": "@taruto1215 may i know what validation pf1,auc and thresholded pf1 you get for each fold? thanks"
            },
            {
              "id": 2134978,
              "postDate": "2023-02-08T11:21:40.770Z",
              "content": "<p><a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> <br>\nI train model 5 epochs.<br>\nthen I choose a model has a best image-level f1 score.<br>\nBut almost of them has the best score at the 5 epoch.</p>",
              "rawMarkdown": "@remekkinas \nI train model 5 epochs.\nthen I choose a model has a best image-level f1 score.\nBut almost of them has the best score at the 5 epoch.",
              "votes": 1
            },
            {
              "id": 2134989,
              "postDate": "2023-02-08T11:36:44.180Z",
              "content": "<p><a href=\"https://www.kaggle.com/mobassir\" target=\"_blank\">@mobassir</a> <br>\nThis is my validation results for each fold.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3830852%2F3e806521970616e694389151d535198b%2FRSNA_results_each_folds.png?generation=1675856167389336&amp;alt=media\" alt=\"\"></p>",
              "rawMarkdown": "@mobassir \nThis is my validation results for each fold.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3830852%2F3e806521970616e694389151d535198b%2FRSNA_results_each_folds.png?generation=1675856167389336&alt=media)",
              "votes": 1
            },
            {
              "id": 2135052,
              "postDate": "2023-02-08T12:21:43.520Z",
              "content": "<p>i ahve been trying another experiment. if you are training too many epoch, then the predicted distribution will be too sharp (i.e. overfitting).</p>\n<p>hence one way is to stop early at, say 5 epoch.<br>\nwe train many of differenet models.</p>\n<p>then we freeze them and use them as feature extractor.<br>\nwe use giba trick (winner of petfinder): <a href=\"https://developer.nvidia.com/blog/fast-fine-tuning-of-ai-transformers-using-rapids-machine-learning/\" target=\"_blank\">https://developer.nvidia.com/blog/fast-fine-tuning-of-ai-transformers-using-rapids-machine-learning/</a><br>\n<a href=\"https://www.kaggle.com/c/petfinder-pawpularity-score/discussion/301686\" target=\"_blank\">https://www.kaggle.com/c/petfinder-pawpularity-score/discussion/301686</a></p>\n<p>but instead, we create ensmble for multi-view model. each image is now treated as a feature map in sequence, and trsnformer is used to fuse the feature of different view, together with other attribute like age, site id.</p>",
              "rawMarkdown": "i ahve been trying another experiment. if you are training too many epoch, then the predicted distribution will be too sharp (i.e. overfitting).\n\nhence one way is to stop early at, say 5 epoch.\nwe train many of differenet models.\n\nthen we freeze them and use them as feature extractor.\nwe use giba trick (winner of petfinder): https://developer.nvidia.com/blog/fast-fine-tuning-of-ai-transformers-using-rapids-machine-learning/\nhttps://www.kaggle.com/c/petfinder-pawpularity-score/discussion/301686\n\nbut instead, we create ensmble for multi-view model. each image is now treated as a feature map in sequence, and trsnformer is used to fuse the feature of different view, together with other attribute like age, site id.",
              "votes": 3
            },
            {
              "id": 2135314,
              "postDate": "2023-02-08T14:59:35.790Z",
              "content": "<p>Thanks I’ll check it<br>\nBut from tomorrow I’ll be very busy at my job . I will not have more time to spend Kaggle than now hahaha😂</p>",
              "rawMarkdown": "Thanks I’ll check it\nBut from tomorrow I’ll be very busy at my job . I will not have more time to spend Kaggle than now hahaha😂"
            }
          ]
        },
        {
          "id": 2132600,
          "postDate": "2023-02-06T22:42:30.270Z",
          "content": "<p>Is it a single fold model or an ensemble ?</p>",
          "rawMarkdown": "Is it a single fold model or an ensemble ?",
          "votes": 1,
          "replies": [
            {
              "id": 2133145,
              "postDate": "2023-02-07T09:15:00.843Z",
              "content": "<p>Number of folds is 3.<br>\nso I used 3 models for submission.</p>",
              "rawMarkdown": "Number of folds is 3.\nso I used 3 models for submission.",
              "votes": 1
            },
            {
              "id": 2141316,
              "postDate": "2023-02-12T16:29:02.437Z",
              "content": "<p>Hi! <a href=\"https://www.kaggle.com/taruto1215\" target=\"_blank\">@taruto1215</a> Can you tell me how do you use 3 models in your inference notebook to make a submission?</p>",
              "rawMarkdown": "Hi! @taruto1215 Can you tell me how do you use 3 models in your inference notebook to make a submission?"
            }
          ]
        },
        {
          "id": 2132927,
          "postDate": "2023-02-07T06:04:44.393Z",
          "content": "<p>5 fold<br>\nCV:0.416<br>\nLB:0.58<br>\nMaybe public leaderboard data are simpler than train data</p>",
          "rawMarkdown": "5 fold\nCV:0.416\nLB:0.58\nMaybe public leaderboard data are simpler than train data",
          "votes": 1,
          "replies": [
            {
              "id": 2133148,
              "postDate": "2023-02-07T09:19:05.653Z",
              "content": "<p>It's the same symptom as mine haha.<br>\nHow about your ROC_AUC for CV??</p>\n<p>I guess so too for the public leaderboard data. <br>\nBut for the private leaderboard data … <br>\nI hope it's same.</p>",
              "rawMarkdown": "It's the same symptom as mine haha.\nHow about your ROC_AUC for CV??\n\nI guess so too for the public leaderboard data. \nBut for the private leaderboard data ... \nI hope it's same."
            },
            {
              "id": 2133176,
              "postDate": "2023-02-07T09:46:15.880Z",
              "content": "<p>ROC_AUC(group by mean) ~0.88</p>",
              "rawMarkdown": "ROC_AUC(group by mean) ~0.88",
              "votes": 1
            },
            {
              "id": 2133293,
              "postDate": "2023-02-07T10:49:20.730Z",
              "content": "<p>thanks for your sharing.<br>\nmy ROC_AUC (group by mean) ~0.896<br>\nI think roc_auc scores has correlation against LB scores.</p>",
              "rawMarkdown": "thanks for your sharing.\nmy ROC_AUC (group by mean) ~0.896\nI think roc_auc scores has correlation against LB scores.",
              "votes": 1
            },
            {
              "id": 2133302,
              "postDate": "2023-02-07T10:55:23.120Z",
              "content": "<p>The same here - ROV_OUC for me is over 0.91 and LB is lower then yours. But I am still talking about single model in my case. All blending solutions failed for me.</p>",
              "rawMarkdown": "The same here - ROV_OUC for me is over 0.91 and LB is lower then yours. But I am still talking about single model in my case. All blending solutions failed for me.",
              "votes": 2
            },
            {
              "id": 2133442,
              "postDate": "2023-02-07T12:18:47.177Z",
              "content": "<p>The ROC_AUC score is also good!<br>\nPerhaps I should suspect that I am overfitting to LB.</p>\n<p>why do you talking about only single model?<br>\nI think we should check oof result.</p>",
              "rawMarkdown": "The ROC_AUC score is also good!\nPerhaps I should suspect that I am overfitting to LB.\n\nwhy do you talking about only single model?\nI think we should check oof result."
            },
            {
              "id": 2133473,
              "postDate": "2023-02-07T12:48:30.570Z",
              "content": "<p>To be honest? I do not know at this time. Fianally I will sub two solution: single model and blend. </p>",
              "rawMarkdown": "To be honest? I do not know at this time. Fianally I will sub two solution: single model and blend. ",
              "votes": 1
            },
            {
              "id": 2142767,
              "postDate": "2023-02-13T18:58:58.997Z",
              "content": "<p><a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> </p>\n<p>I want to follow up on this, my fold 0 results seem good with CV .52 (mean, LB 0.52) similar to the results you posted before but my fold 1 results CV 0.43 (mean), and fold 2 CV 0.44. Did you see similar results on your folds, it seems like your LB is still the same as your single model. Could this be just different folds (or luck as a matter of a fact) fit better on the public LB? Shake up could be possible if others are experiencing this.</p>",
              "rawMarkdown": "@remekkinas \n\nI want to follow up on this, my fold 0 results seem good with CV .52 (mean, LB 0.52) similar to the results you posted before but my fold 1 results CV 0.43 (mean), and fold 2 CV 0.44. Did you see similar results on your folds, it seems like your LB is still the same as your single model. Could this be just different folds (or luck as a matter of a fact) fit better on the public LB? Shake up could be possible if others are experiencing this."
            }
          ]
        }
      ]
    },
    {
      "id": 2121189,
      "postDate": "2023-01-30T05:38:09.993Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/taruto1215\" target=\"_blank\">@taruto1215</a> , how do you get your CV? I mean the LB is calculated by laterality-level, if you train and validate by images, you will get the image-level score. In my experiments, the laterality-level f1 is always much higher than the image-level f1.</p>",
      "rawMarkdown": "Hi @taruto1215 , how do you get your CV? I mean the LB is calculated by laterality-level, if you train and validate by images, you will get the image-level score. In my experiments, the laterality-level f1 is always much higher than the image-level f1.",
      "votes": 5,
      "replies": [
        {
          "id": 2121404,
          "postDate": "2023-01-30T09:42:15.230Z",
          "content": "<p><a href=\"https://www.kaggle.com/forcewithme\" target=\"_blank\">@forcewithme</a> <br>\nThanks for your reply.</p>\n<p>my CV was calculated by laterality-level.<br>\nFor the image-level F1 is ~0.35.</p>",
          "rawMarkdown": "@forcewithme \nThanks for your reply.\n\nmy CV was calculated by laterality-level.\nFor the image-level F1 is ~0.35."
        }
      ]
    },
    {
      "id": 2121526,
      "postDate": "2023-01-30T11:41:48.123Z",
      "content": "<p>i am not very sure if i were correct. but for this competition there may be other ways to reduce shakeup by NOT only relying on LB vs CV.<br>\ne.g. my target is to make a model:</p>\n<ol>\n<li>have recall at least 0.50 for both site1 and site2</li>\n<li>have f1 of at least 0.50 for both site1 and site2 </li>\n<li>threshold not larger than 0.55<br>\n4 . …..</li>\n</ol>\n<p>i am coming out with many such rules based on my observation.</p>\n<p>it is like regularization in machine learning, when you loss is not reliable, you need to think of other \"sensible\" conditions</p>",
      "rawMarkdown": "i am not very sure if i were correct. but for this competition there may be other ways to reduce shakeup by NOT only relying on LB vs CV.\ne.g. my target is to make a model:\n1. have recall at least 0.50 for both site1 and site2\n2. have f1 of at least 0.50 for both site1 and site2 \n3. threshold not larger than 0.55\n4 . .....\n\ni am coming out with many such rules based on my observation.\n\nit is like regularization in machine learning, when you loss is not reliable, you need to think of other \"sensible\" conditions",
      "votes": 3,
      "replies": [
        {
          "id": 2121638,
          "postDate": "2023-01-30T13:09:12.357Z",
          "content": "<p>Thanks for sharing your ideas.<br>\nWe should approach this competition with that perspective for the remaining month.</p>",
          "rawMarkdown": "Thanks for sharing your ideas.\nWe should approach this competition with that perspective for the remaining month.",
          "votes": -1
        }
      ]
    },
    {
      "id": 2121239,
      "postDate": "2023-01-30T06:38:35.973Z",
      "content": "<p>beware if you aggregate by mean. it may be wrong method.</p>\n<p>for example assume</p>\n<p>patient1 (cancer):<br>\nimage1 = 0.9<br>\nimage2 (occluded)=0.1</p>\n<p>patient2 (cancer):<br>\nimage1 = 0.9<br>\nimage2 (occluded)=0.1<br>\nimage3 (occluded)=0.1<br>\nimage4 (occluded)=0.1</p>\n<p>the score for patient 2 is divided out.</p>\n<hr>\n<p>but max is no better<br>\npatient1 (cancer):<br>\nimage1 = 0.7<br>\nimage2 =0.8</p>\n<p>patient1 (no cancer):<br>\nimage1 = 0.1<br>\nimage2 =0.1<br>\nimage3 = 0.7 (noise)<br>\nimage4 =0.1</p>",
      "rawMarkdown": "beware if you aggregate by mean. it may be wrong method.\n\nfor example assume\n\npatient1 (cancer):\nimage1 = 0.9\nimage2 (occluded)=0.1\n\npatient2 (cancer):\nimage1 = 0.9\nimage2 (occluded)=0.1\nimage3 (occluded)=0.1\nimage4 (occluded)=0.1\n\n\nthe score for patient 2 is divided out.\n\n---\nbut max is no better\npatient1 (cancer):\nimage1 = 0.7\nimage2 =0.8\n\npatient1 (no cancer):\nimage1 = 0.1\nimage2 =0.1\nimage3 = 0.7 (noise)\nimage4 =0.1\n\n\n",
      "votes": 4,
      "replies": [
        {
          "id": 2121418,
          "postDate": "2023-01-30T09:54:18.443Z",
          "content": "<p>I put my code to aggregrate.<br>\nI think there is no mistake.</p>\n<p><code>df_merge_groupby = df_oof[['prediction_id','site_id', 'patient_id', 'laterality', 'cancer', 'preds']].groupby(\n            ['patient_id', 'laterality']).mean().reset_index()</code></p>\n<p></p><hr><br>\nyou mean we don't have to use all predictions?<br>\nFor example, in the below case, should we use only the highest 3 cancer prediction  for a patient.<br>\nThanks<p></p>\n<blockquote>\n  <p>patient2 (cancer):<br>\n  image1 = 0.9<br>\n  image2 (occluded)=0.1<br>\n  image3 (occluded)=0.1<br>\n  image4 (occluded)=0.1</p>\n</blockquote>",
          "rawMarkdown": "I put my code to aggregrate.\nI think there is no mistake.\n\n`df_merge_groupby = df_oof[['prediction_id','site_id', 'patient_id', 'laterality', 'cancer', 'preds']].groupby(\n            ['patient_id', 'laterality']).mean().reset_index() `\n\n<hr>\nyou mean we don't have to use all predictions?\nFor example, in the below case, should we use only the highest 3 cancer prediction  for a patient.\nThanks\n\n>patient2 (cancer):\nimage1 = 0.9\nimage2 (occluded)=0.1\nimage3 (occluded)=0.1\nimage4 (occluded)=0.1",
          "replies": [
            {
              "id": 2121511,
              "postDate": "2023-01-30T11:25:30.977Z",
              "content": "<p>you can try mean of top 3.<br>\nbut the correct way is to learn an aggregation model, which is multiple image prediction MIP at my post.</p>\n<hr>\n<p>there is no definite way of mean,max or mean of top3. different aggregation works for different cases, that is why it is better to learn a function.</p>",
              "rawMarkdown": "you can try mean of top 3.\nbut the correct way is to learn an aggregation model, which is multiple image prediction MIP at my post.\n\n---\nthere is no definite way of mean,max or mean of top3. different aggregation works for different cases, that is why it is better to learn a function.\n\n\n",
              "votes": 1
            },
            {
              "id": 2121513,
              "postDate": "2023-01-30T11:27:06.137Z",
              "content": "<p>you can also check <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> discussion at : <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/377654#2113462\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/377654#2113462</a></p>",
              "rawMarkdown": "you can also check @remekkinas discussion at : https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/377654#2113462",
              "votes": 2
            },
            {
              "id": 2121527,
              "postDate": "2023-01-30T11:42:51.167Z",
              "content": "<p>You are right.<br>\nFor me, mean of top 3 boosted my CV at 0.005.</p>\n<p>Thanks for the advice.</p>",
              "rawMarkdown": "You are right.\nFor me, mean of top 3 boosted my CV at 0.005.\n\nThanks for the advice."
            }
          ]
        },
        {
          "id": 2132925,
          "postDate": "2023-02-07T05:54:08.740Z",
          "content": "<p>Hi, I have a doubt regarding the best way to submit.</p>\n<p>Actually I have 4 splits, CV 0.42 and LB 0.52, for every model I have also a best_thr on the mean aggregation.<br>\nFor the prediction of every row I get the mean of the 4 models (trained on the 4 splits):</p>\n<p><code>\np(pid, imageid) = pm = (p1+p2+p3+p4)/4\n</code></p>\n<p>after that I aggregate with the mean using groupby() and after I just predict using a threshold (here the question ) that is the the mean of the 4 thresholds <br>\n<code>\ntm=(t1+t2+t3+t4)/4\n</code></p>\n<p>If I tried a different threshold I get better results on LB (I think that is normal beacuse it's like overfit the LB).</p>\n<p>I tried the same procedure with the max aggregation instead the mean aggregation, with some improvements but depends from the models.</p>\n<p>Do you choose the best threshold in this way? 🤔<br>\nThanks in advance</p>",
          "rawMarkdown": "Hi, I have a doubt regarding the best way to submit.\n\nActually I have 4 splits, CV 0.42 and LB 0.52, for every model I have also a best_thr on the mean aggregation.\nFor the prediction of every row I get the mean of the 4 models (trained on the 4 splits):\n\n`\np(pid, imageid) = pm = (p1+p2+p3+p4)/4\n`\n\nafter that I aggregate with the mean using groupby() and after I just predict using a threshold (here the question ) that is the the mean of the 4 thresholds \n`\ntm=(t1+t2+t3+t4)/4\n`\n\nIf I tried a different threshold I get better results on LB (I think that is normal beacuse it's like overfit the LB).\n\nI tried the same procedure with the max aggregation instead the mean aggregation, with some improvements but depends from the models.\n\nDo you choose the best threshold in this way? 🤔\nThanks in advance",
          "replies": [
            {
              "id": 2133313,
              "postDate": "2023-02-07T11:05:21.283Z",
              "content": "<p>I just choose the threshold that was best for my CV, no special tunning for LB.<br>\nThanks.</p>",
              "rawMarkdown": "I just choose the threshold that was best for my CV, no special tunning for LB.\nThanks.",
              "votes": 3
            }
          ]
        }
      ]
    },
    {
      "id": 2136783,
      "postDate": "2023-02-09T14:38:55.093Z",
      "content": "<p>interesting, LB0.59 is my training score, while my validation score is lower. kaggle LB score is close to my training score.</p>\n<p>could matching train score = LB score be a way to fight generalisation?<br>\n(of course your have to do early exit)</p>\n<p>asuming the top kagglers did not overfit, you can train until train score is around 0.65</p>",
      "rawMarkdown": "interesting, LB0.59 is my training score, while my validation score is lower. kaggle LB score is close to my training score.\n\ncould matching train score = LB score be a way to fight generalisation?\n(of course your have to do early exit)\n\nasuming the top kagglers did not overfit, you can train until train score is around 0.65",
      "votes": 1,
      "replies": [
        {
          "id": 2137447,
          "postDate": "2023-02-10T02:38:46.490Z",
          "content": "<p>on a side note since the server compute probablistic fscore, it is quite easily to overfit public test or probe the public test distribution for some divided bins<br>\ne.g. </p>\n<pre><code>1. submit raw probabiliy and get LB score\n2. for th = 0.1,0.2,0.3 .....\nset probabiliy&gt;th to random and submit,   \nor set set probabiliy&lt;th to random and submit  \nor set set probabiliy&lt;th1 and  set set probabiliy&gt;th2 to random and submit \n</code></pre>",
          "rawMarkdown": "on a side note since the server compute probablistic fscore, it is quite easily to overfit public test or probe the public test distribution for some divided bins\ne.g. \n\n```\n1. submit raw probabiliy and get LB score\n2. for th = 0.1,0.2,0.3 .....\nset probabiliy>th to random and submit,   \nor set set probabiliy<th to random and submit  \nor set set probabiliy<th1 and  set set probabiliy>th2 to random and submit \n\n``` \n\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 2122638,
      "postDate": "2023-01-31T04:22:41.397Z",
      "content": "<p>Hello, Which model did you use?</p>",
      "rawMarkdown": "Hello, Which model did you use?",
      "votes": 1,
      "replies": [
        {
          "id": 2123206,
          "postDate": "2023-01-31T11:34:16.250Z",
          "content": "<p>Hello, I used the efficientnetv2_s</p>",
          "rawMarkdown": "Hello, I used the efficientnetv2_s",
          "votes": 2,
          "replies": [
            {
              "id": 2123496,
              "postDate": "2023-01-31T14:51:48.270Z",
              "content": "<p>Thank you! I tried efficientnet_v2 with size 1024x512 but pf1 only around 0.1-0.15</p>",
              "rawMarkdown": "Thank you! I tried efficientnet_v2 with size 1024x512 but pf1 only around 0.1-0.15",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2121246,
      "postDate": "2023-01-30T06:47:17.847Z",
      "content": "<pre><code>site_id=1 : precision=0.329 recall=0.287 f1=0.307 thres=0.4500 roc_auc=0.637\nsite_id=2 : precision=0.639 recall=0.416 f1=0.504 thres=0.5000 roc_auc=0.706\n</code></pre>\n<p>i am interested in the score of site1 and site2 submitted separately.</p>\n<hr>\n<p>stable f1 if<br>\n1)  precision is close to recall <br>\n2)  precision is very high</p>\n<p>stable LB if ?<br>\n1)site1 close to site2</p>\n<p>high threshold  : maybe overfitting</p>\n<p>note: stable means consistent (i.e. less shakeup) but not necessarily optimum</p>\n<hr>\n<p>CV: 0.395 LB: 0.56</p>\n<p>\"I am very concern about shake down.<br>\nHow is your score gap??\"</p>\n<p>but even if CV close to LB doesn't guarantee  no shakeup </p>",
      "rawMarkdown": "```\nsite_id=1 : precision=0.329 recall=0.287 f1=0.307 thres=0.4500 roc_auc=0.637\nsite_id=2 : precision=0.639 recall=0.416 f1=0.504 thres=0.5000 roc_auc=0.706\n```\n\ni am interested in the score of site1 and site2 submitted separately.\n\n---\nstable f1 if\n1)  precision is close to recall \n2)  precision is very high\n\nstable LB if ?\n1)site1 close to site2\n\nhigh threshold  : maybe overfitting\n\nnote: stable means consistent (i.e. less shakeup) but not necessarily optimum\n\n---\nCV: 0.395 LB: 0.56\n\n\"I am very concern about shake down.\nHow is your score gap??\"\n\nbut even if CV close to LB doesn't guarantee  no shakeup ",
      "votes": 1,
      "replies": [
        {
          "id": 2121420,
          "postDate": "2023-01-30T09:55:10.267Z",
          "content": "<p>Sorry my ROC_AUC was wrong (I made a fundamental code mistake). I wrote up the correct version above.</p>",
          "rawMarkdown": "Sorry my ROC_AUC was wrong (I made a fundamental code mistake). I wrote up the correct version above.\n"
        },
        {
          "id": 2121431,
          "postDate": "2023-01-30T10:04:11.357Z",
          "content": "<p>I used <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/378521\" target=\"_blank\">your code</a> to generate results.<br>\nalso there is a gap between site1 and site2. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3830852%2F0d27d318170899acb50b28639a5f9c79%2FRNSA_results.png?generation=1675072875545524&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3830852%2F96c27c9cc7d78865e123a7cc3345b124%2FRNSA_results2.png?generation=1675072914143213&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "I used [your code](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/378521) to generate results.\nalso there is a gap between site1 and site2. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3830852%2F0d27d318170899acb50b28639a5f9c79%2FRNSA_results.png?generation=1675072875545524&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3830852%2F96c27c9cc7d78865e123a7cc3345b124%2FRNSA_results2.png?generation=1675072914143213&alt=media)"
        }
      ]
    },
    {
      "id": 2121192,
      "postDate": "2023-01-30T05:43:18.480Z",
      "content": "<p>How that possible??<br>\nProbably overfitted to hidden test set?</p>\n<p>I don't think roc_auc about 0.6~0.7 is good performance.</p>",
      "rawMarkdown": "How that possible??\nProbably overfitted to hidden test set?\n\nI don't think roc_auc about 0.6~0.7 is good performance.",
      "votes": 1,
      "replies": [
        {
          "id": 2121400,
          "postDate": "2023-01-30T09:40:04.787Z",
          "content": "<p>I just found I made a fundamental mistakes<br>\nMy CV ROC_AUC is ~0.87.<br>\nThanks.</p>",
          "rawMarkdown": "I just found I made a fundamental mistakes\nMy CV ROC_AUC is ~0.87.\nThanks."
        }
      ]
    },
    {
      "id": 2121017,
      "postDate": "2023-01-30T01:34:10.120Z",
      "content": "<p>me too! cv:0.3 lb:0.46 threshold=0.89 </p>",
      "rawMarkdown": "me too! cv:0.3 lb:0.46 threshold=0.89 ",
      "votes": 1,
      "replies": [
        {
          "id": 2121035,
          "postDate": "2023-01-30T02:27:42.390Z",
          "content": "<p>Thanks for your sharing</p>",
          "rawMarkdown": "Thanks for your sharing",
          "votes": 1
        }
      ]
    },
    {
      "id": 2121011,
      "postDate": "2023-01-30T01:15:29.013Z",
      "content": "<p>did you try same aggregation in your cv (mean, max etc..) as you do for the pred_ids during subm?</p>",
      "rawMarkdown": "did you try same aggregation in your cv (mean, max etc..) as you do for the pred_ids during subm?",
      "votes": 1,
      "replies": [
        {
          "id": 2121029,
          "postDate": "2023-01-30T02:15:46.560Z",
          "content": "<p>Yes, my CV is after mean aggregation.<br>\nHow is your score gap? No gap?<br>\nThanks</p>",
          "rawMarkdown": "Yes, my CV is after mean aggregation.\nHow is your score gap? No gap?\nThanks",
          "replies": [
            {
              "id": 2121033,
              "postDate": "2023-01-30T02:25:45.467Z",
              "content": "<p>CV 0.5, LB 0.53</p>",
              "rawMarkdown": "CV 0.5, LB 0.53",
              "votes": 2
            },
            {
              "id": 2121038,
              "postDate": "2023-01-30T02:31:24.297Z",
              "content": "<p>Good! Almost same scores!<br>\nHow is your CV strategy?<br>\nIs it same as me?<br>\nThanks.</p>",
              "rawMarkdown": "Good! Almost same scores!\nHow is your CV strategy?\nIs it same as me?\nThanks."
            },
            {
              "id": 2121426,
              "postDate": "2023-01-30T09:58:27.640Z",
              "content": "<p>very similar, but 5 splits</p>",
              "rawMarkdown": "very similar, but 5 splits",
              "votes": 1
            },
            {
              "id": 2121505,
              "postDate": "2023-01-30T11:16:50.890Z",
              "content": "<p>Huh, I maybe need more splits.<br>\nThanks for your sharing.</p>",
              "rawMarkdown": "Huh, I maybe need more splits.\nThanks for your sharing.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2132637,
      "postDate": "2023-02-07T00:17:24.230Z",
      "content": "<p>good job!</p>\n<p>currently all team with LB score above 0.60 has good chance of wining.<br>\ndo spend time to:</p>\n<ol>\n<li><p>simulate leader board results. how one or few change of pos samples can change your validation score.<br>\nthis enable you to interprete the LB score correctly and estimate shakeup</p></li>\n<li><p>i still think that mean() aggreataion is not good unless your classifier has very low false positive rate.<br>\nso you may want to spend time to think about how to fuse results.</p></li>\n</ol>\n<p>remember !!!!<br>\nit is the private hidden score that we are interested .you should try to predict the private score for every submission you make !!!</p>",
      "rawMarkdown": "good job!\n\ncurrently all team with LB score above 0.60 has good chance of wining.\ndo spend time to:\n\n1. simulate leader board results. how one or few change of pos samples can change your validation score.\nthis enable you to interprete the LB score correctly and estimate shakeup\n\n2. i still think that mean() aggreataion is not good unless your classifier has very low false positive rate.\nso you may want to spend time to think about how to fuse results.\n\nremember !!!!\nit is the private hidden score that we are interested .you should try to predict the private score for every submission you make !!!\n",
      "votes": 1,
      "replies": [
        {
          "id": 2133161,
          "postDate": "2023-02-07T09:34:24.057Z",
          "content": "<p>Thank you for your advice.<br>\nYour advices help me(us) a lot.<br>\nYes I also am thinking about how to aggregrate.</p>\n<blockquote>\n  <p>1 simulate leader board results. how one or few change of pos samples can change your validation score.<br>\n  this enable you to interprete the LB score correctly and estimate shakeup</p>\n</blockquote>\n<p>You means delete a few pos samples then train model and check the influence (score) ?</p>\n<p>Thanks</p>",
          "rawMarkdown": "Thank you for your advice.\nYour advices help me(us) a lot.\nYes I also am thinking about how to aggregrate.\n\n>1 simulate leader board results. how one or few change of pos samples can change your validation score.\nthis enable you to interprete the LB score correctly and estimate shakeup\n\nYou means delete a few pos samples then train model and check the influence (score) ?\n\nThanks",
          "replies": [
            {
              "id": 2133450,
              "postDate": "2023-02-07T12:25:44.813Z",
              "content": "<p>somtheing like below.</p>\n<p>now you have already make prediction and predicted score and  true label.<br>\nuse 100% of predicted score to compute cv<br>\nuse 90% of predicted score (repeat over a few time)<br>\nuse 80%</p>\n<p>is the change stable?</p>\n<p>ideally, remove of one sample from the predicted score should not changes the CV metric. but is it true?</p>",
              "rawMarkdown": "somtheing like below.\n\nnow you have already make prediction and predicted score and  true label.\nuse 100% of predicted score to compute cv\nuse 90% of predicted score (repeat over a few time)\nuse 80%\n\nis the change stable?\n\nideally, remove of one sample from the predicted score should not changes the CV metric. but is it true?",
              "votes": 1
            },
            {
              "id": 2133456,
              "postDate": "2023-02-07T12:32:03.093Z",
              "content": "<p>I understood.<br>\nI'll try it. <br>\nThanks!</p>",
              "rawMarkdown": "I understood.\nI'll try it. \nThanks!"
            }
          ]
        }
      ]
    },
    {
      "id": 2164140,
      "postDate": "2023-03-01T10:31:36.737Z",
      "content": "<p>I also encounter train and validation set scores of 0.8.<br>\nBut the test set is only 0.02</p>",
      "rawMarkdown": "I also encounter train and validation set scores of 0.8.\nBut the test set is only 0.02"
    },
    {
      "id": 2144025,
      "postDate": "2023-02-14T17:20:51.223Z",
      "content": "<p>My single image CV F1 score is 0.3596 but my LB score is only 0.38. I am using the effnetv2_s on 1536x960 images, what is your CV score for single image right now ?</p>",
      "rawMarkdown": "My single image CV F1 score is 0.3596 but my LB score is only 0.38. I am using the effnetv2_s on 1536x960 images, what is your CV score for single image right now ?",
      "replies": [
        {
          "id": 2145536,
          "postDate": "2023-02-15T07:00:50.407Z",
          "content": "<p><a href=\"https://www.kaggle.com/namansingh2803\" target=\"_blank\">@namansingh2803</a> <br>\nSee this reply.<br>\nmy image level F1 score 0.356 (CV k=3)<br>\nThanks.</p>\n<p><a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/382198#2133432\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/382198#2133432</a></p>",
          "rawMarkdown": "@namansingh2803 \nSee this reply.\nmy image level F1 score 0.356 (CV k=3)\nThanks.\n\nhttps://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/382198#2133432",
          "votes": 1,
          "replies": [
            {
              "id": 2146305,
              "postDate": "2023-02-15T19:09:02.387Z",
              "content": "<p>I see, could you also tell me what was your validation loss ? I observed that some of my models which have a lower CV F1 score but a lower validation loss perform better on the public LB than the models which have a higher F1 and a higher val loss.</p>",
              "rawMarkdown": "I see, could you also tell me what was your validation loss ? I observed that some of my models which have a lower CV F1 score but a lower validation loss perform better on the public LB than the models which have a higher F1 and a higher val loss."
            }
          ]
        }
      ]
    },
    {
      "id": 2138057,
      "postDate": "2023-02-10T14:17:07.460Z",
      "content": "<p>Im curious, whats your single fold LB? </p>",
      "rawMarkdown": "Im curious, whats your single fold LB? "
    },
    {
      "id": 2134791,
      "postDate": "2023-02-08T08:57:20.073Z",
      "content": "<p><a href=\"https://www.kaggle.com/taruto1215\" target=\"_blank\">@taruto1215</a> Sorry if this is not appropriate, could you please share your dataset you used for training? Many thanks</p>",
      "rawMarkdown": "@taruto1215 Sorry if this is not appropriate, could you please share your dataset you used for training? Many thanks",
      "replies": [
        {
          "id": 2134970,
          "postDate": "2023-02-08T11:15:57.680Z",
          "content": "<p>I am sorry.<br>\nit includes one of my unique solution. so I cannot share you.</p>",
          "rawMarkdown": "I am sorry.\nit includes one of my unique solution. so I cannot share you.",
          "votes": 2
        }
      ]
    },
    {
      "id": 2132523,
      "postDate": "2023-02-06T21:16:36.953Z",
      "content": "<p>I think is related to how you split the data. You can have a R-CC on train and a L-CC on validation, R-CC doesn't have cancer, L-CC has. At the same time I expect that there is a high similarity between the 2, so this makes the score on CV smaller.</p>",
      "rawMarkdown": "I think is related to how you split the data. You can have a R-CC on train and a L-CC on validation, R-CC doesn't have cancer, L-CC has. At the same time I expect that there is a high similarity between the 2, so this makes the score on CV smaller.",
      "replies": [
        {
          "id": 2133163,
          "postDate": "2023-02-07T09:37:04.353Z",
          "content": "<p>Thanks<br>\nI split the data as same patients are the same fold.<br>\nso for me I think the problem you mentioned can be ignore.</p>",
          "rawMarkdown": "Thanks\nI split the data as same patients are the same fold.\nso for me I think the problem you mentioned can be ignore."
        }
      ]
    },
    {
      "id": 2130418,
      "postDate": "2023-02-05T12:45:27.553Z",
      "content": "<p><a href=\"https://www.kaggle.com/taruto1215\" target=\"_blank\">@taruto1215</a> Hello, Could you tell me your dataset you use? Did you use VOI_LUT in your dataset? Thanks</p>",
      "rawMarkdown": "@taruto1215 Hello, Could you tell me your dataset you use? Did you use VOI_LUT in your dataset? Thanks",
      "replies": [
        {
          "id": 2131912,
          "postDate": "2023-02-06T13:34:55.653Z",
          "content": "<p>Yes, I did</p>",
          "rawMarkdown": "Yes, I did",
          "votes": 2,
          "replies": [
            {
              "id": 2131951,
              "postDate": "2023-02-06T14:01:25.017Z",
              "content": "<p>I think the main reason why your CV/LB gap maybe in your dataset. Difference in training and inference?</p>",
              "rawMarkdown": "I think the main reason why your CV/LB gap maybe in your dataset. Difference in training and inference?"
            },
            {
              "id": 2133184,
              "postDate": "2023-02-07T09:53:18.987Z",
              "content": "<p>I use exactly same preprocess pipeline… Hmm</p>",
              "rawMarkdown": "I use exactly same preprocess pipeline... Hmm\n"
            }
          ]
        }
      ]
    },
    {
      "id": 2126918,
      "postDate": "2023-02-02T15:04:29.667Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 2127714,
          "postDate": "2023-02-03T05:03:38.933Z",
          "content": "<p>I get my cv score using this cv strategy for all train data.<br>\nn_splits = 3</p>\n<p><a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/382198\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/382198</a></p>",
          "rawMarkdown": "I get my cv score using this cv strategy for all train data.\nn_splits = 3\n\nhttps://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/382198"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2131916,
      "author_name": "taruto",
      "author_url": "",
      "post_date": "2023-02-06T13:37:45.833000",
      "content": "<p>Finally my LB is more than 0.60.<br>\nBut my cv score is 0.42.<br>\nThe gap is too big !!!!!!</p>",
      "votes": 7,
      "replies": [
        {
          "id": 2131925,
          "author_name": "nymfree",
          "author_url": "",
          "post_date": "2023-02-06T13:42:11.230000",
          "content": "<p>any post-processing?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2131955,
              "author_name": "taruto",
              "author_url": "",
              "post_date": "2023-02-06T14:02:18.667000",
              "content": "<p>No post-processing</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2132155,
          "author_name": "Mad_Neil",
          "author_url": "",
          "post_date": "2023-02-06T16:23:15.843000",
          "content": "<p>Your LB score is very impressive. Could you tell me your image size did you use?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2133133,
              "author_name": "taruto",
              "author_url": "",
              "post_date": "2023-02-07T09:04:35.323000",
              "content": "<p>Thanks.<br>\ncurrent my image size is 1536x960</p>",
              "votes": 3,
              "replies": []
            }
          ]
        },
        {
          "id": 2132435,
          "author_name": "Remek Kinas",
          "author_url": "",
          "post_date": "2023-02-06T20:31:11.657000",
          "content": "<p>Amazing result! Congratulations crossing 0.6x.<br>\nDo you still use effnet? </p>\n<p>BTW:<br>\nI have local CV: 0.54 now … and this model is only 0.51 on LB 😂</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2133141,
              "author_name": "taruto",
              "author_url": "",
              "post_date": "2023-02-07T09:11:54.400000",
              "content": "<p>Thank you!<br>\nI still use the effnet.</p>\n<p>You have a good local CV!<br>\nI don't still know why my CV is too low… (it bothers me)</p>\n<p>How is the image level CV score for your model?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2133283,
              "author_name": "Remek Kinas",
              "author_url": "",
              "post_date": "2023-02-07T10:44:32.670000",
              "content": "<p>For single image 0.4.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2133298,
              "author_name": "taruto",
              "author_url": "",
              "post_date": "2023-02-07T10:51:47.230000",
              "content": "<p>it's a also higher score than my image level cv score 0.356.<br>\nthanks for your sharing.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2133300,
              "author_name": "Remek Kinas",
              "author_url": "",
              "post_date": "2023-02-07T10:54:18.857000",
              "content": "<p>If it is possible please share result (validator result). It would be great. Especially I am looking now on predicition dynamics (chart with probability distribution).</p>\n<p>As far as I understand you blend 3 models (3 folds) for final solution?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2133432,
              "author_name": "taruto",
              "author_url": "",
              "post_date": "2023-02-07T12:11:27.647000",
              "content": "<p>I share you my best cv result.<br>\nYes, I just blended 3 models (3folds).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3830852%2F92b091a4351296e585ce247fbd6008c3%2FRSNA_LB063_results2.png?generation=1675771723070708&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3830852%2Fdc4e15caabae71248a9097dfeaebcc08%2FRSNA_LB063_results.png?generation=1675771667026332&amp;alt=media\" alt=\"\"></p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2133443,
              "author_name": "Remek Kinas",
              "author_url": "",
              "post_date": "2023-02-07T12:18:58.940000",
              "content": "<p>Thank you  a lot.👍<br>\nStats you provided are from one model or blended?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2133445,
              "author_name": "taruto",
              "author_url": "",
              "post_date": "2023-02-07T12:22:53.867000",
              "content": "<p>It's blended from 3 models (3folds)</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2134843,
              "author_name": "Remek Kinas",
              "author_url": "",
              "post_date": "2023-02-08T09:32:32.167000",
              "content": "<p><a href=\"https://www.kaggle.com/taruto1215\" target=\"_blank\">@taruto1215</a> how many epochs do you train model? Or from what epoch do you choose models from? </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2134858,
              "author_name": "Mobassir",
              "author_url": "",
              "post_date": "2023-02-08T09:39:15.480000",
              "content": "<p><a href=\"https://www.kaggle.com/taruto1215\" target=\"_blank\">@taruto1215</a> may i know what validation pf1,auc and thresholded pf1 you get for each fold? thanks</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2134978,
              "author_name": "taruto",
              "author_url": "",
              "post_date": "2023-02-08T11:21:40.770000",
              "content": "<p><a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> <br>\nI train model 5 epochs.<br>\nthen I choose a model has a best image-level f1 score.<br>\nBut almost of them has the best score at the 5 epoch.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2134989,
              "author_name": "taruto",
              "author_url": "",
              "post_date": "2023-02-08T11:36:44.180000",
              "content": "<p><a href=\"https://www.kaggle.com/mobassir\" target=\"_blank\">@mobassir</a> <br>\nThis is my validation results for each fold.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3830852%2F3e806521970616e694389151d535198b%2FRSNA_results_each_folds.png?generation=1675856167389336&amp;alt=media\" alt=\"\"></p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2135052,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-02-08T12:21:43.520000",
              "content": "<p>i ahve been trying another experiment. if you are training too many epoch, then the predicted distribution will be too sharp (i.e. overfitting).</p>\n<p>hence one way is to stop early at, say 5 epoch.<br>\nwe train many of differenet models.</p>\n<p>then we freeze them and use them as feature extractor.<br>\nwe use giba trick (winner of petfinder): <a href=\"https://developer.nvidia.com/blog/fast-fine-tuning-of-ai-transformers-using-rapids-machine-learning/\" target=\"_blank\">https://developer.nvidia.com/blog/fast-fine-tuning-of-ai-transformers-using-rapids-machine-learning/</a><br>\n<a href=\"https://www.kaggle.com/c/petfinder-pawpularity-score/discussion/301686\" target=\"_blank\">https://www.kaggle.com/c/petfinder-pawpularity-score/discussion/301686</a></p>\n<p>but instead, we create ensmble for multi-view model. each image is now treated as a feature map in sequence, and trsnformer is used to fuse the feature of different view, together with other attribute like age, site id.</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2135314,
              "author_name": "taruto",
              "author_url": "",
              "post_date": "2023-02-08T14:59:35.790000",
              "content": "<p>Thanks I’ll check it<br>\nBut from tomorrow I’ll be very busy at my job . I will not have more time to spend Kaggle than now hahaha😂</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2132600,
          "author_name": "taghados",
          "author_url": "",
          "post_date": "2023-02-06T22:42:30.270000",
          "content": "<p>Is it a single fold model or an ensemble ?</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2133145,
              "author_name": "taruto",
              "author_url": "",
              "post_date": "2023-02-07T09:15:00.843000",
              "content": "<p>Number of folds is 3.<br>\nso I used 3 models for submission.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2141316,
              "author_name": "Lau2664",
              "author_url": "",
              "post_date": "2023-02-12T16:29:02.437000",
              "content": "<p>Hi! <a href=\"https://www.kaggle.com/taruto1215\" target=\"_blank\">@taruto1215</a> Can you tell me how do you use 3 models in your inference notebook to make a submission?</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2132927,
          "author_name": "fate",
          "author_url": "",
          "post_date": "2023-02-07T06:04:44.393000",
          "content": "<p>5 fold<br>\nCV:0.416<br>\nLB:0.58<br>\nMaybe public leaderboard data are simpler than train data</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2133148,
              "author_name": "taruto",
              "author_url": "",
              "post_date": "2023-02-07T09:19:05.653000",
              "content": "<p>It's the same symptom as mine haha.<br>\nHow about your ROC_AUC for CV??</p>\n<p>I guess so too for the public leaderboard data. <br>\nBut for the private leaderboard data … <br>\nI hope it's same.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2133176,
              "author_name": "fate",
              "author_url": "",
              "post_date": "2023-02-07T09:46:15.880000",
              "content": "<p>ROC_AUC(group by mean) ~0.88</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2133293,
              "author_name": "taruto",
              "author_url": "",
              "post_date": "2023-02-07T10:49:20.730000",
              "content": "<p>thanks for your sharing.<br>\nmy ROC_AUC (group by mean) ~0.896<br>\nI think roc_auc scores has correlation against LB scores.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2133302,
              "author_name": "Remek Kinas",
              "author_url": "",
              "post_date": "2023-02-07T10:55:23.120000",
              "content": "<p>The same here - ROV_OUC for me is over 0.91 and LB is lower then yours. But I am still talking about single model in my case. All blending solutions failed for me.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2133442,
              "author_name": "taruto",
              "author_url": "",
              "post_date": "2023-02-07T12:18:47.177000",
              "content": "<p>The ROC_AUC score is also good!<br>\nPerhaps I should suspect that I am overfitting to LB.</p>\n<p>why do you talking about only single model?<br>\nI think we should check oof result.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2133473,
              "author_name": "Remek Kinas",
              "author_url": "",
              "post_date": "2023-02-07T12:48:30.570000",
              "content": "<p>To be honest? I do not know at this time. Fianally I will sub two solution: single model and blend. </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2142767,
              "author_name": "outwrest",
              "author_url": "",
              "post_date": "2023-02-13T18:58:58.997000",
              "content": "<p><a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> </p>\n<p>I want to follow up on this, my fold 0 results seem good with CV .52 (mean, LB 0.52) similar to the results you posted before but my fold 1 results CV 0.43 (mean), and fold 2 CV 0.44. Did you see similar results on your folds, it seems like your LB is still the same as your single model. Could this be just different folds (or luck as a matter of a fact) fit better on the public LB? Shake up could be possible if others are experiencing this.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2121189,
      "author_name": "ForcewithMe",
      "author_url": "",
      "post_date": "2023-01-30T05:38:09.993000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/taruto1215\" target=\"_blank\">@taruto1215</a> , how do you get your CV? I mean the LB is calculated by laterality-level, if you train and validate by images, you will get the image-level score. In my experiments, the laterality-level f1 is always much higher than the image-level f1.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 2121404,
          "author_name": "taruto",
          "author_url": "",
          "post_date": "2023-01-30T09:42:15.230000",
          "content": "<p><a href=\"https://www.kaggle.com/forcewithme\" target=\"_blank\">@forcewithme</a> <br>\nThanks for your reply.</p>\n<p>my CV was calculated by laterality-level.<br>\nFor the image-level F1 is ~0.35.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2121526,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-01-30T11:41:48.123000",
      "content": "<p>i am not very sure if i were correct. but for this competition there may be other ways to reduce shakeup by NOT only relying on LB vs CV.<br>\ne.g. my target is to make a model:</p>\n<ol>\n<li>have recall at least 0.50 for both site1 and site2</li>\n<li>have f1 of at least 0.50 for both site1 and site2 </li>\n<li>threshold not larger than 0.55<br>\n4 . …..</li>\n</ol>\n<p>i am coming out with many such rules based on my observation.</p>\n<p>it is like regularization in machine learning, when you loss is not reliable, you need to think of other \"sensible\" conditions</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2121638,
          "author_name": "taruto",
          "author_url": "",
          "post_date": "2023-01-30T13:09:12.357000",
          "content": "<p>Thanks for sharing your ideas.<br>\nWe should approach this competition with that perspective for the remaining month.</p>",
          "votes": -1,
          "replies": []
        }
      ]
    },
    {
      "id": 2121239,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-01-30T06:38:35.973000",
      "content": "<p>beware if you aggregate by mean. it may be wrong method.</p>\n<p>for example assume</p>\n<p>patient1 (cancer):<br>\nimage1 = 0.9<br>\nimage2 (occluded)=0.1</p>\n<p>patient2 (cancer):<br>\nimage1 = 0.9<br>\nimage2 (occluded)=0.1<br>\nimage3 (occluded)=0.1<br>\nimage4 (occluded)=0.1</p>\n<p>the score for patient 2 is divided out.</p>\n<hr>\n<p>but max is no better<br>\npatient1 (cancer):<br>\nimage1 = 0.7<br>\nimage2 =0.8</p>\n<p>patient1 (no cancer):<br>\nimage1 = 0.1<br>\nimage2 =0.1<br>\nimage3 = 0.7 (noise)<br>\nimage4 =0.1</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2121418,
          "author_name": "taruto",
          "author_url": "",
          "post_date": "2023-01-30T09:54:18.443000",
          "content": "<p>I put my code to aggregrate.<br>\nI think there is no mistake.</p>\n<p><code>df_merge_groupby = df_oof[['prediction_id','site_id', 'patient_id', 'laterality', 'cancer', 'preds']].groupby(\n            ['patient_id', 'laterality']).mean().reset_index()</code></p>\n<p></p><hr><br>\nyou mean we don't have to use all predictions?<br>\nFor example, in the below case, should we use only the highest 3 cancer prediction  for a patient.<br>\nThanks<p></p>\n<blockquote>\n  <p>patient2 (cancer):<br>\n  image1 = 0.9<br>\n  image2 (occluded)=0.1<br>\n  image3 (occluded)=0.1<br>\n  image4 (occluded)=0.1</p>\n</blockquote>",
          "votes": 0,
          "replies": [
            {
              "id": 2121511,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-01-30T11:25:30.977000",
              "content": "<p>you can try mean of top 3.<br>\nbut the correct way is to learn an aggregation model, which is multiple image prediction MIP at my post.</p>\n<hr>\n<p>there is no definite way of mean,max or mean of top3. different aggregation works for different cases, that is why it is better to learn a function.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2121513,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-01-30T11:27:06.137000",
              "content": "<p>you can also check <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> discussion at : <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/377654#2113462\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/377654#2113462</a></p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2121527,
              "author_name": "taruto",
              "author_url": "",
              "post_date": "2023-01-30T11:42:51.167000",
              "content": "<p>You are right.<br>\nFor me, mean of top 3 boosted my CV at 0.005.</p>\n<p>Thanks for the advice.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2132925,
          "author_name": "AleNic",
          "author_url": "",
          "post_date": "2023-02-07T05:54:08.740000",
          "content": "<p>Hi, I have a doubt regarding the best way to submit.</p>\n<p>Actually I have 4 splits, CV 0.42 and LB 0.52, for every model I have also a best_thr on the mean aggregation.<br>\nFor the prediction of every row I get the mean of the 4 models (trained on the 4 splits):</p>\n<p><code>\np(pid, imageid) = pm = (p1+p2+p3+p4)/4\n</code></p>\n<p>after that I aggregate with the mean using groupby() and after I just predict using a threshold (here the question ) that is the the mean of the 4 thresholds <br>\n<code>\ntm=(t1+t2+t3+t4)/4\n</code></p>\n<p>If I tried a different threshold I get better results on LB (I think that is normal beacuse it's like overfit the LB).</p>\n<p>I tried the same procedure with the max aggregation instead the mean aggregation, with some improvements but depends from the models.</p>\n<p>Do you choose the best threshold in this way? 🤔<br>\nThanks in advance</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2133313,
              "author_name": "taruto",
              "author_url": "",
              "post_date": "2023-02-07T11:05:21.283000",
              "content": "<p>I just choose the threshold that was best for my CV, no special tunning for LB.<br>\nThanks.</p>",
              "votes": 3,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2136783,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-02-09T14:38:55.093000",
      "content": "<p>interesting, LB0.59 is my training score, while my validation score is lower. kaggle LB score is close to my training score.</p>\n<p>could matching train score = LB score be a way to fight generalisation?<br>\n(of course your have to do early exit)</p>\n<p>asuming the top kagglers did not overfit, you can train until train score is around 0.65</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2137447,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2023-02-10T02:38:46.490000",
          "content": "<p>on a side note since the server compute probablistic fscore, it is quite easily to overfit public test or probe the public test distribution for some divided bins<br>\ne.g. </p>\n<pre><code>1. submit raw probabiliy and get LB score\n2. for th = 0.1,0.2,0.3 .....\nset probabiliy&gt;th to random and submit,   \nor set set probabiliy&lt;th to random and submit  \nor set set probabiliy&lt;th1 and  set set probabiliy&gt;th2 to random and submit \n</code></pre>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2122638,
      "author_name": "Mad_Neil",
      "author_url": "",
      "post_date": "2023-01-31T04:22:41.397000",
      "content": "<p>Hello, Which model did you use?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2123206,
          "author_name": "taruto",
          "author_url": "",
          "post_date": "2023-01-31T11:34:16.250000",
          "content": "<p>Hello, I used the efficientnetv2_s</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2123496,
              "author_name": "Mad_Neil",
              "author_url": "",
              "post_date": "2023-01-31T14:51:48.270000",
              "content": "<p>Thank you! I tried efficientnet_v2 with size 1024x512 but pf1 only around 0.1-0.15</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2121246,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-01-30T06:47:17.847000",
      "content": "<pre><code>site_id=1 : precision=0.329 recall=0.287 f1=0.307 thres=0.4500 roc_auc=0.637\nsite_id=2 : precision=0.639 recall=0.416 f1=0.504 thres=0.5000 roc_auc=0.706\n</code></pre>\n<p>i am interested in the score of site1 and site2 submitted separately.</p>\n<hr>\n<p>stable f1 if<br>\n1)  precision is close to recall <br>\n2)  precision is very high</p>\n<p>stable LB if ?<br>\n1)site1 close to site2</p>\n<p>high threshold  : maybe overfitting</p>\n<p>note: stable means consistent (i.e. less shakeup) but not necessarily optimum</p>\n<hr>\n<p>CV: 0.395 LB: 0.56</p>\n<p>\"I am very concern about shake down.<br>\nHow is your score gap??\"</p>\n<p>but even if CV close to LB doesn't guarantee  no shakeup </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2121420,
          "author_name": "taruto",
          "author_url": "",
          "post_date": "2023-01-30T09:55:10.267000",
          "content": "<p>Sorry my ROC_AUC was wrong (I made a fundamental code mistake). I wrote up the correct version above.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2121431,
          "author_name": "taruto",
          "author_url": "",
          "post_date": "2023-01-30T10:04:11.357000",
          "content": "<p>I used <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/378521\" target=\"_blank\">your code</a> to generate results.<br>\nalso there is a gap between site1 and site2. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3830852%2F0d27d318170899acb50b28639a5f9c79%2FRNSA_results.png?generation=1675072875545524&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3830852%2F96c27c9cc7d78865e123a7cc3345b124%2FRNSA_results2.png?generation=1675072914143213&amp;alt=media\" alt=\"\"></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2121192,
      "author_name": "ECO",
      "author_url": "",
      "post_date": "2023-01-30T05:43:18.480000",
      "content": "<p>How that possible??<br>\nProbably overfitted to hidden test set?</p>\n<p>I don't think roc_auc about 0.6~0.7 is good performance.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2121400,
          "author_name": "taruto",
          "author_url": "",
          "post_date": "2023-01-30T09:40:04.787000",
          "content": "<p>I just found I made a fundamental mistakes<br>\nMy CV ROC_AUC is ~0.87.<br>\nThanks.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2121017,
      "author_name": "yoyobar",
      "author_url": "",
      "post_date": "2023-01-30T01:34:10.120000",
      "content": "<p>me too! cv:0.3 lb:0.46 threshold=0.89 </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2121035,
          "author_name": "taruto",
          "author_url": "",
          "post_date": "2023-01-30T02:27:42.390000",
          "content": "<p>Thanks for your sharing</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2121011,
      "author_name": "Eleftherios Fanioudakis",
      "author_url": "",
      "post_date": "2023-01-30T01:15:29.013000",
      "content": "<p>did you try same aggregation in your cv (mean, max etc..) as you do for the pred_ids during subm?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2121029,
          "author_name": "taruto",
          "author_url": "",
          "post_date": "2023-01-30T02:15:46.560000",
          "content": "<p>Yes, my CV is after mean aggregation.<br>\nHow is your score gap? No gap?<br>\nThanks</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2121033,
              "author_name": "Eleftherios Fanioudakis",
              "author_url": "",
              "post_date": "2023-01-30T02:25:45.467000",
              "content": "<p>CV 0.5, LB 0.53</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2121038,
              "author_name": "taruto",
              "author_url": "",
              "post_date": "2023-01-30T02:31:24.297000",
              "content": "<p>Good! Almost same scores!<br>\nHow is your CV strategy?<br>\nIs it same as me?<br>\nThanks.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2121426,
              "author_name": "Eleftherios Fanioudakis",
              "author_url": "",
              "post_date": "2023-01-30T09:58:27.640000",
              "content": "<p>very similar, but 5 splits</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2121505,
              "author_name": "taruto",
              "author_url": "",
              "post_date": "2023-01-30T11:16:50.890000",
              "content": "<p>Huh, I maybe need more splits.<br>\nThanks for your sharing.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2132637,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-02-07T00:17:24.230000",
      "content": "<p>good job!</p>\n<p>currently all team with LB score above 0.60 has good chance of wining.<br>\ndo spend time to:</p>\n<ol>\n<li><p>simulate leader board results. how one or few change of pos samples can change your validation score.<br>\nthis enable you to interprete the LB score correctly and estimate shakeup</p></li>\n<li><p>i still think that mean() aggreataion is not good unless your classifier has very low false positive rate.<br>\nso you may want to spend time to think about how to fuse results.</p></li>\n</ol>\n<p>remember !!!!<br>\nit is the private hidden score that we are interested .you should try to predict the private score for every submission you make !!!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2133161,
          "author_name": "taruto",
          "author_url": "",
          "post_date": "2023-02-07T09:34:24.057000",
          "content": "<p>Thank you for your advice.<br>\nYour advices help me(us) a lot.<br>\nYes I also am thinking about how to aggregrate.</p>\n<blockquote>\n  <p>1 simulate leader board results. how one or few change of pos samples can change your validation score.<br>\n  this enable you to interprete the LB score correctly and estimate shakeup</p>\n</blockquote>\n<p>You means delete a few pos samples then train model and check the influence (score) ?</p>\n<p>Thanks</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2133450,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-02-07T12:25:44.813000",
              "content": "<p>somtheing like below.</p>\n<p>now you have already make prediction and predicted score and  true label.<br>\nuse 100% of predicted score to compute cv<br>\nuse 90% of predicted score (repeat over a few time)<br>\nuse 80%</p>\n<p>is the change stable?</p>\n<p>ideally, remove of one sample from the predicted score should not changes the CV metric. but is it true?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2133456,
              "author_name": "taruto",
              "author_url": "",
              "post_date": "2023-02-07T12:32:03.093000",
              "content": "<p>I understood.<br>\nI'll try it. <br>\nThanks!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2164140,
      "author_name": "Kai-Yi Hsu",
      "author_url": "",
      "post_date": "2023-03-01T10:31:36.737000",
      "content": "<p>I also encounter train and validation set scores of 0.8.<br>\nBut the test set is only 0.02</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2144025,
      "author_name": "Naman Makkar",
      "author_url": "",
      "post_date": "2023-02-14T17:20:51.223000",
      "content": "<p>My single image CV F1 score is 0.3596 but my LB score is only 0.38. I am using the effnetv2_s on 1536x960 images, what is your CV score for single image right now ?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2145536,
          "author_name": "taruto",
          "author_url": "",
          "post_date": "2023-02-15T07:00:50.407000",
          "content": "<p><a href=\"https://www.kaggle.com/namansingh2803\" target=\"_blank\">@namansingh2803</a> <br>\nSee this reply.<br>\nmy image level F1 score 0.356 (CV k=3)<br>\nThanks.</p>\n<p><a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/382198#2133432\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/382198#2133432</a></p>",
          "votes": 1,
          "replies": [
            {
              "id": 2146305,
              "author_name": "Naman Makkar",
              "author_url": "",
              "post_date": "2023-02-15T19:09:02.387000",
              "content": "<p>I see, could you also tell me what was your validation loss ? I observed that some of my models which have a lower CV F1 score but a lower validation loss perform better on the public LB than the models which have a higher F1 and a higher val loss.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2138057,
      "author_name": "Eleftherios Fanioudakis",
      "author_url": "",
      "post_date": "2023-02-10T14:17:07.460000",
      "content": "<p>Im curious, whats your single fold LB? </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2134791,
      "author_name": "Mad_Neil",
      "author_url": "",
      "post_date": "2023-02-08T08:57:20.073000",
      "content": "<p><a href=\"https://www.kaggle.com/taruto1215\" target=\"_blank\">@taruto1215</a> Sorry if this is not appropriate, could you please share your dataset you used for training? Many thanks</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2134970,
          "author_name": "taruto",
          "author_url": "",
          "post_date": "2023-02-08T11:15:57.680000",
          "content": "<p>I am sorry.<br>\nit includes one of my unique solution. so I cannot share you.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2132523,
      "author_name": "FlaviuPaul",
      "author_url": "",
      "post_date": "2023-02-06T21:16:36.953000",
      "content": "<p>I think is related to how you split the data. You can have a R-CC on train and a L-CC on validation, R-CC doesn't have cancer, L-CC has. At the same time I expect that there is a high similarity between the 2, so this makes the score on CV smaller.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2133163,
          "author_name": "taruto",
          "author_url": "",
          "post_date": "2023-02-07T09:37:04.353000",
          "content": "<p>Thanks<br>\nI split the data as same patients are the same fold.<br>\nso for me I think the problem you mentioned can be ignore.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2130418,
      "author_name": "Ơ con lừa!",
      "author_url": "",
      "post_date": "2023-02-05T12:45:27.553000",
      "content": "<p><a href=\"https://www.kaggle.com/taruto1215\" target=\"_blank\">@taruto1215</a> Hello, Could you tell me your dataset you use? Did you use VOI_LUT in your dataset? Thanks</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2131912,
          "author_name": "taruto",
          "author_url": "",
          "post_date": "2023-02-06T13:34:55.653000",
          "content": "<p>Yes, I did</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2131951,
              "author_name": "Ơ con lừa!",
              "author_url": "",
              "post_date": "2023-02-06T14:01:25.017000",
              "content": "<p>I think the main reason why your CV/LB gap maybe in your dataset. Difference in training and inference?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2133184,
              "author_name": "taruto",
              "author_url": "",
              "post_date": "2023-02-07T09:53:18.987000",
              "content": "<p>I use exactly same preprocess pipeline… Hmm</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2126918,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-02-02T15:04:29.667000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 2127714,
          "author_name": "taruto",
          "author_url": "",
          "post_date": "2023-02-03T05:03:38.933000",
          "content": "<p>I get my cv score using this cv strategy for all train data.<br>\nn_splits = 3</p>\n<p><a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/382198\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/382198</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2121006": "there is a huge score gap between my cv and my lb CV: 0.395 (after mean aggregation ) LB: 0.56 (thres=0.20, no strict tunning).\nI am very concern about shake down.\nHow is your score gap??\nthanks.\n\nmy cv strategy is below. (n_splits=3)\nI refer [this notebook](https://www.kaggle.com/code/awsaf49/rsna-bcd-efficientnet-tf-tpu-1vm-train)\n\n \n    strat_cols = [\n        'laterality', 'view', 'biopsy','invasive', 'BIRADS', 'age_bin',\n        'implant', 'density','machine_id', 'difficult_negative_case',\n        'cancer',\n    ]\n\n    df['stratify'] = ''\n    for col in strat_cols:\n        df['stratify'] += df[col].astype(str)\n    \n    kfold = StratifiedGroupKFold(n_splits=Config['SPLITS'],random_state=Config[\"seed\"],shuffle=True)\n    for fold_, (train_idx, valid_idx) in enumerate(kfold.split(df, df['stratify'].values, df['patient_id'].values)):`\n\n**update2 : CV score details (n_splits=3)**\n**all data：precision=0.485 recall=0.333 pf1=0.395 thres=0.5000 roc_auc=0.873**\n**site_id=1 : precision=0.329 recall=0.287 f1=0.307 thres=0.4500 roc_auc=0.826**\n**site_id=2 : precision=0.639 recall=0.416 f1=0.504 thres=0.5000 roc_auc=0.906**",
    "2131916": "Finally my LB is more than 0.60.\nBut my cv score is 0.42.\nThe gap is too big !!!!!!",
    "2121189": "Hi @taruto1215 , how do you get your CV? I mean the LB is calculated by laterality-level, if you train and validate by images, you will get the image-level score. In my experiments, the laterality-level f1 is always much higher than the image-level f1.",
    "2121526": "i am not very sure if i were correct. but for this competition there may be other ways to reduce shakeup by NOT only relying on LB vs CV.\ne.g. my target is to make a model:\n1. have recall at least 0.50 for both site1 and site2\n2. have f1 of at least 0.50 for both site1 and site2 \n3. threshold not larger than 0.55\n4 . .....\n\ni am coming out with many such rules based on my observation.\n\nit is like regularization in machine learning, when you loss is not reliable, you need to think of other \"sensible\" conditions",
    "2121239": "beware if you aggregate by mean. it may be wrong method.\n\nfor example assume\n\npatient1 (cancer):\nimage1 = 0.9\nimage2 (occluded)=0.1\n\npatient2 (cancer):\nimage1 = 0.9\nimage2 (occluded)=0.1\nimage3 (occluded)=0.1\nimage4 (occluded)=0.1\n\n\nthe score for patient 2 is divided out.\n\n---\nbut max is no better\npatient1 (cancer):\nimage1 = 0.7\nimage2 =0.8\n\npatient1 (no cancer):\nimage1 = 0.1\nimage2 =0.1\nimage3 = 0.7 (noise)\nimage4 =0.1\n\n\n",
    "2136783": "interesting, LB0.59 is my training score, while my validation score is lower. kaggle LB score is close to my training score.\n\ncould matching train score = LB score be a way to fight generalisation?\n(of course your have to do early exit)\n\nasuming the top kagglers did not overfit, you can train until train score is around 0.65",
    "2122638": "Hello, Which model did you use?",
    "2121246": "```\nsite_id=1 : precision=0.329 recall=0.287 f1=0.307 thres=0.4500 roc_auc=0.637\nsite_id=2 : precision=0.639 recall=0.416 f1=0.504 thres=0.5000 roc_auc=0.706\n```\n\ni am interested in the score of site1 and site2 submitted separately.\n\n---\nstable f1 if\n1)  precision is close to recall \n2)  precision is very high\n\nstable LB if ?\n1)site1 close to site2\n\nhigh threshold  : maybe overfitting\n\nnote: stable means consistent (i.e. less shakeup) but not necessarily optimum\n\n---\nCV: 0.395 LB: 0.56\n\n\"I am very concern about shake down.\nHow is your score gap??\"\n\nbut even if CV close to LB doesn't guarantee  no shakeup ",
    "2121192": "How that possible??\nProbably overfitted to hidden test set?\n\nI don't think roc_auc about 0.6~0.7 is good performance.",
    "2121017": "me too! cv:0.3 lb:0.46 threshold=0.89 ",
    "2121011": "did you try same aggregation in your cv (mean, max etc..) as you do for the pred_ids during subm?",
    "2132637": "good job!\n\ncurrently all team with LB score above 0.60 has good chance of wining.\ndo spend time to:\n\n1. simulate leader board results. how one or few change of pos samples can change your validation score.\nthis enable you to interprete the LB score correctly and estimate shakeup\n\n2. i still think that mean() aggreataion is not good unless your classifier has very low false positive rate.\nso you may want to spend time to think about how to fuse results.\n\nremember !!!!\nit is the private hidden score that we are interested .you should try to predict the private score for every submission you make !!!\n",
    "2164140": "I also encounter train and validation set scores of 0.8.\nBut the test set is only 0.02",
    "2144025": "My single image CV F1 score is 0.3596 but my LB score is only 0.38. I am using the effnetv2_s on 1536x960 images, what is your CV score for single image right now ?",
    "2138057": "Im curious, whats your single fold LB? ",
    "2134791": "@taruto1215 Sorry if this is not appropriate, could you please share your dataset you used for training? Many thanks",
    "2132523": "I think is related to how you split the data. You can have a R-CC on train and a L-CC on validation, R-CC doesn't have cancer, L-CC has. At the same time I expect that there is a high similarity between the 2, so this makes the score on CV smaller.",
    "2130418": "@taruto1215 Hello, Could you tell me your dataset you use? Did you use VOI_LUT in your dataset? Thanks",
    "2126918": ""
  }
}