{
  "id": 371185,
  "title": "RSNA SMBCD CV vs LB",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/371185",
  "author_name": "Martin Kovacevic Buvinic",
  "post_date": "2022-12-08T13:23:22.678000",
  "votes": 27,
  "comment_count": 28,
  "views": 0,
  "content": "<p>Classic discussion of each competition.</p>\n<p>Update: Run and submitted 5 folds model, cv is based on 100% of the train data and the threshold is computed using the concatenation of the 5 validations sets.</p>\n<p>CV: 0.3884, LB: 0.48 -&gt; My 4 fold model gave better LB 😲</p>\n<p>In my case I have the following scores:</p>\n<p>1 Fold model (CV per patient, laterality)<br>\nCV: 0.4096, LB: 0.44</p>\n<p>2 Fold model (CV per patient, laterality) - CV (average of the folds)<br>\nCV: 0.4317, LB: 0.50</p>\n<p>3 Fold model (CV per patient, laterality) - CV (average of the folds)<br>\nCV: 0.4264, LB: 0.47</p>\n<p>4 Fold model (CV per patient, laterality) - CV (average of the folds)<br>\nCV: 0.4089, LB: 0.51</p>\n<p>Notice that for 2 folds model cv is the best, this is because the cv for those 2 folds is bigger, therefore the difference of CV between folds is not small in my case. For example:</p>\n<p>CV Fold2: 0.45<br>\nCV Fold4: 0.36</p>\n<p>Big difference</p>\n<p>I tried to run the last fold but when the model is validating I got this error.</p>\n<p>all the input arrays must have same number of dimensions, but the array at index 0 has 1 dimension(s) and the array at index 684 has 0 dimension(s)</p>\n<p>My best guess is that there is an observation in the validation where the model can't predict and return an empty value, anyone experienced the same issue?</p>",
  "messages": [
    {
      "id": 2059084,
      "postDate": "2022-12-08T13:23:22.680Z",
      "content": "<p>Classic discussion of each competition.</p>\n<p>Update: Run and submitted 5 folds model, cv is based on 100% of the train data and the threshold is computed using the concatenation of the 5 validations sets.</p>\n<p>CV: 0.3884, LB: 0.48 -&gt; My 4 fold model gave better LB 😲</p>\n<p>In my case I have the following scores:</p>\n<p>1 Fold model (CV per patient, laterality)<br>\nCV: 0.4096, LB: 0.44</p>\n<p>2 Fold model (CV per patient, laterality) - CV (average of the folds)<br>\nCV: 0.4317, LB: 0.50</p>\n<p>3 Fold model (CV per patient, laterality) - CV (average of the folds)<br>\nCV: 0.4264, LB: 0.47</p>\n<p>4 Fold model (CV per patient, laterality) - CV (average of the folds)<br>\nCV: 0.4089, LB: 0.51</p>\n<p>Notice that for 2 folds model cv is the best, this is because the cv for those 2 folds is bigger, therefore the difference of CV between folds is not small in my case. For example:</p>\n<p>CV Fold2: 0.45<br>\nCV Fold4: 0.36</p>\n<p>Big difference</p>\n<p>I tried to run the last fold but when the model is validating I got this error.</p>\n<p>all the input arrays must have same number of dimensions, but the array at index 0 has 1 dimension(s) and the array at index 684 has 0 dimension(s)</p>\n<p>My best guess is that there is an observation in the validation where the model can't predict and return an empty value, anyone experienced the same issue?</p>",
      "rawMarkdown": "Classic discussion of each competition.\n\nUpdate: Run and submitted 5 folds model, cv is based on 100% of the train data and the threshold is computed using the concatenation of the 5 validations sets.\n\nCV: 0.3884, LB: 0.48 -> My 4 fold model gave better LB 😲\n\nIn my case I have the following scores:\n\n1 Fold model (CV per patient, laterality)\nCV: 0.4096, LB: 0.44\n\n2 Fold model (CV per patient, laterality) - CV (average of the folds)\nCV: 0.4317, LB: 0.50\n\n3 Fold model (CV per patient, laterality) - CV (average of the folds)\nCV: 0.4264, LB: 0.47\n\n4 Fold model (CV per patient, laterality) - CV (average of the folds)\nCV: 0.4089, LB: 0.51\n\nNotice that for 2 folds model cv is the best, this is because the cv for those 2 folds is bigger, therefore the difference of CV between folds is not small in my case. For example:\n\nCV Fold2: 0.45\nCV Fold4: 0.36\n\nBig difference\n\nI tried to run the last fold but when the model is validating I got this error.\n\nall the input arrays must have same number of dimensions, but the array at index 0 has 1 dimension(s) and the array at index 684 has 0 dimension(s)\n\nMy best guess is that there is an observation in the validation where the model can't predict and return an empty value, anyone experienced the same issue?",
      "votes": 27
    },
    {
      "id": 2060717,
      "postDate": "2022-12-10T10:42:30.077Z",
      "content": "<ul>\n<li>CV : 0.3880 - LB 0.42  (reference model)</li>\n<li>CV <strong>0.3985</strong> - LB 0.45  (reference model, improved CV)</li>\n<li>CV 0.3869 - LB <strong>0.49</strong>  (bigger model)</li>\n</ul>\n<p>My correlation does not look so good. All models use image size 1024</p>\n<p>Also, 3 of my subs gave a score of 0. Anyone encountered this ?<br>\nI'm guessing that's because the models predicted no 1, but that makes no sense to me.</p>\n<p>EDIT : Fixed by decreasing batch size, somehow code was crashing during inference while working fine in kaggle notebooks.</p>",
      "rawMarkdown": "- CV : 0.3880 - LB 0.42  (reference model)\n- CV **0.3985** - LB 0.45  (reference model, improved CV)\n- CV 0.3869 - LB **0.49**  (bigger model)\n\nMy correlation does not look so good. All models use image size 1024\n\nAlso, 3 of my subs gave a score of 0. Anyone encountered this ?\nI'm guessing that's because the models predicted no 1, but that makes no sense to me.\n\nEDIT : Fixed by decreasing batch size, somehow code was crashing during inference while working fine in kaggle notebooks.",
      "votes": 5,
      "replies": [
        {
          "id": 2060827,
          "postDate": "2022-12-10T13:41:08.793Z",
          "content": "<p>That is strange, it was exactly the same code, the only thing that change where the weights of the model?</p>\n<p>I also don't have great correlation between cv and lb, public lb is not so big, and we are using thresholds, maybe for that specific chunk of the public lb you may move your threshold a little and have a better lb but a worst private leaderboard. My best guess is that you will have correlations only if you get big improvements on CV, smaller ones could not correlated so well</p>",
          "rawMarkdown": "That is strange, it was exactly the same code, the only thing that change where the weights of the model?\n\nI also don't have great correlation between cv and lb, public lb is not so big, and we are using thresholds, maybe for that specific chunk of the public lb you may move your threshold a little and have a better lb but a worst private leaderboard. My best guess is that you will have correlations only if you get big improvements on CV, smaller ones could not correlated so well",
          "votes": 2
        },
        {
          "id": 2060961,
          "postDate": "2022-12-10T15:55:01.080Z",
          "content": "<p>Changed the weights and the threshold only yup. I'm doing a few subs for debugging right now.</p>",
          "rawMarkdown": "Changed the weights and the threshold only yup. I'm doing a few subs for debugging right now.",
          "votes": 1
        },
        {
          "id": 2086821,
          "postDate": "2023-01-05T04:27:52.570Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> Which split do you use?</p>",
          "rawMarkdown": "Hi @theoviel Which split do you use?"
        }
      ]
    },
    {
      "id": 2059106,
      "postDate": "2022-12-08T13:39:57.920Z",
      "content": "<p>If you want to assess your CV score reliably, you should compute the pfbeta on the whole data. Main reason is that <code>mean_over_all_data(pfbeta) != pfbeta(all_data)</code>.<br>\nThis matters even more here because we are tweaking thresholds.</p>\n<p><strong>Pseudo code :</strong></p>\n<pre><code>pred_oof = np.concatenate([pred_val for pred_val in pred_vals])\ncv = pfbeta(pred_oof &gt; thresh, y)\n</code></pre>\n<p><strong>My CV scores :</strong> </p>\n<p>Fold 0 : 0.4022<br>\nFold 1 : 0.3684<br>\nFold 2 : 0.4086<br>\nFold 3 : 0.3902<br>\nAverage : 0.3925<br>\nCV : 0.3880<br>\nLB : 0.40 +/- 0.02  (quite sensitive to thresholding)</p>",
      "rawMarkdown": "If you want to assess your CV score reliably, you should compute the pfbeta on the whole data. Main reason is that `mean_over_all_data(pfbeta) != pfbeta(all_data)`.\nThis matters even more here because we are tweaking thresholds.\n\n**Pseudo code :**\n```\npred_oof = np.concatenate([pred_val for pred_val in pred_vals])\ncv = pfbeta(pred_oof > thresh, y)\n```\n\n**My CV scores :** \n\nFold 0 : 0.4022\nFold 1 : 0.3684\nFold 2 : 0.4086\nFold 3 : 0.3902\nAverage : 0.3925\nCV : 0.3880\nLB : 0.40 +/- 0.02  (quite sensitive to thresholding)",
      "votes": 5,
      "replies": [
        {
          "id": 2059120,
          "postDate": "2022-12-08T13:50:05.470Z",
          "content": "<p>Test set is build for each image, laterality combinations, if you compute your CV on the raw data without any aggregations you will not be aligned to the test. At least this is how I am thinking the problem, maybe I am wrong 😅</p>",
          "rawMarkdown": "Test set is build for each image, laterality combinations, if you compute your CV on the raw data without any aggregations you will not be aligned to the test. At least this is how I am thinking the problem, maybe I am wrong 😅"
        },
        {
          "id": 2059134,
          "postDate": "2022-12-08T14:02:27.383Z",
          "content": "<p>You are right on this, I am discussing the way you aggregate fold predictions, not patient_laterality :)<br>\nUnless I misunderstood the \"CV (average of the folds)\" thing.</p>",
          "rawMarkdown": "You are right on this, I am discussing the way you aggregate fold predictions, not patient_laterality :)\nUnless I misunderstood the \"CV (average of the folds)\" thing.",
          "votes": 4
        },
        {
          "id": 2059148,
          "postDate": "2022-12-08T14:10:25.400Z",
          "content": "<blockquote>\n  <p>Average : 0.3925<br>\n  CV : 0.3880</p>\n</blockquote>\n<p>The difference between the two is what Theo wants to explain.</p>",
          "rawMarkdown": "> Average : 0.3925\nCV : 0.3880\n\nThe difference between the two is what Theo wants to explain.",
          "votes": 3
        },
        {
          "id": 2059150,
          "postDate": "2022-12-08T14:11:39.580Z",
          "content": "<p>what input size are you using? 1024?</p>",
          "rawMarkdown": "what input size are you using? 1024?\n",
          "votes": 1
        },
        {
          "id": 2059151,
          "postDate": "2022-12-08T14:11:55.963Z",
          "content": "<p>Ohh I think you misunderstood, what I meant is that for example if I want to compute de CV of 2 folds I do:</p>\n<p>(CV fold1 + CV fold2) / 2.</p>\n<p>It is just a lazy way to estimate real out of folds CV, to compute correctly 5 folds training would be needed and then you can validate on 100% of the data</p>",
          "rawMarkdown": "Ohh I think you misunderstood, what I meant is that for example if I want to compute de CV of 2 folds I do:\n\n(CV fold1 + CV fold2) / 2.\n\nIt is just a lazy way to estimate real out of folds CV, to compute correctly 5 folds training would be needed and then you can validate on 100% of the data"
        },
        {
          "id": 2059152,
          "postDate": "2022-12-08T14:14:36.463Z",
          "content": "<p>Yep 1024, I tried 1280 but I had hardware constraints, I could run a training loop but with a small batch size, and CV was worst. Nevertheless this does not mean that 1024 is better than 1280, maybe 1280 trained correctly will provide better results </p>",
          "rawMarkdown": "Yep 1024, I tried 1280 but I had hardware constraints, I could run a training loop but with a small batch size, and CV was worst. Nevertheless this does not mean that 1024 is better than 1280, maybe 1280 trained correctly will provide better results ",
          "votes": 1
        },
        {
          "id": 2059153,
          "postDate": "2022-12-08T14:14:42.033Z",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> I use 1024 as well<br>\n<a href=\"https://www.kaggle.com/ragnar123\" target=\"_blank\">@ragnar123</a> I understood, what I'm saying is that it's better to compute the CV on the concatenation of the predictions of your folds than average your two validation scores</p>",
          "rawMarkdown": "@hengck23 I use 1024 as well\n@ragnar123 I understood, what I'm saying is that it's better to compute the CV on the concatenation of the predictions of your folds than average your two validation scores",
          "votes": 1
        },
        {
          "id": 2059154,
          "postDate": "2022-12-08T14:15:40.353Z",
          "content": "<p>Yes that is correct, actually to get threshold also the concatenation would be the best. </p>\n<p>Concatenation = out of folds CV score</p>\n<p>This are initial test, actually I don't know if submission time will allow the predictions of 3 or 5 folds based on that we have only 9 hours, 1 fold run all good for me, I want to check how many folds (models) we can use for inference based on the 9 hours we have</p>",
          "rawMarkdown": "Yes that is correct, actually to get threshold also the concatenation would be the best. \n\nConcatenation = out of folds CV score\n\nThis are initial test, actually I don't know if submission time will allow the predictions of 3 or 5 folds based on that we have only 9 hours, 1 fold run all good for me, I want to check how many folds (models) we can use for inference based on the 9 hours we have"
        }
      ]
    },
    {
      "id": 2082353,
      "postDate": "2023-01-01T13:14:10.363Z",
      "content": "<p>Single fold:<br>\nLocal validation=0.44 but LB goes from 0.35 to 0.45 depending on +-0.025 threshold.</p>\n<p><strong>Update:</strong> 4 folds:<br>\nCV=0.378, LB goes  from 0.50 to 0.52 depending on +-0.025 threshold.</p>",
      "rawMarkdown": "Single fold:\nLocal validation=0.44 but LB goes from 0.35 to 0.45 depending on +-0.025 threshold.\n\n**Update:** 4 folds:\nCV=0.378, LB goes  from 0.50 to 0.52 depending on +-0.025 threshold.",
      "votes": 4,
      "replies": [
        {
          "id": 2091715,
          "postDate": "2023-01-08T17:12:25.763Z",
          "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> Thank you for sharing. BTW, Which split do you use?</p>",
          "rawMarkdown": "@mpware Thank you for sharing. BTW, Which split do you use?",
          "votes": 1,
          "replies": [
            {
              "id": 2091719,
              "postDate": "2023-01-08T17:16:38.087Z",
              "content": "<p>StratifiedGroupKFold with groups=PATIENT_ID and CANCER to stratify.<br>\nThere is a probing discussion here: <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/370341\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/370341</a><br>\nAnd we know there is no PATIENT_ID overlap between train/test.</p>\n<p>I'm wondering if binarizing predictions based on threshold would lead to huge shakeup at the end.</p>",
              "rawMarkdown": "StratifiedGroupKFold with groups=PATIENT_ID and CANCER to stratify.\nThere is a probing discussion here: https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/370341\nAnd we know there is no PATIENT_ID overlap between train/test.\n\nI'm wondering if binarizing predictions based on threshold would lead to huge shakeup at the end.",
              "votes": 5
            },
            {
              "id": 2091806,
              "postDate": "2023-01-08T18:48:24.023Z",
              "content": "<p>Thank you. I think the shake will not happen with the top teams</p>",
              "rawMarkdown": "Thank you. I think the shake will not happen with the top teams",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2059584,
      "postDate": "2022-12-09T03:41:58.743Z",
      "content": "<p>interesting results here<br>\n<a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/370333#2059570\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/370333#2059570</a></p>\n<p>2 fold 2048 input for efficientnet-B2<br>\ncv = 0.369942,0.37333<br>\nlb = 0.51</p>",
      "rawMarkdown": "interesting results here\nhttps://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/370333#2059570\n\n2 fold 2048 input for efficientnet-B2\ncv = 0.369942,0.37333\nlb = 0.51\n ",
      "votes": 3,
      "replies": [
        {
          "id": 2060163,
          "postDate": "2022-12-09T15:45:05.453Z",
          "content": "<p>Interesting, 2 folds using 1024 and efficientnetb3 had a better CV in my case and LB 0.50. Maybe 1024 is big enough and going from 1024 to 2048 is not a big difference (does not provide more signal).</p>\n<p>On the other hand image size 2048 will take more time in inference so searching for a nice image size that has nice signal and that can run in inference fast is going to be one of the keys</p>\n<p>What hardware are you using for 2048 image size?</p>",
          "rawMarkdown": "Interesting, 2 folds using 1024 and efficientnetb3 had a better CV in my case and LB 0.50. Maybe 1024 is big enough and going from 1024 to 2048 is not a big difference (does not provide more signal).\n\nOn the other hand image size 2048 will take more time in inference so searching for a nice image size that has nice signal and that can run in inference fast is going to be one of the keys\n\nWhat hardware are you using for 2048 image size?",
          "votes": 3
        }
      ]
    },
    {
      "id": 2059523,
      "postDate": "2022-12-09T01:06:44.243Z",
      "content": "<p>thanks for the post. there are two ways to improve score of a single model:</p>\n<p>[1] larger image<br>\n[2] stronger model.</p>\n<p>i have been focusing on [1] (e.g. 2048 input). But i my last submissions make me realize that this is difficult and i got submission timeout and out of memory. There is more work here … (maybe i need to distill the results to smaller models etc)</p>\n<p>this post makes me realized  that [2] is a better option unless I can solved the the issues above.<br>\nThis post and my experiments also confirms that LB-CV gap decreases when you have better results in the range of LB 0.5 (i.e. your generalization improves)</p>",
      "rawMarkdown": "thanks for the post. there are two ways to improve score of a single model:\n\n[1] larger image\n[2] stronger model.\n\ni have been focusing on [1] (e.g. 2048 input). But i my last submissions make me realize that this is difficult and i got submission timeout and out of memory. There is more work here ... (maybe i need to distill the results to smaller models etc)\n\nthis post makes me realized  that [2] is a better option unless I can solved the the issues above.\nThis post and my experiments also confirms that LB-CV gap decreases when you have better results in the range of LB 0.5 (i.e. your generalization improves)",
      "votes": 4,
      "replies": [
        {
          "id": 2060239,
          "postDate": "2022-12-09T17:00:47.847Z",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> </p>\n<p>How about </p>\n<p>[3] Multiple models</p>\n<p>?</p>",
          "rawMarkdown": "@hengck23 \n\nHow about \n\n[3] Multiple models\n\n?\n"
        }
      ]
    },
    {
      "id": 2064008,
      "postDate": "2022-12-13T13:07:34.603Z",
      "content": "<p><a href=\"https://www.kaggle.com/ragnar123\" target=\"_blank\">@ragnar123</a> <br>\nHow did you compute CV for each fold? I tried to use <a href=\"https://www.kaggle.com/code/sohier/probabilistic-f-score\" target=\"_blank\">https://www.kaggle.com/code/sohier/probabilistic-f-score</a> but the train pf1 score and validation pf1 score always end up with nearly 0.03 ~ 0.08. </p>",
      "rawMarkdown": "@ragnar123 \nHow did you compute CV for each fold? I tried to use https://www.kaggle.com/code/sohier/probabilistic-f-score but the train pf1 score and validation pf1 score always end up with nearly 0.03 ~ 0.08. ",
      "votes": 2,
      "replies": [
        {
          "id": 2064100,
          "postDate": "2022-12-13T14:20:47.777Z",
          "content": "<p>You need to truncate your predictions and transform them to 1 or 0. This is a postprocess that helps a lot with the metric of the competition.</p>\n<p>For traditional f1 score is the same, you need to use a threshold and transform predictions into 1 or 0.</p>",
          "rawMarkdown": "You need to truncate your predictions and transform them to 1 or 0. This is a postprocess that helps a lot with the metric of the competition.\n\nFor traditional f1 score is the same, you need to use a threshold and transform predictions into 1 or 0.\n\n",
          "votes": 1
        },
        {
          "id": 2064211,
          "postDate": "2022-12-13T15:23:22.877Z",
          "content": "<p>Thanks for the reply. Could you please elaborate on truncation and transformation? Any sample code?</p>",
          "rawMarkdown": "Thanks for the reply. Could you please elaborate on truncation and transformation? Any sample code?"
        },
        {
          "id": 2064245,
          "postDate": "2022-12-13T16:08:19.037Z",
          "content": "<p>With the following code you will transform your predictions to 1 or 0, 0.20 is the threshold and you can find it by trying a lot of threshold and pick the one that gives the best cv<br>\nnp.where(predictions &gt; 0.20, 1, 0)</p>",
          "rawMarkdown": "With the following code you will transform your predictions to 1 or 0, 0.20 is the threshold and you can find it by trying a lot of threshold and pick the one that gives the best cv\nnp.where(predictions > 0.20, 1, 0)",
          "votes": 2
        },
        {
          "id": 2064265,
          "postDate": "2022-12-13T16:26:13.090Z",
          "content": "<p>Uh, got it. And one last thing, do you know why the metric is called <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/overview/evaluation\" target=\"_blank\">probabilistic f-score</a> rather than simply f1-score? </p>",
          "rawMarkdown": "Uh, got it. And one last thing, do you know why the metric is called [probabilistic f-score](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/overview/evaluation) rather than simply f1-score? ",
          "votes": 1
        },
        {
          "id": 2064318,
          "postDate": "2022-12-13T17:00:17.043Z",
          "content": "<p>There is a paper for probabilitic f-score that explains better what is the difference compared to f1-score.</p>",
          "rawMarkdown": "There is a paper for probabilitic f-score that explains better what is the difference compared to f1-score.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2069446,
      "postDate": "2022-12-19T01:59:26.050Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/ragnar123\" target=\"_blank\">@ragnar123</a> I don't understand that you split your data (per patient, laterality). Is it GroupKFold with patient? But laterality is? Thanks</p>",
      "rawMarkdown": "Hi @ragnar123 I don't understand that you split your data (per patient, laterality). Is it GroupKFold with patient? But laterality is? Thanks"
    }
  ],
  "comments": [
    {
      "id": 2060717,
      "author_name": "Theo Viel",
      "author_url": "",
      "post_date": "2022-12-10T10:42:30.077000",
      "content": "<ul>\n<li>CV : 0.3880 - LB 0.42  (reference model)</li>\n<li>CV <strong>0.3985</strong> - LB 0.45  (reference model, improved CV)</li>\n<li>CV 0.3869 - LB <strong>0.49</strong>  (bigger model)</li>\n</ul>\n<p>My correlation does not look so good. All models use image size 1024</p>\n<p>Also, 3 of my subs gave a score of 0. Anyone encountered this ?<br>\nI'm guessing that's because the models predicted no 1, but that makes no sense to me.</p>\n<p>EDIT : Fixed by decreasing batch size, somehow code was crashing during inference while working fine in kaggle notebooks.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 2060827,
          "author_name": "Martin Kovacevic Buvinic",
          "author_url": "",
          "post_date": "2022-12-10T13:41:08.793000",
          "content": "<p>That is strange, it was exactly the same code, the only thing that change where the weights of the model?</p>\n<p>I also don't have great correlation between cv and lb, public lb is not so big, and we are using thresholds, maybe for that specific chunk of the public lb you may move your threshold a little and have a better lb but a worst private leaderboard. My best guess is that you will have correlations only if you get big improvements on CV, smaller ones could not correlated so well</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2060961,
          "author_name": "Theo Viel",
          "author_url": "",
          "post_date": "2022-12-10T15:55:01.080000",
          "content": "<p>Changed the weights and the threshold only yup. I'm doing a few subs for debugging right now.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2086821,
          "author_name": "Mad_Neil",
          "author_url": "",
          "post_date": "2023-01-05T04:27:52.570000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> Which split do you use?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2059106,
      "author_name": "Theo Viel",
      "author_url": "",
      "post_date": "2022-12-08T13:39:57.920000",
      "content": "<p>If you want to assess your CV score reliably, you should compute the pfbeta on the whole data. Main reason is that <code>mean_over_all_data(pfbeta) != pfbeta(all_data)</code>.<br>\nThis matters even more here because we are tweaking thresholds.</p>\n<p><strong>Pseudo code :</strong></p>\n<pre><code>pred_oof = np.concatenate([pred_val for pred_val in pred_vals])\ncv = pfbeta(pred_oof &gt; thresh, y)\n</code></pre>\n<p><strong>My CV scores :</strong> </p>\n<p>Fold 0 : 0.4022<br>\nFold 1 : 0.3684<br>\nFold 2 : 0.4086<br>\nFold 3 : 0.3902<br>\nAverage : 0.3925<br>\nCV : 0.3880<br>\nLB : 0.40 +/- 0.02  (quite sensitive to thresholding)</p>",
      "votes": 5,
      "replies": [
        {
          "id": 2059120,
          "author_name": "Martin Kovacevic Buvinic",
          "author_url": "",
          "post_date": "2022-12-08T13:50:05.470000",
          "content": "<p>Test set is build for each image, laterality combinations, if you compute your CV on the raw data without any aggregations you will not be aligned to the test. At least this is how I am thinking the problem, maybe I am wrong 😅</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2059134,
          "author_name": "Theo Viel",
          "author_url": "",
          "post_date": "2022-12-08T14:02:27.383000",
          "content": "<p>You are right on this, I am discussing the way you aggregate fold predictions, not patient_laterality :)<br>\nUnless I misunderstood the \"CV (average of the folds)\" thing.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 2059148,
          "author_name": "Leon",
          "author_url": "",
          "post_date": "2022-12-08T14:10:25.400000",
          "content": "<blockquote>\n  <p>Average : 0.3925<br>\n  CV : 0.3880</p>\n</blockquote>\n<p>The difference between the two is what Theo wants to explain.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 2059150,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2022-12-08T14:11:39.580000",
          "content": "<p>what input size are you using? 1024?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2059151,
          "author_name": "Martin Kovacevic Buvinic",
          "author_url": "",
          "post_date": "2022-12-08T14:11:55.963000",
          "content": "<p>Ohh I think you misunderstood, what I meant is that for example if I want to compute de CV of 2 folds I do:</p>\n<p>(CV fold1 + CV fold2) / 2.</p>\n<p>It is just a lazy way to estimate real out of folds CV, to compute correctly 5 folds training would be needed and then you can validate on 100% of the data</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2059152,
          "author_name": "Martin Kovacevic Buvinic",
          "author_url": "",
          "post_date": "2022-12-08T14:14:36.463000",
          "content": "<p>Yep 1024, I tried 1280 but I had hardware constraints, I could run a training loop but with a small batch size, and CV was worst. Nevertheless this does not mean that 1024 is better than 1280, maybe 1280 trained correctly will provide better results </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2059153,
          "author_name": "Theo Viel",
          "author_url": "",
          "post_date": "2022-12-08T14:14:42.033000",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> I use 1024 as well<br>\n<a href=\"https://www.kaggle.com/ragnar123\" target=\"_blank\">@ragnar123</a> I understood, what I'm saying is that it's better to compute the CV on the concatenation of the predictions of your folds than average your two validation scores</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2059154,
          "author_name": "Martin Kovacevic Buvinic",
          "author_url": "",
          "post_date": "2022-12-08T14:15:40.353000",
          "content": "<p>Yes that is correct, actually to get threshold also the concatenation would be the best. </p>\n<p>Concatenation = out of folds CV score</p>\n<p>This are initial test, actually I don't know if submission time will allow the predictions of 3 or 5 folds based on that we have only 9 hours, 1 fold run all good for me, I want to check how many folds (models) we can use for inference based on the 9 hours we have</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2082353,
      "author_name": "MPWARE",
      "author_url": "",
      "post_date": "2023-01-01T13:14:10.363000",
      "content": "<p>Single fold:<br>\nLocal validation=0.44 but LB goes from 0.35 to 0.45 depending on +-0.025 threshold.</p>\n<p><strong>Update:</strong> 4 folds:<br>\nCV=0.378, LB goes  from 0.50 to 0.52 depending on +-0.025 threshold.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2091715,
          "author_name": "Mad_Neil",
          "author_url": "",
          "post_date": "2023-01-08T17:12:25.763000",
          "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> Thank you for sharing. BTW, Which split do you use?</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2091719,
              "author_name": "MPWARE",
              "author_url": "",
              "post_date": "2023-01-08T17:16:38.087000",
              "content": "<p>StratifiedGroupKFold with groups=PATIENT_ID and CANCER to stratify.<br>\nThere is a probing discussion here: <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/370341\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/370341</a><br>\nAnd we know there is no PATIENT_ID overlap between train/test.</p>\n<p>I'm wondering if binarizing predictions based on threshold would lead to huge shakeup at the end.</p>",
              "votes": 5,
              "replies": []
            },
            {
              "id": 2091806,
              "author_name": "Mad_Neil",
              "author_url": "",
              "post_date": "2023-01-08T18:48:24.023000",
              "content": "<p>Thank you. I think the shake will not happen with the top teams</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2059584,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2022-12-09T03:41:58.743000",
      "content": "<p>interesting results here<br>\n<a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/370333#2059570\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/370333#2059570</a></p>\n<p>2 fold 2048 input for efficientnet-B2<br>\ncv = 0.369942,0.37333<br>\nlb = 0.51</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2060163,
          "author_name": "Martin Kovacevic Buvinic",
          "author_url": "",
          "post_date": "2022-12-09T15:45:05.453000",
          "content": "<p>Interesting, 2 folds using 1024 and efficientnetb3 had a better CV in my case and LB 0.50. Maybe 1024 is big enough and going from 1024 to 2048 is not a big difference (does not provide more signal).</p>\n<p>On the other hand image size 2048 will take more time in inference so searching for a nice image size that has nice signal and that can run in inference fast is going to be one of the keys</p>\n<p>What hardware are you using for 2048 image size?</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 2059523,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2022-12-09T01:06:44.243000",
      "content": "<p>thanks for the post. there are two ways to improve score of a single model:</p>\n<p>[1] larger image<br>\n[2] stronger model.</p>\n<p>i have been focusing on [1] (e.g. 2048 input). But i my last submissions make me realize that this is difficult and i got submission timeout and out of memory. There is more work here … (maybe i need to distill the results to smaller models etc)</p>\n<p>this post makes me realized  that [2] is a better option unless I can solved the the issues above.<br>\nThis post and my experiments also confirms that LB-CV gap decreases when you have better results in the range of LB 0.5 (i.e. your generalization improves)</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2060239,
          "author_name": "@kaggleqrdl",
          "author_url": "",
          "post_date": "2022-12-09T17:00:47.847000",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> </p>\n<p>How about </p>\n<p>[3] Multiple models</p>\n<p>?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2064008,
      "author_name": "Simon Alerdic",
      "author_url": "",
      "post_date": "2022-12-13T13:07:34.603000",
      "content": "<p><a href=\"https://www.kaggle.com/ragnar123\" target=\"_blank\">@ragnar123</a> <br>\nHow did you compute CV for each fold? I tried to use <a href=\"https://www.kaggle.com/code/sohier/probabilistic-f-score\" target=\"_blank\">https://www.kaggle.com/code/sohier/probabilistic-f-score</a> but the train pf1 score and validation pf1 score always end up with nearly 0.03 ~ 0.08. </p>",
      "votes": 2,
      "replies": [
        {
          "id": 2064100,
          "author_name": "Martin Kovacevic Buvinic",
          "author_url": "",
          "post_date": "2022-12-13T14:20:47.777000",
          "content": "<p>You need to truncate your predictions and transform them to 1 or 0. This is a postprocess that helps a lot with the metric of the competition.</p>\n<p>For traditional f1 score is the same, you need to use a threshold and transform predictions into 1 or 0.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2064211,
          "author_name": "Simon Alerdic",
          "author_url": "",
          "post_date": "2022-12-13T15:23:22.877000",
          "content": "<p>Thanks for the reply. Could you please elaborate on truncation and transformation? Any sample code?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2064245,
          "author_name": "Martin Kovacevic Buvinic",
          "author_url": "",
          "post_date": "2022-12-13T16:08:19.037000",
          "content": "<p>With the following code you will transform your predictions to 1 or 0, 0.20 is the threshold and you can find it by trying a lot of threshold and pick the one that gives the best cv<br>\nnp.where(predictions &gt; 0.20, 1, 0)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2064265,
          "author_name": "Simon Alerdic",
          "author_url": "",
          "post_date": "2022-12-13T16:26:13.090000",
          "content": "<p>Uh, got it. And one last thing, do you know why the metric is called <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/overview/evaluation\" target=\"_blank\">probabilistic f-score</a> rather than simply f1-score? </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2064318,
          "author_name": "Martin Kovacevic Buvinic",
          "author_url": "",
          "post_date": "2022-12-13T17:00:17.043000",
          "content": "<p>There is a paper for probabilitic f-score that explains better what is the difference compared to f1-score.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2069446,
      "author_name": "Hưng Nguyễn Khánh",
      "author_url": "",
      "post_date": "2022-12-19T01:59:26.050000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/ragnar123\" target=\"_blank\">@ragnar123</a> I don't understand that you split your data (per patient, laterality). Is it GroupKFold with patient? But laterality is? Thanks</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2059084": "Classic discussion of each competition.\n\nUpdate: Run and submitted 5 folds model, cv is based on 100% of the train data and the threshold is computed using the concatenation of the 5 validations sets.\n\nCV: 0.3884, LB: 0.48 -> My 4 fold model gave better LB 😲\n\nIn my case I have the following scores:\n\n1 Fold model (CV per patient, laterality)\nCV: 0.4096, LB: 0.44\n\n2 Fold model (CV per patient, laterality) - CV (average of the folds)\nCV: 0.4317, LB: 0.50\n\n3 Fold model (CV per patient, laterality) - CV (average of the folds)\nCV: 0.4264, LB: 0.47\n\n4 Fold model (CV per patient, laterality) - CV (average of the folds)\nCV: 0.4089, LB: 0.51\n\nNotice that for 2 folds model cv is the best, this is because the cv for those 2 folds is bigger, therefore the difference of CV between folds is not small in my case. For example:\n\nCV Fold2: 0.45\nCV Fold4: 0.36\n\nBig difference\n\nI tried to run the last fold but when the model is validating I got this error.\n\nall the input arrays must have same number of dimensions, but the array at index 0 has 1 dimension(s) and the array at index 684 has 0 dimension(s)\n\nMy best guess is that there is an observation in the validation where the model can't predict and return an empty value, anyone experienced the same issue?",
    "2060717": "- CV : 0.3880 - LB 0.42  (reference model)\n- CV **0.3985** - LB 0.45  (reference model, improved CV)\n- CV 0.3869 - LB **0.49**  (bigger model)\n\nMy correlation does not look so good. All models use image size 1024\n\nAlso, 3 of my subs gave a score of 0. Anyone encountered this ?\nI'm guessing that's because the models predicted no 1, but that makes no sense to me.\n\nEDIT : Fixed by decreasing batch size, somehow code was crashing during inference while working fine in kaggle notebooks.",
    "2059106": "If you want to assess your CV score reliably, you should compute the pfbeta on the whole data. Main reason is that `mean_over_all_data(pfbeta) != pfbeta(all_data)`.\nThis matters even more here because we are tweaking thresholds.\n\n**Pseudo code :**\n```\npred_oof = np.concatenate([pred_val for pred_val in pred_vals])\ncv = pfbeta(pred_oof > thresh, y)\n```\n\n**My CV scores :** \n\nFold 0 : 0.4022\nFold 1 : 0.3684\nFold 2 : 0.4086\nFold 3 : 0.3902\nAverage : 0.3925\nCV : 0.3880\nLB : 0.40 +/- 0.02  (quite sensitive to thresholding)",
    "2082353": "Single fold:\nLocal validation=0.44 but LB goes from 0.35 to 0.45 depending on +-0.025 threshold.\n\n**Update:** 4 folds:\nCV=0.378, LB goes  from 0.50 to 0.52 depending on +-0.025 threshold.",
    "2059584": "interesting results here\nhttps://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/370333#2059570\n\n2 fold 2048 input for efficientnet-B2\ncv = 0.369942,0.37333\nlb = 0.51\n ",
    "2059523": "thanks for the post. there are two ways to improve score of a single model:\n\n[1] larger image\n[2] stronger model.\n\ni have been focusing on [1] (e.g. 2048 input). But i my last submissions make me realize that this is difficult and i got submission timeout and out of memory. There is more work here ... (maybe i need to distill the results to smaller models etc)\n\nthis post makes me realized  that [2] is a better option unless I can solved the the issues above.\nThis post and my experiments also confirms that LB-CV gap decreases when you have better results in the range of LB 0.5 (i.e. your generalization improves)",
    "2064008": "@ragnar123 \nHow did you compute CV for each fold? I tried to use https://www.kaggle.com/code/sohier/probabilistic-f-score but the train pf1 score and validation pf1 score always end up with nearly 0.03 ~ 0.08. ",
    "2069446": "Hi @ragnar123 I don't understand that you split your data (per patient, laterality). Is it GroupKFold with patient? But laterality is? Thanks"
  }
}