{
  "id": 413153,
  "title": "Single Model CV-LB Thread",
  "url": "/competitions/google-research-identify-contrails-reduce-global-warming/discussion/413153",
  "author_name": "Dracarys",
  "post_date": "2023-05-27T08:06:14.322000",
  "votes": 47,
  "comment_count": 86,
  "views": 0,
  "content": "<table>\n<thead>\n<tr>\n<th>CV</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.65</td>\n<td>0.536</td>\n</tr>\n<tr>\n<td>0.71</td>\n<td>0.578</td>\n</tr>\n<tr>\n<td>0.76</td>\n<td>0.621</td>\n</tr>\n<tr>\n<td>0.77</td>\n<td>0.633</td>\n</tr>\n</tbody>\n</table>\n<blockquote>\n  <p>So far CV-LB seems perfectly correlated. What are your results?</p>\n</blockquote>",
  "messages": [
    {
      "id": 2276804,
      "postDate": "2023-05-27T08:06:14.323Z",
      "content": "<table>\n<thead>\n<tr>\n<th>CV</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.65</td>\n<td>0.536</td>\n</tr>\n<tr>\n<td>0.71</td>\n<td>0.578</td>\n</tr>\n<tr>\n<td>0.76</td>\n<td>0.621</td>\n</tr>\n<tr>\n<td>0.77</td>\n<td>0.633</td>\n</tr>\n</tbody>\n</table>\n<blockquote>\n  <p>So far CV-LB seems perfectly correlated. What are your results?</p>\n</blockquote>",
      "rawMarkdown": "\n| CV | LB |\n| --- | --- |\n| 0.65 | 0.536 |\n|0.71  | 0.578 |\n|0.76  | 0.621 |\n|0.77  | 0.633 |\n\n> So far CV-LB seems perfectly correlated. What are your results?\n",
      "votes": 45
    },
    {
      "id": 2326006,
      "postDate": "2023-07-01T19:11:02.290Z",
      "content": "<p>No folds. Train folder vs validation folder.<br>\nThreshold was selected by iterating thru<code>np.arange(0.01, 0.51, 0.01)</code>.</p>\n<table>\n<thead>\n<tr>\n<th>CV</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.633</td>\n<td>0.636</td>\n</tr>\n<tr>\n<td>0.641</td>\n<td>0.642</td>\n</tr>\n<tr>\n<td>0.656</td>\n<td>0.662</td>\n</tr>\n<tr>\n<td>0.658</td>\n<td>0.666</td>\n</tr>\n<tr>\n<td>0.663</td>\n<td>0.671</td>\n</tr>\n<tr>\n<td>0.664</td>\n<td>0.673</td>\n</tr>\n<tr>\n<td>0.666</td>\n<td>0.679</td>\n</tr>\n<tr>\n<td>0.668</td>\n<td>0.681</td>\n</tr>\n<tr>\n<td><strong>0.673</strong></td>\n<td><strong>0.687</strong></td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "No folds. Train folder vs validation folder.\nThreshold was selected by iterating thru` np.arange(0.01, 0.51, 0.01)`.\n\n\n| CV     | LB     |\n|-------|--------|\n| 0.633 | 0.636  |\n| 0.641  | 0.642  |\n| 0.656  | 0.662  |\n| 0.658  | 0.666  |\n| 0.663  | 0.671  |\n| 0.664  | 0.673  |\n| 0.666  | 0.679  |\n| 0.668  | 0.681  |\n| **0.673**  | **0.687**  |",
      "votes": 16
    },
    {
      "id": 2322598,
      "postDate": "2023-06-29T11:16:48.610Z",
      "content": "<p>Simple Train/Validation Split (No folds)</p>\n<p>All scores are from single model.</p>\n<p>CV：<strong>0.676</strong><br>\nLB：<strong>0.689</strong> (Threshold <strong>0.5</strong> not tuned)</p>\n<p><strong>Compilation of my results so far:</strong></p>\n<table>\n<thead>\n<tr>\n<th>CV</th>\n<th>Public LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.645</td>\n<td>0.640</td>\n</tr>\n<tr>\n<td>0.647</td>\n<td>0.661</td>\n</tr>\n<tr>\n<td>0.654</td>\n<td>0.665</td>\n</tr>\n<tr>\n<td>0.659</td>\n<td>0.671</td>\n</tr>\n<tr>\n<td>0.660</td>\n<td>0.666</td>\n</tr>\n<tr>\n<td>0.664</td>\n<td>0.677</td>\n</tr>\n<tr>\n<td>0.665</td>\n<td>0.663</td>\n</tr>\n<tr>\n<td>0.668</td>\n<td>0.684</td>\n</tr>\n<tr>\n<td>0.676</td>\n<td>0.689</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "Simple Train/Validation Split (No folds)\n\nAll scores are from single model.\n\nCV：**0.676**\nLB：**0.689** (Threshold **0.5** not tuned)\n\n**Compilation of my results so far:**\n| CV | Public LB |\n| --- | --- |\n| 0.645 | 0.640 |\n| 0.647 | 0.661  |\n| 0.654 | 0.665 |\n| 0.659 | 0.671  |\n| 0.660 | 0.666 |\n| 0.664 | 0.677 |\n| 0.665 | 0.663 |\n| 0.668 | 0.684 |\n| 0.676 | 0.689 |\n\n",
      "votes": 15,
      "replies": [
        {
          "id": 2324558,
          "postDate": "2023-06-30T17:18:48.900Z",
          "content": "<p>Pretty good correlation, any hints about the split? 😏</p>",
          "rawMarkdown": "Pretty good correlation, any hints about the split? 😏",
          "votes": 1,
          "replies": [
            {
              "id": 2324829,
              "postDate": "2023-06-30T22:50:23.243Z",
              "content": "<p>These results are using the split provided by hosts ;) In my experiments, correlation deteriorates when I shift between architectures or image size. So, I tried to keep them constant until the last stage of the competition. </p>",
              "rawMarkdown": "These results are using the split provided by hosts ;) In my experiments, correlation deteriorates when I shift between architectures or image size. So, I tried to keep them constant until the last stage of the competition. ",
              "votes": 6
            },
            {
              "id": 2328719,
              "postDate": "2023-07-03T20:10:07.667Z",
              "content": "<p>Hello <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a>, <br>\nCorrect me if I'm wrong but a correlation that doesn't stick with diff models and image size is a very fragile one.<br>\nI had similar issues in the past but now with a 4-fold Split the scores are perfectely correlated even with a image size/architecture shift, I suggest you trouble shoot your pipeline so that you'r not going to be unpleasantly surprised at the end of the comp..</p>",
              "rawMarkdown": "Hello @nischaydnk, \nCorrect me if I'm wrong but a correlation that doesn't stick with diff models and image size is a very fragile one.\nI had similar issues in the past but now with a 4-fold Split the scores are perfectely correlated even with a image size/architecture shift, I suggest you trouble shoot your pipeline so that you'r not going to be unpleasantly surprised at the end of the comp..",
              "votes": 3
            },
            {
              "id": 2328986,
              "postDate": "2023-07-04T02:54:29.070Z",
              "content": "<p>Do you use train + validation while doing 4 fold split ?</p>",
              "rawMarkdown": "Do you use train + validation while doing 4 fold split ?"
            },
            {
              "id": 2329831,
              "postDate": "2023-07-04T14:35:22.830Z",
              "content": "<p>Hello <a href=\"https://www.kaggle.com/phoenix9032\" target=\"_blank\">@phoenix9032</a>,<br>\nNo I keep validation as a holdout.</p>",
              "rawMarkdown": "Hello @phoenix9032,\nNo I keep validation as a holdout.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2314964,
      "postDate": "2023-06-23T17:49:27.320Z",
      "content": "<p>This is what I have accumulated so far.<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2126325%2Fa9862998a0b606ccf17b4aaec52ca8bd%2FCV_LB.png?generation=1687542535053652&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "This is what I have accumulated so far.![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2126325%2Fa9862998a0b606ccf17b4aaec52ca8bd%2FCV_LB.png?generation=1687542535053652&alt=media)",
      "votes": 15,
      "replies": [
        {
          "id": 2315071,
          "postDate": "2023-06-23T19:42:00.590Z",
          "content": "<p>You achieved a very good correlation. Having checked on your work, I can infer that you used validation folder inside the cross-validation loop. May I ask the question, how did you adjust the dice metric like this without inflating it? Though, it really looks like you did not use the validation folder during the training.</p>",
          "rawMarkdown": "You achieved a very good correlation. Having checked on your work, I can infer that you used validation folder inside the cross-validation loop. May I ask the question, how did you adjust the dice metric like this without inflating it? Though, it really looks like you did not use the validation folder during the training.",
          "votes": 1,
          "replies": [
            {
              "id": 2315117,
              "postDate": "2023-06-23T20:32:26.023Z",
              "content": "<p>That does not look correlated, above .684, the higher he scores in CV, the lower he scores on LB</p>",
              "rawMarkdown": "That does not look correlated, above .684, the higher he scores in CV, the lower he scores on LB",
              "votes": 1
            },
            {
              "id": 2315142,
              "postDate": "2023-06-23T20:57:09.390Z",
              "content": "<p>Generally, it looks correlated.<br>\nIt would look more correlated if x and y values were set to [0, 1].<br>\nStatistical graphs can lie. </p>",
              "rawMarkdown": "Generally, it looks correlated.\nIt would look more correlated if x and y values were set to [0, 1].\nStatistical graphs can lie. ",
              "votes": 1
            },
            {
              "id": 2315150,
              "postDate": "2023-06-23T21:05:01.060Z",
              "content": "<p>Well it does. There is a variance for sure but \"not correlated\" means some of us needs to revisit what was considered a good correlation until this moment. <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2Fa966293d9733610f507131c38a65e4e9%2Fcorrelated.png?generation=1687554457285576&amp;alt=media\" alt=\"\"></p>",
              "rawMarkdown": "Well it does. There is a variance for sure but \"not correlated\" means some of us needs to revisit what was considered a good correlation until this moment. ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2Fa966293d9733610f507131c38a65e4e9%2Fcorrelated.png?generation=1687554457285576&alt=media)",
              "votes": 2
            },
            {
              "id": 2315182,
              "postDate": "2023-06-23T22:06:32.707Z",
              "content": "<p>On kaggle, we push to squeeze out the last bit of performance out of the models on CV and if they start randomly crashing a large amount (by LB standards), even though they are \"technically\" co-related, this correlation is not of use, LB not correlating CV is not a new thing on kaggle, but, your \"strongly\" correlated has destroyed many potential golds of countless people on kaggle ;)</p>",
              "rawMarkdown": "On kaggle, we push to squeeze out the last bit of performance out of the models on CV and if they start randomly crashing a large amount (by LB standards), even though they are \"technically\" co-related, this correlation is not of use, LB not correlating CV is not a new thing on kaggle, but, your \"strongly\" correlated has destroyed many potential golds of countless people on kaggle ;)",
              "votes": 1
            },
            {
              "id": 2315243,
              "postDate": "2023-06-24T00:16:41.673Z",
              "content": "<p>How did you calculate “many” out of countless? LoL</p>",
              "rawMarkdown": "How did you calculate “many” out of countless? LoL",
              "votes": 5
            },
            {
              "id": 2315245,
              "postDate": "2023-06-24T00:23:46.037Z",
              "content": "<p>LoL, Sorry, english is not my native language, I meant to say countless people have lost many of their own golds because of not good enough co-relation, including me 5 months ago</p>",
              "rawMarkdown": "LoL, Sorry, english is not my native language, I meant to say countless people have lost many of their own golds because of not good enough co-relation, including me 5 months ago",
              "votes": 5
            },
            {
              "id": 2315819,
              "postDate": "2023-06-24T12:53:55.133Z",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/sergiosaharovskiy\" target=\"_blank\">@sergiosaharovskiy</a> <br>\nI am not doing anything special with dice metrics at the moment.<br>\nYes, I train with a combined train/val dataset. But I think it is actually a good idea to have the val dataset separately as a holdout, e.g. to do experiments with threshold optimization.</p>",
              "rawMarkdown": "Hi @sergiosaharovskiy \nI am not doing anything special with dice metrics at the moment.\nYes, I train with a combined train/val dataset. But I think it is actually a good idea to have the val dataset separately as a holdout, e.g. to do experiments with threshold optimization.",
              "votes": 4
            },
            {
              "id": 2320929,
              "postDate": "2023-06-28T06:59:42.050Z",
              "content": "<p>Hi Egor, May I ask if you're using dice loss for training &amp; dice coefficient for score? My local CV &amp; LB score have huge gap and the correlation is not stable..</p>\n<p>Thank you.</p>",
              "rawMarkdown": "Hi Egor, May I ask if you're using dice loss for training & dice coefficient for score? My local CV & LB score have huge gap and the correlation is not stable..\n\nThank you."
            }
          ]
        },
        {
          "id": 2360109,
          "postDate": "2023-07-26T14:58:58.687Z",
          "content": "<p>Hi, Egor, May I know your currently CV score</p>",
          "rawMarkdown": "Hi, Egor, May I know your currently CV score"
        }
      ]
    },
    {
      "id": 2294891,
      "postDate": "2023-06-10T12:29:46.980Z",
      "content": "<p>CV5 single model:</p>\n<ul>\n<li>Global Dice: 0.6785</li>\n<li>Dice per Image: 0.7757</li>\n<li>AUCPR: 0.7555</li>\n</ul>\n<p>LB: 0.680</p>",
      "rawMarkdown": "CV5 single model:\n- Global Dice: 0.6785\n- Dice per Image: 0.7757\n- AUCPR: 0.7555\n\nLB: 0.680",
      "votes": 10,
      "replies": [
        {
          "id": 2295073,
          "postDate": "2023-06-10T15:08:44.217Z",
          "content": "<p>Great Global Dice result  …is it for a threshold? </p>",
          "rawMarkdown": "Great Global Dice result  ...is it for a threshold? ",
          "votes": 2,
          "replies": [
            {
              "id": 2295169,
              "postDate": "2023-06-10T16:39:27.207Z",
              "content": "<p>Threshold on inference is around 0.40</p>",
              "rawMarkdown": "Threshold on inference is around 0.40",
              "votes": 2
            }
          ]
        },
        {
          "id": 2306511,
          "postDate": "2023-06-17T11:25:16.560Z",
          "content": "<p><strong>Updates</strong>: </p>\n<p>Same model but improved training procedure. No post-processing.</p>\n<ul>\n<li>Global Dice: 0.6819</li>\n<li>Dice per Image: 0.7763</li>\n<li>AUCPR: 0.7526</li>\n</ul>\n<p>LB: 0.691 (Best threshold around 0.4)</p>",
          "rawMarkdown": "**Updates**: \n\nSame model but improved training procedure. No post-processing.\n\n- Global Dice: 0.6819\n- Dice per Image: 0.7763\n- AUCPR: 0.7526\n\nLB: 0.691 (Best threshold around 0.4)",
          "votes": 4,
          "replies": [
            {
              "id": 2306538,
              "postDate": "2023-06-17T11:47:35.670Z",
              "content": "<p>Impressive score. Do you use validation to calulate cv or KFold?.I already get 0.68lb when my cv is 0.66….but now my cv get ~0.695, my lb still is 0.68+. My lb has been standing still for the past few weeks. This even makes me wonder if I'm overfitting the local validation dataset😂</p>",
              "rawMarkdown": "Impressive score. Do you use validation to calulate cv or KFold?.I already get 0.68lb when my cv is 0.66....but now my cv get ~0.695, my lb still is 0.68+. My lb has been standing still for the past few weeks. This even makes me wonder if I'm overfitting the local validation dataset😂",
              "votes": 3
            },
            {
              "id": 2306601,
              "postDate": "2023-06-17T12:32:53.570Z",
              "content": "<p>KFold, I'm not using validation/ folder for now. Maybe you've an issue with your global dice metric which is not the same one as Kaggle.</p>\n<p>One example for one of my fold (finetuning): Global dice (dice_coef), dice, AUCPR and LR.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2F37d09273dd79429fa7a0c22d724cbf87%2Fstage2.png?generation=1687007010478502&amp;alt=media\" alt=\"\"></p>",
              "rawMarkdown": "KFold, I'm not using validation/ folder for now. Maybe you've an issue with your global dice metric which is not the same one as Kaggle.\n\nOne example for one of my fold (finetuning): Global dice (dice_coef), dice, AUCPR and LR.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2F37d09273dd79429fa7a0c22d724cbf87%2Fstage2.png?generation=1687007010478502&alt=media)\n",
              "votes": 8
            },
            {
              "id": 2306756,
              "postDate": "2023-06-17T14:47:06.397Z",
              "content": "<p>thanks for your reply! maybe I need check my code to confirm</p>",
              "rawMarkdown": "thanks for your reply! maybe I need check my code to confirm",
              "votes": 1
            },
            {
              "id": 2306780,
              "postDate": "2023-06-17T15:10:44.607Z",
              "content": "<p><a href=\"https://www.kaggle.com/zhuwanglju\" target=\"_blank\">@zhuwanglju</a> I'm having the same issue as you have, and this with a SKF CV and with the Train/valid original split… I feel like using the train/valid is a massive bait. </p>",
              "rawMarkdown": "@zhuwanglju I'm having the same issue as you have, and this with a SKF CV and with the Train/valid original split... I feel like using the train/valid is a massive bait. ",
              "votes": 1
            },
            {
              "id": 2307233,
              "postDate": "2023-06-18T03:05:36.323Z",
              "content": "<blockquote>\n  <p>KFold, I'm not using validation/ folder for now. Maybe you've an issue with your global dice metric which is not the same one as Kaggle.</p>\n  <p>One example for one of my fold (finetuning): Global dice (dice_coef), dice, AUCPR and LR.</p>\n  <p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2F37d09273dd79429fa7a0c22d724cbf87%2Fstage2.png?generation=1687007010478502&amp;alt=media\" alt=\"\"></p>\n</blockquote>\n<p>Thank you for posting the initial results.<br>\nAchieving a score of 0.68 after just the second epoch is a promising start. Though, waiting for another 70 epochs to get +0.019, god you are tough man! It is so Kaggle :)</p>\n<p>P.s. I cannot really see, but it looks the warmup is 3-4 and the first two epochs score is around 0 or the enumeration starts from 1?</p>",
              "rawMarkdown": "> KFold, I'm not using validation/ folder for now. Maybe you've an issue with your global dice metric which is not the same one as Kaggle.\n> \n> One example for one of my fold (finetuning): Global dice (dice_coef), dice, AUCPR and LR.\n> \n> ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2F37d09273dd79429fa7a0c22d724cbf87%2Fstage2.png?generation=1687007010478502&alt=media)\n\nThank you for posting the initial results.\nAchieving a score of 0.68 after just the second epoch is a promising start. Though, waiting for another 70 epochs to get +0.019, god you are tough man! It is so Kaggle :)\n\n P.s. I cannot really see, but it looks the warmup is 3-4 and the first two epochs score is around 0 or the enumeration starts from 1?",
              "votes": 3
            },
            {
              "id": 2307452,
              "postDate": "2023-06-18T07:48:54.683Z",
              "content": "<p>It's the finetuning part as I mentioned that's why it starts very high. The initial training part (not posted here) starts very low. And yes I've LR warmup and display starts from epoch 1.</p>",
              "rawMarkdown": "It's the finetuning part as I mentioned that's why it starts very high. The initial training part (not posted here) starts very low. And yes I've LR warmup and display starts from epoch 1.",
              "votes": 6
            },
            {
              "id": 2307613,
              "postDate": "2023-06-18T10:19:37.890Z",
              "content": "<p>now makes sense - thanks for sharing your results. Just to make sure we are comparing apples with apples the scores you report above are the average across all folds, right?<br>\nEDIT: or from a single fold?</p>",
              "rawMarkdown": "now makes sense - thanks for sharing your results. Just to make sure we are comparing apples with apples the scores you report above are the average across all folds, right?\nEDIT: or from a single fold?",
              "votes": 1
            },
            {
              "id": 2307634,
              "postDate": "2023-06-18T10:47:03.917Z",
              "content": "<p><a href=\"https://www.kaggle.com/imeintanis\" target=\"_blank\">@imeintanis</a> <br>\nThat should answer your question:</p>\n<blockquote>\n  <p>One example for <strong>one of my fold</strong> (finetuning): Global dice (dice_coef), dice, AUCPR and LR.</p>\n</blockquote>",
              "rawMarkdown": "@imeintanis \nThat should answer your question:\n>One example for **one of my fold** (finetuning): Global dice (dice_coef), dice, AUCPR and LR.",
              "votes": 2
            },
            {
              "id": 2311987,
              "postDate": "2023-06-21T15:27:51.147Z",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a>,<br>\nIs your lb score the ensemble of the 5 folds or is it just a single fold submission?<br>\nThanks :)</p>",
              "rawMarkdown": "Hi @mpware,\nIs your lb score the ensemble of the 5 folds or is it just a single fold submission?\nThanks :)",
              "votes": 1
            },
            {
              "id": 2312257,
              "postDate": "2023-06-21T19:21:46.390Z",
              "content": "<p><a href=\"https://www.kaggle.com/shashwatraman\" target=\"_blank\">@shashwatraman</a> Ensemble of 5 folds. We've have lot of room/time on inference so I've not even tried any single fold for now. But I will do and post my results here.</p>\n<p>And here is what looks like the search for best threshold:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2F5d894e8837c5ccb49551ae48a67bc84f%2Fthreshold.png?generation=1687386508197030&amp;alt=media\" alt=\"\"></p>",
              "rawMarkdown": "@shashwatraman Ensemble of 5 folds. We've have lot of room/time on inference so I've not even tried any single fold for now. But I will do and post my results here.\n\nAnd here is what looks like the search for best threshold:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2F5d894e8837c5ccb49551ae48a67bc84f%2Fthreshold.png?generation=1687386508197030&alt=media)\n",
              "votes": 3
            },
            {
              "id": 2312581,
              "postDate": "2023-06-22T03:12:32.417Z",
              "content": "<p>Thank you very much for your reply <a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a>, it's really helpful. I'm still struggling to find a good cross validation strategy. Cv and Lb are not correlated for me rn.</p>",
              "rawMarkdown": "Thank you very much for your reply @mpware, it's really helpful. I'm still struggling to find a good cross validation strategy. Cv and Lb are not correlated for me rn.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2314341,
      "postDate": "2023-06-23T11:07:38.580Z",
      "content": "<p>training by 5-fold cross validaition:</p>\n<table>\n<thead>\n<tr>\n<th>train folder(oof pred)</th>\n<th>validation folder(5-fold avg)</th>\n<th>LB(5-fold avg)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.658</td>\n<td>0.635</td>\n<td>0.650</td>\n</tr>\n<tr>\n<td>0.659</td>\n<td>0.642</td>\n<td>0.657</td>\n</tr>\n</tbody>\n</table>\n<p>　<br>\nIt seems that \"validation\" and LB have correlation🤔</p>",
      "rawMarkdown": "training by 5-fold cross validaition:\n\n| train folder(oof pred) | validation folder(5-fold avg) | LB(5-fold avg) |\n|:------:|:------:|:------:|\n| 0.658 | 0.635  | 0.650 |\n| 0.659 | 0.642  | 0.657  |\n\n　\nIt seems that \"validation\" and LB have correlation🤔",
      "votes": 7,
      "replies": [
        {
          "id": 2314665,
          "postDate": "2023-06-23T14:21:28.510Z",
          "content": "<p>same here , only improvements in validation-folder score boosts LB.</p>",
          "rawMarkdown": "same here , only improvements in validation-folder score boosts LB.",
          "votes": 2,
          "replies": [
            {
              "id": 2314751,
              "postDate": "2023-06-23T15:25:06.370Z",
              "content": "<p>I have validation-folder scoring .655 and .668 global dice and both scores .678 on LB :(</p>",
              "rawMarkdown": "I have validation-folder scoring .655 and .668 global dice and both scores .678 on LB :(",
              "votes": 4
            },
            {
              "id": 2314788,
              "postDate": "2023-06-23T15:44:26.313Z",
              "content": "<p>In my experiments (not many) Lb and validation-folder perfectly correlates , maybe it is a threshold issue (i always use cv-threshold, usually in range(0.4,0.5) )</p>",
              "rawMarkdown": "In my experiments (not many) Lb and validation-folder perfectly correlates , maybe it is a threshold issue (i always use cv-threshold, usually in range(0.4,0.5) )",
              "votes": 2
            },
            {
              "id": 2314859,
              "postDate": "2023-06-23T16:33:48.397Z",
              "content": "<p>Using threshold=0.4 I have .678 LB but I was using threshold=0.5 when doing validation-folder cv, when I do threshold=0.4 for cv, I get .670, so that's another increase on validation-folder</p>",
              "rawMarkdown": "Using threshold=0.4 I have .678 LB but I was using threshold=0.5 when doing validation-folder cv, when I do threshold=0.4 for cv, I get .670, so that's another increase on validation-folder"
            },
            {
              "id": 2314879,
              "postDate": "2023-06-23T16:48:50.033Z",
              "content": "<p>Yes, the same thing happens for me. Some models which have a better global dice score on 5 fold cv and validation folder have a lower score on the lb. <br>\nMaybe this happens because the Public Test set is very small?</p>",
              "rawMarkdown": "Yes, the same thing happens for me. Some models which have a better global dice score on 5 fold cv and validation folder have a lower score on the lb. \nMaybe this happens because the Public Test set is very small?",
              "votes": 1
            },
            {
              "id": 2315122,
              "postDate": "2023-06-23T20:36:40.580Z",
              "content": "<p>same here, but history told us it was a bad idea to trust LB blindly..</p>",
              "rawMarkdown": "same here, but history told us it was a bad idea to trust LB blindly..",
              "votes": 2
            }
          ]
        },
        {
          "id": 2320934,
          "postDate": "2023-06-28T07:04:50.133Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a> When you say \"train folder(oof pred)\", are you splitting the train folder into K folds to perform a cross-validation? Then use the trained k models to test on the validation folder? </p>\n<p>For my cross-validation, I use the validation folder for each of K fold while training, which is different.<br>\nThank you.</p>",
          "rawMarkdown": "Hi @ttahara When you say \"train folder(oof pred)\", are you splitting the train folder into K folds to perform a cross-validation? Then use the trained k models to test on the validation folder? \n\nFor my cross-validation, I use the validation folder for each of K fold while training, which is different.\nThank you.",
          "votes": 1,
          "replies": [
            {
              "id": 2321450,
              "postDate": "2023-06-28T15:28:21.167Z",
              "content": "<p>Yes, you are right.<br>\nI trained k models by K-folds cross-validation, then use these k models to test on the validation folder.</p>",
              "rawMarkdown": "Yes, you are right.\nI trained k models by K-folds cross-validation, then use these k models to test on the validation folder."
            },
            {
              "id": 2327472,
              "postDate": "2023-07-03T01:40:16.917Z",
              "content": "<p>Thank you for your information. I take the same way as you did and it helps. </p>",
              "rawMarkdown": "Thank you for your information. I take the same way as you did and it helps. "
            }
          ]
        }
      ]
    },
    {
      "id": 2277264,
      "postDate": "2023-05-27T15:21:55.800Z",
      "content": "<p>I  just submited twice. cv:587,lb:589. cv:607 lb:604. it seems cv and lb have good trend👀</p>",
      "rawMarkdown": "I  just submited twice. cv:587,lb:589. cv:607 lb:604. it seems cv and lb have good trend👀",
      "votes": 5,
      "replies": [
        {
          "id": 2277334,
          "postDate": "2023-05-27T16:12:32.010Z",
          "content": "<p>Nice! In your case the LB and CV seem even closer. Do you also use the training folder for training and validation for CV?</p>",
          "rawMarkdown": "Nice! In your case the LB and CV seem even closer. Do you also use the training folder for training and validation for CV?",
          "votes": 1,
          "replies": [
            {
              "id": 2277808,
              "postDate": "2023-05-28T04:44:44.537Z",
              "content": "<p>Yes, I rewrote the function of dice. If  dice_score is similar to the following code(from <a href=\"https://www.kaggle.com/code/lupin11/u-net-baseline-training\" target=\"_blank\">this</a>), which will make the local CV higher because many images do not have masks, and their dice score is 1. But it doesn't seem to matter, its score and lb trend are also very consistent.😀</p>\n<pre><code>def dice_score(y_p, y_t, smooth=1e-6):\n    i = torch.sum(y_p * y_t, dim=(2, 3))\n    u = torch.sum(y_p, dim=(2, 3)) + torch.sum(y_t, dim=(2, 3))\n    score = (2 * i + smooth)/(u + smooth)\n    return torch.mean(score)\n</code></pre>",
              "rawMarkdown": "Yes, I rewrote the function of dice. If  dice_score is similar to the following code(from [this](https://www.kaggle.com/code/lupin11/u-net-baseline-training)), which will make the local CV higher because many images do not have masks, and their dice score is 1. But it doesn't seem to matter, its score and lb trend are also very consistent.😀\n\n```\ndef dice_score(y_p, y_t, smooth=1e-6):\n    i = torch.sum(y_p * y_t, dim=(2, 3))\n    u = torch.sum(y_p, dim=(2, 3)) + torch.sum(y_t, dim=(2, 3))\n    score = (2 * i + smooth)/(u + smooth)\n    return torch.mean(score)\n```",
              "votes": 10
            }
          ]
        }
      ]
    },
    {
      "id": 2309146,
      "postDate": "2023-06-19T12:34:08.297Z",
      "content": "<p>5 folds single model : </p>\n<ul>\n<li>global dice : 0.6418 (mean of cv scores)</li>\n<li>LB : 0.655 </li>\n</ul>",
      "rawMarkdown": "5 folds single model : \n- global dice : 0.6418 (mean of cv scores)\n- LB : 0.655 ",
      "votes": 3
    },
    {
      "id": 2289988,
      "postDate": "2023-06-06T13:34:20.390Z",
      "content": "<p>CV - 0.637 LB - 0.662</p>",
      "rawMarkdown": "CV - 0.637 LB - 0.662",
      "votes": 3
    },
    {
      "id": 2289893,
      "postDate": "2023-06-06T12:14:28.450Z",
      "content": "<p>Current best:</p>\n<table>\n<thead>\n<tr>\n<th>CV</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.606</td>\n<td>0.646</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "Current best:\n| CV | LB |\n| --- | --- |\n|0.606  | 0.646 |\n",
      "votes": 3
    },
    {
      "id": 2312467,
      "postDate": "2023-06-21T23:33:40.800Z",
      "content": "<p>5 Fold training using \"train\" folder only</p>\n<p>Fold-wise average CV Global Dice: .69</p>\n<p>5-Fold ensemble on \"validation\" folder:<br>\nGlobal Dice: .668<br>\nImage-wise Dice: .830</p>\n<p>LB: .678</p>\n<p>I have many models in CV .68 to .69 and they score from .66 to .678 on LB, I am unable to find correlation as .687 CV scores .66 LB and .685 also scores .678 the same as .690 CV</p>",
      "rawMarkdown": "5 Fold training using \"train\" folder only\n\nFold-wise average CV Global Dice: .69\n\n5-Fold ensemble on \"validation\" folder:\nGlobal Dice: .668\nImage-wise Dice: .830\n\nLB: .678\n\nI have many models in CV .68 to .69 and they score from .66 to .678 on LB, I am unable to find correlation as .687 CV scores .66 LB and .685 also scores .678 the same as .690 CV",
      "votes": 4,
      "replies": [
        {
          "id": 2314029,
          "postDate": "2023-06-23T05:48:14.957Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 2316761,
          "postDate": "2023-06-25T08:11:17.810Z",
          "content": "<p>Update:</p>\n<p>5 Fold average CV Global Dice: .6916<br>\nLB: .657</p>\n<p>Each one of these 5 folds scores over .69 on their respective CV, very stable</p>",
          "rawMarkdown": "Update:\n\n5 Fold average CV Global Dice: .6916\nLB: .657\n\nEach one of these 5 folds scores over .69 on their respective CV, very stable",
          "replies": [
            {
              "id": 2316965,
              "postDate": "2023-06-25T10:55:02.710Z",
              "content": "<p>we have similiar problem… my cv is 0.70+, and lb is 0.67+….Now, I choose to believe my cv….</p>",
              "rawMarkdown": "we have similiar problem... my cv is 0.70+, and lb is 0.67+....Now, I choose to believe my cv....",
              "votes": 2
            },
            {
              "id": 2338132,
              "postDate": "2023-07-10T14:53:44.263Z",
              "content": "<p>I have just started with a few experiments and my correlation is terrible</p>\n<table>\n<thead>\n<tr>\n<th>exp</th>\n<th>type</th>\n<th>CV</th>\n<th>Validation</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>1 fold</td>\n<td>0.6824</td>\n<td>0.650</td>\n<td>0.662</td>\n</tr>\n<tr>\n<td>2</td>\n<td>5 fold</td>\n<td>0.6775</td>\n<td>0.6575</td>\n<td>0.656</td>\n</tr>\n<tr>\n<td>3</td>\n<td>5 fold</td>\n<td>0.6822</td>\n<td>0.6687</td>\n<td>0.656</td>\n</tr>\n<tr>\n<td>4</td>\n<td>1 fold</td>\n<td>Nan</td>\n<td>0.6701</td>\n<td>0.641</td>\n</tr>\n<tr>\n<td>5</td>\n<td>5 fold</td>\n<td>0.6867</td>\n<td>0.6724</td>\n<td>0.659</td>\n</tr>\n</tbody>\n</table>",
              "rawMarkdown": "I have just started with a few experiments and my correlation is terrible\n\n| exp | type   | CV     | Validation | LB    |\n|-----|--------|--------|------------|-------|\n| 1   | 1 fold | 0.6824 | 0.650      | 0.662 |\n| 2   | 5 fold | 0.6775 | 0.6575     | 0.656 |\n| 3   | 5 fold | 0.6822 | 0.6687     | 0.656 |\n| 4   | 1 fold | Nan    | 0.6701     | 0.641 |\n| 5   | 5 fold | 0.6867 | 0.6724     | 0.659 |",
              "votes": 2
            },
            {
              "id": 2338372,
              "postDate": "2023-07-10T17:57:58.687Z",
              "content": "<p>Are you using some threshold?</p>",
              "rawMarkdown": "Are you using some threshold?"
            },
            {
              "id": 2339508,
              "postDate": "2023-07-10T20:20:53.587Z",
              "content": "<p>I'm giving here best thresholded scores for validation.</p>\n<p>From my observations, single models predict hard distributions of 1s and 0s mostly so in validation score is the same across a wide range of thresholds, for 5 fold CV averaging on validation : best threshold is around 0.3.</p>\n<p>For the same single fold, I've seen 0.641LB for threshold 0.95, and 0.626LB for threshold 0.80. So I don't observe the same behavior between validation and LB. Anyone can relate ?</p>",
              "rawMarkdown": "I'm giving here best thresholded scores for validation.\n\nFrom my observations, single models predict hard distributions of 1s and 0s mostly so in validation score is the same across a wide range of thresholds, for 5 fold CV averaging on validation : best threshold is around 0.3.\n\nFor the same single fold, I've seen 0.641LB for threshold 0.95, and 0.626LB for threshold 0.80. So I don't observe the same behavior between validation and LB. Anyone can relate ?"
            }
          ]
        },
        {
          "id": 2320939,
          "postDate": "2023-06-28T07:08:47.427Z",
          "content": "<p>Hello, <a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a> may I ask your image size? Your local CV score is crazy…</p>",
          "rawMarkdown": "Hello, @harshitsheoran may I ask your image size? Your local CV score is crazy..."
        }
      ]
    },
    {
      "id": 2289868,
      "postDate": "2023-06-06T11:48:54.170Z",
      "content": "<p>CV on validation data - 0.634 LB - 644</p>",
      "rawMarkdown": "CV on validation data - 0.634 LB - 644",
      "votes": 4,
      "replies": [
        {
          "id": 2311491,
          "postDate": "2023-06-21T07:31:41.393Z",
          "content": "<p>CV on 4 kfolds - 0.677 LB - 0.668</p>",
          "rawMarkdown": "CV on 4 kfolds - 0.677 LB - 0.668",
          "votes": 2,
          "replies": [
            {
              "id": 2324294,
              "postDate": "2023-06-30T14:08:56.907Z",
              "content": "<p>CV - 688<br>\nLB - 693</p>",
              "rawMarkdown": "CV - 688\nLB - 693",
              "votes": 1
            },
            {
              "id": 2382592,
              "postDate": "2023-08-09T21:02:31.100Z",
              "content": "<p>This one is nice is it an ensemble ?</p>",
              "rawMarkdown": "This one is nice is it an ensemble ?"
            }
          ]
        }
      ]
    },
    {
      "id": 2339675,
      "postDate": "2023-07-11T02:24:37.170Z",
      "content": "<table>\n<thead>\n<tr>\n<th>CV (Single Model Threshold 0.5)</th>\n<th>LB (Threshold 0.5 not tuned)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.680</td>\n<td>0.646</td>\n</tr>\n<tr>\n<td>0.673 (<strong>-0.007</strong>)</td>\n<td>0.638 (<strong>-0.008</strong>)</td>\n</tr>\n</tbody>\n</table>\n<p>Will compare the cv correlation with few more submissions</p>",
      "rawMarkdown": "| CV (Single Model Threshold 0.5) | LB (Threshold 0.5 not tuned) |\n| --- | --- |\n| 0.680 | 0.646 |\n| 0.673 (**-0.007**) | 0.638 (**-0.008**) |\n\nWill compare the cv correlation with few more submissions",
      "votes": 1
    },
    {
      "id": 2317041,
      "postDate": "2023-06-25T12:05:44.657Z",
      "content": "<p>How many epochs do you typically train your model for each fold? I would like to perform cross-validation in order to obtain a more accurate estimation of my model's performance. However, running 5 * 20 epochs would require several days. I am training the model on my own decent PC (Nvidia 3090ti, 24GB VRAM), but performing cross-validation for all my models is not feasible due to the significant time it would take. Do you utilize any external resources? Kaggle notebooks tend to be unstable and often shut down after a few hours. Google Colab, even with Collab Pro, is quite slow when it comes to loading data, and the credits may run out before completing the cross-validation process.</p>",
      "rawMarkdown": "How many epochs do you typically train your model for each fold? I would like to perform cross-validation in order to obtain a more accurate estimation of my model's performance. However, running 5 * 20 epochs would require several days. I am training the model on my own decent PC (Nvidia 3090ti, 24GB VRAM), but performing cross-validation for all my models is not feasible due to the significant time it would take. Do you utilize any external resources? Kaggle notebooks tend to be unstable and often shut down after a few hours. Google Colab, even with Collab Pro, is quite slow when it comes to loading data, and the credits may run out before completing the cross-validation process.",
      "votes": 1,
      "replies": [
        {
          "id": 2317977,
          "postDate": "2023-06-26T05:39:57.300Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/louisweiss\" target=\"_blank\">@louisweiss</a>,<br>\nI am training my models for about 30-40 epochs. More epochs help here.<br>\nNow for cross validation, it is generally better to use a 5 fold cv, however you can go for 3 or 4 folds. Just don't go below that.<br>\nYes I am using external recources. Kaggle notebooks are great. They are very stable if you Save and Run the notebook. Just the thing is the 30 hour limit is too less. You can try JarvisLabs, PaperSpace, etc and see what suits you.<br>\nTraining a medium sized model with 5 Folds and 30 epochs each takes around 3 hours for me using a single A5000 GPU.</p>",
          "rawMarkdown": "Hi @louisweiss,\nI am training my models for about 30-40 epochs. More epochs help here.\nNow for cross validation, it is generally better to use a 5 fold cv, however you can go for 3 or 4 folds. Just don't go below that.\nYes I am using external recources. Kaggle notebooks are great. They are very stable if you Save and Run the notebook. Just the thing is the 30 hour limit is too less. You can try JarvisLabs, PaperSpace, etc and see what suits you.\nTraining a medium sized model with 5 Folds and 30 epochs each takes around 3 hours for me using a single A5000 GPU.",
          "votes": 2,
          "replies": [
            {
              "id": 2318465,
              "postDate": "2023-06-26T11:38:09.860Z",
              "content": "<p><a href=\"https://www.kaggle.com/shashwatraman\" target=\"_blank\">@shashwatraman</a> what is your python and pytorch version? It takes 15 min per epoch for me 512x512 size. I wonder if you use resnet50 like architecture.</p>",
              "rawMarkdown": "@shashwatraman what is your python and pytorch version? It takes 15 min per epoch for me 512x512 size. I wonder if you use resnet50 like architecture.",
              "votes": 1
            },
            {
              "id": 2318472,
              "postDate": "2023-06-26T11:49:38.220Z",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/sergiosaharovskiy\" target=\"_blank\">@sergiosaharovskiy</a>,<br>\nPython Version - 3.10.11<br>\nPytorch Version - 2.0.1<br>\nWhen using 512 sized images, it takes me around 5 mins per epoch to train. I use an EfficientNetB3!</p>",
              "rawMarkdown": "Hi @sergiosaharovskiy,\nPython Version - 3.10.11\nPytorch Version - 2.0.1\nWhen using 512 sized images, it takes me around 5 mins per epoch to train. I use an EfficientNetB3!",
              "votes": 1
            },
            {
              "id": 2318477,
              "postDate": "2023-06-26T11:53:04.903Z",
              "content": "<p>Takes me around 3 minutes per epoch for efficientnetv2 L at a 512*512 size for 5 fold training with 2x4090. I use FP16 with torch.compile </p>",
              "rawMarkdown": "Takes me around 3 minutes per epoch for efficientnetv2 L at a 512*512 size for 5 fold training with 2x4090. I use FP16 with torch.compile ",
              "votes": 2
            },
            {
              "id": 2318485,
              "postDate": "2023-06-26T12:01:48.743Z",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/mithilsalunkhe\" target=\"_blank\">@mithilsalunkhe</a>,<br>\nThat is really amazing. Currently I am using a single A5000 Gpu :)</p>",
              "rawMarkdown": "Hi @mithilsalunkhe,\nThat is really amazing. Currently I am using a single A5000 Gpu :)"
            },
            {
              "id": 2319268,
              "postDate": "2023-06-27T03:01:45.887Z",
              "content": "<p>torch.compile is such a pain. Kaggle kernel throws an error for P100 since it is not supported. T4 x 2 throw the error with CUDA. Once you switch to x1 it says Cannot find a working triton installation. Either it sucks being still in Beta Version or me sucks, no alternatives. </p>",
              "rawMarkdown": "torch.compile is such a pain. Kaggle kernel throws an error for P100 since it is not supported. T4 x 2 throw the error with CUDA. Once you switch to x1 it says Cannot find a working triton installation. Either it sucks being still in Beta Version or me sucks, no alternatives. ",
              "votes": 1
            },
            {
              "id": 2319333,
              "postDate": "2023-06-27T04:31:24.257Z",
              "content": "<p>I use the NGC container  given by nvidia  , only had a few issues with it. I still have to try using it in inference as the speedup would be minimal  </p>",
              "rawMarkdown": "I use the NGC container  given by nvidia  , only had a few issues with it. I still have to try using it in inference as the speedup would be minimal  ",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2292643,
      "postDate": "2023-06-08T14:24:15.080Z",
      "content": "<p>CV: 0.4931 LB: 0.636<br>\nI'm using the mDice metric (so not averaging over the images but over all pixels in val set) and validating on provided val dataset, only on positive images. If keeping track of negative images, my mDice would fall at ~0.20 in CV. Don't see particular strong correlation with my CV/LB. Any idea? </p>",
      "rawMarkdown": "CV: 0.4931 LB: 0.636\nI'm using the mDice metric (so not averaging over the images but over all pixels in val set) and validating on provided val dataset, only on positive images. If keeping track of negative images, my mDice would fall at ~0.20 in CV. Don't see particular strong correlation with my CV/LB. Any idea? ",
      "votes": 1,
      "replies": [
        {
          "id": 2293054,
          "postDate": "2023-06-08T23:43:28.810Z",
          "content": "<p>As you said, you don't use the \"right\" metric, maybe changing that would be the firsy thing to try but then my guess is that you did already.<br>\nif that doesn't work, I suggest you take a look at the Dice metric in pytroch metrics<br>\nPytorch : <a href=\"https://torchmetrics.readthedocs.io/en/v0.8.2/classification/dice_score.html\" target=\"_blank\">https://torchmetrics.readthedocs.io/en/v0.8.2/classification/dice_score.html</a> (the one I use, I certify it works well)<br>\nyou meed need a few unsqueeze or squeeze to match the input of this function, but the fact that you can choose the threshold as a parameter is very nice for 0.01 optimisations. Good luck</p>",
          "rawMarkdown": "As you said, you don't use the \"right\" metric, maybe changing that would be the firsy thing to try but then my guess is that you did already.\nif that doesn't work, I suggest you take a look at the Dice metric in pytroch metrics\nPytorch : https://torchmetrics.readthedocs.io/en/v0.8.2/classification/dice_score.html (the one I use, I certify it works well)\nyou meed need a few unsqueeze or squeeze to match the input of this function, but the fact that you can choose the threshold as a parameter is very nice for 0.01 optimisations. Good luck",
          "votes": 2,
          "replies": [
            {
              "id": 2293368,
              "postDate": "2023-06-09T06:07:43.767Z",
              "content": "<p>Thank you for your answer! </p>\n<p>I decided to use the mDice because they said they'd use it for evaluation, and I was thinking than I have this discrepancy between CV/LB because the LB has only positive images with many contrails. In particular, I noticed that the models that in my local with higher precision (and even lower recall) works better in this LB. However, having 0.2 vs 0.63 is a big leap. 😅 </p>\n<p>I don't know if I'm being wrong.</p>",
              "rawMarkdown": "Thank you for your answer! \n\nI decided to use the mDice because they said they'd use it for evaluation, and I was thinking than I have this discrepancy between CV/LB because the LB has only positive images with many contrails. In particular, I noticed that the models that in my local with higher precision (and even lower recall) works better in this LB. However, having 0.2 vs 0.63 is a big leap. 😅 \n\nI don't know if I'm being wrong.",
              "votes": 2
            },
            {
              "id": 2295522,
              "postDate": "2023-06-11T04:09:25.900Z",
              "content": "<blockquote>\n  <p>LB has only positive images with many contrails</p>\n</blockquote>\n<p>This is a dangerous assumption for your final LB score, <br>\ntake a look here:<br>\n<a href=\"https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/discussion/415764#2291944\" target=\"_blank\">https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/discussion/415764#2291944</a></p>\n<p>what you'r saying might be correct for the public LB, but don't trust this at all, public LB is very small, get CV up not this.</p>",
              "rawMarkdown": ">LB has only positive images with many contrails\n\nThis is a dangerous assumption for your final LB score, \ntake a look here:\nhttps://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/discussion/415764#2291944\n\nwhat you'r saying might be correct for the public LB, but don't trust this at all, public LB is very small, get CV up not this.",
              "votes": 2
            },
            {
              "id": 2296451,
              "postDate": "2023-06-11T20:56:15.587Z",
              "content": "<p>Sorry, I explained myself wrong. I was saying that I'm possibly having this discrepancy because this particular LB has many positive example. My assumption is that I'll have to trust my CV despite the LB scores. :)</p>",
              "rawMarkdown": "Sorry, I explained myself wrong. I was saying that I'm possibly having this discrepancy because this particular LB has many positive example. My assumption is that I'll have to trust my CV despite the LB scores. :)",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2277210,
      "postDate": "2023-05-27T14:38:39.317Z",
      "content": "<p>Very strong CV results! Which data do you use for CV? Do you use the \"train\" folder for train / validation and the \"validation\" data for the test?</p>",
      "rawMarkdown": "Very strong CV results! Which data do you use for CV? Do you use the \"train\" folder for train / validation and the \"validation\" data for the test?",
      "votes": 1,
      "replies": [
        {
          "id": 2277222,
          "postDate": "2023-05-27T14:49:16.303Z",
          "content": "<p>Model is trained on train folder and validated on validation folder.  </p>",
          "rawMarkdown": "Model is trained on train folder and validated on validation folder.  ",
          "votes": 2,
          "replies": [
            {
              "id": 2277239,
              "postDate": "2023-05-27T14:59:14Z",
              "content": "<p>Alright, so thr CV refers to the validation dice coef during / at the end of training?</p>",
              "rawMarkdown": "Alright, so thr CV refers to the validation dice coef during / at the end of training?",
              "votes": 1
            },
            {
              "id": 2277245,
              "postDate": "2023-05-27T15:04:58.040Z",
              "content": "<p>yes, CV is calculated on validation data at the end of each epoch.</p>",
              "rawMarkdown": "yes, CV is calculated on validation data at the end of each epoch.",
              "votes": 3
            }
          ]
        }
      ]
    },
    {
      "id": 2317391,
      "postDate": "2023-06-25T17:14:29.897Z",
      "content": "<p>How are you splitting train data into 5 folds? </p>\n<p>I was reading the preprint and I noticed these two paragraphs:</p>\n<blockquote>\n  <p>To further boot the number of positives in the dataset, we also included some GOES-16 ABI imagery at locations in the US where Google Street View images of the sky contained contrails. To define when a Street View image of the sky contained a contrail, we used 64-dimensional image feature vectors derived from image-text data, created with an approach similar to that used by Juan et al. [26]. We applied a threshold to the cosine similarity of the Street View image feature vector and that of a seed image of a contrail taken from the ground; if it was similar enough, GOES-16 imagery at that time and location were sampled for human labeling of contrails. These additional labeled images are only in the training set: because Street View cars operate on days with sunnier weather, it may be easier than usual to identify contrails in the GOES16 imagery of those locations.</p>\n</blockquote>\n<p>And also this:</p>\n<blockquote>\n  <p>The full dataset contains 20,544 examples in the train set and 1,866 examples in the validation set. The examples are randomly partitioned except for the satellites scenes that were identified as likely to have contrails by Google Street View, which are only included in the training set. 9,283 of the training examples contain at least one annotated contrail. About 1.2% of the pixels in the training set are labeled as contrails. The dataset contains a wide variety of times and locations, as shown in Figure 5 and 6. The examples are not uniformly distributed in space and time as the images are sampled to include more contrail examples as described above.\"</p>\n</blockquote>\n<p>If the same applies to our train/validation/test data, I would deduce that provided validation data could be more similar to the test set than creating 5 folds from the train ourselves (well, depending on how you create the folds, I guess). Maybe it is a good idea to create the folds using the same proportion of contrails as in provided validation data?</p>",
      "rawMarkdown": "How are you splitting train data into 5 folds? \n\nI was reading the preprint and I noticed these two paragraphs:\n\n>To further boot the number of positives in the dataset, we also included some GOES-16 ABI imagery at locations in the US where Google Street View images of the sky contained contrails. To define when a Street View image of the sky contained a contrail, we used 64-dimensional image feature vectors derived from image-text data, created with an approach similar to that used by Juan et al. [26]. We applied a threshold to the cosine similarity of the Street View image feature vector and that of a seed image of a contrail taken from the ground; if it was similar enough, GOES-16 imagery at that time and location were sampled for human labeling of contrails. These additional labeled images are only in the training set: because Street View cars operate on days with sunnier weather, it may be easier than usual to identify contrails in the GOES16 imagery of those locations.\n\nAnd also this:\n\n>The full dataset contains 20,544 examples in the train set and 1,866 examples in the validation set. The examples are randomly partitioned except for the satellites scenes that were identified as likely to have contrails by Google Street View, which are only included in the training set. 9,283 of the training examples contain at least one annotated contrail. About 1.2% of the pixels in the training set are labeled as contrails. The dataset contains a wide variety of times and locations, as shown in Figure 5 and 6. The examples are not uniformly distributed in space and time as the images are sampled to include more contrail examples as described above.\"\n\nIf the same applies to our train/validation/test data, I would deduce that provided validation data could be more similar to the test set than creating 5 folds from the train ourselves (well, depending on how you create the folds, I guess). Maybe it is a good idea to create the folds using the same proportion of contrails as in provided validation data?",
      "votes": 2,
      "replies": [
        {
          "id": 2318722,
          "postDate": "2023-06-26T14:33:58.600Z",
          "content": "<p>Here is some more info on the split. I have been using a random 5-fold split with data in the training folder for now. </p>\n<p>Current plan is to determine the optimal prediction threshold using the validation folder.</p>\n<table>\n<thead>\n<tr>\n<th>Split</th>\n<th>Percent w/ Contrails</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Train</td>\n<td>45.10</td>\n</tr>\n<tr>\n<td>Valid</td>\n<td>29.74</td>\n</tr>\n</tbody>\n</table>",
          "rawMarkdown": "Here is some more info on the split. I have been using a random 5-fold split with data in the training folder for now. \n\nCurrent plan is to determine the optimal prediction threshold using the validation folder.\n\n| Split | Percent w/ Contrails |\n| --- | --- |\n| Train | 45.10 |\n| Valid | 29.74 |",
          "votes": 2
        }
      ]
    },
    {
      "id": 2314307,
      "postDate": "2023-06-23T10:35:23.917Z",
      "content": "<p>SingleModel<br>\nTTSplit(OriginalDataset \"train\"/\"validation\")</p>\n<ul>\n<li>CV：0.649(Threshold 0.05)</li>\n<li>LB：0.661(Threshold 0.35)</li>\n</ul>",
      "rawMarkdown": "SingleModel\nTTSplit(OriginalDataset \"train\"/\"validation\")\n* CV：0.649(Threshold 0.05)\n* LB：0.661(Threshold 0.35)",
      "votes": 2,
      "replies": [
        {
          "id": 2320935,
          "postDate": "2023-06-28T07:06:02.263Z",
          "content": "<p><a href=\"https://www.kaggle.com/hideyukizushi\" target=\"_blank\">@hideyukizushi</a> why your threshold of CV is so low?</p>",
          "rawMarkdown": "@hideyukizushi why your threshold of CV is so low?",
          "votes": 1
        }
      ]
    },
    {
      "id": 2382587,
      "postDate": "2023-08-09T21:01:35.227Z",
      "content": "<p><code>CV: 0.6861001545292715 LB : 0.686</code> on an ensemble , others cv more but lb less 😑 the usual confusion again .</p>",
      "rawMarkdown": "`CV: 0.6861001545292715 LB : 0.686` on an ensemble , others cv more but lb less 😑 the usual confusion again ."
    },
    {
      "id": 2333072,
      "postDate": "2023-07-06T16:22:05.437Z",
      "content": "<p>single model local dice0.671 lb0.687</p>",
      "rawMarkdown": "single model local dice0.671 lb0.687"
    },
    {
      "id": 2328577,
      "postDate": "2023-07-03T17:32:24.517Z",
      "rawMarkdown": "",
      "votes": 2,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2326006,
      "author_name": "SSS",
      "author_url": "",
      "post_date": "2023-07-01T19:11:02.290000",
      "content": "<p>No folds. Train folder vs validation folder.<br>\nThreshold was selected by iterating thru<code>np.arange(0.01, 0.51, 0.01)</code>.</p>\n<table>\n<thead>\n<tr>\n<th>CV</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.633</td>\n<td>0.636</td>\n</tr>\n<tr>\n<td>0.641</td>\n<td>0.642</td>\n</tr>\n<tr>\n<td>0.656</td>\n<td>0.662</td>\n</tr>\n<tr>\n<td>0.658</td>\n<td>0.666</td>\n</tr>\n<tr>\n<td>0.663</td>\n<td>0.671</td>\n</tr>\n<tr>\n<td>0.664</td>\n<td>0.673</td>\n</tr>\n<tr>\n<td>0.666</td>\n<td>0.679</td>\n</tr>\n<tr>\n<td>0.668</td>\n<td>0.681</td>\n</tr>\n<tr>\n<td><strong>0.673</strong></td>\n<td><strong>0.687</strong></td>\n</tr>\n</tbody>\n</table>",
      "votes": 16,
      "replies": []
    },
    {
      "id": 2322598,
      "author_name": "Nischay Dhankhar",
      "author_url": "",
      "post_date": "2023-06-29T11:16:48.610000",
      "content": "<p>Simple Train/Validation Split (No folds)</p>\n<p>All scores are from single model.</p>\n<p>CV：<strong>0.676</strong><br>\nLB：<strong>0.689</strong> (Threshold <strong>0.5</strong> not tuned)</p>\n<p><strong>Compilation of my results so far:</strong></p>\n<table>\n<thead>\n<tr>\n<th>CV</th>\n<th>Public LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.645</td>\n<td>0.640</td>\n</tr>\n<tr>\n<td>0.647</td>\n<td>0.661</td>\n</tr>\n<tr>\n<td>0.654</td>\n<td>0.665</td>\n</tr>\n<tr>\n<td>0.659</td>\n<td>0.671</td>\n</tr>\n<tr>\n<td>0.660</td>\n<td>0.666</td>\n</tr>\n<tr>\n<td>0.664</td>\n<td>0.677</td>\n</tr>\n<tr>\n<td>0.665</td>\n<td>0.663</td>\n</tr>\n<tr>\n<td>0.668</td>\n<td>0.684</td>\n</tr>\n<tr>\n<td>0.676</td>\n<td>0.689</td>\n</tr>\n</tbody>\n</table>",
      "votes": 15,
      "replies": [
        {
          "id": 2324558,
          "author_name": "delai50",
          "author_url": "",
          "post_date": "2023-06-30T17:18:48.900000",
          "content": "<p>Pretty good correlation, any hints about the split? 😏</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2324829,
              "author_name": "Nischay Dhankhar",
              "author_url": "",
              "post_date": "2023-06-30T22:50:23.243000",
              "content": "<p>These results are using the split provided by hosts ;) In my experiments, correlation deteriorates when I shift between architectures or image size. So, I tried to keep them constant until the last stage of the competition. </p>",
              "votes": 6,
              "replies": []
            },
            {
              "id": 2328719,
              "author_name": "JEANMPIA",
              "author_url": "",
              "post_date": "2023-07-03T20:10:07.667000",
              "content": "<p>Hello <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a>, <br>\nCorrect me if I'm wrong but a correlation that doesn't stick with diff models and image size is a very fragile one.<br>\nI had similar issues in the past but now with a 4-fold Split the scores are perfectely correlated even with a image size/architecture shift, I suggest you trouble shoot your pipeline so that you'r not going to be unpleasantly surprised at the end of the comp..</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2328986,
              "author_name": "Nirjhar Roy",
              "author_url": "",
              "post_date": "2023-07-04T02:54:29.070000",
              "content": "<p>Do you use train + validation while doing 4 fold split ?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2329831,
              "author_name": "JEANMPIA",
              "author_url": "",
              "post_date": "2023-07-04T14:35:22.830000",
              "content": "<p>Hello <a href=\"https://www.kaggle.com/phoenix9032\" target=\"_blank\">@phoenix9032</a>,<br>\nNo I keep validation as a holdout.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2314964,
      "author_name": "Egor Trushin",
      "author_url": "",
      "post_date": "2023-06-23T17:49:27.320000",
      "content": "<p>This is what I have accumulated so far.<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2126325%2Fa9862998a0b606ccf17b4aaec52ca8bd%2FCV_LB.png?generation=1687542535053652&amp;alt=media\" alt=\"\"></p>",
      "votes": 15,
      "replies": [
        {
          "id": 2315071,
          "author_name": "SSS",
          "author_url": "",
          "post_date": "2023-06-23T19:42:00.590000",
          "content": "<p>You achieved a very good correlation. Having checked on your work, I can infer that you used validation folder inside the cross-validation loop. May I ask the question, how did you adjust the dice metric like this without inflating it? Though, it really looks like you did not use the validation folder during the training.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2315117,
              "author_name": "Harshit Sheoran",
              "author_url": "",
              "post_date": "2023-06-23T20:32:26.023000",
              "content": "<p>That does not look correlated, above .684, the higher he scores in CV, the lower he scores on LB</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2315142,
              "author_name": "june",
              "author_url": "",
              "post_date": "2023-06-23T20:57:09.390000",
              "content": "<p>Generally, it looks correlated.<br>\nIt would look more correlated if x and y values were set to [0, 1].<br>\nStatistical graphs can lie. </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2315150,
              "author_name": "SSS",
              "author_url": "",
              "post_date": "2023-06-23T21:05:01.060000",
              "content": "<p>Well it does. There is a variance for sure but \"not correlated\" means some of us needs to revisit what was considered a good correlation until this moment. <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2Fa966293d9733610f507131c38a65e4e9%2Fcorrelated.png?generation=1687554457285576&amp;alt=media\" alt=\"\"></p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2315182,
              "author_name": "Harshit Sheoran",
              "author_url": "",
              "post_date": "2023-06-23T22:06:32.707000",
              "content": "<p>On kaggle, we push to squeeze out the last bit of performance out of the models on CV and if they start randomly crashing a large amount (by LB standards), even though they are \"technically\" co-related, this correlation is not of use, LB not correlating CV is not a new thing on kaggle, but, your \"strongly\" correlated has destroyed many potential golds of countless people on kaggle ;)</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2315243,
              "author_name": "SSS",
              "author_url": "",
              "post_date": "2023-06-24T00:16:41.673000",
              "content": "<p>How did you calculate “many” out of countless? LoL</p>",
              "votes": 5,
              "replies": []
            },
            {
              "id": 2315245,
              "author_name": "Harshit Sheoran",
              "author_url": "",
              "post_date": "2023-06-24T00:23:46.037000",
              "content": "<p>LoL, Sorry, english is not my native language, I meant to say countless people have lost many of their own golds because of not good enough co-relation, including me 5 months ago</p>",
              "votes": 5,
              "replies": []
            },
            {
              "id": 2315819,
              "author_name": "Egor Trushin",
              "author_url": "",
              "post_date": "2023-06-24T12:53:55.133000",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/sergiosaharovskiy\" target=\"_blank\">@sergiosaharovskiy</a> <br>\nI am not doing anything special with dice metrics at the moment.<br>\nYes, I train with a combined train/val dataset. But I think it is actually a good idea to have the val dataset separately as a holdout, e.g. to do experiments with threshold optimization.</p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 2320929,
              "author_name": "豆柴金鯱",
              "author_url": "",
              "post_date": "2023-06-28T06:59:42.050000",
              "content": "<p>Hi Egor, May I ask if you're using dice loss for training &amp; dice coefficient for score? My local CV &amp; LB score have huge gap and the correlation is not stable..</p>\n<p>Thank you.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2360109,
          "author_name": "HZM",
          "author_url": "",
          "post_date": "2023-07-26T14:58:58.687000",
          "content": "<p>Hi, Egor, May I know your currently CV score</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2294891,
      "author_name": "MPWARE",
      "author_url": "",
      "post_date": "2023-06-10T12:29:46.980000",
      "content": "<p>CV5 single model:</p>\n<ul>\n<li>Global Dice: 0.6785</li>\n<li>Dice per Image: 0.7757</li>\n<li>AUCPR: 0.7555</li>\n</ul>\n<p>LB: 0.680</p>",
      "votes": 10,
      "replies": [
        {
          "id": 2295073,
          "author_name": "Nirjhar Roy",
          "author_url": "",
          "post_date": "2023-06-10T15:08:44.217000",
          "content": "<p>Great Global Dice result  …is it for a threshold? </p>",
          "votes": 2,
          "replies": [
            {
              "id": 2295169,
              "author_name": "MPWARE",
              "author_url": "",
              "post_date": "2023-06-10T16:39:27.207000",
              "content": "<p>Threshold on inference is around 0.40</p>",
              "votes": 2,
              "replies": []
            }
          ]
        },
        {
          "id": 2306511,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2023-06-17T11:25:16.560000",
          "content": "<p><strong>Updates</strong>: </p>\n<p>Same model but improved training procedure. No post-processing.</p>\n<ul>\n<li>Global Dice: 0.6819</li>\n<li>Dice per Image: 0.7763</li>\n<li>AUCPR: 0.7526</li>\n</ul>\n<p>LB: 0.691 (Best threshold around 0.4)</p>",
          "votes": 4,
          "replies": [
            {
              "id": 2306538,
              "author_name": "lyu",
              "author_url": "",
              "post_date": "2023-06-17T11:47:35.670000",
              "content": "<p>Impressive score. Do you use validation to calulate cv or KFold?.I already get 0.68lb when my cv is 0.66….but now my cv get ~0.695, my lb still is 0.68+. My lb has been standing still for the past few weeks. This even makes me wonder if I'm overfitting the local validation dataset😂</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2306601,
              "author_name": "MPWARE",
              "author_url": "",
              "post_date": "2023-06-17T12:32:53.570000",
              "content": "<p>KFold, I'm not using validation/ folder for now. Maybe you've an issue with your global dice metric which is not the same one as Kaggle.</p>\n<p>One example for one of my fold (finetuning): Global dice (dice_coef), dice, AUCPR and LR.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2F37d09273dd79429fa7a0c22d724cbf87%2Fstage2.png?generation=1687007010478502&amp;alt=media\" alt=\"\"></p>",
              "votes": 8,
              "replies": []
            },
            {
              "id": 2306756,
              "author_name": "lyu",
              "author_url": "",
              "post_date": "2023-06-17T14:47:06.397000",
              "content": "<p>thanks for your reply! maybe I need check my code to confirm</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2306780,
              "author_name": "JEANMPIA",
              "author_url": "",
              "post_date": "2023-06-17T15:10:44.607000",
              "content": "<p><a href=\"https://www.kaggle.com/zhuwanglju\" target=\"_blank\">@zhuwanglju</a> I'm having the same issue as you have, and this with a SKF CV and with the Train/valid original split… I feel like using the train/valid is a massive bait. </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2307233,
              "author_name": "SSS",
              "author_url": "",
              "post_date": "2023-06-18T03:05:36.323000",
              "content": "<blockquote>\n  <p>KFold, I'm not using validation/ folder for now. Maybe you've an issue with your global dice metric which is not the same one as Kaggle.</p>\n  <p>One example for one of my fold (finetuning): Global dice (dice_coef), dice, AUCPR and LR.</p>\n  <p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2F37d09273dd79429fa7a0c22d724cbf87%2Fstage2.png?generation=1687007010478502&amp;alt=media\" alt=\"\"></p>\n</blockquote>\n<p>Thank you for posting the initial results.<br>\nAchieving a score of 0.68 after just the second epoch is a promising start. Though, waiting for another 70 epochs to get +0.019, god you are tough man! It is so Kaggle :)</p>\n<p>P.s. I cannot really see, but it looks the warmup is 3-4 and the first two epochs score is around 0 or the enumeration starts from 1?</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2307452,
              "author_name": "MPWARE",
              "author_url": "",
              "post_date": "2023-06-18T07:48:54.683000",
              "content": "<p>It's the finetuning part as I mentioned that's why it starts very high. The initial training part (not posted here) starts very low. And yes I've LR warmup and display starts from epoch 1.</p>",
              "votes": 6,
              "replies": []
            },
            {
              "id": 2307613,
              "author_name": "Ioannis M",
              "author_url": "",
              "post_date": "2023-06-18T10:19:37.890000",
              "content": "<p>now makes sense - thanks for sharing your results. Just to make sure we are comparing apples with apples the scores you report above are the average across all folds, right?<br>\nEDIT: or from a single fold?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2307634,
              "author_name": "JEANMPIA",
              "author_url": "",
              "post_date": "2023-06-18T10:47:03.917000",
              "content": "<p><a href=\"https://www.kaggle.com/imeintanis\" target=\"_blank\">@imeintanis</a> <br>\nThat should answer your question:</p>\n<blockquote>\n  <p>One example for <strong>one of my fold</strong> (finetuning): Global dice (dice_coef), dice, AUCPR and LR.</p>\n</blockquote>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2311987,
              "author_name": "Shashwat Raman",
              "author_url": "",
              "post_date": "2023-06-21T15:27:51.147000",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a>,<br>\nIs your lb score the ensemble of the 5 folds or is it just a single fold submission?<br>\nThanks :)</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2312257,
              "author_name": "MPWARE",
              "author_url": "",
              "post_date": "2023-06-21T19:21:46.390000",
              "content": "<p><a href=\"https://www.kaggle.com/shashwatraman\" target=\"_blank\">@shashwatraman</a> Ensemble of 5 folds. We've have lot of room/time on inference so I've not even tried any single fold for now. But I will do and post my results here.</p>\n<p>And here is what looks like the search for best threshold:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2F5d894e8837c5ccb49551ae48a67bc84f%2Fthreshold.png?generation=1687386508197030&amp;alt=media\" alt=\"\"></p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2312581,
              "author_name": "Shashwat Raman",
              "author_url": "",
              "post_date": "2023-06-22T03:12:32.417000",
              "content": "<p>Thank you very much for your reply <a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a>, it's really helpful. I'm still struggling to find a good cross validation strategy. Cv and Lb are not correlated for me rn.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2314341,
      "author_name": "Tawara",
      "author_url": "",
      "post_date": "2023-06-23T11:07:38.580000",
      "content": "<p>training by 5-fold cross validaition:</p>\n<table>\n<thead>\n<tr>\n<th>train folder(oof pred)</th>\n<th>validation folder(5-fold avg)</th>\n<th>LB(5-fold avg)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.658</td>\n<td>0.635</td>\n<td>0.650</td>\n</tr>\n<tr>\n<td>0.659</td>\n<td>0.642</td>\n<td>0.657</td>\n</tr>\n</tbody>\n</table>\n<p>　<br>\nIt seems that \"validation\" and LB have correlation🤔</p>",
      "votes": 7,
      "replies": [
        {
          "id": 2314665,
          "author_name": "Reacher",
          "author_url": "",
          "post_date": "2023-06-23T14:21:28.510000",
          "content": "<p>same here , only improvements in validation-folder score boosts LB.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2314751,
              "author_name": "Harshit Sheoran",
              "author_url": "",
              "post_date": "2023-06-23T15:25:06.370000",
              "content": "<p>I have validation-folder scoring .655 and .668 global dice and both scores .678 on LB :(</p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 2314788,
              "author_name": "Reacher",
              "author_url": "",
              "post_date": "2023-06-23T15:44:26.313000",
              "content": "<p>In my experiments (not many) Lb and validation-folder perfectly correlates , maybe it is a threshold issue (i always use cv-threshold, usually in range(0.4,0.5) )</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2314859,
              "author_name": "Harshit Sheoran",
              "author_url": "",
              "post_date": "2023-06-23T16:33:48.397000",
              "content": "<p>Using threshold=0.4 I have .678 LB but I was using threshold=0.5 when doing validation-folder cv, when I do threshold=0.4 for cv, I get .670, so that's another increase on validation-folder</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2314879,
              "author_name": "Shashwat Raman",
              "author_url": "",
              "post_date": "2023-06-23T16:48:50.033000",
              "content": "<p>Yes, the same thing happens for me. Some models which have a better global dice score on 5 fold cv and validation folder have a lower score on the lb. <br>\nMaybe this happens because the Public Test set is very small?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2315122,
              "author_name": "JEANMPIA",
              "author_url": "",
              "post_date": "2023-06-23T20:36:40.580000",
              "content": "<p>same here, but history told us it was a bad idea to trust LB blindly..</p>",
              "votes": 2,
              "replies": []
            }
          ]
        },
        {
          "id": 2320934,
          "author_name": "豆柴金鯱",
          "author_url": "",
          "post_date": "2023-06-28T07:04:50.133000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a> When you say \"train folder(oof pred)\", are you splitting the train folder into K folds to perform a cross-validation? Then use the trained k models to test on the validation folder? </p>\n<p>For my cross-validation, I use the validation folder for each of K fold while training, which is different.<br>\nThank you.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2321450,
              "author_name": "Tawara",
              "author_url": "",
              "post_date": "2023-06-28T15:28:21.167000",
              "content": "<p>Yes, you are right.<br>\nI trained k models by K-folds cross-validation, then use these k models to test on the validation folder.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2327472,
              "author_name": "豆柴金鯱",
              "author_url": "",
              "post_date": "2023-07-03T01:40:16.917000",
              "content": "<p>Thank you for your information. I take the same way as you did and it helps. </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2277264,
      "author_name": "lyu",
      "author_url": "",
      "post_date": "2023-05-27T15:21:55.800000",
      "content": "<p>I  just submited twice. cv:587,lb:589. cv:607 lb:604. it seems cv and lb have good trend👀</p>",
      "votes": 5,
      "replies": [
        {
          "id": 2277334,
          "author_name": "Jan H",
          "author_url": "",
          "post_date": "2023-05-27T16:12:32.010000",
          "content": "<p>Nice! In your case the LB and CV seem even closer. Do you also use the training folder for training and validation for CV?</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2277808,
              "author_name": "lyu",
              "author_url": "",
              "post_date": "2023-05-28T04:44:44.537000",
              "content": "<p>Yes, I rewrote the function of dice. If  dice_score is similar to the following code(from <a href=\"https://www.kaggle.com/code/lupin11/u-net-baseline-training\" target=\"_blank\">this</a>), which will make the local CV higher because many images do not have masks, and their dice score is 1. But it doesn't seem to matter, its score and lb trend are also very consistent.😀</p>\n<pre><code>def dice_score(y_p, y_t, smooth=1e-6):\n    i = torch.sum(y_p * y_t, dim=(2, 3))\n    u = torch.sum(y_p, dim=(2, 3)) + torch.sum(y_t, dim=(2, 3))\n    score = (2 * i + smooth)/(u + smooth)\n    return torch.mean(score)\n</code></pre>",
              "votes": 10,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2309146,
      "author_name": "Reacher",
      "author_url": "",
      "post_date": "2023-06-19T12:34:08.297000",
      "content": "<p>5 folds single model : </p>\n<ul>\n<li>global dice : 0.6418 (mean of cv scores)</li>\n<li>LB : 0.655 </li>\n</ul>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2289988,
      "author_name": "liuzhangzhen",
      "author_url": "",
      "post_date": "2023-06-06T13:34:20.390000",
      "content": "<p>CV - 0.637 LB - 0.662</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2289893,
      "author_name": "JEANMPIA",
      "author_url": "",
      "post_date": "2023-06-06T12:14:28.450000",
      "content": "<p>Current best:</p>\n<table>\n<thead>\n<tr>\n<th>CV</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.606</td>\n<td>0.646</td>\n</tr>\n</tbody>\n</table>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2312467,
      "author_name": "Harshit Sheoran",
      "author_url": "",
      "post_date": "2023-06-21T23:33:40.800000",
      "content": "<p>5 Fold training using \"train\" folder only</p>\n<p>Fold-wise average CV Global Dice: .69</p>\n<p>5-Fold ensemble on \"validation\" folder:<br>\nGlobal Dice: .668<br>\nImage-wise Dice: .830</p>\n<p>LB: .678</p>\n<p>I have many models in CV .68 to .69 and they score from .66 to .678 on LB, I am unable to find correlation as .687 CV scores .66 LB and .685 also scores .678 the same as .690 CV</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2314029,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-06-23T05:48:14.957000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2316761,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2023-06-25T08:11:17.810000",
          "content": "<p>Update:</p>\n<p>5 Fold average CV Global Dice: .6916<br>\nLB: .657</p>\n<p>Each one of these 5 folds scores over .69 on their respective CV, very stable</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2316965,
              "author_name": "lyu",
              "author_url": "",
              "post_date": "2023-06-25T10:55:02.710000",
              "content": "<p>we have similiar problem… my cv is 0.70+, and lb is 0.67+….Now, I choose to believe my cv….</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2338132,
              "author_name": "Optimo",
              "author_url": "",
              "post_date": "2023-07-10T14:53:44.263000",
              "content": "<p>I have just started with a few experiments and my correlation is terrible</p>\n<table>\n<thead>\n<tr>\n<th>exp</th>\n<th>type</th>\n<th>CV</th>\n<th>Validation</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>1 fold</td>\n<td>0.6824</td>\n<td>0.650</td>\n<td>0.662</td>\n</tr>\n<tr>\n<td>2</td>\n<td>5 fold</td>\n<td>0.6775</td>\n<td>0.6575</td>\n<td>0.656</td>\n</tr>\n<tr>\n<td>3</td>\n<td>5 fold</td>\n<td>0.6822</td>\n<td>0.6687</td>\n<td>0.656</td>\n</tr>\n<tr>\n<td>4</td>\n<td>1 fold</td>\n<td>Nan</td>\n<td>0.6701</td>\n<td>0.641</td>\n</tr>\n<tr>\n<td>5</td>\n<td>5 fold</td>\n<td>0.6867</td>\n<td>0.6724</td>\n<td>0.659</td>\n</tr>\n</tbody>\n</table>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2338372,
              "author_name": "MPWARE",
              "author_url": "",
              "post_date": "2023-07-10T17:57:58.687000",
              "content": "<p>Are you using some threshold?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2339508,
              "author_name": "Optimo",
              "author_url": "",
              "post_date": "2023-07-10T20:20:53.587000",
              "content": "<p>I'm giving here best thresholded scores for validation.</p>\n<p>From my observations, single models predict hard distributions of 1s and 0s mostly so in validation score is the same across a wide range of thresholds, for 5 fold CV averaging on validation : best threshold is around 0.3.</p>\n<p>For the same single fold, I've seen 0.641LB for threshold 0.95, and 0.626LB for threshold 0.80. So I don't observe the same behavior between validation and LB. Anyone can relate ?</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2320939,
          "author_name": "豆柴金鯱",
          "author_url": "",
          "post_date": "2023-06-28T07:08:47.427000",
          "content": "<p>Hello, <a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a> may I ask your image size? Your local CV score is crazy…</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2289868,
      "author_name": "Mithil Salunkhe",
      "author_url": "",
      "post_date": "2023-06-06T11:48:54.170000",
      "content": "<p>CV on validation data - 0.634 LB - 644</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2311491,
          "author_name": "Mithil Salunkhe",
          "author_url": "",
          "post_date": "2023-06-21T07:31:41.393000",
          "content": "<p>CV on 4 kfolds - 0.677 LB - 0.668</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2324294,
              "author_name": "Mithil Salunkhe",
              "author_url": "",
              "post_date": "2023-06-30T14:08:56.907000",
              "content": "<p>CV - 688<br>\nLB - 693</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2382592,
              "author_name": "Gaurav Rawat",
              "author_url": "",
              "post_date": "2023-08-09T21:02:31.100000",
              "content": "<p>This one is nice is it an ensemble ?</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2339675,
      "author_name": "SeshuRaju 🧘‍♂️",
      "author_url": "",
      "post_date": "2023-07-11T02:24:37.170000",
      "content": "<table>\n<thead>\n<tr>\n<th>CV (Single Model Threshold 0.5)</th>\n<th>LB (Threshold 0.5 not tuned)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.680</td>\n<td>0.646</td>\n</tr>\n<tr>\n<td>0.673 (<strong>-0.007</strong>)</td>\n<td>0.638 (<strong>-0.008</strong>)</td>\n</tr>\n</tbody>\n</table>\n<p>Will compare the cv correlation with few more submissions</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2317041,
      "author_name": "Louis Weiss",
      "author_url": "",
      "post_date": "2023-06-25T12:05:44.657000",
      "content": "<p>How many epochs do you typically train your model for each fold? I would like to perform cross-validation in order to obtain a more accurate estimation of my model's performance. However, running 5 * 20 epochs would require several days. I am training the model on my own decent PC (Nvidia 3090ti, 24GB VRAM), but performing cross-validation for all my models is not feasible due to the significant time it would take. Do you utilize any external resources? Kaggle notebooks tend to be unstable and often shut down after a few hours. Google Colab, even with Collab Pro, is quite slow when it comes to loading data, and the credits may run out before completing the cross-validation process.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2317977,
          "author_name": "Shashwat Raman",
          "author_url": "",
          "post_date": "2023-06-26T05:39:57.300000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/louisweiss\" target=\"_blank\">@louisweiss</a>,<br>\nI am training my models for about 30-40 epochs. More epochs help here.<br>\nNow for cross validation, it is generally better to use a 5 fold cv, however you can go for 3 or 4 folds. Just don't go below that.<br>\nYes I am using external recources. Kaggle notebooks are great. They are very stable if you Save and Run the notebook. Just the thing is the 30 hour limit is too less. You can try JarvisLabs, PaperSpace, etc and see what suits you.<br>\nTraining a medium sized model with 5 Folds and 30 epochs each takes around 3 hours for me using a single A5000 GPU.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2318465,
              "author_name": "SSS",
              "author_url": "",
              "post_date": "2023-06-26T11:38:09.860000",
              "content": "<p><a href=\"https://www.kaggle.com/shashwatraman\" target=\"_blank\">@shashwatraman</a> what is your python and pytorch version? It takes 15 min per epoch for me 512x512 size. I wonder if you use resnet50 like architecture.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2318472,
              "author_name": "Shashwat Raman",
              "author_url": "",
              "post_date": "2023-06-26T11:49:38.220000",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/sergiosaharovskiy\" target=\"_blank\">@sergiosaharovskiy</a>,<br>\nPython Version - 3.10.11<br>\nPytorch Version - 2.0.1<br>\nWhen using 512 sized images, it takes me around 5 mins per epoch to train. I use an EfficientNetB3!</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2318477,
              "author_name": "Mithil Salunkhe",
              "author_url": "",
              "post_date": "2023-06-26T11:53:04.903000",
              "content": "<p>Takes me around 3 minutes per epoch for efficientnetv2 L at a 512*512 size for 5 fold training with 2x4090. I use FP16 with torch.compile </p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2318485,
              "author_name": "Shashwat Raman",
              "author_url": "",
              "post_date": "2023-06-26T12:01:48.743000",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/mithilsalunkhe\" target=\"_blank\">@mithilsalunkhe</a>,<br>\nThat is really amazing. Currently I am using a single A5000 Gpu :)</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2319268,
              "author_name": "SSS",
              "author_url": "",
              "post_date": "2023-06-27T03:01:45.887000",
              "content": "<p>torch.compile is such a pain. Kaggle kernel throws an error for P100 since it is not supported. T4 x 2 throw the error with CUDA. Once you switch to x1 it says Cannot find a working triton installation. Either it sucks being still in Beta Version or me sucks, no alternatives. </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2319333,
              "author_name": "Mithil Salunkhe",
              "author_url": "",
              "post_date": "2023-06-27T04:31:24.257000",
              "content": "<p>I use the NGC container  given by nvidia  , only had a few issues with it. I still have to try using it in inference as the speedup would be minimal  </p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2292643,
      "author_name": "Giovanni Cavallin",
      "author_url": "",
      "post_date": "2023-06-08T14:24:15.080000",
      "content": "<p>CV: 0.4931 LB: 0.636<br>\nI'm using the mDice metric (so not averaging over the images but over all pixels in val set) and validating on provided val dataset, only on positive images. If keeping track of negative images, my mDice would fall at ~0.20 in CV. Don't see particular strong correlation with my CV/LB. Any idea? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2293054,
          "author_name": "JEANMPIA",
          "author_url": "",
          "post_date": "2023-06-08T23:43:28.810000",
          "content": "<p>As you said, you don't use the \"right\" metric, maybe changing that would be the firsy thing to try but then my guess is that you did already.<br>\nif that doesn't work, I suggest you take a look at the Dice metric in pytroch metrics<br>\nPytorch : <a href=\"https://torchmetrics.readthedocs.io/en/v0.8.2/classification/dice_score.html\" target=\"_blank\">https://torchmetrics.readthedocs.io/en/v0.8.2/classification/dice_score.html</a> (the one I use, I certify it works well)<br>\nyou meed need a few unsqueeze or squeeze to match the input of this function, but the fact that you can choose the threshold as a parameter is very nice for 0.01 optimisations. Good luck</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2293368,
              "author_name": "Giovanni Cavallin",
              "author_url": "",
              "post_date": "2023-06-09T06:07:43.767000",
              "content": "<p>Thank you for your answer! </p>\n<p>I decided to use the mDice because they said they'd use it for evaluation, and I was thinking than I have this discrepancy between CV/LB because the LB has only positive images with many contrails. In particular, I noticed that the models that in my local with higher precision (and even lower recall) works better in this LB. However, having 0.2 vs 0.63 is a big leap. 😅 </p>\n<p>I don't know if I'm being wrong.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2295522,
              "author_name": "JEANMPIA",
              "author_url": "",
              "post_date": "2023-06-11T04:09:25.900000",
              "content": "<blockquote>\n  <p>LB has only positive images with many contrails</p>\n</blockquote>\n<p>This is a dangerous assumption for your final LB score, <br>\ntake a look here:<br>\n<a href=\"https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/discussion/415764#2291944\" target=\"_blank\">https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/discussion/415764#2291944</a></p>\n<p>what you'r saying might be correct for the public LB, but don't trust this at all, public LB is very small, get CV up not this.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2296451,
              "author_name": "Giovanni Cavallin",
              "author_url": "",
              "post_date": "2023-06-11T20:56:15.587000",
              "content": "<p>Sorry, I explained myself wrong. I was saying that I'm possibly having this discrepancy because this particular LB has many positive example. My assumption is that I'll have to trust my CV despite the LB scores. :)</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2277210,
      "author_name": "Jan H",
      "author_url": "",
      "post_date": "2023-05-27T14:38:39.317000",
      "content": "<p>Very strong CV results! Which data do you use for CV? Do you use the \"train\" folder for train / validation and the \"validation\" data for the test?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2277222,
          "author_name": "Dracarys",
          "author_url": "",
          "post_date": "2023-05-27T14:49:16.303000",
          "content": "<p>Model is trained on train folder and validated on validation folder.  </p>",
          "votes": 2,
          "replies": [
            {
              "id": 2277239,
              "author_name": "Jan H",
              "author_url": "",
              "post_date": "2023-05-27T14:59:14",
              "content": "<p>Alright, so thr CV refers to the validation dice coef during / at the end of training?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2277245,
              "author_name": "Dracarys",
              "author_url": "",
              "post_date": "2023-05-27T15:04:58.040000",
              "content": "<p>yes, CV is calculated on validation data at the end of each epoch.</p>",
              "votes": 3,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2317391,
      "author_name": "delai50",
      "author_url": "",
      "post_date": "2023-06-25T17:14:29.897000",
      "content": "<p>How are you splitting train data into 5 folds? </p>\n<p>I was reading the preprint and I noticed these two paragraphs:</p>\n<blockquote>\n  <p>To further boot the number of positives in the dataset, we also included some GOES-16 ABI imagery at locations in the US where Google Street View images of the sky contained contrails. To define when a Street View image of the sky contained a contrail, we used 64-dimensional image feature vectors derived from image-text data, created with an approach similar to that used by Juan et al. [26]. We applied a threshold to the cosine similarity of the Street View image feature vector and that of a seed image of a contrail taken from the ground; if it was similar enough, GOES-16 imagery at that time and location were sampled for human labeling of contrails. These additional labeled images are only in the training set: because Street View cars operate on days with sunnier weather, it may be easier than usual to identify contrails in the GOES16 imagery of those locations.</p>\n</blockquote>\n<p>And also this:</p>\n<blockquote>\n  <p>The full dataset contains 20,544 examples in the train set and 1,866 examples in the validation set. The examples are randomly partitioned except for the satellites scenes that were identified as likely to have contrails by Google Street View, which are only included in the training set. 9,283 of the training examples contain at least one annotated contrail. About 1.2% of the pixels in the training set are labeled as contrails. The dataset contains a wide variety of times and locations, as shown in Figure 5 and 6. The examples are not uniformly distributed in space and time as the images are sampled to include more contrail examples as described above.\"</p>\n</blockquote>\n<p>If the same applies to our train/validation/test data, I would deduce that provided validation data could be more similar to the test set than creating 5 folds from the train ourselves (well, depending on how you create the folds, I guess). Maybe it is a good idea to create the folds using the same proportion of contrails as in provided validation data?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2318722,
          "author_name": "Bartley",
          "author_url": "",
          "post_date": "2023-06-26T14:33:58.600000",
          "content": "<p>Here is some more info on the split. I have been using a random 5-fold split with data in the training folder for now. </p>\n<p>Current plan is to determine the optimal prediction threshold using the validation folder.</p>\n<table>\n<thead>\n<tr>\n<th>Split</th>\n<th>Percent w/ Contrails</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Train</td>\n<td>45.10</td>\n</tr>\n<tr>\n<td>Valid</td>\n<td>29.74</td>\n</tr>\n</tbody>\n</table>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2314307,
      "author_name": "yukiZ",
      "author_url": "",
      "post_date": "2023-06-23T10:35:23.917000",
      "content": "<p>SingleModel<br>\nTTSplit(OriginalDataset \"train\"/\"validation\")</p>\n<ul>\n<li>CV：0.649(Threshold 0.05)</li>\n<li>LB：0.661(Threshold 0.35)</li>\n</ul>",
      "votes": 2,
      "replies": [
        {
          "id": 2320935,
          "author_name": "豆柴金鯱",
          "author_url": "",
          "post_date": "2023-06-28T07:06:02.263000",
          "content": "<p><a href=\"https://www.kaggle.com/hideyukizushi\" target=\"_blank\">@hideyukizushi</a> why your threshold of CV is so low?</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2382587,
      "author_name": "Gaurav Rawat",
      "author_url": "",
      "post_date": "2023-08-09T21:01:35.227000",
      "content": "<p><code>CV: 0.6861001545292715 LB : 0.686</code> on an ensemble , others cv more but lb less 😑 the usual confusion again .</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2333072,
      "author_name": "pky",
      "author_url": "",
      "post_date": "2023-07-06T16:22:05.437000",
      "content": "<p>single model local dice0.671 lb0.687</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2328577,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-07-03T17:32:24.517000",
      "content": "",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2276804": "\n| CV | LB |\n| --- | --- |\n| 0.65 | 0.536 |\n|0.71  | 0.578 |\n|0.76  | 0.621 |\n|0.77  | 0.633 |\n\n> So far CV-LB seems perfectly correlated. What are your results?\n",
    "2326006": "No folds. Train folder vs validation folder.\nThreshold was selected by iterating thru` np.arange(0.01, 0.51, 0.01)`.\n\n\n| CV     | LB     |\n|-------|--------|\n| 0.633 | 0.636  |\n| 0.641  | 0.642  |\n| 0.656  | 0.662  |\n| 0.658  | 0.666  |\n| 0.663  | 0.671  |\n| 0.664  | 0.673  |\n| 0.666  | 0.679  |\n| 0.668  | 0.681  |\n| **0.673**  | **0.687**  |",
    "2322598": "Simple Train/Validation Split (No folds)\n\nAll scores are from single model.\n\nCV：**0.676**\nLB：**0.689** (Threshold **0.5** not tuned)\n\n**Compilation of my results so far:**\n| CV | Public LB |\n| --- | --- |\n| 0.645 | 0.640 |\n| 0.647 | 0.661  |\n| 0.654 | 0.665 |\n| 0.659 | 0.671  |\n| 0.660 | 0.666 |\n| 0.664 | 0.677 |\n| 0.665 | 0.663 |\n| 0.668 | 0.684 |\n| 0.676 | 0.689 |\n\n",
    "2314964": "This is what I have accumulated so far.![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2126325%2Fa9862998a0b606ccf17b4aaec52ca8bd%2FCV_LB.png?generation=1687542535053652&alt=media)",
    "2294891": "CV5 single model:\n- Global Dice: 0.6785\n- Dice per Image: 0.7757\n- AUCPR: 0.7555\n\nLB: 0.680",
    "2314341": "training by 5-fold cross validaition:\n\n| train folder(oof pred) | validation folder(5-fold avg) | LB(5-fold avg) |\n|:------:|:------:|:------:|\n| 0.658 | 0.635  | 0.650 |\n| 0.659 | 0.642  | 0.657  |\n\n　\nIt seems that \"validation\" and LB have correlation🤔",
    "2277264": "I  just submited twice. cv:587,lb:589. cv:607 lb:604. it seems cv and lb have good trend👀",
    "2309146": "5 folds single model : \n- global dice : 0.6418 (mean of cv scores)\n- LB : 0.655 ",
    "2289988": "CV - 0.637 LB - 0.662",
    "2289893": "Current best:\n| CV | LB |\n| --- | --- |\n|0.606  | 0.646 |\n",
    "2312467": "5 Fold training using \"train\" folder only\n\nFold-wise average CV Global Dice: .69\n\n5-Fold ensemble on \"validation\" folder:\nGlobal Dice: .668\nImage-wise Dice: .830\n\nLB: .678\n\nI have many models in CV .68 to .69 and they score from .66 to .678 on LB, I am unable to find correlation as .687 CV scores .66 LB and .685 also scores .678 the same as .690 CV",
    "2289868": "CV on validation data - 0.634 LB - 644",
    "2339675": "| CV (Single Model Threshold 0.5) | LB (Threshold 0.5 not tuned) |\n| --- | --- |\n| 0.680 | 0.646 |\n| 0.673 (**-0.007**) | 0.638 (**-0.008**) |\n\nWill compare the cv correlation with few more submissions",
    "2317041": "How many epochs do you typically train your model for each fold? I would like to perform cross-validation in order to obtain a more accurate estimation of my model's performance. However, running 5 * 20 epochs would require several days. I am training the model on my own decent PC (Nvidia 3090ti, 24GB VRAM), but performing cross-validation for all my models is not feasible due to the significant time it would take. Do you utilize any external resources? Kaggle notebooks tend to be unstable and often shut down after a few hours. Google Colab, even with Collab Pro, is quite slow when it comes to loading data, and the credits may run out before completing the cross-validation process.",
    "2292643": "CV: 0.4931 LB: 0.636\nI'm using the mDice metric (so not averaging over the images but over all pixels in val set) and validating on provided val dataset, only on positive images. If keeping track of negative images, my mDice would fall at ~0.20 in CV. Don't see particular strong correlation with my CV/LB. Any idea? ",
    "2277210": "Very strong CV results! Which data do you use for CV? Do you use the \"train\" folder for train / validation and the \"validation\" data for the test?",
    "2317391": "How are you splitting train data into 5 folds? \n\nI was reading the preprint and I noticed these two paragraphs:\n\n>To further boot the number of positives in the dataset, we also included some GOES-16 ABI imagery at locations in the US where Google Street View images of the sky contained contrails. To define when a Street View image of the sky contained a contrail, we used 64-dimensional image feature vectors derived from image-text data, created with an approach similar to that used by Juan et al. [26]. We applied a threshold to the cosine similarity of the Street View image feature vector and that of a seed image of a contrail taken from the ground; if it was similar enough, GOES-16 imagery at that time and location were sampled for human labeling of contrails. These additional labeled images are only in the training set: because Street View cars operate on days with sunnier weather, it may be easier than usual to identify contrails in the GOES16 imagery of those locations.\n\nAnd also this:\n\n>The full dataset contains 20,544 examples in the train set and 1,866 examples in the validation set. The examples are randomly partitioned except for the satellites scenes that were identified as likely to have contrails by Google Street View, which are only included in the training set. 9,283 of the training examples contain at least one annotated contrail. About 1.2% of the pixels in the training set are labeled as contrails. The dataset contains a wide variety of times and locations, as shown in Figure 5 and 6. The examples are not uniformly distributed in space and time as the images are sampled to include more contrail examples as described above.\"\n\nIf the same applies to our train/validation/test data, I would deduce that provided validation data could be more similar to the test set than creating 5 folds from the train ourselves (well, depending on how you create the folds, I guess). Maybe it is a good idea to create the folds using the same proportion of contrails as in provided validation data?",
    "2314307": "SingleModel\nTTSplit(OriginalDataset \"train\"/\"validation\")\n* CV：0.649(Threshold 0.05)\n* LB：0.661(Threshold 0.35)",
    "2382587": "`CV: 0.6861001545292715 LB : 0.686` on an ensemble , others cv more but lb less 😑 the usual confusion again .",
    "2333072": "single model local dice0.671 lb0.687",
    "2328577": ""
  }
}