{
  "id": 498194,
  "title": "Let's share CV-LB Insights  here 🔥",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/498194",
  "author_name": "Ulrich G.",
  "post_date": "2024-04-27T08:53:31.541000",
  "votes": 22,
  "comment_count": 21,
  "views": 0,
  "content": "<p>Hi Everyone. I hope you're doing well in this competition. I'm a little bit curious to see how things goes for you concerning CV and LB gaps. Feel free to share here.</p>\n<p>I start first by sharing two experiments:</p>\n<ul>\n<li>A simple Xgboost : CV=0.43x / LB = 0.44x</li>\n<li>A simple FCN : CV=0.51x / LB=0.52x</li>\n</ul>\n<p>So basically, I noticed that there is a <strong>big correlation between CV and LB</strong>. In my experiments for CV &lt; 0.55 the gap is around 0.01x and when CV&gt;0.60 the gap is around 0.03x.</p>\n<p><strong>Happy Kaggling !!!</strong></p>",
  "messages": [
    {
      "id": 2778613,
      "postDate": "2024-04-27T08:53:31.540Z",
      "content": "<p>Hi Everyone. I hope you're doing well in this competition. I'm a little bit curious to see how things goes for you concerning CV and LB gaps. Feel free to share here.</p>\n<p>I start first by sharing two experiments:</p>\n<ul>\n<li>A simple Xgboost : CV=0.43x / LB = 0.44x</li>\n<li>A simple FCN : CV=0.51x / LB=0.52x</li>\n</ul>\n<p>So basically, I noticed that there is a <strong>big correlation between CV and LB</strong>. In my experiments for CV &lt; 0.55 the gap is around 0.01x and when CV&gt;0.60 the gap is around 0.03x.</p>\n<p><strong>Happy Kaggling !!!</strong></p>",
      "rawMarkdown": "Hi Everyone. I hope you're doing well in this competition. I'm a little bit curious to see how things goes for you concerning CV and LB gaps. Feel free to share here.\n\nI start first by sharing two experiments:\n* A simple Xgboost : CV=0.43x / LB = 0.44x\n* A simple FCN : CV=0.51x / LB=0.52x\n\nSo basically, I noticed that there is a **big correlation between CV and LB**. In my experiments for CV < 0.55 the gap is around 0.01x and when CV>0.60 the gap is around 0.03x.\n\n\n**Happy Kaggling !!!**",
      "votes": 22
    },
    {
      "id": 2778710,
      "postDate": "2024-04-27T09:42:40.823Z",
      "content": "<p>For constant or unpredictable targets I assign R2=0, for the rest I compute per-class R2 score and then take average of all scores,<br>\nthe validation is done on last 2* million samples.</p>\n<table>\n<thead>\n<tr>\n<th>CV</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.38753</td>\n<td>0.56053</td>\n</tr>\n<tr>\n<td>0.41259</td>\n<td>0.58585</td>\n</tr>\n<tr>\n<td>0.43521</td>\n<td>0.60716</td>\n</tr>\n<tr>\n<td>0.45459</td>\n<td>0.62662</td>\n</tr>\n</tbody>\n</table>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5487737%2Fee1d48213af364532dc2ffdb67357c94%2Fleap_corr.png?generation=1714210918935331&amp;alt=media\"></p>\n<p>I expect this competition have little to none shakeup.</p>",
      "rawMarkdown": "For constant or unpredictable targets I assign R2=0, for the rest I compute per-class R2 score and then take average of all scores,\nthe validation is done on last 2* million samples.\n\n| CV | LB |\n| --- | --- |\n| 0.38753  | 0.56053 |\n| 0.41259 | 0.58585 |\n| 0.43521 | 0.60716  | \n| 0.45459| 0.62662 |\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5487737%2Fee1d48213af364532dc2ffdb67357c94%2Fleap_corr.png?generation=1714210918935331&alt=media)\n\nI expect this competition have little to none shakeup.",
      "votes": 7,
      "replies": [
        {
          "id": 2778721,
          "postDate": "2024-04-27T09:46:34.183Z",
          "content": "<p>Yes you're right. Actual R2 is inflated because for the constant var, predicting the same constant yiels an R2= 1 just inflating the computations.</p>",
          "rawMarkdown": "Yes you're right. Actual R2 is inflated because for the constant var, predicting the same constant yiels an R2= 1 just inflating the computations.",
          "votes": 5
        },
        {
          "id": 2778756,
          "postDate": "2024-04-27T10:27:15.437Z",
          "content": "<p>validation on the last 10M samples or 1M samples? Thanks for sharing this insightful result.</p>",
          "rawMarkdown": "validation on the last 10M samples or 1M samples? Thanks for sharing this insightful result.\n",
          "votes": 3,
          "replies": [
            {
              "id": 2778759,
              "postDate": "2024-04-27T10:28:16.270Z",
              "content": "<p>Validation on 500 000 rows</p>",
              "rawMarkdown": "Validation on 500 000 rows",
              "votes": 4
            }
          ]
        },
        {
          "id": 2778931,
          "postDate": "2024-04-27T13:04:05.413Z",
          "content": "<p><a href=\"https://www.kaggle.com/martynoveduard\" target=\"_blank\">@martynoveduard</a>, please what is meant by unpredictable targets, the constant targets is perhaps those with 0 as sample_submission(scale) ?<br>\nThanks</p>",
          "rawMarkdown": "@martynoveduard, please what is meant by unpredictable targets, the constant targets is perhaps those with 0 as sample_submission(scale) ?\nThanks",
          "replies": [
            {
              "id": 2778992,
              "postDate": "2024-04-27T13:35:02.110Z",
              "content": "<p>I didn't dig deep yet, but I found some targets that are neither constant or predictable (i.e. for these columns it's impossible to obtain R2 score &gt; 0), I called them unpredictable </p>",
              "rawMarkdown": "I didn't dig deep yet, but I found some targets that are neither constant or predictable (i.e. for these columns it's impossible to obtain R2 score > 0), I called them unpredictable "
            }
          ]
        }
      ]
    },
    {
      "id": 2779144,
      "postDate": "2024-04-27T14:43:53.600Z",
      "content": "<p>In my experiments CV-LB are well correlated. </p>\n<table>\n<thead>\n<tr>\n<th>CV</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.6274</td>\n<td>0.6241</td>\n</tr>\n<tr>\n<td>0.6197</td>\n<td>0.6164</td>\n</tr>\n<tr>\n<td>0.6075</td>\n<td>0.6081</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "In my experiments CV-LB are well correlated. \n| CV | LB |\n| --- | --- |\n| 0.6274 | 0.6241 |\n| 0.6197 | 0.6164 |\n| 0.6075 | 0.6081 |",
      "votes": 6
    },
    {
      "id": 2826676,
      "postDate": "2024-05-21T04:34:27.743Z",
      "content": "<p>Use the last 650,000 records as validation.<br>\nCV: 0.765<br>\nLB: 0.756</p>",
      "rawMarkdown": "Use the last 650,000 records as validation.\nCV: 0.765\nLB: 0.756",
      "votes": 3
    },
    {
      "id": 2828263,
      "postDate": "2024-05-22T02:16:39.737Z",
      "content": "<p>How is target with value ==0 calculated? R2 equals 1 for them? It seems to be impossible to predict ptend_q0002_15~ptend_q0002_26, and they make the cv score very bad …</p>",
      "rawMarkdown": "How is target with value ==0 calculated? R2 equals 1 for them? It seems to be impossible to predict ptend_q0002_15~ptend_q0002_26, and they make the cv score very bad ...",
      "votes": 1,
      "replies": [
        {
          "id": 2828383,
          "postDate": "2024-05-22T04:52:53.913Z",
          "content": "<p>You can have perfect scores for those:<br>\n<a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/502484\" target=\"_blank\">https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/502484</a></p>",
          "rawMarkdown": "You can have perfect scores for those:\nhttps://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/502484",
          "replies": [
            {
              "id": 2828432,
              "postDate": "2024-05-22T05:19:27.700Z",
              "content": "<p>Thank you for the info, from the metric calculation, it seems that we get 1 for 0s. But the correlation between local CV and LB seems to be perfect, and my CV skips the columns with all 0. </p>",
              "rawMarkdown": "Thank you for the info, from the metric calculation, it seems that we get 1 for 0s. But the correlation between local CV and LB seems to be perfect, and my CV skips the columns with all 0. "
            }
          ]
        }
      ]
    },
    {
      "id": 2778889,
      "postDate": "2024-04-27T12:29:35.387Z",
      "content": "<p>Excuse the newbie question: I know LB stands for Leadboard [score], and the competition brief explains how that is calculated, but what exactly is CV in this context? Thanks!</p>",
      "rawMarkdown": "Excuse the newbie question: I know LB stands for Leadboard [score], and the competition brief explains how that is calculated, but what exactly is CV in this context? Thanks!",
      "votes": 1,
      "replies": [
        {
          "id": 2779492,
          "postDate": "2024-04-27T17:11:37.537Z",
          "content": "<p>The CV is the score on your validation dagta. It is supposed to be <em>proper cross validation score</em>, but sometimes it is the score on one validation fold. I hope it mates it clear</p>",
          "rawMarkdown": "The CV is the score on your validation dagta. It is supposed to be *proper cross validation score*, but sometimes it is the score on one validation fold. I hope it mates it clear",
          "votes": 7,
          "replies": [
            {
              "id": 2779538,
              "postDate": "2024-04-27T17:29:27.430Z",
              "content": "<p>Ah thanks, Googling it seemed to be related to cross-validation, but as that was a process not a metric I thought it must be something else! (Any idea what metric \"CV\" is then, anyone, out of curiosity? -- I don't need to know, was just plugging my lack of knowledge)</p>",
              "rawMarkdown": "Ah thanks, Googling it seemed to be related to cross-validation, but as that was a process not a metric I thought it must be something else! (Any idea what metric \"CV\" is then, anyone, out of curiosity? -- I don't need to know, was just plugging my lack of knowledge)",
              "votes": 1
            },
            {
              "id": 2781041,
              "postDate": "2024-04-28T14:21:04.960Z",
              "content": "<p>\"CV score\" means \"score parsed on the validation data\".</p>",
              "rawMarkdown": "\"CV score\" means \"score parsed on the validation data\"."
            },
            {
              "id": 2781524,
              "postDate": "2024-04-28T20:27:10.070Z",
              "content": "<p>Thanks -- but what does 'score' mean mathematically?</p>",
              "rawMarkdown": "Thanks -- but what does 'score' mean mathematically?"
            },
            {
              "id": 2781636,
              "postDate": "2024-04-28T22:27:35.533Z",
              "content": "<p>Same 'score' (metric) as LB…</p>",
              "rawMarkdown": "Same 'score' (metric) as LB..."
            }
          ]
        },
        {
          "id": 2779501,
          "postDate": "2024-04-27T17:15:27.510Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 2830030,
      "postDate": "2024-05-23T00:57:44.230Z",
      "content": "<p>current my best model:</p>\n<p>CV: 0.7657<br>\nLB: 0.76360</p>",
      "rawMarkdown": "current my best model:\n\nCV: 0.7657\nLB: 0.76360",
      "votes": 2
    },
    {
      "id": 2789401,
      "postDate": "2024-05-02T16:54:14.153Z",
      "content": "<p>Good to know</p>",
      "rawMarkdown": "Good to know"
    },
    {
      "id": 2788978,
      "postDate": "2024-05-02T13:52:03.647Z",
      "content": "<p>I've got quiet the deflated LB score compared to my CV score. CV around ~.6 dropping to ~.5 regularly. </p>\n<p>*for reference, i'll define:</p>\n<ul>\n<li>original scale \"_o\"</li>\n<li>normalized variables \"_n\" </li>\n<li>scaled by submission weights \"_s\"</li>\n</ul>\n<p>In summary, what I'm doing is…</p>\n<ol>\n<li>y_o -&gt; y_n</li>\n<li>model(y_n) -&gt; y_pred_n</li>\n<li>y_pred_n -&gt; y_pred_o</li>\n<li>r2_score for (y_o, y_pred_o)</li>\n<li>replace y_pred_o with mean of y_o where r2 is negative<br>\n    * This is where I perform cross val </li>\n<li>y_pred_o -&gt; y_pred_s</li>\n</ol>\n<p>Should I be replacing y_pred_n with zero, or calculating y_train_s means and replacing y_pred_s at that point?</p>",
      "rawMarkdown": "I've got quiet the deflated LB score compared to my CV score. CV around ~.6 dropping to ~.5 regularly. \n\n*for reference, i'll define:\n - original scale \"_o\"\n - normalized variables \"_n\" \n - scaled by submission weights \"_s\"\n\nIn summary, what I'm doing is...\n1. y_o -> y_n\n2. model(y_n) -> y_pred_n\n3. y_pred_n -> y_pred_o\n4. r2_score for (y_o, y_pred_o)\n5. replace y_pred_o with mean of y_o where r2 is negative\n        * This is where I perform cross val \n6. y_pred_o -> y_pred_s\n\nShould I be replacing y_pred_n with zero, or calculating y_train_s means and replacing y_pred_s at that point?"
    }
  ],
  "comments": [
    {
      "id": 2778710,
      "author_name": "slime",
      "author_url": "",
      "post_date": "2024-04-27T09:42:40.823000",
      "content": "<p>For constant or unpredictable targets I assign R2=0, for the rest I compute per-class R2 score and then take average of all scores,<br>\nthe validation is done on last 2* million samples.</p>\n<table>\n<thead>\n<tr>\n<th>CV</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.38753</td>\n<td>0.56053</td>\n</tr>\n<tr>\n<td>0.41259</td>\n<td>0.58585</td>\n</tr>\n<tr>\n<td>0.43521</td>\n<td>0.60716</td>\n</tr>\n<tr>\n<td>0.45459</td>\n<td>0.62662</td>\n</tr>\n</tbody>\n</table>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5487737%2Fee1d48213af364532dc2ffdb67357c94%2Fleap_corr.png?generation=1714210918935331&amp;alt=media\"></p>\n<p>I expect this competition have little to none shakeup.</p>",
      "votes": 7,
      "replies": [
        {
          "id": 2778721,
          "author_name": "Ulrich G.",
          "author_url": "",
          "post_date": "2024-04-27T09:46:34.183000",
          "content": "<p>Yes you're right. Actual R2 is inflated because for the constant var, predicting the same constant yiels an R2= 1 just inflating the computations.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 2778756,
          "author_name": "HungryLearner",
          "author_url": "",
          "post_date": "2024-04-27T10:27:15.437000",
          "content": "<p>validation on the last 10M samples or 1M samples? Thanks for sharing this insightful result.</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2778759,
              "author_name": "Ulrich G.",
              "author_url": "",
              "post_date": "2024-04-27T10:28:16.270000",
              "content": "<p>Validation on 500 000 rows</p>",
              "votes": 4,
              "replies": []
            }
          ]
        },
        {
          "id": 2778931,
          "author_name": "HungryLearner",
          "author_url": "",
          "post_date": "2024-04-27T13:04:05.413000",
          "content": "<p><a href=\"https://www.kaggle.com/martynoveduard\" target=\"_blank\">@martynoveduard</a>, please what is meant by unpredictable targets, the constant targets is perhaps those with 0 as sample_submission(scale) ?<br>\nThanks</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2778992,
              "author_name": "slime",
              "author_url": "",
              "post_date": "2024-04-27T13:35:02.110000",
              "content": "<p>I didn't dig deep yet, but I found some targets that are neither constant or predictable (i.e. for these columns it's impossible to obtain R2 score &gt; 0), I called them unpredictable </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2779144,
      "author_name": "Giba",
      "author_url": "",
      "post_date": "2024-04-27T14:43:53.600000",
      "content": "<p>In my experiments CV-LB are well correlated. </p>\n<table>\n<thead>\n<tr>\n<th>CV</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.6274</td>\n<td>0.6241</td>\n</tr>\n<tr>\n<td>0.6197</td>\n<td>0.6164</td>\n</tr>\n<tr>\n<td>0.6075</td>\n<td>0.6081</td>\n</tr>\n</tbody>\n</table>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 2826676,
      "author_name": "phalanx",
      "author_url": "",
      "post_date": "2024-05-21T04:34:27.743000",
      "content": "<p>Use the last 650,000 records as validation.<br>\nCV: 0.765<br>\nLB: 0.756</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2828263,
      "author_name": "yuanzhe zhou",
      "author_url": "",
      "post_date": "2024-05-22T02:16:39.737000",
      "content": "<p>How is target with value ==0 calculated? R2 equals 1 for them? It seems to be impossible to predict ptend_q0002_15~ptend_q0002_26, and they make the cv score very bad …</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2828383,
          "author_name": "DennisSakva",
          "author_url": "",
          "post_date": "2024-05-22T04:52:53.913000",
          "content": "<p>You can have perfect scores for those:<br>\n<a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/502484\" target=\"_blank\">https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/502484</a></p>",
          "votes": 0,
          "replies": [
            {
              "id": 2828432,
              "author_name": "yuanzhe zhou",
              "author_url": "",
              "post_date": "2024-05-22T05:19:27.700000",
              "content": "<p>Thank you for the info, from the metric calculation, it seems that we get 1 for 0s. But the correlation between local CV and LB seems to be perfect, and my CV skips the columns with all 0. </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2778889,
      "author_name": "Charlie Wartnaby",
      "author_url": "",
      "post_date": "2024-04-27T12:29:35.387000",
      "content": "<p>Excuse the newbie question: I know LB stands for Leadboard [score], and the competition brief explains how that is calculated, but what exactly is CV in this context? Thanks!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2779492,
          "author_name": "Ulrich G.",
          "author_url": "",
          "post_date": "2024-04-27T17:11:37.537000",
          "content": "<p>The CV is the score on your validation dagta. It is supposed to be <em>proper cross validation score</em>, but sometimes it is the score on one validation fold. I hope it mates it clear</p>",
          "votes": 7,
          "replies": [
            {
              "id": 2779538,
              "author_name": "Charlie Wartnaby",
              "author_url": "",
              "post_date": "2024-04-27T17:29:27.430000",
              "content": "<p>Ah thanks, Googling it seemed to be related to cross-validation, but as that was a process not a metric I thought it must be something else! (Any idea what metric \"CV\" is then, anyone, out of curiosity? -- I don't need to know, was just plugging my lack of knowledge)</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2781041,
              "author_name": "Catadanna",
              "author_url": "",
              "post_date": "2024-04-28T14:21:04.960000",
              "content": "<p>\"CV score\" means \"score parsed on the validation data\".</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2781524,
              "author_name": "Charlie Wartnaby",
              "author_url": "",
              "post_date": "2024-04-28T20:27:10.070000",
              "content": "<p>Thanks -- but what does 'score' mean mathematically?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2781636,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2024-04-28T22:27:35.533000",
              "content": "<p>Same 'score' (metric) as LB…</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2779501,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-04-27T17:15:27.510000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2830030,
      "author_name": "Bilzard",
      "author_url": "",
      "post_date": "2024-05-23T00:57:44.230000",
      "content": "<p>current my best model:</p>\n<p>CV: 0.7657<br>\nLB: 0.76360</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2789401,
      "author_name": "CARMEN.AI",
      "author_url": "",
      "post_date": "2024-05-02T16:54:14.153000",
      "content": "<p>Good to know</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2788978,
      "author_name": "Jekasm19",
      "author_url": "",
      "post_date": "2024-05-02T13:52:03.647000",
      "content": "<p>I've got quiet the deflated LB score compared to my CV score. CV around ~.6 dropping to ~.5 regularly. </p>\n<p>*for reference, i'll define:</p>\n<ul>\n<li>original scale \"_o\"</li>\n<li>normalized variables \"_n\" </li>\n<li>scaled by submission weights \"_s\"</li>\n</ul>\n<p>In summary, what I'm doing is…</p>\n<ol>\n<li>y_o -&gt; y_n</li>\n<li>model(y_n) -&gt; y_pred_n</li>\n<li>y_pred_n -&gt; y_pred_o</li>\n<li>r2_score for (y_o, y_pred_o)</li>\n<li>replace y_pred_o with mean of y_o where r2 is negative<br>\n    * This is where I perform cross val </li>\n<li>y_pred_o -&gt; y_pred_s</li>\n</ol>\n<p>Should I be replacing y_pred_n with zero, or calculating y_train_s means and replacing y_pred_s at that point?</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2778613": "Hi Everyone. I hope you're doing well in this competition. I'm a little bit curious to see how things goes for you concerning CV and LB gaps. Feel free to share here.\n\nI start first by sharing two experiments:\n* A simple Xgboost : CV=0.43x / LB = 0.44x\n* A simple FCN : CV=0.51x / LB=0.52x\n\nSo basically, I noticed that there is a **big correlation between CV and LB**. In my experiments for CV < 0.55 the gap is around 0.01x and when CV>0.60 the gap is around 0.03x.\n\n\n**Happy Kaggling !!!**",
    "2778710": "For constant or unpredictable targets I assign R2=0, for the rest I compute per-class R2 score and then take average of all scores,\nthe validation is done on last 2* million samples.\n\n| CV | LB |\n| --- | --- |\n| 0.38753  | 0.56053 |\n| 0.41259 | 0.58585 |\n| 0.43521 | 0.60716  | \n| 0.45459| 0.62662 |\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5487737%2Fee1d48213af364532dc2ffdb67357c94%2Fleap_corr.png?generation=1714210918935331&alt=media)\n\nI expect this competition have little to none shakeup.",
    "2779144": "In my experiments CV-LB are well correlated. \n| CV | LB |\n| --- | --- |\n| 0.6274 | 0.6241 |\n| 0.6197 | 0.6164 |\n| 0.6075 | 0.6081 |",
    "2826676": "Use the last 650,000 records as validation.\nCV: 0.765\nLB: 0.756",
    "2828263": "How is target with value ==0 calculated? R2 equals 1 for them? It seems to be impossible to predict ptend_q0002_15~ptend_q0002_26, and they make the cv score very bad ...",
    "2778889": "Excuse the newbie question: I know LB stands for Leadboard [score], and the competition brief explains how that is calculated, but what exactly is CV in this context? Thanks!",
    "2830030": "current my best model:\n\nCV: 0.7657\nLB: 0.76360",
    "2789401": "Good to know",
    "2788978": "I've got quiet the deflated LB score compared to my CV score. CV around ~.6 dropping to ~.5 regularly. \n\n*for reference, i'll define:\n - original scale \"_o\"\n - normalized variables \"_n\" \n - scaled by submission weights \"_s\"\n\nIn summary, what I'm doing is...\n1. y_o -> y_n\n2. model(y_n) -> y_pred_n\n3. y_pred_n -> y_pred_o\n4. r2_score for (y_o, y_pred_o)\n5. replace y_pred_o with mean of y_o where r2 is negative\n        * This is where I perform cross val \n6. y_pred_o -> y_pred_s\n\nShould I be replacing y_pred_n with zero, or calculating y_train_s means and replacing y_pred_s at that point?"
  }
}