{
  "id": 497749,
  "title": "'Multiply by sample_submission'",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/497749",
  "author_name": "Jekasm19",
  "post_date": "2024-04-25T15:01:46.916000",
  "votes": 2,
  "comment_count": 4,
  "views": 0,
  "content": "<p>First of all, thank you to anyone who offers help!<br>\nI'm getting the massive negative score, and I can't seem to diagnose the problem.<br>\nI've tried simply multiplying element wise (outputs * sample_submission).to_csv('submission.csv')  , but to no avail.</p>\n<p>Also, what is meant by this?</p>\n<blockquote>\n  <p>This weighting can also be calculated without downloading this file. The top 12 levels (0-11 inclusive) for 'ptend_q0001', 'ptend_q0002', 'ptend_q0003', 'ptend_u', and 'ptend_v' can be ignored. For 'ptend_q0002', indices 12-14 (inclusive) can be ignored as well. These values are zeroed out by the prediction weightings in sample_submission.csv.&gt;</p>\n</blockquote>\n<p>I'm so unclear on this I don't even know where to start.<br>\nThank you!</p>\n<p>Update:<br>\nOk, so I found that the problem columns (from scoring within my notebook on validation set) are indeed those top 12 levels. I get this… but practically speaking what does it mean to \"ignore\" them? They are being graded in my final score and wrecking it, clearly not being zeroed out.</p>",
  "messages": [
    {
      "id": 2775888,
      "postDate": "2024-04-25T21:27:18.143Z",
      "content": "<p>Hi! </p>\n<p>Some targets are either constant or unpredictable, in order to identify them you must compute per-class R2 score on your validation set, once you do it and handle them (e.g. replace with their train-mean values), you should be able to get a positive score on LB</p>",
      "rawMarkdown": "Hi! \n\nSome targets are either constant or unpredictable, in order to identify them you must compute per-class R2 score on your validation set, once you do it and handle them (e.g. replace with their train-mean values), you should be able to get a positive score on LB",
      "votes": 1,
      "replies": [
        {
          "id": 2775963,
          "postDate": "2024-04-25T23:09:00.987Z",
          "content": "<p>Thanks! That’s the solution I ended up going with to get my ~.2 </p>",
          "rawMarkdown": "Thanks! That’s the solution I ended up going with to get my ~.2 ",
          "replies": [
            {
              "id": 2827492,
              "postDate": "2024-05-21T14:05:14.360Z",
              "content": "<p>I understand some targets are constants (probably authors have used mean values to replace nans), but what do you mean by \"unpredictable\"?</p>",
              "rawMarkdown": "I understand some targets are constants (probably authors have used mean values to replace nans), but what do you mean by \"unpredictable\"?\n"
            },
            {
              "id": 2827557,
              "postDate": "2024-05-21T14:49:09.857Z",
              "content": "<p>It seems to me that some target values in train.csv are bad, they have all values close to zero for some target variables. So I'll simply drop these lines</p>",
              "rawMarkdown": "It seems to me that some target values in train.csv are bad, they have all values close to zero for some target variables. So I'll simply drop these lines"
            }
          ]
        }
      ]
    },
    {
      "id": 2775239,
      "postDate": "2024-04-25T15:01:46.917Z",
      "content": "<p>First of all, thank you to anyone who offers help!<br>\nI'm getting the massive negative score, and I can't seem to diagnose the problem.<br>\nI've tried simply multiplying element wise (outputs * sample_submission).to_csv('submission.csv')  , but to no avail.</p>\n<p>Also, what is meant by this?</p>\n<blockquote>\n  <p>This weighting can also be calculated without downloading this file. The top 12 levels (0-11 inclusive) for 'ptend_q0001', 'ptend_q0002', 'ptend_q0003', 'ptend_u', and 'ptend_v' can be ignored. For 'ptend_q0002', indices 12-14 (inclusive) can be ignored as well. These values are zeroed out by the prediction weightings in sample_submission.csv.&gt;</p>\n</blockquote>\n<p>I'm so unclear on this I don't even know where to start.<br>\nThank you!</p>\n<p>Update:<br>\nOk, so I found that the problem columns (from scoring within my notebook on validation set) are indeed those top 12 levels. I get this… but practically speaking what does it mean to \"ignore\" them? They are being graded in my final score and wrecking it, clearly not being zeroed out.</p>",
      "rawMarkdown": "First of all, thank you to anyone who offers help!\nI'm getting the massive negative score, and I can't seem to diagnose the problem.\nI've tried simply multiplying element wise (outputs * sample_submission).to_csv('submission.csv')  , but to no avail.\n\nAlso, what is meant by this?\n\n>This weighting can also be calculated without downloading this file. The top 12 levels (0-11 inclusive) for 'ptend_q0001', 'ptend_q0002', 'ptend_q0003', 'ptend_u', and 'ptend_v' can be ignored. For 'ptend_q0002', indices 12-14 (inclusive) can be ignored as well. These values are zeroed out by the prediction weightings in sample_submission.csv.>\n\nI'm so unclear on this I don't even know where to start.\nThank you!\n\nUpdate:\nOk, so I found that the problem columns (from scoring within my notebook on validation set) are indeed those top 12 levels. I get this... but practically speaking what does it mean to \"ignore\" them? They are being graded in my final score and wrecking it, clearly not being zeroed out.",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 2775888,
      "author_name": "slime",
      "author_url": "",
      "post_date": "2024-04-25T21:27:18.143000",
      "content": "<p>Hi! </p>\n<p>Some targets are either constant or unpredictable, in order to identify them you must compute per-class R2 score on your validation set, once you do it and handle them (e.g. replace with their train-mean values), you should be able to get a positive score on LB</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2775963,
          "author_name": "Jekasm19",
          "author_url": "",
          "post_date": "2024-04-25T23:09:00.987000",
          "content": "<p>Thanks! That’s the solution I ended up going with to get my ~.2 </p>",
          "votes": 0,
          "replies": [
            {
              "id": 2827492,
              "author_name": "Markus Kaukonen",
              "author_url": "",
              "post_date": "2024-05-21T14:05:14.360000",
              "content": "<p>I understand some targets are constants (probably authors have used mean values to replace nans), but what do you mean by \"unpredictable\"?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2827557,
              "author_name": "Markus Kaukonen",
              "author_url": "",
              "post_date": "2024-05-21T14:49:09.857000",
              "content": "<p>It seems to me that some target values in train.csv are bad, they have all values close to zero for some target variables. So I'll simply drop these lines</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2775888": "Hi! \n\nSome targets are either constant or unpredictable, in order to identify them you must compute per-class R2 score on your validation set, once you do it and handle them (e.g. replace with their train-mean values), you should be able to get a positive score on LB",
    "2775239": "First of all, thank you to anyone who offers help!\nI'm getting the massive negative score, and I can't seem to diagnose the problem.\nI've tried simply multiplying element wise (outputs * sample_submission).to_csv('submission.csv')  , but to no avail.\n\nAlso, what is meant by this?\n\n>This weighting can also be calculated without downloading this file. The top 12 levels (0-11 inclusive) for 'ptend_q0001', 'ptend_q0002', 'ptend_q0003', 'ptend_u', and 'ptend_v' can be ignored. For 'ptend_q0002', indices 12-14 (inclusive) can be ignored as well. These values are zeroed out by the prediction weightings in sample_submission.csv.>\n\nI'm so unclear on this I don't even know where to start.\nThank you!\n\nUpdate:\nOk, so I found that the problem columns (from scoring within my notebook on validation set) are indeed those top 12 levels. I get this... but practically speaking what does it mean to \"ignore\" them? They are being graded in my final score and wrecking it, clearly not being zeroed out."
  }
}