{
  "id": 514800,
  "title": "Target preprocessing before training",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/514800",
  "author_name": "Federico Peccia",
  "post_date": "2024-06-25T14:48:49.865000",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I am interested in other ideas to preprocess the targets before the training procedure. Until now, I have been using just a standard normalization (as proposed in <a href=\"https://www.kaggle.com/code/abiolatti/keras-baseline-seq2seq\" target=\"_blank\">this notebook</a> from <a href=\"https://www.kaggle.com/abiolatti\" target=\"_blank\">@abiolatti</a> ). I also experimented with other clipping values for the std (the original notebook proposed 1e-10), but all other values gave worse results. I also experimented with the PowerTransformer and QuantileTransformer from sklearn, but they also did not worked very well.</p>\n<p>My hunch is that there must be a better transformation, that helps the model better learn the outliers in the high-risk target variables.</p>\n<p>I am interested in your thoughts on this!</p>\n<p>EDIT1: I am aware of the trick to replace variables q0002_12 to _27. Do you still use them in your model? Or do you just ignore them, and replace them afterwards?</p>\n<p>EDIT2: of course, the same question applies to the input variables. Are you using something different than a standard normalization?</p>",
  "messages": [
    {
      "id": 2895076,
      "postDate": "2024-06-28T22:51:40.023Z",
      "content": "<p>Mooers et al. 2021 used range normalization. 😉</p>\n<p><a href=\"https://agupubs.onlinelibrary.wiley.com/doi/full/10.1029/2020MS002385\" target=\"_blank\">https://agupubs.onlinelibrary.wiley.com/doi/full/10.1029/2020MS002385</a></p>",
      "rawMarkdown": "Mooers et al. 2021 used range normalization. 😉\n\nhttps://agupubs.onlinelibrary.wiley.com/doi/full/10.1029/2020MS002385",
      "votes": 1
    },
    {
      "id": 2895374,
      "postDate": "2024-06-29T06:29:03.267Z",
      "content": "<p>I tried gauss rank, z score, rms, min max and all of them on log scale but rms was the best.</p>",
      "rawMarkdown": "I tried gauss rank, z score, rms, min max and all of them on log scale but rms was the best.",
      "votes": 2
    },
    {
      "id": 2889541,
      "postDate": "2024-06-25T14:48:49.867Z",
      "content": "<p>I am interested in other ideas to preprocess the targets before the training procedure. Until now, I have been using just a standard normalization (as proposed in <a href=\"https://www.kaggle.com/code/abiolatti/keras-baseline-seq2seq\" target=\"_blank\">this notebook</a> from <a href=\"https://www.kaggle.com/abiolatti\" target=\"_blank\">@abiolatti</a> ). I also experimented with other clipping values for the std (the original notebook proposed 1e-10), but all other values gave worse results. I also experimented with the PowerTransformer and QuantileTransformer from sklearn, but they also did not worked very well.</p>\n<p>My hunch is that there must be a better transformation, that helps the model better learn the outliers in the high-risk target variables.</p>\n<p>I am interested in your thoughts on this!</p>\n<p>EDIT1: I am aware of the trick to replace variables q0002_12 to _27. Do you still use them in your model? Or do you just ignore them, and replace them afterwards?</p>\n<p>EDIT2: of course, the same question applies to the input variables. Are you using something different than a standard normalization?</p>",
      "rawMarkdown": "I am interested in other ideas to preprocess the targets before the training procedure. Until now, I have been using just a standard normalization (as proposed in [this notebook](https://www.kaggle.com/code/abiolatti/keras-baseline-seq2seq) from @abiolatti ). I also experimented with other clipping values for the std (the original notebook proposed 1e-10), but all other values gave worse results. I also experimented with the PowerTransformer and QuantileTransformer from sklearn, but they also did not worked very well.\n\nMy hunch is that there must be a better transformation, that helps the model better learn the outliers in the high-risk target variables.\n\nI am interested in your thoughts on this!\n\nEDIT1: I am aware of the trick to replace variables q0002_12 to _27. Do you still use them in your model? Or do you just ignore them, and replace them afterwards?\n\nEDIT2: of course, the same question applies to the input variables. Are you using something different than a standard normalization?",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 2895076,
      "author_name": "Jerry Lin",
      "author_url": "",
      "post_date": "2024-06-28T22:51:40.023000",
      "content": "<p>Mooers et al. 2021 used range normalization. 😉</p>\n<p><a href=\"https://agupubs.onlinelibrary.wiley.com/doi/full/10.1029/2020MS002385\" target=\"_blank\">https://agupubs.onlinelibrary.wiley.com/doi/full/10.1029/2020MS002385</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2895374,
      "author_name": "Gunes Evitan",
      "author_url": "",
      "post_date": "2024-06-29T06:29:03.267000",
      "content": "<p>I tried gauss rank, z score, rms, min max and all of them on log scale but rms was the best.</p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2895076": "Mooers et al. 2021 used range normalization. 😉\n\nhttps://agupubs.onlinelibrary.wiley.com/doi/full/10.1029/2020MS002385",
    "2895374": "I tried gauss rank, z score, rms, min max and all of them on log scale but rms was the best.",
    "2889541": "I am interested in other ideas to preprocess the targets before the training procedure. Until now, I have been using just a standard normalization (as proposed in [this notebook](https://www.kaggle.com/code/abiolatti/keras-baseline-seq2seq) from @abiolatti ). I also experimented with other clipping values for the std (the original notebook proposed 1e-10), but all other values gave worse results. I also experimented with the PowerTransformer and QuantileTransformer from sklearn, but they also did not worked very well.\n\nMy hunch is that there must be a better transformation, that helps the model better learn the outliers in the high-risk target variables.\n\nI am interested in your thoughts on this!\n\nEDIT1: I am aware of the trick to replace variables q0002_12 to _27. Do you still use them in your model? Or do you just ignore them, and replace them afterwards?\n\nEDIT2: of course, the same question applies to the input variables. Are you using something different than a standard normalization?"
  }
}