{
  "id": 495301,
  "title": "Many columns with one unique value in the training set.",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/495301",
  "author_name": "Catadanna",
  "post_date": "2024-04-20T13:38:53.858000",
  "votes": 11,
  "comment_count": 7,
  "views": 0,
  "content": "<p>It seems that in the training set there are a lot of features with only one unique value. <br>\nI made a notebook with an example code which illustrates that for one feature.<br>\nYou can find it <a href=\"https://www.kaggle.com/code/catadanna/leap-get-features-with-one-unique-value\" target=\"_blank\">here</a>.</p>\n<p>These columns should be deleted for training in order to get a better result.</p>\n<p><strong>UPDATE</strong></p>\n<p>I add <a href=\"https://www.kaggle.com/code/catadanna/columns-to-delete/notebook\" target=\"_blank\">here</a> a notebook which selects the columns with only one unique value.</p>",
  "messages": [
    {
      "id": 2763345,
      "postDate": "2024-04-20T13:38:53.860Z",
      "content": "<p>It seems that in the training set there are a lot of features with only one unique value. <br>\nI made a notebook with an example code which illustrates that for one feature.<br>\nYou can find it <a href=\"https://www.kaggle.com/code/catadanna/leap-get-features-with-one-unique-value\" target=\"_blank\">here</a>.</p>\n<p>These columns should be deleted for training in order to get a better result.</p>\n<p><strong>UPDATE</strong></p>\n<p>I add <a href=\"https://www.kaggle.com/code/catadanna/columns-to-delete/notebook\" target=\"_blank\">here</a> a notebook which selects the columns with only one unique value.</p>",
      "rawMarkdown": "It seems that in the training set there are a lot of features with only one unique value. \nI made a notebook with an example code which illustrates that for one feature.\nYou can find it [here](https://www.kaggle.com/code/catadanna/leap-get-features-with-one-unique-value).\n\nThese columns should be deleted for training in order to get a better result.\n\n**UPDATE**\n\nI add [here](https://www.kaggle.com/code/catadanna/columns-to-delete/notebook) a notebook which selects the columns with only one unique value.",
      "votes": 10
    },
    {
      "id": 2764101,
      "postDate": "2024-04-20T22:31:46.203Z",
      "content": "<p>i think all 60 values of each vector input should be interpreted as a series. If some elements of the series are always 0, they still should be kept for processing like CNN or RNN. Otherwise they could be dropped.</p>",
      "rawMarkdown": "i think all 60 values of each vector input should be interpreted as a series. If some elements of the series are always 0, they still should be kept for processing like CNN or RNN. Otherwise they could be dropped.",
      "votes": 7,
      "replies": [
        {
          "id": 2764111,
          "postDate": "2024-04-20T23:02:40.990Z",
          "content": "<p>Good point.</p>",
          "rawMarkdown": "Good point."
        },
        {
          "id": 2765506,
          "postDate": "2024-04-21T07:01:14.287Z",
          "content": "<p>I agree that they should be treated as series, maybe some kind of sequential modelling could be useful. I think we can still drop them if they always have the same value within different samples.</p>",
          "rawMarkdown": "I agree that they should be treated as series, maybe some kind of sequential modelling could be useful. I think we can still drop them if they always have the same value within different samples.",
          "votes": 2
        }
      ]
    },
    {
      "id": 2763748,
      "postDate": "2024-04-20T18:04:57.763Z",
      "content": "<p>I noticed that too. Maybe this is happening due to 32-bit float downcasting.</p>",
      "rawMarkdown": "I noticed that too. Maybe this is happening due to 32-bit float downcasting.",
      "replies": [
        {
          "id": 2764110,
          "postDate": "2024-04-20T23:01:13.607Z",
          "content": "<p>Was there a downcast though?</p>",
          "rawMarkdown": "Was there a downcast though?",
          "replies": [
            {
              "id": 2765284,
              "postDate": "2024-04-21T04:46:04.373Z",
              "content": "<p>I'm not able to load all columns as 64-bit floats so I do it myself and I assumed that she would also do that.</p>",
              "rawMarkdown": "I'm not able to load all columns as 64-bit floats so I do it myself and I assumed that she would also do that.",
              "votes": 1
            }
          ]
        },
        {
          "id": 2765739,
          "postDate": "2024-04-21T10:27:33.547Z",
          "content": "<p>In the related notebook (see the link) there is no downcasting so we have float64.</p>",
          "rawMarkdown": "In the related notebook (see the link) there is no downcasting so we have float64."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2764101,
      "author_name": "Youri Matiounine",
      "author_url": "",
      "post_date": "2024-04-20T22:31:46.203000",
      "content": "<p>i think all 60 values of each vector input should be interpreted as a series. If some elements of the series are always 0, they still should be kept for processing like CNN or RNN. Otherwise they could be dropped.</p>",
      "votes": 7,
      "replies": [
        {
          "id": 2764111,
          "author_name": "greySnow",
          "author_url": "",
          "post_date": "2024-04-20T23:02:40.990000",
          "content": "<p>Good point.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2765506,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2024-04-21T07:01:14.287000",
          "content": "<p>I agree that they should be treated as series, maybe some kind of sequential modelling could be useful. I think we can still drop them if they always have the same value within different samples.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2763748,
      "author_name": "Gunes Evitan",
      "author_url": "",
      "post_date": "2024-04-20T18:04:57.763000",
      "content": "<p>I noticed that too. Maybe this is happening due to 32-bit float downcasting.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2764110,
          "author_name": "greySnow",
          "author_url": "",
          "post_date": "2024-04-20T23:01:13.607000",
          "content": "<p>Was there a downcast though?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2765284,
              "author_name": "Gunes Evitan",
              "author_url": "",
              "post_date": "2024-04-21T04:46:04.373000",
              "content": "<p>I'm not able to load all columns as 64-bit floats so I do it myself and I assumed that she would also do that.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 2765739,
          "author_name": "Catadanna",
          "author_url": "",
          "post_date": "2024-04-21T10:27:33.547000",
          "content": "<p>In the related notebook (see the link) there is no downcasting so we have float64.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2763345": "It seems that in the training set there are a lot of features with only one unique value. \nI made a notebook with an example code which illustrates that for one feature.\nYou can find it [here](https://www.kaggle.com/code/catadanna/leap-get-features-with-one-unique-value).\n\nThese columns should be deleted for training in order to get a better result.\n\n**UPDATE**\n\nI add [here](https://www.kaggle.com/code/catadanna/columns-to-delete/notebook) a notebook which selects the columns with only one unique value.",
    "2764101": "i think all 60 values of each vector input should be interpreted as a series. If some elements of the series are always 0, they still should be kept for processing like CNN or RNN. Otherwise they could be dropped.",
    "2763748": "I noticed that too. Maybe this is happening due to 32-bit float downcasting."
  }
}