{
  "id": 496558,
  "title": "should we drop single value cols ?",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/496558",
  "author_name": "Satej Raste",
  "post_date": "2024-04-21T15:14:34.376000",
  "votes": 4,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Should we drop columns with only 1 unique value if not what is relevance of that column for example many columns in pbuf series are same for train and test then why should i even bother to add them since it only increase my data size </p>",
  "messages": [
    {
      "id": 2777683,
      "postDate": "2024-04-26T19:07:52.213Z",
      "content": "<p>If you consider them as sequences of 60, you should not delete them. If you consider the data as tabular, one should delete them. </p>",
      "rawMarkdown": "If you consider them as sequences of 60, you should not delete them. If you consider the data as tabular, one should delete them. ",
      "votes": 4
    },
    {
      "id": 2766252,
      "postDate": "2024-04-21T15:14:34.377Z",
      "content": "<p>Should we drop columns with only 1 unique value if not what is relevance of that column for example many columns in pbuf series are same for train and test then why should i even bother to add them since it only increase my data size </p>",
      "rawMarkdown": "Should we drop columns with only 1 unique value if not what is relevance of that column for example many columns in pbuf series are same for train and test then why should i even bother to add them since it only increase my data size ",
      "votes": 4
    },
    {
      "id": 2767236,
      "postDate": "2024-04-22T07:34:40.923Z",
      "content": "<p>I think columns can be dropped  if same unique value is present in both test and train column.</p>",
      "rawMarkdown": "I think columns can be dropped  if same unique value is present in both test and train column.",
      "votes": 1,
      "replies": [
        {
          "id": 2767243,
          "postDate": "2024-04-22T07:42:26.197Z",
          "content": "<p>i tried that in some iterations but the performance drops for neural network i am still trying to figure the rational behind it </p>",
          "rawMarkdown": "i tried that in some iterations but the performance drops for neural network i am still trying to figure the rational behind it \n",
          "replies": [
            {
              "id": 2767347,
              "postDate": "2024-04-22T09:15:26.193Z",
              "content": "<p>Is there a specific reason this column is included? Intuitively, it seems like it might not be contributing much.</p>",
              "rawMarkdown": "Is there a specific reason this column is included? Intuitively, it seems like it might not be contributing much.",
              "votes": 1
            },
            {
              "id": 2767414,
              "postDate": "2024-04-22T09:52:15.573Z",
              "content": "<p>not sure about this but i did some further analysis and found this many columns you can check my notebook here <br>\n<a href=\"https://www.kaggle.com/code/starcs2001/data-reduction-notebook\" target=\"_blank\">https://www.kaggle.com/code/starcs2001/data-reduction-notebook</a></p>",
              "rawMarkdown": "not sure about this but i did some further analysis and found this many columns you can check my notebook here \nhttps://www.kaggle.com/code/starcs2001/data-reduction-notebook\n"
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2777683,
      "author_name": "Catadanna",
      "author_url": "",
      "post_date": "2024-04-26T19:07:52.213000",
      "content": "<p>If you consider them as sequences of 60, you should not delete them. If you consider the data as tabular, one should delete them. </p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 2767236,
      "author_name": "Chinmaya",
      "author_url": "",
      "post_date": "2024-04-22T07:34:40.923000",
      "content": "<p>I think columns can be dropped  if same unique value is present in both test and train column.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2767243,
          "author_name": "Satej Raste",
          "author_url": "",
          "post_date": "2024-04-22T07:42:26.197000",
          "content": "<p>i tried that in some iterations but the performance drops for neural network i am still trying to figure the rational behind it </p>",
          "votes": 0,
          "replies": [
            {
              "id": 2767347,
              "author_name": "Chinmaya",
              "author_url": "",
              "post_date": "2024-04-22T09:15:26.193000",
              "content": "<p>Is there a specific reason this column is included? Intuitively, it seems like it might not be contributing much.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2767414,
              "author_name": "Satej Raste",
              "author_url": "",
              "post_date": "2024-04-22T09:52:15.573000",
              "content": "<p>not sure about this but i did some further analysis and found this many columns you can check my notebook here <br>\n<a href=\"https://www.kaggle.com/code/starcs2001/data-reduction-notebook\" target=\"_blank\">https://www.kaggle.com/code/starcs2001/data-reduction-notebook</a></p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2777683": "If you consider them as sequences of 60, you should not delete them. If you consider the data as tabular, one should delete them. ",
    "2766252": "Should we drop columns with only 1 unique value if not what is relevance of that column for example many columns in pbuf series are same for train and test then why should i even bother to add them since it only increase my data size ",
    "2767236": "I think columns can be dropped  if same unique value is present in both test and train column."
  }
}