{
  "id": 495316,
  "title": "Sharing some starter artefacts for convenience",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/495316",
  "author_name": "Ravi Ramakrishnan",
  "post_date": "2024-04-20T14:24:14.585000",
  "votes": 5,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hello everyone,</p>\n<p>I am sure that this competition is likely to posit significant challenges for data collation and wrangling considering the large size of the training data. I have thereby prepared 7.5 million rows of data as below-</p>\n<ol>\n<li>I collated a dataset with batches of 5000 rows per chunk and 2.5 million records per folder</li>\n<li>I created 2 folders in the dataset, totally amounting to 5 million records in the dataset <a href=\"https://www.kaggle.com/datasets/ravi20076/climsim2024starterv1\" target=\"_blank\">here</a></li>\n</ol>\n<p>One may simply concatenate the requisite number of batches/ chunks as per his/ her requirements and proceed as a starter material. Please note that all the chunks preserve the train data structure with some datatype shrinkages. </p>\n<p>I also augmented the dataset with a <a href=\"https://www.kaggle.com/code/ravi20076/climsim2024-starterdata-v1\" target=\"_blank\">starter kernel</a> with additional 2.5 million records amounting to 7.5 million record collection, all in batches of 5000 per chunk. All files are named as <strong>Train_Batch{i}.parquet</strong>,  *i ranging from 0-499 per folder/ kernel output. <br>\nEach folder is approximately <strong>6.2 GB large</strong>, so use with caution. </p>\n<p>I also extended the idea discussed in <a href=\"https://www.kaggle.com/code/asarvazyan/leap-predict-the-mean\" target=\"_blank\">this kernel</a> to include this using polars and on my 7.5 million train data subset to yield a starter result. The submitted kernel is <a href=\"https://www.kaggle.com/code/ravi20076/climsim-starter-v2\" target=\"_blank\">here</a> for perusal. This kernel scores <strong>0.17117</strong> on the leaderboard. </p>\n<p>I purposely used polars for wrangling to illustrate the efficacy of the library herewith. </p>\n<p>Additionally, I have saved the sample submission file in parquet in my starter dataset, reducing the time taken to import and use it too hereby. Please feel free to use these elements in your pipelines too!</p>\n<p>Best wishes!</p>",
  "messages": [
    {
      "id": 2763410,
      "postDate": "2024-04-20T14:24:14.587Z",
      "content": "<p>Hello everyone,</p>\n<p>I am sure that this competition is likely to posit significant challenges for data collation and wrangling considering the large size of the training data. I have thereby prepared 7.5 million rows of data as below-</p>\n<ol>\n<li>I collated a dataset with batches of 5000 rows per chunk and 2.5 million records per folder</li>\n<li>I created 2 folders in the dataset, totally amounting to 5 million records in the dataset <a href=\"https://www.kaggle.com/datasets/ravi20076/climsim2024starterv1\" target=\"_blank\">here</a></li>\n</ol>\n<p>One may simply concatenate the requisite number of batches/ chunks as per his/ her requirements and proceed as a starter material. Please note that all the chunks preserve the train data structure with some datatype shrinkages. </p>\n<p>I also augmented the dataset with a <a href=\"https://www.kaggle.com/code/ravi20076/climsim2024-starterdata-v1\" target=\"_blank\">starter kernel</a> with additional 2.5 million records amounting to 7.5 million record collection, all in batches of 5000 per chunk. All files are named as <strong>Train_Batch{i}.parquet</strong>,  *i ranging from 0-499 per folder/ kernel output. <br>\nEach folder is approximately <strong>6.2 GB large</strong>, so use with caution. </p>\n<p>I also extended the idea discussed in <a href=\"https://www.kaggle.com/code/asarvazyan/leap-predict-the-mean\" target=\"_blank\">this kernel</a> to include this using polars and on my 7.5 million train data subset to yield a starter result. The submitted kernel is <a href=\"https://www.kaggle.com/code/ravi20076/climsim-starter-v2\" target=\"_blank\">here</a> for perusal. This kernel scores <strong>0.17117</strong> on the leaderboard. </p>\n<p>I purposely used polars for wrangling to illustrate the efficacy of the library herewith. </p>\n<p>Additionally, I have saved the sample submission file in parquet in my starter dataset, reducing the time taken to import and use it too hereby. Please feel free to use these elements in your pipelines too!</p>\n<p>Best wishes!</p>",
      "rawMarkdown": "Hello everyone,\n\nI am sure that this competition is likely to posit significant challenges for data collation and wrangling considering the large size of the training data. I have thereby prepared 7.5 million rows of data as below-\n1. I collated a dataset with batches of 5000 rows per chunk and 2.5 million records per folder\n2. I created 2 folders in the dataset, totally amounting to 5 million records in the dataset [here](https://www.kaggle.com/datasets/ravi20076/climsim2024starterv1)\n\nOne may simply concatenate the requisite number of batches/ chunks as per his/ her requirements and proceed as a starter material. Please note that all the chunks preserve the train data structure with some datatype shrinkages. \n\nI also augmented the dataset with a [starter kernel](https://www.kaggle.com/code/ravi20076/climsim2024-starterdata-v1) with additional 2.5 million records amounting to 7.5 million record collection, all in batches of 5000 per chunk. All files are named as **Train_Batch{i}.parquet**,  *i ranging from 0-499 per folder/ kernel output. \nEach folder is approximately **6.2 GB large**, so use with caution. \n\nI also extended the idea discussed in [this kernel](https://www.kaggle.com/code/asarvazyan/leap-predict-the-mean) to include this using polars and on my 7.5 million train data subset to yield a starter result. The submitted kernel is [here](https://www.kaggle.com/code/ravi20076/climsim-starter-v2) for perusal. This kernel scores **0.17117** on the leaderboard. \n\nI purposely used polars for wrangling to illustrate the efficacy of the library herewith. \n\nAdditionally, I have saved the sample submission file in parquet in my starter dataset, reducing the time taken to import and use it too hereby. Please feel free to use these elements in your pipelines too!\n\nBest wishes!",
      "votes": 5
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2763410": "Hello everyone,\n\nI am sure that this competition is likely to posit significant challenges for data collation and wrangling considering the large size of the training data. I have thereby prepared 7.5 million rows of data as below-\n1. I collated a dataset with batches of 5000 rows per chunk and 2.5 million records per folder\n2. I created 2 folders in the dataset, totally amounting to 5 million records in the dataset [here](https://www.kaggle.com/datasets/ravi20076/climsim2024starterv1)\n\nOne may simply concatenate the requisite number of batches/ chunks as per his/ her requirements and proceed as a starter material. Please note that all the chunks preserve the train data structure with some datatype shrinkages. \n\nI also augmented the dataset with a [starter kernel](https://www.kaggle.com/code/ravi20076/climsim2024-starterdata-v1) with additional 2.5 million records amounting to 7.5 million record collection, all in batches of 5000 per chunk. All files are named as **Train_Batch{i}.parquet**,  *i ranging from 0-499 per folder/ kernel output. \nEach folder is approximately **6.2 GB large**, so use with caution. \n\nI also extended the idea discussed in [this kernel](https://www.kaggle.com/code/asarvazyan/leap-predict-the-mean) to include this using polars and on my 7.5 million train data subset to yield a starter result. The submitted kernel is [here](https://www.kaggle.com/code/ravi20076/climsim-starter-v2) for perusal. This kernel scores **0.17117** on the leaderboard. \n\nI purposely used polars for wrangling to illustrate the efficacy of the library herewith. \n\nAdditionally, I have saved the sample submission file in parquet in my starter dataset, reducing the time taken to import and use it too hereby. Please feel free to use these elements in your pipelines too!\n\nBest wishes!"
  }
}