{
  "id": 519184,
  "title": "Hack the Planet (Simulator)",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/519184",
  "author_name": "bethegoose",
  "post_date": "2024-07-10T04:19:55.859000",
  "votes": 34,
  "comment_count": 46,
  "views": 0,
  "content": "<p>tl;dr you can perfectly assign the simulation tick and location for all test samples</p>\n<p>A number of features in the ClimSim dataset are fixed for a given (location, timestamp).<br>\nRegardless of simulation year, the value for that (location, timestamp, feature) is always the same.<br>\n<code>pbuf_ozone_2</code> is such a feature that is both fixed and a unique fingerprint value: there are no repeated <code>pbuf_ozone_2</code> measurements within a simulation year, but they always appear at the same time and place.<br>\nWhich means that it is possible to construct a mapping table of <code>pbuf_ozone_2</code> -&gt; (location, timestamp) and unblind all of the test samples.<br>\nThis requires 10,091,520 <code>pbuf_ozone_2</code> readings (384 locations * 365 days * 72ticks/day).</p>\n<h2>Praise the Sun</h2>\n<p>The sun is important to the weather system (insert citation here).<br>\nLuckily for us, <code>pbuf_SOLIN</code> (solar insolation) is also fixed for a simulation (location, timestamp).<br>\nWhich means that once we have derived the (location, timestamp) from <code>pbuf_ozone_2</code>, we now know the value of solar insolation at the previous tick.<br>\nAnd the next tick.<br>\nAnd the tick a week from now.</p>\n<p>(unlike <code>pbuf_ozone_2</code>, <code>pbuf_SOLIN</code> has repeated 0 values at night, so it cannot be used as a fingerprint feature)</p>\n<h2>Test Set Fun Facts</h2>\n<p>Original testset was sampled at 10 hour increments (Jan-01 06:00, Jan-01 16:00, Jan-01 02:00, Jan 02 12:00, etc).<br>\nReplacement testset was sampled at 8 hour increments (Jan-01 00:00, Jan-01 08:00, Jan-01 16:00, Jan-02 00:00, etc).</p>\n<p>This arrangement means that the original testset ~randomly sampled each location throughout the day (every even hour was hit).<br>\nThe new testset only touched each location at three of the possible ticks per day.</p>\n<h2>Fun Things to Try</h2>\n<ul>\n<li>Targeted training set. Why bother with those irrelevant timestamps? Only utilize 00:00, 08:00, and 16:00.</li>\n<li>Train by region. Latitudes appear to be well correlated.</li>\n<li>Global balance. For half of the year (some inference required) it is possible to build a global map incorporating the 384 locations.</li>\n<li>Timeseries - can align a feature for the present test value and the next one 8 hours in the future. By incorporating the original testset, could even have moments where the future location tick is only two hours in the future.</li>\n<li>Solar oracle. Incorporate additional features for the previous/next solar input tick as a sort of solar momentum</li>\n</ul>\n<h2>Script to Spot Check</h2>\n<p>To spot check a few simulation ticks can pull a few files from the low_res dataset.</p>\n<pre><code> datetime\n pathlib  Path\n\n netCDF4 \n pandas  pd\n requests\n xarray  xr\n\nTICK_NUMBER =  \nHUGGING_ROOT = \nDST_DIR = Path()\nDST_DIR.mkdir(exist_ok=)\n\n ():\n    \n    initial = datetime.datetime(year=, month=, day=)\n    ts = initial + datetime.timedelta(seconds=tick_number*(*))\n    \n    second_of_day = (ts.hour**) + (ts.minute*) + (ts.second)\n\n    \n    fname = \n    url = \n    dst = dst_dir / fname\n      dst.exists():\n        resp = requests.get(url, params={:})\n         (dst, )  fd:\n            fd.write(resp.content)\n\n\ndownload_tick(, TICK_NUMBER, DST_DIR)\ndownload_tick(, TICK_NUMBER, DST_DIR)\ndownload_tick(, TICK_NUMBER, DST_DIR)\n\n\nparts = []\n fpath  DST_DIR.glob():\n    part = xr.load_dataset(fpath).to_dataframe().reset_index(names=[, ])\n    part.insert(, , fpath.name)\n    sim_year = (fpath.name.split()[].split(, )[])\n    part.insert(, , sim_year)\n    parts.append(part)\ndf = pd.concat(parts)\ndf.shape\n\n\ndf = df.sort_values([, , , ])\n\n\nlev2 = df[df[]==]\n\nlev2.groupby([]).pbuf_ozone.nunique()\n\n\nlev2.pbuf_ozone.nunique()\n\n\ny7 = lev2[lev2[] == ]\ny8 = lev2[lev2[] == ]\n\nozone_matches = y7[] == y8[]\nozone_matches.value_counts()\n</code></pre>\n<h2>Script to Extract Fingerprint IDs</h2>\n<p>Given a directory containing the ClimSim low res dataset, can extract out the mapping values using some variation of the below.<br>\nThe ClimSim low res repository contain a definition file mapping location id to (latitude, longitude).</p>\n<pre><code> datetime\n pathlib  Path\n\n netCDF4 \n pandas  pd\n tqdm\n xarray  xr\n\n\nDIR_CLIMSIM = Path()\n\nparts = []\n year  [, ]: \n     month  (, ):\n        dir_month = DIR_CLIMSIM / \n         fpath  tqdm.tqdm(dir_month.glob()):\n            part = xr.load_dataset(fpath).to_dataframe().reset_index()\n            part[] = year\n            part[] = fpath.name\n            \n            bidx = part[].isin((, ))\n            parts.append(part[bidx][[, , , , , , , ]])\ndf = pd.concat(parts)\ndf = df.rename(columns={:, :, :})\ndf.shape\n\n\nsim_date = pd.to_datetime( + df[].astype().[-:], =)\ndf[] = sim_date + df[].apply( x: datetime.timedelta(seconds=x))\ndf[] = df.day_seconds // (*)\n\ndf = df.sort_values([, , , ])\ndf\n\ny7 = df[df[] == ]\ny8 = df[df[] == ]\ny7.shape, y8.shape\n\n\nperfect_matches = y7[y7[] == ][] == y8[y8[] == ][]\nperfect_matches.value_counts()\n\n\n\nperfect_matches = y7[y7[] == ][] == y8[y8[] == ][]\nperfect_matches.value_counts()\n\n\n\ny7_lev2_oz_uniques = y7[y7[] == ][].nunique()\ny8_lev2_oz_uniques = y8[y8[] == ][].nunique()\ny7_lev2_oz_uniques, y8_lev2_oz_uniques, y7[y7[] == ].shape[]\n\n\nchosen_day = df[df[] == datetime.datetime(year=, month=, day=, hour=)]\nchosen_day\n\n\n\nchosen_day[chosen_day[] == ][].nunique()\n\n\n\nmapping_table = df[df[] == ][[, , , ]].drop_duplicates()\nmapping_table.shape\n\n\nmapping_table.to_csv()\n</code></pre>",
  "messages": [
    {
      "id": 2914599,
      "postDate": "2024-07-10T04:19:55.860Z",
      "content": "<p>tl;dr you can perfectly assign the simulation tick and location for all test samples</p>\n<p>A number of features in the ClimSim dataset are fixed for a given (location, timestamp).<br>\nRegardless of simulation year, the value for that (location, timestamp, feature) is always the same.<br>\n<code>pbuf_ozone_2</code> is such a feature that is both fixed and a unique fingerprint value: there are no repeated <code>pbuf_ozone_2</code> measurements within a simulation year, but they always appear at the same time and place.<br>\nWhich means that it is possible to construct a mapping table of <code>pbuf_ozone_2</code> -&gt; (location, timestamp) and unblind all of the test samples.<br>\nThis requires 10,091,520 <code>pbuf_ozone_2</code> readings (384 locations * 365 days * 72ticks/day).</p>\n<h2>Praise the Sun</h2>\n<p>The sun is important to the weather system (insert citation here).<br>\nLuckily for us, <code>pbuf_SOLIN</code> (solar insolation) is also fixed for a simulation (location, timestamp).<br>\nWhich means that once we have derived the (location, timestamp) from <code>pbuf_ozone_2</code>, we now know the value of solar insolation at the previous tick.<br>\nAnd the next tick.<br>\nAnd the tick a week from now.</p>\n<p>(unlike <code>pbuf_ozone_2</code>, <code>pbuf_SOLIN</code> has repeated 0 values at night, so it cannot be used as a fingerprint feature)</p>\n<h2>Test Set Fun Facts</h2>\n<p>Original testset was sampled at 10 hour increments (Jan-01 06:00, Jan-01 16:00, Jan-01 02:00, Jan 02 12:00, etc).<br>\nReplacement testset was sampled at 8 hour increments (Jan-01 00:00, Jan-01 08:00, Jan-01 16:00, Jan-02 00:00, etc).</p>\n<p>This arrangement means that the original testset ~randomly sampled each location throughout the day (every even hour was hit).<br>\nThe new testset only touched each location at three of the possible ticks per day.</p>\n<h2>Fun Things to Try</h2>\n<ul>\n<li>Targeted training set. Why bother with those irrelevant timestamps? Only utilize 00:00, 08:00, and 16:00.</li>\n<li>Train by region. Latitudes appear to be well correlated.</li>\n<li>Global balance. For half of the year (some inference required) it is possible to build a global map incorporating the 384 locations.</li>\n<li>Timeseries - can align a feature for the present test value and the next one 8 hours in the future. By incorporating the original testset, could even have moments where the future location tick is only two hours in the future.</li>\n<li>Solar oracle. Incorporate additional features for the previous/next solar input tick as a sort of solar momentum</li>\n</ul>\n<h2>Script to Spot Check</h2>\n<p>To spot check a few simulation ticks can pull a few files from the low_res dataset.</p>\n<pre><code> datetime\n pathlib  Path\n\n netCDF4 \n pandas  pd\n requests\n xarray  xr\n\nTICK_NUMBER =  \nHUGGING_ROOT = \nDST_DIR = Path()\nDST_DIR.mkdir(exist_ok=)\n\n ():\n    \n    initial = datetime.datetime(year=, month=, day=)\n    ts = initial + datetime.timedelta(seconds=tick_number*(*))\n    \n    second_of_day = (ts.hour**) + (ts.minute*) + (ts.second)\n\n    \n    fname = \n    url = \n    dst = dst_dir / fname\n      dst.exists():\n        resp = requests.get(url, params={:})\n         (dst, )  fd:\n            fd.write(resp.content)\n\n\ndownload_tick(, TICK_NUMBER, DST_DIR)\ndownload_tick(, TICK_NUMBER, DST_DIR)\ndownload_tick(, TICK_NUMBER, DST_DIR)\n\n\nparts = []\n fpath  DST_DIR.glob():\n    part = xr.load_dataset(fpath).to_dataframe().reset_index(names=[, ])\n    part.insert(, , fpath.name)\n    sim_year = (fpath.name.split()[].split(, )[])\n    part.insert(, , sim_year)\n    parts.append(part)\ndf = pd.concat(parts)\ndf.shape\n\n\ndf = df.sort_values([, , , ])\n\n\nlev2 = df[df[]==]\n\nlev2.groupby([]).pbuf_ozone.nunique()\n\n\nlev2.pbuf_ozone.nunique()\n\n\ny7 = lev2[lev2[] == ]\ny8 = lev2[lev2[] == ]\n\nozone_matches = y7[] == y8[]\nozone_matches.value_counts()\n</code></pre>\n<h2>Script to Extract Fingerprint IDs</h2>\n<p>Given a directory containing the ClimSim low res dataset, can extract out the mapping values using some variation of the below.<br>\nThe ClimSim low res repository contain a definition file mapping location id to (latitude, longitude).</p>\n<pre><code> datetime\n pathlib  Path\n\n netCDF4 \n pandas  pd\n tqdm\n xarray  xr\n\n\nDIR_CLIMSIM = Path()\n\nparts = []\n year  [, ]: \n     month  (, ):\n        dir_month = DIR_CLIMSIM / \n         fpath  tqdm.tqdm(dir_month.glob()):\n            part = xr.load_dataset(fpath).to_dataframe().reset_index()\n            part[] = year\n            part[] = fpath.name\n            \n            bidx = part[].isin((, ))\n            parts.append(part[bidx][[, , , , , , , ]])\ndf = pd.concat(parts)\ndf = df.rename(columns={:, :, :})\ndf.shape\n\n\nsim_date = pd.to_datetime( + df[].astype().[-:], =)\ndf[] = sim_date + df[].apply( x: datetime.timedelta(seconds=x))\ndf[] = df.day_seconds // (*)\n\ndf = df.sort_values([, , , ])\ndf\n\ny7 = df[df[] == ]\ny8 = df[df[] == ]\ny7.shape, y8.shape\n\n\nperfect_matches = y7[y7[] == ][] == y8[y8[] == ][]\nperfect_matches.value_counts()\n\n\n\nperfect_matches = y7[y7[] == ][] == y8[y8[] == ][]\nperfect_matches.value_counts()\n\n\n\ny7_lev2_oz_uniques = y7[y7[] == ][].nunique()\ny8_lev2_oz_uniques = y8[y8[] == ][].nunique()\ny7_lev2_oz_uniques, y8_lev2_oz_uniques, y7[y7[] == ].shape[]\n\n\nchosen_day = df[df[] == datetime.datetime(year=, month=, day=, hour=)]\nchosen_day\n\n\n\nchosen_day[chosen_day[] == ][].nunique()\n\n\n\nmapping_table = df[df[] == ][[, , , ]].drop_duplicates()\nmapping_table.shape\n\n\nmapping_table.to_csv()\n</code></pre>",
      "rawMarkdown": "tl;dr you can perfectly assign the simulation tick and location for all test samples\n\nA number of features in the ClimSim dataset are fixed for a given (location, timestamp).\nRegardless of simulation year, the value for that (location, timestamp, feature) is always the same.\n``pbuf_ozone_2`` is such a feature that is both fixed and a unique fingerprint value: there are no repeated ``pbuf_ozone_2`` measurements within a simulation year, but they always appear at the same time and place.\nWhich means that it is possible to construct a mapping table of ``pbuf_ozone_2`` -> (location, timestamp) and unblind all of the test samples.\nThis requires 10,091,520 ``pbuf_ozone_2`` readings (384 locations * 365 days * 72ticks/day).\n\n## Praise the Sun\n\nThe sun is important to the weather system (insert citation here).\nLuckily for us, ``pbuf_SOLIN`` (solar insolation) is also fixed for a simulation (location, timestamp).\nWhich means that once we have derived the (location, timestamp) from ``pbuf_ozone_2``, we now know the value of solar insolation at the previous tick.\nAnd the next tick.\nAnd the tick a week from now.\n\n(unlike ``pbuf_ozone_2``, ``pbuf_SOLIN`` has repeated 0 values at night, so it cannot be used as a fingerprint feature)\n\n## Test Set Fun Facts\n\nOriginal testset was sampled at 10 hour increments (Jan-01 06:00, Jan-01 16:00, Jan-01 02:00, Jan 02 12:00, etc).\nReplacement testset was sampled at 8 hour increments (Jan-01 00:00, Jan-01 08:00, Jan-01 16:00, Jan-02 00:00, etc).\n\nThis arrangement means that the original testset ~randomly sampled each location throughout the day (every even hour was hit).\nThe new testset only touched each location at three of the possible ticks per day.\n\n## Fun Things to Try\n\n- Targeted training set. Why bother with those irrelevant timestamps? Only utilize 00:00, 08:00, and 16:00.\n- Train by region. Latitudes appear to be well correlated.\n- Global balance. For half of the year (some inference required) it is possible to build a global map incorporating the 384 locations.\n- Timeseries - can align a feature for the present test value and the next one 8 hours in the future. By incorporating the original testset, could even have moments where the future location tick is only two hours in the future.\n- Solar oracle. Incorporate additional features for the previous/next solar input tick as a sort of solar momentum\n\n## Script to Spot Check \n\nTo spot check a few simulation ticks can pull a few files from the low_res dataset.\n\n```python\nimport datetime\nfrom pathlib import Path\n\nimport netCDF4 # implicitly required by xarray\nimport pandas as pd\nimport requests\nimport xarray as xr\n\nTICK_NUMBER = 12_345 # whatever you want to check. tick is every 20 minutes; 72/day\nHUGGING_ROOT = \"https://huggingface.co/datasets/LEAP/ClimSim_low-res/resolve/main/train/\"\nDST_DIR = Path(\"climsim_ticks\")\nDST_DIR.mkdir(exist_ok=True)\n\ndef download_tick(sim_year:int, tick_number:int, dst_dir:Path):\n    # 2021 is arbitrary, but want non-leapyear\n    initial = datetime.datetime(year=2021, month=1, day=1)\n    ts = initial + datetime.timedelta(seconds=tick_number*(60*20))\n    ## seconds are always zero, but ¯\\_ (ツ)_/¯ \n    second_of_day = (ts.hour*60*60) + (ts.minute*60) + (ts.second)\n\n    # `E3SM-MMF.mli.0008-05-01-03600.nc`\n    fname = f\"E3SM-MMF.mli.{sim_year:04d}-{ts.month:02d}-{ts.day:02d}-{second_of_day:05d}.nc\"\n    url = f\"{HUGGING_ROOT}{sim_year:04d}-{ts.month:02d}/{fname}\"\n    dst = dst_dir / fname\n    if not dst.exists():\n        resp = requests.get(url, params={\"download\":\"true\"})\n        with open(dst, 'wb') as fd:\n            fd.write(resp.content)\n\n## examine same simulation tick across different simulation years\ndownload_tick(6, TICK_NUMBER, DST_DIR)\ndownload_tick(7, TICK_NUMBER, DST_DIR)\ndownload_tick(8, TICK_NUMBER, DST_DIR)\n\n## Build up the three samples\nparts = []\nfor fpath in DST_DIR.glob(\"*.nc\"):\n    part = xr.load_dataset(fpath).to_dataframe().reset_index(names=[\"climsim_location_id\", \"level\"])\n    part.insert(0, \"origin\", fpath.name)\n    sim_year = int(fpath.name.split(\".mli.\")[1].split(\"-\", 1)[0])\n    part.insert(1, \"simulation_year\", sim_year)\n    parts.append(part)\ndf = pd.concat(parts)\ndf.shape\n\n## ensure ordering\ndf = df.sort_values(['simulation_year', 'ymd', 'tod', 'level'])\n\n## grab just level 2\nlev2 = df[df[\"level\"]==2]\n## how many uniques per simulation year?\nlev2.groupby(['simulation_year']).pbuf_ozone.nunique()\n\n## how many uniques total?\nlev2.pbuf_ozone.nunique()\n\n## look at years specifically\ny7 = lev2[lev2[\"simulation_year\"] == 7]\ny8 = lev2[lev2[\"simulation_year\"] == 8]\n\nozone_matches = y7[\"pbuf_ozone\"] == y8[\"pbuf_ozone\"]\nozone_matches.value_counts()\n```\n\n## Script to Extract Fingerprint IDs\n\nGiven a directory containing the ClimSim low res dataset, can extract out the mapping values using some variation of the below.\nThe ClimSim low res repository contain a definition file mapping location id to (latitude, longitude).\n\n```python\nimport datetime\nfrom pathlib import Path\n\nimport netCDF4 # implicitly required by xarray to load .nc\nimport pandas as pd\nimport tqdm\nimport xarray as xr\n\n# https://huggingface.co/datasets/LEAP/ClimSim_low-res/ \nDIR_CLIMSIM = Path(\"ClimSim_low-res/train/\")\n\nparts = []\nfor year in [7, 8]: # available sim years are 01-Jan to 09-Jan\n    for month in range(1, 13):\n        dir_month = DIR_CLIMSIM / f\"{year:04d}-{month:02d}\"\n        for fpath in tqdm.tqdm(dir_month.glob(\"*mli*.nc\")):\n            part = xr.load_dataset(fpath).to_dataframe().reset_index()\n            part[\"simulation_year\"] = year\n            part[\"origin\"] = fpath.name\n            ## level 2 is what we want, grab 53 as random level comparison\n            bidx = part[\"lev\"].isin((2, 53))\n            parts.append(part[bidx][[\"origin\", \"simulation_year\", \"ncol\", \"lev\", \"ymd\", \"tod\", \"pbuf_ozone\", \"pbuf_SOLIN\"]])\ndf = pd.concat(parts)\ndf = df.rename(columns={\"ncol\":\"climsim_location_id\", \"lev\":\"level\", \"tod\":\"day_seconds\"})\ndf.shape\n\n## arbitrary \"2021\" as year - just need something without a leap year\nsim_date = pd.to_datetime(\"2021\" + df[\"ymd\"].astype(str).str[-4:], format=\"%Y%m%d\")\ndf[\"simulation_ts\"] = sim_date + df[\"day_seconds\"].apply(lambda x: datetime.timedelta(seconds=x))\ndf[\"day_tick\"] = df.day_seconds // (20*60)\n\ndf = df.sort_values([\"simulation_year\", \"simulation_ts\", \"climsim_location_id\", \"level\"])\ndf\n\ny7 = df[df[\"simulation_year\"] == 7]\ny8 = df[df[\"simulation_year\"] == 8]\ny7.shape, y8.shape\n\n## does the entire year align at level 2?\nperfect_matches = y7[y7[\"level\"] == 2][\"pbuf_ozone\"] == y8[y8[\"level\"] == 2][\"pbuf_ozone\"]\nperfect_matches.value_counts()\n# it's perfect!\n\n## does the entire year align at level 53?\nperfect_matches = y7[y7[\"level\"] == 53][\"pbuf_ozone\"] == y8[y8[\"level\"] == 53][\"pbuf_ozone\"]\nperfect_matches.value_counts()\n## definitely not\n\n## no repeated ozone values within a year\ny7_lev2_oz_uniques = y7[y7[\"level\"] == 2][\"pbuf_ozone\"].nunique()\ny8_lev2_oz_uniques = y8[y8[\"level\"] == 2][\"pbuf_ozone\"].nunique()\ny7_lev2_oz_uniques, y8_lev2_oz_uniques, y7[y7[\"level\"] == 2].shape[0]\n\n## maybe we just got lucky for the entire year comparison, let's pick a random date: April 18th at 8am\nchosen_day = df[df[\"simulation_ts\"] == datetime.datetime(year=2021, month=4, day=18, hour=8)]\nchosen_day\n\n## two years of data @ level 2, how many unique ozone values do we have?\n## this tick represents 384 locations * 1 tick, so expect 384 distinct values\nchosen_day[chosen_day[\"level\"] == 2][\"pbuf_ozone\"].nunique()\n\n## to save out the mapping table, just grab level 2\n## just for fun, drop duplicates instead of subsetting on a year, will equal the number of values for either 7 or 8\nmapping_table = df[df[\"level\"] == 2][[\"simulation_ts\", \"climsim_location_id\", \"pbuf_ozone\", \"pbuf_SOLIN\"]].drop_duplicates()\nmapping_table.shape\n\n# was paranoid about round-tripping the float value for merging, but seems fine\nmapping_table.to_csv(\"pbuf_ozone_2_mapping.csv.gz\")\n```\n",
      "votes": 34
    },
    {
      "id": 2916316,
      "postDate": "2024-07-10T22:31:48.160Z",
      "content": "<p>(Reposting from another thread)</p>\n<p>I am waiting to confirm this with the Kaggle staff, but it may be safe to assume for now that we will disqualify any submissions that take advantage of any leak or use a multicolumn (e.g. 2D to 1D or 2D to 2D) approach for cash prizes and gold medals.</p>",
      "rawMarkdown": "(Reposting from another thread)\n\nI am waiting to confirm this with the Kaggle staff, but it may be safe to assume for now that we will disqualify any submissions that take advantage of any leak or use a multicolumn (e.g. 2D to 1D or 2D to 2D) approach for cash prizes and gold medals.",
      "votes": 17,
      "replies": [
        {
          "id": 2917438,
          "postDate": "2024-07-11T16:24:32.407Z",
          "content": "<p>(reposting from another thread)</p>\n<p>From Section B.10 :</p>\n<p>\"Competition Sponsor reserves the right to disqualify any participant from the Competition if the Competition Sponsor reasonably believes that the participant has attempted to undermine the legitimate operation of the Competition by cheating, deception, or other unfair playing practices or abuses, threatens or harasses any other participants, Competition Sponsor or Kaggle.</p>\n<p>A disqualified participant may be removed from the Competition leaderboard, at Kaggle's sole discretion. If a Participant is removed from the Competition Leaderboard, additional winning features associated with the Kaggle competition platform, for example Kaggle points or medals, may also not be awarded.\"</p>\n<p>We will disqualify leak-based solutions for cash prizes and gold medal awards. To be eligible for a cash prize or gold medal, Kaggle teams will have to verify their training and inference code with the competition hosts.</p>\n<p>We will not be extending the deadline.</p>",
          "rawMarkdown": "(reposting from another thread)\n\nFrom Section B.10 :\n\n\"Competition Sponsor reserves the right to disqualify any participant from the Competition if the Competition Sponsor reasonably believes that the participant has attempted to undermine the legitimate operation of the Competition by cheating, deception, or other unfair playing practices or abuses, threatens or harasses any other participants, Competition Sponsor or Kaggle.\n\nA disqualified participant may be removed from the Competition leaderboard, at Kaggle's sole discretion. If a Participant is removed from the Competition Leaderboard, additional winning features associated with the Kaggle competition platform, for example Kaggle points or medals, may also not be awarded.\"\n\nWe will disqualify leak-based solutions for cash prizes and gold medal awards. To be eligible for a cash prize or gold medal, Kaggle teams will have to verify their training and inference code with the competition hosts.\n\nWe will not be extending the deadline.",
          "votes": 3,
          "replies": [
            {
              "id": 2917475,
              "postDate": "2024-07-11T16:40:46.490Z",
              "rawMarkdown": "",
              "votes": 1,
              "isDeleted": true
            },
            {
              "id": 2917521,
              "postDate": "2024-07-11T16:54:19.117Z",
              "content": "<blockquote>\n  <p>We will not be extending the deadline.</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a> Can you at least keep the submission/scoring engine up and running for at least a few weeks after Jul 15?</p>\n<p>I've been building a fundamentally novel learning algorithm. The math and optimization problems I managed to solve are truly eye-opening. (I say this as a matter of fact, without an iota of self-conceit.) This huge-scale multi-target regression problem is simply perfect as a testing ground.</p>\n<p>I don't care about winning but I'd rather die than not seeing how my learning algorithm performs eventually. I implore you to please keep the scoring engine alive after the deadline if at all possible <a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a> </p>",
              "rawMarkdown": ">We will not be extending the deadline.\n\n@jerrylin96 Can you at least keep the submission/scoring engine up and running for at least a few weeks after Jul 15?\n\nI've been building a fundamentally novel learning algorithm. The math and optimization problems I managed to solve are truly eye-opening. (I say this as a matter of fact, without an iota of self-conceit.) This huge-scale multi-target regression problem is simply perfect as a testing ground.\n\nI don't care about winning but I'd rather die than not seeing how my learning algorithm performs eventually. I implore you to please keep the scoring engine alive after the deadline if at all possible @jerrylin96 ",
              "votes": 1
            },
            {
              "id": 2917526,
              "postDate": "2024-07-11T16:57:02.537Z",
              "content": "<p>scoring engine is kept on by default on kaggle, you can submit to past competition if such need arise</p>",
              "rawMarkdown": "scoring engine is kept on by default on kaggle, you can submit to past competition if such need arise",
              "votes": 2
            },
            {
              "id": 2918045,
              "postDate": "2024-07-12T01:40:44.173Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    },
    {
      "id": 2914865,
      "postDate": "2024-07-10T08:07:55.347Z",
      "content": "<p>Good finding!<br>\nWe now understand that time series prediction comp always need a time-series API.😅<br>\nI don't mind that host post a new dataset and extend 15 days again, i am not tired.</p>",
      "rawMarkdown": "Good finding!\nWe now understand that time series prediction comp always need a time-series API.😅\nI don't mind that host post a new dataset and extend 15 days again, i am not tired.\n\n\n",
      "votes": 14,
      "replies": [
        {
          "id": 2914980,
          "postDate": "2024-07-10T09:20:03.280Z",
          "content": "<p>Evil. Evil is good 😁</p>",
          "rawMarkdown": "Evil. Evil is good 😁",
          "votes": 1,
          "replies": [
            {
              "id": 2915016,
              "postDate": "2024-07-10T09:55:27.810Z",
              "content": "<p>I am also competing, not making fun of this😅.</p>",
              "rawMarkdown": "I am also competing, not making fun of this😅.",
              "votes": 1
            }
          ]
        },
        {
          "id": 2914981,
          "postDate": "2024-07-10T09:20:05.260Z",
          "content": "<p>Don’t take the risk of making submissions on last day of the competition in case of API 😉</p>",
          "rawMarkdown": "Don’t take the risk of making submissions on last day of the competition in case of API 😉",
          "votes": 1,
          "replies": [
            {
              "id": 2915007,
              "postDate": "2024-07-10T09:49:40.013Z",
              "content": "<p>🤣I will submit right now if so.</p>",
              "rawMarkdown": "🤣I will submit right now if so."
            }
          ]
        },
        {
          "id": 2915075,
          "postDate": "2024-07-10T10:16:44.860Z",
          "content": "<p>i would much prefer for the hosts to end the competition today or tomorrow than another extend :P</p>",
          "rawMarkdown": "i would much prefer for the hosts to end the competition today or tomorrow than another extend :P",
          "votes": 1
        },
        {
          "id": 2915320,
          "postDate": "2024-07-10T12:47:53.710Z",
          "content": "<p><a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a> <br>\n15 days is not enough imho. Let's make it 1 month. <br>\n3 suggestions: <br>\n1 - remove those columns from test dataset (new one ofc)<br>\n2- add some noise  to those columns and be sure no-one can exploit it.<br>\n3- Switch to high resolution data and give location and time information and make test dataset time period like  6h  or more (time difference between consecutive instance for same location).  So, Leak won't be an issue again.</p>\n<p>3rd one is solid. </p>",
          "rawMarkdown": "@jerrylin96 \n15 days is not enough imho. Let's make it 1 month. \n3 suggestions: \n1 - remove those columns from test dataset (new one ofc)\n2- add some noise  to those columns and be sure no-one can exploit it.\n3- Switch to high resolution data and give location and time information and make test dataset time period like  6h  or more (time difference between consecutive instance for same location).  So, Leak won't be an issue again.\n\n3rd one is solid. ",
          "votes": -2,
          "replies": [
            {
              "id": 2915324,
              "postDate": "2024-07-10T12:52:58.407Z",
              "content": "<p>I also suggest we make it 1 year so we can train on 50Tb data &amp; 21,600 seq length</p>",
              "rawMarkdown": "I also suggest we make it 1 year so we can train on 50Tb data & 21,600 seq length",
              "votes": 10
            },
            {
              "id": 2915330,
              "postDate": "2024-07-10T12:57:58.783Z",
              "content": "<p>Great idea! :D  I was actually joking about suggestions. it was a mistake that host did not give the location and time info  at the beginning.  It is Fluid Dynamic / Thermodynamic problem. location and  spatial neighborhood is important in the calculations.  </p>",
              "rawMarkdown": "Great idea! :D  I was actually joking about suggestions. it was a mistake that host did not give the location and time info  at the beginning.  It is Fluid Dynamic / Thermodynamic problem. location and  spatial neighborhood is important in the calculations.  ",
              "votes": 2
            },
            {
              "id": 2915332,
              "postDate": "2024-07-10T12:58:19.143Z",
              "content": "<p>does high res data improve score?</p>",
              "rawMarkdown": "does high res data improve score?"
            },
            {
              "id": 2915735,
              "postDate": "2024-07-10T16:12:33.373Z",
              "content": "<p>Don't forget to add a new set of target weights!</p>",
              "rawMarkdown": "Don't forget to add a new set of target weights!",
              "votes": 1
            },
            {
              "id": 2915749,
              "postDate": "2024-07-10T16:20:43.190Z",
              "content": "<p>One month is too long!<br>\nand i suggest more prizes for top 10 but not top 5, we spend too much time! </p>",
              "rawMarkdown": "One month is too long!\nand i suggest more prizes for top 10 but not top 5, we spend too much time! ",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2914810,
      "postDate": "2024-07-10T07:12:16.293Z",
      "content": "<p>Maybe the host <a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a> needs to clarify that how will they treat winners that utilizing time\\position information (Since they don't want solution that using time\\position information).</p>\n<ol>\n<li>Keep both of the rank and prize.</li>\n<li>Disqualified the prize but keep the rank.</li>\n<li>Disqualified both of the prize and the rank.</li>\n</ol>",
      "rawMarkdown": "Maybe the host @jerrylin96 needs to clarify that how will they treat winners that utilizing time\\position information (Since they don't want solution that using time\\position information).\n\n1. Keep both of the rank and prize.\n2. Disqualified the prize but keep the rank.\n3. Disqualified both of the prize and the rank.",
      "votes": 10,
      "replies": [
        {
          "id": 2914823,
          "postDate": "2024-07-10T07:22:18.083Z",
          "content": "<p>I've never seen a team getting removed because they used a leak. I don't think it makes sense to do it since it is a grey area. No one knows how much models can learn time information. I guess there isn't much to do here.</p>",
          "rawMarkdown": "I've never seen a team getting removed because they used a leak. I don't think it makes sense to do it since it is a grey area. No one knows how much models can learn time information. I guess there isn't much to do here.",
          "votes": 1,
          "replies": [
            {
              "id": 2914837,
              "postDate": "2024-07-10T07:34:03.180Z",
              "content": "<p>In DeepFake Challenge, the original top1 team became top7, and the original top2 team was removed from the leaderboard. But the situation is different, they used the data without license by mistake. In fact, the hosts of the DeepFake did not particularly emphasize restrictions on external data, so I think their handling was perhaps overly strict in that regard. </p>\n<p>However, the LEAP hosts made it clear right from the start that time and location data were not to be used. They even went to the extent of changing the test set specifically to prevent such issues and delayed the competition by two weeks. </p>\n<p>In summary, if the organizers have repeatedly highlighted prohibited actions, competitors should abide by these rules in good faith</p>",
              "rawMarkdown": "In DeepFake Challenge, the original top1 team became top7, and the original top2 team was removed from the leaderboard. But the situation is different, they used the data without license by mistake. In fact, the hosts of the DeepFake did not particularly emphasize restrictions on external data, so I think their handling was perhaps overly strict in that regard. \n\nHowever, the LEAP hosts made it clear right from the start that time and location data were not to be used. They even went to the extent of changing the test set specifically to prevent such issues and delayed the competition by two weeks. \n\nIn summary, if the organizers have repeatedly highlighted prohibited actions, competitors should abide by these rules in good faith",
              "votes": 9
            },
            {
              "id": 2914843,
              "postDate": "2024-07-10T07:40:21.810Z",
              "content": "<p>I don't think its explicitly written in the rules right so in that sense I think it is a grey zone. </p>\n<p>I also think in this case it would be very difficult to detect…especially if using the information implicitly (e.g. train on a chosen subset) vs explicitly (multi-col prediction).</p>",
              "rawMarkdown": "I don't think its explicitly written in the rules right so in that sense I think it is a grey zone. \n\nI also think in this case it would be very difficult to detect...especially if using the information implicitly (e.g. train on a chosen subset) vs explicitly (multi-col prediction).\n",
              "votes": 3
            },
            {
              "id": 2914884,
              "postDate": "2024-07-10T08:23:04.623Z",
              "content": "<p>They can't detect whether to use the leak automatically. But the winners(Top 5 teams in this competition) have to share the code and the solution after the competition if they want to get the prize. So the host can at least audit the winner solution to forbid someone using this leak to win.</p>",
              "rawMarkdown": "They can't detect whether to use the leak automatically. But the winners(Top 5 teams in this competition) have to share the code and the solution after the competition if they want to get the prize. So the host can at least audit the winner solution to forbid someone using this leak to win."
            },
            {
              "id": 2914911,
              "postDate": "2024-07-10T08:31:54.040Z",
              "content": "<p>But what about those who are not in the bonus zone but in the gold medal zone? Do they need to submit code for review to get a gold medal? And what about those silver zone and bronze zone or just someone to get some kaggle points?</p>",
              "rawMarkdown": "But what about those who are not in the bonus zone but in the gold medal zone? Do they need to submit code for review to get a gold medal? And what about those silver zone and bronze zone or just someone to get some kaggle points?"
            },
            {
              "id": 2915142,
              "postDate": "2024-07-10T11:09:35.370Z",
              "content": "<p>Sadly, Kaggle can't detect those teams not in prize zone about whether they used the leak.</p>",
              "rawMarkdown": "Sadly, Kaggle can't detect those teams not in prize zone about whether they used the leak."
            }
          ]
        },
        {
          "id": 2914828,
          "postDate": "2024-07-10T07:26:29.163Z",
          "content": "<p>Considering that the organizers have repeatedly emphasized that time and location information should not be used, I believe that if the winning team hacked this information, they should at least be disqualified from the prize, and it would even be acceptable to disqualified the rank.</p>\n<p>I love this competition and do not want it to become another Home Credit.</p>",
          "rawMarkdown": "Considering that the organizers have repeatedly emphasized that time and location information should not be used, I believe that if the winning team hacked this information, they should at least be disqualified from the prize, and it would even be acceptable to disqualified the rank.\n\nI love this competition and do not want it to become another Home Credit.",
          "votes": 5,
          "replies": [
            {
              "id": 2915001,
              "postDate": "2024-07-10T09:44:19.100Z",
              "content": "<p>Agreed, it also make all our (teams that not utilize leaks, which is most of us) efforts meaningless.</p>",
              "rawMarkdown": "Agreed, it also make all our (teams that not utilize leaks, which is most of us) efforts meaningless.",
              "votes": 2
            },
            {
              "id": 2921847,
              "postDate": "2024-07-14T17:35:36.727Z",
              "content": "<p>Agree! It makes no sense otherwise for other participants.</p>",
              "rawMarkdown": "Agree! It makes no sense otherwise for other participants."
            }
          ]
        }
      ]
    },
    {
      "id": 2914672,
      "postDate": "2024-07-10T05:39:49.630Z",
      "content": "<p>What a great leak (or finding). If the fact this article saying is true, all of our model (more or less) implicitly utilizes this trick by these targets.</p>",
      "rawMarkdown": "What a great leak (or finding). If the fact this article saying is true, all of our model (more or less) implicitly utilizes this trick by these targets.",
      "votes": 3,
      "replies": [
        {
          "id": 2914844,
          "postDate": "2024-07-10T07:40:57.653Z",
          "content": "<p>just checked: on train and old test set holds true</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F214989%2F8f5f1db9cc22f2eeecae977c82084686%2Fa.png?generation=1720597382059388&amp;alt=media\"></p>",
          "rawMarkdown": "just checked: on train and old test set holds true\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F214989%2F8f5f1db9cc22f2eeecae977c82084686%2Fa.png?generation=1720597382059388&alt=media)",
          "votes": 1
        },
        {
          "id": 2915005,
          "postDate": "2024-07-10T09:47:09.900Z",
          "content": "<p>Yes but we utilize this information (implicitly) in 1D-&gt;1D. Turning this to 2D-1D is another story.</p>",
          "rawMarkdown": "Yes but we utilize this information (implicitly) in 1D->1D. Turning this to 2D-1D is another story.",
          "votes": 3
        }
      ]
    },
    {
      "id": 2915384,
      "postDate": "2024-07-10T13:31:49.507Z",
      "content": "<p>TL;DR</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F214989%2F8a191d674c9c6aa6c6bbcb213b50d7e0%2Fa.jpg?generation=1720618256605003&amp;alt=media\"></p>",
      "rawMarkdown": "TL;DR\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F214989%2F8a191d674c9c6aa6c6bbcb213b50d7e0%2Fa.jpg?generation=1720618256605003&alt=media)",
      "votes": 1
    },
    {
      "id": 2914781,
      "postDate": "2024-07-10T06:46:59.953Z",
      "content": "<p>I don't know if we can use this information. The host specifically said that we couldn't use any timestamp and location information during the discussion before changing the dataset. Also, I don't see the creator of the post in the Leaderboard… </p>",
      "rawMarkdown": "I don't know if we can use this information. The host specifically said that we couldn't use any timestamp and location information during the discussion before changing the dataset. Also, I don't see the creator of the post in the Leaderboard... ",
      "votes": 1
    },
    {
      "id": 2914684,
      "postDate": "2024-07-10T05:57:12.613Z",
      "content": "<p>I don't really understand why there are so many downvotes for  without any explanation  for this discussion</p>",
      "rawMarkdown": "I don't really understand why there are so many downvotes for  without any explanation  for this discussion",
      "replies": [
        {
          "id": 2914691,
          "postDate": "2024-07-10T05:59:14.727Z",
          "content": "<p>Some people worked for 3 month on this comp :D</p>",
          "rawMarkdown": "Some people worked for 3 month on this comp :D",
          "votes": 5,
          "replies": [
            {
              "id": 2914725,
              "postDate": "2024-07-10T06:13:27.337Z",
              "content": "<p>I understand it perfectly, but this is not in the spirit of Kaggle! You can't reject an analysis just because it doesn't align with your interests.</p>",
              "rawMarkdown": "I understand it perfectly, but this is not in the spirit of Kaggle! You can't reject an analysis just because it doesn't align with your interests.",
              "votes": 3
            },
            {
              "id": 2914736,
              "postDate": "2024-07-10T06:25:36.043Z",
              "content": "<p>i'm tired :D (most likely not just me!~)</p>",
              "rawMarkdown": "i'm tired :D (most likely not just me!~)",
              "votes": 16
            },
            {
              "id": 2914753,
              "postDate": "2024-07-10T06:33:20.147Z",
              "content": "<p>tired, too.</p>",
              "rawMarkdown": "tired, too.",
              "votes": 6
            },
            {
              "id": 2914785,
              "postDate": "2024-07-10T06:50:05.883Z",
              "content": "<p>Get ready to be tired for another 2 weeks.🫠</p>",
              "rawMarkdown": "Get ready to be tired for another 2 weeks.🫠",
              "votes": 1
            },
            {
              "id": 2914821,
              "postDate": "2024-07-10T07:20:07.940Z",
              "content": "<p>I can understand that it it not the spirit of Kaggle. But since the host emphasized that <code>it would be a column to column regression problem, not a multi-column to column regression problem</code>, what if someone use this leak to get prize, is it allowable? </p>\n<p>Are there some cases before?</p>",
              "rawMarkdown": "I can understand that it it not the spirit of Kaggle. But since the host emphasized that `it would be a column to column regression problem, not a multi-column to column regression problem`, what if someone use this leak to get prize, is it allowable? \n\nAre there some cases before?",
              "votes": 2
            },
            {
              "id": 2915243,
              "postDate": "2024-07-10T11:52:18.003Z",
              "content": "<blockquote>\n  <p>Get ready to be tired for another 2 weeks.🫠  </p>\n</blockquote>\n<p>Abd then someone find a leak in another column 😭</p>",
              "rawMarkdown": "> Get ready to be tired for another 2 weeks.🫠  \n\nAbd then someone find a leak in another column 😭",
              "votes": 3
            },
            {
              "id": 2917043,
              "postDate": "2024-07-11T12:50:04.480Z",
              "content": "<p>given past leak, most likely they will extend deadline</p>",
              "rawMarkdown": "given past leak, most likely they will extend deadline"
            }
          ]
        },
        {
          "id": 2914735,
          "postDate": "2024-07-10T06:23:20.787Z",
          "content": "<p>I think it would be better if he writes this much earlier or later, not now only 6 days before the deadline. What's more, it's very strange that this account was not active in the past one year and suddenly writes this discussion. It seems someone uses multiple accounts.</p>",
          "rawMarkdown": "I think it would be better if he writes this much earlier or later, not now only 6 days before the deadline. What's more, it's very strange that this account was not active in the past one year and suddenly writes this discussion. It seems someone uses multiple accounts.",
          "votes": 19,
          "replies": [
            {
              "id": 2914740,
              "postDate": "2024-07-10T06:27:54.180Z",
              "content": "<p>Agree - quite strange… no submissions/posts and then comes out with this - extracting the low-res set (which many are presumably using - but definitely not the first thing to investigate…)</p>",
              "rawMarkdown": "Agree - quite strange... no submissions/posts and then comes out with this - extracting the low-res set (which many are presumably using - but definitely not the first thing to investigate...)",
              "votes": 5
            }
          ]
        },
        {
          "id": 2914742,
          "postDate": "2024-07-10T06:29:49.833Z",
          "content": "<p>The host try a lot to let us not use the time or location information. It's totally against the host's intention. We've already been extended by 14 days and changed dataset once.</p>",
          "rawMarkdown": "The host try a lot to let us not use the time or location information. It's totally against the host's intention. We've already been extended by 14 days and changed dataset once.",
          "votes": 6
        }
      ]
    },
    {
      "id": 2921251,
      "postDate": "2024-07-14T08:48:05.453Z",
      "content": "<p>Thank you for the sharing.<br>\nJust to make sure, does the value of pbuf_ozone_2 in the CSV refer to normalized values?</p>",
      "rawMarkdown": "Thank you for the sharing.\nJust to make sure, does the value of pbuf_ozone_2 in the CSV refer to normalized values?"
    }
  ],
  "comments": [
    {
      "id": 2916316,
      "author_name": "Jerry Lin",
      "author_url": "",
      "post_date": "2024-07-10T22:31:48.160000",
      "content": "<p>(Reposting from another thread)</p>\n<p>I am waiting to confirm this with the Kaggle staff, but it may be safe to assume for now that we will disqualify any submissions that take advantage of any leak or use a multicolumn (e.g. 2D to 1D or 2D to 2D) approach for cash prizes and gold medals.</p>",
      "votes": 17,
      "replies": [
        {
          "id": 2917438,
          "author_name": "Jerry Lin",
          "author_url": "",
          "post_date": "2024-07-11T16:24:32.407000",
          "content": "<p>(reposting from another thread)</p>\n<p>From Section B.10 :</p>\n<p>\"Competition Sponsor reserves the right to disqualify any participant from the Competition if the Competition Sponsor reasonably believes that the participant has attempted to undermine the legitimate operation of the Competition by cheating, deception, or other unfair playing practices or abuses, threatens or harasses any other participants, Competition Sponsor or Kaggle.</p>\n<p>A disqualified participant may be removed from the Competition leaderboard, at Kaggle's sole discretion. If a Participant is removed from the Competition Leaderboard, additional winning features associated with the Kaggle competition platform, for example Kaggle points or medals, may also not be awarded.\"</p>\n<p>We will disqualify leak-based solutions for cash prizes and gold medal awards. To be eligible for a cash prize or gold medal, Kaggle teams will have to verify their training and inference code with the competition hosts.</p>\n<p>We will not be extending the deadline.</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2917475,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-07-11T16:40:46.490000",
              "content": "",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2917521,
              "author_name": "Truth Seeker",
              "author_url": "",
              "post_date": "2024-07-11T16:54:19.117000",
              "content": "<blockquote>\n  <p>We will not be extending the deadline.</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a> Can you at least keep the submission/scoring engine up and running for at least a few weeks after Jul 15?</p>\n<p>I've been building a fundamentally novel learning algorithm. The math and optimization problems I managed to solve are truly eye-opening. (I say this as a matter of fact, without an iota of self-conceit.) This huge-scale multi-target regression problem is simply perfect as a testing ground.</p>\n<p>I don't care about winning but I'd rather die than not seeing how my learning algorithm performs eventually. I implore you to please keep the scoring engine alive after the deadline if at all possible <a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a> </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2917526,
              "author_name": "slime",
              "author_url": "",
              "post_date": "2024-07-11T16:57:02.537000",
              "content": "<p>scoring engine is kept on by default on kaggle, you can submit to past competition if such need arise</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2918045,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-07-12T01:40:44.173000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2914865,
      "author_name": "hyd",
      "author_url": "",
      "post_date": "2024-07-10T08:07:55.347000",
      "content": "<p>Good finding!<br>\nWe now understand that time series prediction comp always need a time-series API.😅<br>\nI don't mind that host post a new dataset and extend 15 days again, i am not tired.</p>",
      "votes": 14,
      "replies": [
        {
          "id": 2914980,
          "author_name": "nymfree",
          "author_url": "",
          "post_date": "2024-07-10T09:20:03.280000",
          "content": "<p>Evil. Evil is good 😁</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2915016,
              "author_name": "hyd",
              "author_url": "",
              "post_date": "2024-07-10T09:55:27.810000",
              "content": "<p>I am also competing, not making fun of this😅.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 2914981,
          "author_name": "Nischay Dhankhar",
          "author_url": "",
          "post_date": "2024-07-10T09:20:05.260000",
          "content": "<p>Don’t take the risk of making submissions on last day of the competition in case of API 😉</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2915007,
              "author_name": "hyd",
              "author_url": "",
              "post_date": "2024-07-10T09:49:40.013000",
              "content": "<p>🤣I will submit right now if so.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2915075,
          "author_name": "Vasilis",
          "author_url": "",
          "post_date": "2024-07-10T10:16:44.860000",
          "content": "<p>i would much prefer for the hosts to end the competition today or tomorrow than another extend :P</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2915320,
          "author_name": "Davut Polat",
          "author_url": "",
          "post_date": "2024-07-10T12:47:53.710000",
          "content": "<p><a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a> <br>\n15 days is not enough imho. Let's make it 1 month. <br>\n3 suggestions: <br>\n1 - remove those columns from test dataset (new one ofc)<br>\n2- add some noise  to those columns and be sure no-one can exploit it.<br>\n3- Switch to high resolution data and give location and time information and make test dataset time period like  6h  or more (time difference between consecutive instance for same location).  So, Leak won't be an issue again.</p>\n<p>3rd one is solid. </p>",
          "votes": -2,
          "replies": [
            {
              "id": 2915324,
              "author_name": "slime",
              "author_url": "",
              "post_date": "2024-07-10T12:52:58.407000",
              "content": "<p>I also suggest we make it 1 year so we can train on 50Tb data &amp; 21,600 seq length</p>",
              "votes": 10,
              "replies": []
            },
            {
              "id": 2915330,
              "author_name": "Davut Polat",
              "author_url": "",
              "post_date": "2024-07-10T12:57:58.783000",
              "content": "<p>Great idea! :D  I was actually joking about suggestions. it was a mistake that host did not give the location and time info  at the beginning.  It is Fluid Dynamic / Thermodynamic problem. location and  spatial neighborhood is important in the calculations.  </p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2915332,
              "author_name": "yuanzhe zhou",
              "author_url": "",
              "post_date": "2024-07-10T12:58:19.143000",
              "content": "<p>does high res data improve score?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2915735,
              "author_name": "Yusef A.",
              "author_url": "",
              "post_date": "2024-07-10T16:12:33.373000",
              "content": "<p>Don't forget to add a new set of target weights!</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2915749,
              "author_name": "hyd",
              "author_url": "",
              "post_date": "2024-07-10T16:20:43.190000",
              "content": "<p>One month is too long!<br>\nand i suggest more prizes for top 10 but not top 5, we spend too much time! </p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2914810,
      "author_name": "ForcewithMe",
      "author_url": "",
      "post_date": "2024-07-10T07:12:16.293000",
      "content": "<p>Maybe the host <a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a> needs to clarify that how will they treat winners that utilizing time\\position information (Since they don't want solution that using time\\position information).</p>\n<ol>\n<li>Keep both of the rank and prize.</li>\n<li>Disqualified the prize but keep the rank.</li>\n<li>Disqualified both of the prize and the rank.</li>\n</ol>",
      "votes": 10,
      "replies": [
        {
          "id": 2914823,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2024-07-10T07:22:18.083000",
          "content": "<p>I've never seen a team getting removed because they used a leak. I don't think it makes sense to do it since it is a grey area. No one knows how much models can learn time information. I guess there isn't much to do here.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2914837,
              "author_name": "ForcewithMe",
              "author_url": "",
              "post_date": "2024-07-10T07:34:03.180000",
              "content": "<p>In DeepFake Challenge, the original top1 team became top7, and the original top2 team was removed from the leaderboard. But the situation is different, they used the data without license by mistake. In fact, the hosts of the DeepFake did not particularly emphasize restrictions on external data, so I think their handling was perhaps overly strict in that regard. </p>\n<p>However, the LEAP hosts made it clear right from the start that time and location data were not to be used. They even went to the extent of changing the test set specifically to prevent such issues and delayed the competition by two weeks. </p>\n<p>In summary, if the organizers have repeatedly highlighted prohibited actions, competitors should abide by these rules in good faith</p>",
              "votes": 9,
              "replies": []
            },
            {
              "id": 2914843,
              "author_name": "Yusef A.",
              "author_url": "",
              "post_date": "2024-07-10T07:40:21.810000",
              "content": "<p>I don't think its explicitly written in the rules right so in that sense I think it is a grey zone. </p>\n<p>I also think in this case it would be very difficult to detect…especially if using the information implicitly (e.g. train on a chosen subset) vs explicitly (multi-col prediction).</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2914884,
              "author_name": "ForcewithMe",
              "author_url": "",
              "post_date": "2024-07-10T08:23:04.623000",
              "content": "<p>They can't detect whether to use the leak automatically. But the winners(Top 5 teams in this competition) have to share the code and the solution after the competition if they want to get the prize. So the host can at least audit the winner solution to forbid someone using this leak to win.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2914911,
              "author_name": "heng",
              "author_url": "",
              "post_date": "2024-07-10T08:31:54.040000",
              "content": "<p>But what about those who are not in the bonus zone but in the gold medal zone? Do they need to submit code for review to get a gold medal? And what about those silver zone and bronze zone or just someone to get some kaggle points?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2915142,
              "author_name": "ForcewithMe",
              "author_url": "",
              "post_date": "2024-07-10T11:09:35.370000",
              "content": "<p>Sadly, Kaggle can't detect those teams not in prize zone about whether they used the leak.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2914828,
          "author_name": "ForcewithMe",
          "author_url": "",
          "post_date": "2024-07-10T07:26:29.163000",
          "content": "<p>Considering that the organizers have repeatedly emphasized that time and location information should not be used, I believe that if the winning team hacked this information, they should at least be disqualified from the prize, and it would even be acceptable to disqualified the rank.</p>\n<p>I love this competition and do not want it to become another Home Credit.</p>",
          "votes": 5,
          "replies": [
            {
              "id": 2915001,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2024-07-10T09:44:19.100000",
              "content": "<p>Agreed, it also make all our (teams that not utilize leaks, which is most of us) efforts meaningless.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2921847,
              "author_name": "Linus X",
              "author_url": "",
              "post_date": "2024-07-14T17:35:36.727000",
              "content": "<p>Agree! It makes no sense otherwise for other participants.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2914672,
      "author_name": "Bilzard",
      "author_url": "",
      "post_date": "2024-07-10T05:39:49.630000",
      "content": "<p>What a great leak (or finding). If the fact this article saying is true, all of our model (more or less) implicitly utilizes this trick by these targets.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2914844,
          "author_name": "steubk",
          "author_url": "",
          "post_date": "2024-07-10T07:40:57.653000",
          "content": "<p>just checked: on train and old test set holds true</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F214989%2F8f5f1db9cc22f2eeecae977c82084686%2Fa.png?generation=1720597382059388&amp;alt=media\"></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2915005,
          "author_name": "greySnow",
          "author_url": "",
          "post_date": "2024-07-10T09:47:09.900000",
          "content": "<p>Yes but we utilize this information (implicitly) in 1D-&gt;1D. Turning this to 2D-1D is another story.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 2915384,
      "author_name": "steubk",
      "author_url": "",
      "post_date": "2024-07-10T13:31:49.507000",
      "content": "<p>TL;DR</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F214989%2F8a191d674c9c6aa6c6bbcb213b50d7e0%2Fa.jpg?generation=1720618256605003&amp;alt=media\"></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2914781,
      "author_name": "Federico Peccia",
      "author_url": "",
      "post_date": "2024-07-10T06:46:59.953000",
      "content": "<p>I don't know if we can use this information. The host specifically said that we couldn't use any timestamp and location information during the discussion before changing the dataset. Also, I don't see the creator of the post in the Leaderboard… </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2914684,
      "author_name": "steubk",
      "author_url": "",
      "post_date": "2024-07-10T05:57:12.613000",
      "content": "<p>I don't really understand why there are so many downvotes for  without any explanation  for this discussion</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2914691,
          "author_name": "slime",
          "author_url": "",
          "post_date": "2024-07-10T05:59:14.727000",
          "content": "<p>Some people worked for 3 month on this comp :D</p>",
          "votes": 5,
          "replies": [
            {
              "id": 2914725,
              "author_name": "steubk",
              "author_url": "",
              "post_date": "2024-07-10T06:13:27.337000",
              "content": "<p>I understand it perfectly, but this is not in the spirit of Kaggle! You can't reject an analysis just because it doesn't align with your interests.</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2914736,
              "author_name": "slime",
              "author_url": "",
              "post_date": "2024-07-10T06:25:36.043000",
              "content": "<p>i'm tired :D (most likely not just me!~)</p>",
              "votes": 16,
              "replies": []
            },
            {
              "id": 2914753,
              "author_name": "ADAM.",
              "author_url": "",
              "post_date": "2024-07-10T06:33:20.147000",
              "content": "<p>tired, too.</p>",
              "votes": 6,
              "replies": []
            },
            {
              "id": 2914785,
              "author_name": "Yusef A.",
              "author_url": "",
              "post_date": "2024-07-10T06:50:05.883000",
              "content": "<p>Get ready to be tired for another 2 weeks.🫠</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2914821,
              "author_name": "ADAM.",
              "author_url": "",
              "post_date": "2024-07-10T07:20:07.940000",
              "content": "<p>I can understand that it it not the spirit of Kaggle. But since the host emphasized that <code>it would be a column to column regression problem, not a multi-column to column regression problem</code>, what if someone use this leak to get prize, is it allowable? </p>\n<p>Are there some cases before?</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2915243,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2024-07-10T11:52:18.003000",
              "content": "<blockquote>\n  <p>Get ready to be tired for another 2 weeks.🫠  </p>\n</blockquote>\n<p>Abd then someone find a leak in another column 😭</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2917043,
              "author_name": "VITALIY",
              "author_url": "",
              "post_date": "2024-07-11T12:50:04.480000",
              "content": "<p>given past leak, most likely they will extend deadline</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2914735,
          "author_name": "嘴爷",
          "author_url": "",
          "post_date": "2024-07-10T06:23:20.787000",
          "content": "<p>I think it would be better if he writes this much earlier or later, not now only 6 days before the deadline. What's more, it's very strange that this account was not active in the past one year and suddenly writes this discussion. It seems someone uses multiple accounts.</p>",
          "votes": 19,
          "replies": [
            {
              "id": 2914740,
              "author_name": "Yusef A.",
              "author_url": "",
              "post_date": "2024-07-10T06:27:54.180000",
              "content": "<p>Agree - quite strange… no submissions/posts and then comes out with this - extracting the low-res set (which many are presumably using - but definitely not the first thing to investigate…)</p>",
              "votes": 5,
              "replies": []
            }
          ]
        },
        {
          "id": 2914742,
          "author_name": "ADAM.",
          "author_url": "",
          "post_date": "2024-07-10T06:29:49.833000",
          "content": "<p>The host try a lot to let us not use the time or location information. It's totally against the host's intention. We've already been extended by 14 days and changed dataset once.</p>",
          "votes": 6,
          "replies": []
        }
      ]
    },
    {
      "id": 2921251,
      "author_name": "yelim421",
      "author_url": "",
      "post_date": "2024-07-14T08:48:05.453000",
      "content": "<p>Thank you for the sharing.<br>\nJust to make sure, does the value of pbuf_ozone_2 in the CSV refer to normalized values?</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2914599": "tl;dr you can perfectly assign the simulation tick and location for all test samples\n\nA number of features in the ClimSim dataset are fixed for a given (location, timestamp).\nRegardless of simulation year, the value for that (location, timestamp, feature) is always the same.\n``pbuf_ozone_2`` is such a feature that is both fixed and a unique fingerprint value: there are no repeated ``pbuf_ozone_2`` measurements within a simulation year, but they always appear at the same time and place.\nWhich means that it is possible to construct a mapping table of ``pbuf_ozone_2`` -> (location, timestamp) and unblind all of the test samples.\nThis requires 10,091,520 ``pbuf_ozone_2`` readings (384 locations * 365 days * 72ticks/day).\n\n## Praise the Sun\n\nThe sun is important to the weather system (insert citation here).\nLuckily for us, ``pbuf_SOLIN`` (solar insolation) is also fixed for a simulation (location, timestamp).\nWhich means that once we have derived the (location, timestamp) from ``pbuf_ozone_2``, we now know the value of solar insolation at the previous tick.\nAnd the next tick.\nAnd the tick a week from now.\n\n(unlike ``pbuf_ozone_2``, ``pbuf_SOLIN`` has repeated 0 values at night, so it cannot be used as a fingerprint feature)\n\n## Test Set Fun Facts\n\nOriginal testset was sampled at 10 hour increments (Jan-01 06:00, Jan-01 16:00, Jan-01 02:00, Jan 02 12:00, etc).\nReplacement testset was sampled at 8 hour increments (Jan-01 00:00, Jan-01 08:00, Jan-01 16:00, Jan-02 00:00, etc).\n\nThis arrangement means that the original testset ~randomly sampled each location throughout the day (every even hour was hit).\nThe new testset only touched each location at three of the possible ticks per day.\n\n## Fun Things to Try\n\n- Targeted training set. Why bother with those irrelevant timestamps? Only utilize 00:00, 08:00, and 16:00.\n- Train by region. Latitudes appear to be well correlated.\n- Global balance. For half of the year (some inference required) it is possible to build a global map incorporating the 384 locations.\n- Timeseries - can align a feature for the present test value and the next one 8 hours in the future. By incorporating the original testset, could even have moments where the future location tick is only two hours in the future.\n- Solar oracle. Incorporate additional features for the previous/next solar input tick as a sort of solar momentum\n\n## Script to Spot Check \n\nTo spot check a few simulation ticks can pull a few files from the low_res dataset.\n\n```python\nimport datetime\nfrom pathlib import Path\n\nimport netCDF4 # implicitly required by xarray\nimport pandas as pd\nimport requests\nimport xarray as xr\n\nTICK_NUMBER = 12_345 # whatever you want to check. tick is every 20 minutes; 72/day\nHUGGING_ROOT = \"https://huggingface.co/datasets/LEAP/ClimSim_low-res/resolve/main/train/\"\nDST_DIR = Path(\"climsim_ticks\")\nDST_DIR.mkdir(exist_ok=True)\n\ndef download_tick(sim_year:int, tick_number:int, dst_dir:Path):\n    # 2021 is arbitrary, but want non-leapyear\n    initial = datetime.datetime(year=2021, month=1, day=1)\n    ts = initial + datetime.timedelta(seconds=tick_number*(60*20))\n    ## seconds are always zero, but ¯\\_ (ツ)_/¯ \n    second_of_day = (ts.hour*60*60) + (ts.minute*60) + (ts.second)\n\n    # `E3SM-MMF.mli.0008-05-01-03600.nc`\n    fname = f\"E3SM-MMF.mli.{sim_year:04d}-{ts.month:02d}-{ts.day:02d}-{second_of_day:05d}.nc\"\n    url = f\"{HUGGING_ROOT}{sim_year:04d}-{ts.month:02d}/{fname}\"\n    dst = dst_dir / fname\n    if not dst.exists():\n        resp = requests.get(url, params={\"download\":\"true\"})\n        with open(dst, 'wb') as fd:\n            fd.write(resp.content)\n\n## examine same simulation tick across different simulation years\ndownload_tick(6, TICK_NUMBER, DST_DIR)\ndownload_tick(7, TICK_NUMBER, DST_DIR)\ndownload_tick(8, TICK_NUMBER, DST_DIR)\n\n## Build up the three samples\nparts = []\nfor fpath in DST_DIR.glob(\"*.nc\"):\n    part = xr.load_dataset(fpath).to_dataframe().reset_index(names=[\"climsim_location_id\", \"level\"])\n    part.insert(0, \"origin\", fpath.name)\n    sim_year = int(fpath.name.split(\".mli.\")[1].split(\"-\", 1)[0])\n    part.insert(1, \"simulation_year\", sim_year)\n    parts.append(part)\ndf = pd.concat(parts)\ndf.shape\n\n## ensure ordering\ndf = df.sort_values(['simulation_year', 'ymd', 'tod', 'level'])\n\n## grab just level 2\nlev2 = df[df[\"level\"]==2]\n## how many uniques per simulation year?\nlev2.groupby(['simulation_year']).pbuf_ozone.nunique()\n\n## how many uniques total?\nlev2.pbuf_ozone.nunique()\n\n## look at years specifically\ny7 = lev2[lev2[\"simulation_year\"] == 7]\ny8 = lev2[lev2[\"simulation_year\"] == 8]\n\nozone_matches = y7[\"pbuf_ozone\"] == y8[\"pbuf_ozone\"]\nozone_matches.value_counts()\n```\n\n## Script to Extract Fingerprint IDs\n\nGiven a directory containing the ClimSim low res dataset, can extract out the mapping values using some variation of the below.\nThe ClimSim low res repository contain a definition file mapping location id to (latitude, longitude).\n\n```python\nimport datetime\nfrom pathlib import Path\n\nimport netCDF4 # implicitly required by xarray to load .nc\nimport pandas as pd\nimport tqdm\nimport xarray as xr\n\n# https://huggingface.co/datasets/LEAP/ClimSim_low-res/ \nDIR_CLIMSIM = Path(\"ClimSim_low-res/train/\")\n\nparts = []\nfor year in [7, 8]: # available sim years are 01-Jan to 09-Jan\n    for month in range(1, 13):\n        dir_month = DIR_CLIMSIM / f\"{year:04d}-{month:02d}\"\n        for fpath in tqdm.tqdm(dir_month.glob(\"*mli*.nc\")):\n            part = xr.load_dataset(fpath).to_dataframe().reset_index()\n            part[\"simulation_year\"] = year\n            part[\"origin\"] = fpath.name\n            ## level 2 is what we want, grab 53 as random level comparison\n            bidx = part[\"lev\"].isin((2, 53))\n            parts.append(part[bidx][[\"origin\", \"simulation_year\", \"ncol\", \"lev\", \"ymd\", \"tod\", \"pbuf_ozone\", \"pbuf_SOLIN\"]])\ndf = pd.concat(parts)\ndf = df.rename(columns={\"ncol\":\"climsim_location_id\", \"lev\":\"level\", \"tod\":\"day_seconds\"})\ndf.shape\n\n## arbitrary \"2021\" as year - just need something without a leap year\nsim_date = pd.to_datetime(\"2021\" + df[\"ymd\"].astype(str).str[-4:], format=\"%Y%m%d\")\ndf[\"simulation_ts\"] = sim_date + df[\"day_seconds\"].apply(lambda x: datetime.timedelta(seconds=x))\ndf[\"day_tick\"] = df.day_seconds // (20*60)\n\ndf = df.sort_values([\"simulation_year\", \"simulation_ts\", \"climsim_location_id\", \"level\"])\ndf\n\ny7 = df[df[\"simulation_year\"] == 7]\ny8 = df[df[\"simulation_year\"] == 8]\ny7.shape, y8.shape\n\n## does the entire year align at level 2?\nperfect_matches = y7[y7[\"level\"] == 2][\"pbuf_ozone\"] == y8[y8[\"level\"] == 2][\"pbuf_ozone\"]\nperfect_matches.value_counts()\n# it's perfect!\n\n## does the entire year align at level 53?\nperfect_matches = y7[y7[\"level\"] == 53][\"pbuf_ozone\"] == y8[y8[\"level\"] == 53][\"pbuf_ozone\"]\nperfect_matches.value_counts()\n## definitely not\n\n## no repeated ozone values within a year\ny7_lev2_oz_uniques = y7[y7[\"level\"] == 2][\"pbuf_ozone\"].nunique()\ny8_lev2_oz_uniques = y8[y8[\"level\"] == 2][\"pbuf_ozone\"].nunique()\ny7_lev2_oz_uniques, y8_lev2_oz_uniques, y7[y7[\"level\"] == 2].shape[0]\n\n## maybe we just got lucky for the entire year comparison, let's pick a random date: April 18th at 8am\nchosen_day = df[df[\"simulation_ts\"] == datetime.datetime(year=2021, month=4, day=18, hour=8)]\nchosen_day\n\n## two years of data @ level 2, how many unique ozone values do we have?\n## this tick represents 384 locations * 1 tick, so expect 384 distinct values\nchosen_day[chosen_day[\"level\"] == 2][\"pbuf_ozone\"].nunique()\n\n## to save out the mapping table, just grab level 2\n## just for fun, drop duplicates instead of subsetting on a year, will equal the number of values for either 7 or 8\nmapping_table = df[df[\"level\"] == 2][[\"simulation_ts\", \"climsim_location_id\", \"pbuf_ozone\", \"pbuf_SOLIN\"]].drop_duplicates()\nmapping_table.shape\n\n# was paranoid about round-tripping the float value for merging, but seems fine\nmapping_table.to_csv(\"pbuf_ozone_2_mapping.csv.gz\")\n```\n",
    "2916316": "(Reposting from another thread)\n\nI am waiting to confirm this with the Kaggle staff, but it may be safe to assume for now that we will disqualify any submissions that take advantage of any leak or use a multicolumn (e.g. 2D to 1D or 2D to 2D) approach for cash prizes and gold medals.",
    "2914865": "Good finding!\nWe now understand that time series prediction comp always need a time-series API.😅\nI don't mind that host post a new dataset and extend 15 days again, i am not tired.\n\n\n",
    "2914810": "Maybe the host @jerrylin96 needs to clarify that how will they treat winners that utilizing time\\position information (Since they don't want solution that using time\\position information).\n\n1. Keep both of the rank and prize.\n2. Disqualified the prize but keep the rank.\n3. Disqualified both of the prize and the rank.",
    "2914672": "What a great leak (or finding). If the fact this article saying is true, all of our model (more or less) implicitly utilizes this trick by these targets.",
    "2915384": "TL;DR\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F214989%2F8a191d674c9c6aa6c6bbcb213b50d7e0%2Fa.jpg?generation=1720618256605003&alt=media)",
    "2914781": "I don't know if we can use this information. The host specifically said that we couldn't use any timestamp and location information during the discussion before changing the dataset. Also, I don't see the creator of the post in the Leaderboard... ",
    "2914684": "I don't really understand why there are so many downvotes for  without any explanation  for this discussion",
    "2921251": "Thank you for the sharing.\nJust to make sure, does the value of pbuf_ozone_2 in the CSV refer to normalized values?"
  }
}