{
  "id": 519575,
  "title": "Questions about the criteria for leaks",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/519575",
  "author_name": "Ryota",
  "post_date": "2024-07-11T19:10:55.302000",
  "votes": 8,
  "comment_count": 25,
  "views": 0,
  "content": "<p><a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a> </p>\n<p>However, I have already mentioned <a href=\"url\" target=\"_blank\">here</a>, I am concerned that it may get lost in the thread and the host might not find it, so I have decided to raise this discussion myself. Could you please answer the following questions?</p>\n<blockquote>\n  <p>Thank you for clarifying the measures to be taken when leaks are used. However, the criteria for what falls under a leak are still unclear, and I would like this to be clarified.<br>\n  From your intentions, we understand that using explicit leaks present in the test data (such as pbuf_ozone_2) to predict a single row from multiple rows is prohibited.<br>\n  However, for example:</p>\n  <ol>\n  <li>What about cases where location and timestamp information are inferred and restored using training data? This does not use the direct leak from pbuf_ozone_2. (The approach of inferring and restoring these information was known to be possible even before the leak involving pbuf_ozone_2 was discovered, and it was not prohibited.)</li>\n  <li>Additionally, what happens if the restored data is used as features for a single row without using information from multiple rows? (e.g. month, day, hour, lat, lon, etc.)</li>\n  </ol>\n</blockquote>",
  "messages": [
    {
      "id": 2918339,
      "postDate": "2024-07-12T06:50:23.093Z",
      "content": "<p><a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a> <br>\nThank you for providing the guidelines. If you don’t mind, I would like to inquire further about the use of location and timestamp information when not intentionally analyzing the test set.</p>\n<p>When training with the provided low-resolution data, we can use the location and timestamp information from the training set. While we cannot use these features as inputs to the model without analyzing the test set, there are other ways to utilize them:</p>\n<ul>\n<li><strong>2D pretrain to 1D fine-tuning</strong>: For example, we can build a 2D to 1D model using location and timestamp in the train set (Of course, timestamp/location information of test data is not used. ), and then fine-tune it as a 1D to 1D model. This doesn’t utilize leakage from the test set, but it’s an example where part of the train pipeline includes a 2D to 1D component.</li>\n<li><strong>Auxiliary Loss</strong>: This technique involves adding a classification head to predict location and timestamp during model training. Auxiliary loss is a commonly used method in machine learning. The advantage of this method is that these meta-information features are not required during inference on test data or when deploying the model without actual location and timestamp data.</li>\n<li><strong>Difference from Statistics</strong>: In this method, we calculate statistical values for each location and timestamp within the training set. For example, we calculate the mean value of features for each of the 384 locations in the training set. During model training, we use the difference between these 384 mean values and the input features (1D). This method doesn’t require analyzing the test set and performs 1D to 1D prediction. However, from some perspectives, it might be considered as utilizing other data within the training set. We would appreciate your thoughts on this method.</li>\n</ul>\n<p>Additionally, there might be ways to implicitly utilize location information without explicitly using the above information. For example:</p>\n<ul>\n<li><strong>Clustering Techniques</strong>: We could apply clustering techniques to the combined train and test sets, build models for each cluster, or use cluster information as input to the model or as auxiliary loss. This method might unintentionally create clusters based on location information. There’s a concern that participants might unknowingly use this cluster information. In data analysis, it’s common to perform cluster analysis on provided data and treat each cluster separately (e.g., normalizing each cluster) if mixed distributions are found. We haven’t tried this, but if this method generates clusters containing location information and improves R2 score, how would it be judged in light of the competition rules?</li>\n</ul>\n<p>The current leakage issue is different from the typical target label leakage often seen on Kaggle. It relates to the inherent spatial and temporal nature of the data. Considering this, it might be difficult to completely exclude location and timestamp information from modeling, whether intentional or not. Furthermore, setting clear disqualification criteria might be challenging, though not impossible with extensive expertise and time for verification.</p>\n<p>Once again, thank you for the guidance provided. We would appreciate your explanation on the above points to minimize the risk of participants being unintentionally disqualified.</p>",
      "rawMarkdown": "@jerrylin96 \nThank you for providing the guidelines. If you don’t mind, I would like to inquire further about the use of location and timestamp information when not intentionally analyzing the test set.\n\nWhen training with the provided low-resolution data, we can use the location and timestamp information from the training set. While we cannot use these features as inputs to the model without analyzing the test set, there are other ways to utilize them:\n- **2D pretrain to 1D fine-tuning**: For example, we can build a 2D to 1D model using location and timestamp in the train set (Of course, timestamp/location information of test data is not used. ), and then fine-tune it as a 1D to 1D model. This doesn’t utilize leakage from the test set, but it’s an example where part of the train pipeline includes a 2D to 1D component.\n- **Auxiliary Loss**: This technique involves adding a classification head to predict location and timestamp during model training. Auxiliary loss is a commonly used method in machine learning. The advantage of this method is that these meta-information features are not required during inference on test data or when deploying the model without actual location and timestamp data.\n- **Difference from Statistics**: In this method, we calculate statistical values for each location and timestamp within the training set. For example, we calculate the mean value of features for each of the 384 locations in the training set. During model training, we use the difference between these 384 mean values and the input features (1D). This method doesn’t require analyzing the test set and performs 1D to 1D prediction. However, from some perspectives, it might be considered as utilizing other data within the training set. We would appreciate your thoughts on this method.\n\nAdditionally, there might be ways to implicitly utilize location information without explicitly using the above information. For example:\n- **Clustering Techniques**: We could apply clustering techniques to the combined train and test sets, build models for each cluster, or use cluster information as input to the model or as auxiliary loss. This method might unintentionally create clusters based on location information. There’s a concern that participants might unknowingly use this cluster information. In data analysis, it’s common to perform cluster analysis on provided data and treat each cluster separately (e.g., normalizing each cluster) if mixed distributions are found. We haven’t tried this, but if this method generates clusters containing location information and improves R2 score, how would it be judged in light of the competition rules?\n\nThe current leakage issue is different from the typical target label leakage often seen on Kaggle. It relates to the inherent spatial and temporal nature of the data. Considering this, it might be difficult to completely exclude location and timestamp information from modeling, whether intentional or not. Furthermore, setting clear disqualification criteria might be challenging, though not impossible with extensive expertise and time for verification.\n\nOnce again, thank you for the guidance provided. We would appreciate your explanation on the above points to minimize the risk of participants being unintentionally disqualified.",
      "votes": 11,
      "replies": [
        {
          "id": 2918346,
          "postDate": "2024-07-12T07:02:17.827Z",
          "content": "<p>I use auxiliary spacetime loss, it helps a little. It does not utilize any leak (data is readily available in HF low-res dataset) so I saw no reason not to</p>",
          "rawMarkdown": "I use auxiliary spacetime loss, it helps a little. It does not utilize any leak (data is readily available in HF low-res dataset) so I saw no reason not to",
          "votes": 3,
          "replies": [
            {
              "id": 2918382,
              "postDate": "2024-07-12T07:49:24.683Z",
              "content": "<p>if u don't use multiple rows together, I think it is fine.  If you use HF data to predict any location/time information \"to combine multiple instances together or calculate features from multiple rows\".  it is %100 leaky  and  the solution should be disqualified.  That is why the extension was made. </p>",
              "rawMarkdown": "if u don't use multiple rows together, I think it is fine.  If you use HF data to predict any location/time information \"to combine multiple instances together or calculate features from multiple rows\".  it is %100 leaky  and  the solution should be disqualified.  That is why the extension was made. ",
              "votes": 2
            },
            {
              "id": 2918422,
              "postDate": "2024-07-12T08:36:41.370Z",
              "content": "<p>No combination. It's a simple 1d--&gt;1d model. Auxiliary loss was used only because it helped the loss a bit, did not even though about complex pseudo-labeling and combination to 2D since I was sure that time stamp prediction would be far far from accurate enough</p>",
              "rawMarkdown": "No combination. It's a simple 1d-->1d model. Auxiliary loss was used only because it helped the loss a bit, did not even though about complex pseudo-labeling and combination to 2D since I was sure that time stamp prediction would be far far from accurate enough",
              "votes": -1
            },
            {
              "id": 2918582,
              "postDate": "2024-07-12T11:06:43.133Z",
              "content": "<p>indeed Auxiliary Loss is a 1D -&gt; 2D technique 😀</p>",
              "rawMarkdown": "indeed Auxiliary Loss is a 1D -> 2D technique 😀",
              "votes": 1
            },
            {
              "id": 2918606,
              "postDate": "2024-07-12T11:27:55.657Z",
              "content": "<p>someone can hide the leak so it can be read as 1D--&gt;1D . reorder the test data so bring neighbor locations/timestamps instances together then use for some layer with batch_first = False to extract information from other rows. Theoretically it is possible. Community/kaggle authorities also should check if test data is reordered.  </p>",
              "rawMarkdown": "someone can hide the leak so it can be read as 1D-->1D . reorder the test data so bring neighbor locations/timestamps instances together then use for some layer with batch_first = False to extract information from other rows. Theoretically it is possible. Community/kaggle authorities also should check if test data is reordered.  "
            },
            {
              "id": 2918707,
              "postDate": "2024-07-12T12:46:33.747Z",
              "content": "<blockquote>\n  <p>indeed Auxiliary Loss is a 1D -&gt; 2D technique 😀</p>\n</blockquote>\n<p>How exactly 🤣 I see only one row both in training and inference.<br>\nI did not even thought it would be in question, but now I start to worry that host decide that it is considered a leak too lol. Started to train without auxiliary anyway. I'm pretty sure I can still reach 0.793+ without, question is if remining 3 days are enough time for sufficient ensemble…</p>",
              "rawMarkdown": ">indeed Auxiliary Loss is a 1D -> 2D technique 😀\n\nHow exactly 🤣 I see only one row both in training and inference.\nI did not even thought it would be in question, but now I start to worry that host decide that it is considered a leak too lol. Started to train without auxiliary anyway. I'm pretty sure I can still reach 0.793+ without, question is if remining 3 days are enough time for sufficient ensemble...\n",
              "votes": -2
            },
            {
              "id": 2918733,
              "postDate": "2024-07-12T13:21:53.613Z",
              "content": "<p>Don't go into the grey zone here! :D</p>\n<p>All I can say is that since leak is intrinsic in the competition dataset, we are all using the leak, but some more transparent than another, since this is how NNs work.</p>",
              "rawMarkdown": "Don't go into the grey zone here! :D\n\nAll I can say is that since leak is intrinsic in the competition dataset, we are all using the leak, but some more transparent than another, since this is how NNs work.",
              "votes": -1
            },
            {
              "id": 2918763,
              "postDate": "2024-07-12T13:35:40.257Z",
              "content": "<blockquote>\n  <p>Don't go into the grey zone here! :D</p>\n</blockquote>\n<p>Auxiliary loss is a common practice and we had explicit timespace data for training set. It was not some secret. Moreover, host never said anything about it. You can hardly call it a grey zone… If it become forbidden now, 3 days before comp' end, I think I will justly feel robbed 😂 especially considering it was not THAT important, just another thing that gave a small bump so stayed in my model… but host is king. </p>",
              "rawMarkdown": ">Don't go into the grey zone here! :D\n\nAuxiliary loss is a common practice and we had explicit timespace data for training set. It was not some secret. Moreover, host never said anything about it. You can hardly call it a grey zone... If it become forbidden now, 3 days before comp' end, I think I will justly feel robbed 😂 especially considering it was not THAT important, just another thing that gave a small bump so stayed in my model... but host is king. "
            },
            {
              "id": 2918922,
              "postDate": "2024-07-12T15:41:21.157Z",
              "content": "<p>I agree that aux loss is legit, however there can be more tricks than that :D </p>\n<p>Anyways, will be waiting for your solution, super excited :]</p>",
              "rawMarkdown": "I agree that aux loss is legit, however there can be more tricks than that :D \n\nAnyways, will be waiting for your solution, super excited :]"
            }
          ]
        }
      ]
    },
    {
      "id": 2917758,
      "postDate": "2024-07-11T19:10:55.303Z",
      "content": "<p><a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a> </p>\n<p>However, I have already mentioned <a href=\"url\" target=\"_blank\">here</a>, I am concerned that it may get lost in the thread and the host might not find it, so I have decided to raise this discussion myself. Could you please answer the following questions?</p>\n<blockquote>\n  <p>Thank you for clarifying the measures to be taken when leaks are used. However, the criteria for what falls under a leak are still unclear, and I would like this to be clarified.<br>\n  From your intentions, we understand that using explicit leaks present in the test data (such as pbuf_ozone_2) to predict a single row from multiple rows is prohibited.<br>\n  However, for example:</p>\n  <ol>\n  <li>What about cases where location and timestamp information are inferred and restored using training data? This does not use the direct leak from pbuf_ozone_2. (The approach of inferring and restoring these information was known to be possible even before the leak involving pbuf_ozone_2 was discovered, and it was not prohibited.)</li>\n  <li>Additionally, what happens if the restored data is used as features for a single row without using information from multiple rows? (e.g. month, day, hour, lat, lon, etc.)</li>\n  </ol>\n</blockquote>",
      "rawMarkdown": "@jerrylin96 \n\nHowever, I have already mentioned [here](url), I am concerned that it may get lost in the thread and the host might not find it, so I have decided to raise this discussion myself. Could you please answer the following questions?\n\n>Thank you for clarifying the measures to be taken when leaks are used. However, the criteria for what falls under a leak are still unclear, and I would like this to be clarified.\nFrom your intentions, we understand that using explicit leaks present in the test data (such as pbuf_ozone_2) to predict a single row from multiple rows is prohibited.\nHowever, for example:\n1. What about cases where location and timestamp information are inferred and restored using training data? This does not use the direct leak from pbuf_ozone_2. (The approach of inferring and restoring these information was known to be possible even before the leak involving pbuf_ozone_2 was discovered, and it was not prohibited.)\n2. Additionally, what happens if the restored data is used as features for a single row without using information from multiple rows? (e.g. month, day, hour, lat, lon, etc.)",
      "votes": 9
    },
    {
      "id": 2918189,
      "postDate": "2024-07-12T04:18:43.783Z",
      "content": "<p>Using multiple atmospheric columns corresponding to the same timestep to predict tendencies for an atmospheric column (e.g. 2D -&gt; 1D regression or 2D -&gt; 2D regression) will be disqualified. This is virtually impossible to do unintentionally.</p>",
      "rawMarkdown": "Using multiple atmospheric columns corresponding to the same timestep to predict tendencies for an atmospheric column (e.g. 2D -> 1D regression or 2D -> 2D regression) will be disqualified. This is virtually impossible to do unintentionally.",
      "votes": 3,
      "replies": [
        {
          "id": 2918195,
          "postDate": "2024-07-12T04:23:04.740Z",
          "content": "<p>Thank you for your very quick response!<br>\nSo, does this mean that cases in which information linked to a single record (e.g. month, lat-lon) is included as just features are not disqualified?</p>",
          "rawMarkdown": "Thank you for your very quick response!\nSo, does this mean that cases in which information linked to a single record (e.g. month, lat-lon) is included as just features are not disqualified?",
          "replies": [
            {
              "id": 2918225,
              "postDate": "2024-07-12T05:13:56.907Z",
              "content": "<p>I don't see how you could infer timestamp without abusing some kind of leak we would not allow.</p>",
              "rawMarkdown": "I don't see how you could infer timestamp without abusing some kind of leak we would not allow."
            },
            {
              "id": 2918235,
              "postDate": "2024-07-12T05:24:20.540Z",
              "content": "<p>OK, I understand your intentions.<br>\nTo be honest, I was using features linked to the restored location_id after inferring it using the train data(The train data is ordered by location-timestamp), but I will stop using this and retrain a \"clean\" model.<br>\nThank you for your response.</p>",
              "rawMarkdown": "OK, I understand your intentions.\nTo be honest, I was using features linked to the restored location_id after inferring it using the train data(The train data is ordered by location-timestamp), but I will stop using this and retrain a \"clean\" model.\nThank you for your response.",
              "votes": 1
            },
            {
              "id": 2918386,
              "postDate": "2024-07-12T07:56:07.877Z",
              "content": "<p>If I understand correctly, we can use location and timestamp information on huggingface datasets since they are explicitly given. For example I can pretrain a model with location and timestamp heads using aux losses on that dataset, but I can't use timestamp or location information explicitly on competition dataset.</p>",
              "rawMarkdown": "If I understand correctly, we can use location and timestamp information on huggingface datasets since they are explicitly given. For example I can pretrain a model with location and timestamp heads using aux losses on that dataset, but I can't use timestamp or location information explicitly on competition dataset.",
              "votes": 2
            }
          ]
        },
        {
          "id": 2918592,
          "postDate": "2024-07-12T11:16:26.207Z",
          "content": "<p><a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a> <br>\nthat is why a clean and solid extension is needed here :D </p>",
          "rawMarkdown": "@jerrylin96 \nthat is why a clean and solid extension is needed here :D "
        }
      ]
    },
    {
      "id": 2918079,
      "postDate": "2024-07-12T03:02:39.637Z",
      "content": "<p>IMO, both of your examples uses the leak, so both of them will lead to a disqualification. Honestly, if someone wants to use the leak, there are hundreds of implications. Just don’t use any of those information, then there is no any risk.</p>",
      "rawMarkdown": "IMO, both of your examples uses the leak, so both of them will lead to a disqualification. Honestly, if someone wants to use the leak, there are hundreds of implications. Just don’t use any of those information, then there is no any risk.",
      "votes": 3,
      "replies": [
        {
          "id": 2918086,
          "postDate": "2024-07-12T03:16:12.227Z",
          "content": "<p>even batchnorm uses the leak. should models with batchnorm be disqualified?</p>",
          "rawMarkdown": "even batchnorm uses the leak. should models with batchnorm be disqualified?",
          "votes": -4,
          "replies": [
            {
              "id": 2918105,
              "postDate": "2024-07-12T03:33:46.937Z",
              "content": "<p>Batchnorm don’t use the leak. It normalize with the data distribution. But it didn’t try to recover time or position of each row.</p>",
              "rawMarkdown": "Batchnorm don’t use the leak. It normalize with the data distribution. But it didn’t try to recover time or position of each row.",
              "votes": 2
            },
            {
              "id": 2918136,
              "postDate": "2024-07-12T03:45:46.417Z",
              "content": "<p>In training, data is shuffled, all the rows in a batch has no correlation on time and position. In inference, it uses the mean and std got in training time. So BN has zero correlation with leak.</p>",
              "rawMarkdown": "In training, data is shuffled, all the rows in a batch has no correlation on time and position. In inference, it uses the mean and std got in training time. So BN has zero correlation with leak.",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2918094,
      "postDate": "2024-07-12T03:25:18.010Z",
      "content": "<p>IMHO,</p>\n<ol>\n<li>Don't explicitly input month, day, hour, lat, lon, location_id, etc. [But I guess actually model can learn it well, just a guess. I didn't do a Ablation experiment so far]</li>\n<li>use 1d -&gt; 1d model, do not use multi-rows inputs.</li>\n</ol>",
      "rawMarkdown": "IMHO,\n1. Don't explicitly input month, day, hour, lat, lon, location_id, etc. [But I guess actually model can learn it well, just a guess. I didn't do a Ablation experiment so far]\n2. use 1d -> 1d model, do not use multi-rows inputs.",
      "votes": 4
    },
    {
      "id": 2921024,
      "postDate": "2024-07-14T03:06:02.293Z",
      "content": "<p>IMO, both of your examples uses the leak, so both of them will lead to a disqualification. Honestly, if someone wants to use the leak, there are hundreds of implications. Just don’t use any of those information, then there is no any risk.</p>\n<p>Reply</p>\n<p>React</p>",
      "rawMarkdown": "IMO, both of your examples uses the leak, so both of them will lead to a disqualification. Honestly, if someone wants to use the leak, there are hundreds of implications. Just don’t use any of those information, then there is no any risk.\n\n\nReply\n\nReact",
      "votes": -1
    },
    {
      "id": 2918190,
      "postDate": "2024-07-12T04:19:01.350Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2918017,
      "postDate": "2024-07-12T00:23:34.630Z",
      "content": "<p>Since the organizers have repeatedly emphasized avoiding the use of data leaks and have even changed the test.csv .etc for this purpose, we shouldn't have to exploit the loopholes in the rules, and wouldn't it be nice to make this contest a little purer?</p>",
      "rawMarkdown": "Since the organizers have repeatedly emphasized avoiding the use of data leaks and have even changed the test.csv .etc for this purpose, we shouldn't have to exploit the loopholes in the rules, and wouldn't it be nice to make this contest a little purer?\n",
      "votes": 3,
      "isDeleted": true,
      "replies": [
        {
          "id": 2918049,
          "postDate": "2024-07-12T02:06:46.297Z",
          "content": "<p>I agree with your opinion. I do not wish to further explore or exploit leaks. Since the criteria for leaks are ambiguous, I do not want to unintentionally disqualify my team. If there were more time before the deadline, we could retrain the model to ensure its safety. However, with only 4 days left until the deadline, I would like to reconfirm the criteria for leaks.</p>",
          "rawMarkdown": "I agree with your opinion. I do not wish to further explore or exploit leaks. Since the criteria for leaks are ambiguous, I do not want to unintentionally disqualify my team. If there were more time before the deadline, we could retrain the model to ensure its safety. However, with only 4 days left until the deadline, I would like to reconfirm the criteria for leaks.",
          "votes": 4
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2918339,
      "author_name": "ishikei",
      "author_url": "",
      "post_date": "2024-07-12T06:50:23.093000",
      "content": "<p><a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a> <br>\nThank you for providing the guidelines. If you don’t mind, I would like to inquire further about the use of location and timestamp information when not intentionally analyzing the test set.</p>\n<p>When training with the provided low-resolution data, we can use the location and timestamp information from the training set. While we cannot use these features as inputs to the model without analyzing the test set, there are other ways to utilize them:</p>\n<ul>\n<li><strong>2D pretrain to 1D fine-tuning</strong>: For example, we can build a 2D to 1D model using location and timestamp in the train set (Of course, timestamp/location information of test data is not used. ), and then fine-tune it as a 1D to 1D model. This doesn’t utilize leakage from the test set, but it’s an example where part of the train pipeline includes a 2D to 1D component.</li>\n<li><strong>Auxiliary Loss</strong>: This technique involves adding a classification head to predict location and timestamp during model training. Auxiliary loss is a commonly used method in machine learning. The advantage of this method is that these meta-information features are not required during inference on test data or when deploying the model without actual location and timestamp data.</li>\n<li><strong>Difference from Statistics</strong>: In this method, we calculate statistical values for each location and timestamp within the training set. For example, we calculate the mean value of features for each of the 384 locations in the training set. During model training, we use the difference between these 384 mean values and the input features (1D). This method doesn’t require analyzing the test set and performs 1D to 1D prediction. However, from some perspectives, it might be considered as utilizing other data within the training set. We would appreciate your thoughts on this method.</li>\n</ul>\n<p>Additionally, there might be ways to implicitly utilize location information without explicitly using the above information. For example:</p>\n<ul>\n<li><strong>Clustering Techniques</strong>: We could apply clustering techniques to the combined train and test sets, build models for each cluster, or use cluster information as input to the model or as auxiliary loss. This method might unintentionally create clusters based on location information. There’s a concern that participants might unknowingly use this cluster information. In data analysis, it’s common to perform cluster analysis on provided data and treat each cluster separately (e.g., normalizing each cluster) if mixed distributions are found. We haven’t tried this, but if this method generates clusters containing location information and improves R2 score, how would it be judged in light of the competition rules?</li>\n</ul>\n<p>The current leakage issue is different from the typical target label leakage often seen on Kaggle. It relates to the inherent spatial and temporal nature of the data. Considering this, it might be difficult to completely exclude location and timestamp information from modeling, whether intentional or not. Furthermore, setting clear disqualification criteria might be challenging, though not impossible with extensive expertise and time for verification.</p>\n<p>Once again, thank you for the guidance provided. We would appreciate your explanation on the above points to minimize the risk of participants being unintentionally disqualified.</p>",
      "votes": 11,
      "replies": [
        {
          "id": 2918346,
          "author_name": "greySnow",
          "author_url": "",
          "post_date": "2024-07-12T07:02:17.827000",
          "content": "<p>I use auxiliary spacetime loss, it helps a little. It does not utilize any leak (data is readily available in HF low-res dataset) so I saw no reason not to</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2918382,
              "author_name": "Davut Polat",
              "author_url": "",
              "post_date": "2024-07-12T07:49:24.683000",
              "content": "<p>if u don't use multiple rows together, I think it is fine.  If you use HF data to predict any location/time information \"to combine multiple instances together or calculate features from multiple rows\".  it is %100 leaky  and  the solution should be disqualified.  That is why the extension was made. </p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2918422,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2024-07-12T08:36:41.370000",
              "content": "<p>No combination. It's a simple 1d--&gt;1d model. Auxiliary loss was used only because it helped the loss a bit, did not even though about complex pseudo-labeling and combination to 2D since I was sure that time stamp prediction would be far far from accurate enough</p>",
              "votes": -1,
              "replies": []
            },
            {
              "id": 2918582,
              "author_name": "steubk",
              "author_url": "",
              "post_date": "2024-07-12T11:06:43.133000",
              "content": "<p>indeed Auxiliary Loss is a 1D -&gt; 2D technique 😀</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2918606,
              "author_name": "Davut Polat",
              "author_url": "",
              "post_date": "2024-07-12T11:27:55.657000",
              "content": "<p>someone can hide the leak so it can be read as 1D--&gt;1D . reorder the test data so bring neighbor locations/timestamps instances together then use for some layer with batch_first = False to extract information from other rows. Theoretically it is possible. Community/kaggle authorities also should check if test data is reordered.  </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2918707,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2024-07-12T12:46:33.747000",
              "content": "<blockquote>\n  <p>indeed Auxiliary Loss is a 1D -&gt; 2D technique 😀</p>\n</blockquote>\n<p>How exactly 🤣 I see only one row both in training and inference.<br>\nI did not even thought it would be in question, but now I start to worry that host decide that it is considered a leak too lol. Started to train without auxiliary anyway. I'm pretty sure I can still reach 0.793+ without, question is if remining 3 days are enough time for sufficient ensemble…</p>",
              "votes": -2,
              "replies": []
            },
            {
              "id": 2918733,
              "author_name": "slime",
              "author_url": "",
              "post_date": "2024-07-12T13:21:53.613000",
              "content": "<p>Don't go into the grey zone here! :D</p>\n<p>All I can say is that since leak is intrinsic in the competition dataset, we are all using the leak, but some more transparent than another, since this is how NNs work.</p>",
              "votes": -1,
              "replies": []
            },
            {
              "id": 2918763,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2024-07-12T13:35:40.257000",
              "content": "<blockquote>\n  <p>Don't go into the grey zone here! :D</p>\n</blockquote>\n<p>Auxiliary loss is a common practice and we had explicit timespace data for training set. It was not some secret. Moreover, host never said anything about it. You can hardly call it a grey zone… If it become forbidden now, 3 days before comp' end, I think I will justly feel robbed 😂 especially considering it was not THAT important, just another thing that gave a small bump so stayed in my model… but host is king. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2918922,
              "author_name": "slime",
              "author_url": "",
              "post_date": "2024-07-12T15:41:21.157000",
              "content": "<p>I agree that aux loss is legit, however there can be more tricks than that :D </p>\n<p>Anyways, will be waiting for your solution, super excited :]</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2918189,
      "author_name": "Jerry Lin",
      "author_url": "",
      "post_date": "2024-07-12T04:18:43.783000",
      "content": "<p>Using multiple atmospheric columns corresponding to the same timestep to predict tendencies for an atmospheric column (e.g. 2D -&gt; 1D regression or 2D -&gt; 2D regression) will be disqualified. This is virtually impossible to do unintentionally.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2918195,
          "author_name": "Ryota",
          "author_url": "",
          "post_date": "2024-07-12T04:23:04.740000",
          "content": "<p>Thank you for your very quick response!<br>\nSo, does this mean that cases in which information linked to a single record (e.g. month, lat-lon) is included as just features are not disqualified?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2918225,
              "author_name": "Jerry Lin",
              "author_url": "",
              "post_date": "2024-07-12T05:13:56.907000",
              "content": "<p>I don't see how you could infer timestamp without abusing some kind of leak we would not allow.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2918235,
              "author_name": "Ryota",
              "author_url": "",
              "post_date": "2024-07-12T05:24:20.540000",
              "content": "<p>OK, I understand your intentions.<br>\nTo be honest, I was using features linked to the restored location_id after inferring it using the train data(The train data is ordered by location-timestamp), but I will stop using this and retrain a \"clean\" model.<br>\nThank you for your response.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2918386,
              "author_name": "Gunes Evitan",
              "author_url": "",
              "post_date": "2024-07-12T07:56:07.877000",
              "content": "<p>If I understand correctly, we can use location and timestamp information on huggingface datasets since they are explicitly given. For example I can pretrain a model with location and timestamp heads using aux losses on that dataset, but I can't use timestamp or location information explicitly on competition dataset.</p>",
              "votes": 2,
              "replies": []
            }
          ]
        },
        {
          "id": 2918592,
          "author_name": "Davut Polat",
          "author_url": "",
          "post_date": "2024-07-12T11:16:26.207000",
          "content": "<p><a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a> <br>\nthat is why a clean and solid extension is needed here :D </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2918079,
      "author_name": "ForcewithMe",
      "author_url": "",
      "post_date": "2024-07-12T03:02:39.637000",
      "content": "<p>IMO, both of your examples uses the leak, so both of them will lead to a disqualification. Honestly, if someone wants to use the leak, there are hundreds of implications. Just don’t use any of those information, then there is no any risk.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2918086,
          "author_name": "nymfree",
          "author_url": "",
          "post_date": "2024-07-12T03:16:12.227000",
          "content": "<p>even batchnorm uses the leak. should models with batchnorm be disqualified?</p>",
          "votes": -4,
          "replies": [
            {
              "id": 2918105,
              "author_name": "ForcewithMe",
              "author_url": "",
              "post_date": "2024-07-12T03:33:46.937000",
              "content": "<p>Batchnorm don’t use the leak. It normalize with the data distribution. But it didn’t try to recover time or position of each row.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2918136,
              "author_name": "ForcewithMe",
              "author_url": "",
              "post_date": "2024-07-12T03:45:46.417000",
              "content": "<p>In training, data is shuffled, all the rows in a batch has no correlation on time and position. In inference, it uses the mean and std got in training time. So BN has zero correlation with leak.</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2918094,
      "author_name": "ADAM.",
      "author_url": "",
      "post_date": "2024-07-12T03:25:18.010000",
      "content": "<p>IMHO,</p>\n<ol>\n<li>Don't explicitly input month, day, hour, lat, lon, location_id, etc. [But I guess actually model can learn it well, just a guess. I didn't do a Ablation experiment so far]</li>\n<li>use 1d -&gt; 1d model, do not use multi-rows inputs.</li>\n</ol>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 2921024,
      "author_name": "Ankit Pandey",
      "author_url": "",
      "post_date": "2024-07-14T03:06:02.293000",
      "content": "<p>IMO, both of your examples uses the leak, so both of them will lead to a disqualification. Honestly, if someone wants to use the leak, there are hundreds of implications. Just don’t use any of those information, then there is no any risk.</p>\n<p>Reply</p>\n<p>React</p>",
      "votes": -1,
      "replies": []
    },
    {
      "id": 2918190,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-07-12T04:19:01.350000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2918017,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-07-12T00:23:34.630000",
      "content": "<p>Since the organizers have repeatedly emphasized avoiding the use of data leaks and have even changed the test.csv .etc for this purpose, we shouldn't have to exploit the loopholes in the rules, and wouldn't it be nice to make this contest a little purer?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2918049,
          "author_name": "Ryota",
          "author_url": "",
          "post_date": "2024-07-12T02:06:46.297000",
          "content": "<p>I agree with your opinion. I do not wish to further explore or exploit leaks. Since the criteria for leaks are ambiguous, I do not want to unintentionally disqualify my team. If there were more time before the deadline, we could retrain the model to ensure its safety. However, with only 4 days left until the deadline, I would like to reconfirm the criteria for leaks.</p>",
          "votes": 4,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2918339": "@jerrylin96 \nThank you for providing the guidelines. If you don’t mind, I would like to inquire further about the use of location and timestamp information when not intentionally analyzing the test set.\n\nWhen training with the provided low-resolution data, we can use the location and timestamp information from the training set. While we cannot use these features as inputs to the model without analyzing the test set, there are other ways to utilize them:\n- **2D pretrain to 1D fine-tuning**: For example, we can build a 2D to 1D model using location and timestamp in the train set (Of course, timestamp/location information of test data is not used. ), and then fine-tune it as a 1D to 1D model. This doesn’t utilize leakage from the test set, but it’s an example where part of the train pipeline includes a 2D to 1D component.\n- **Auxiliary Loss**: This technique involves adding a classification head to predict location and timestamp during model training. Auxiliary loss is a commonly used method in machine learning. The advantage of this method is that these meta-information features are not required during inference on test data or when deploying the model without actual location and timestamp data.\n- **Difference from Statistics**: In this method, we calculate statistical values for each location and timestamp within the training set. For example, we calculate the mean value of features for each of the 384 locations in the training set. During model training, we use the difference between these 384 mean values and the input features (1D). This method doesn’t require analyzing the test set and performs 1D to 1D prediction. However, from some perspectives, it might be considered as utilizing other data within the training set. We would appreciate your thoughts on this method.\n\nAdditionally, there might be ways to implicitly utilize location information without explicitly using the above information. For example:\n- **Clustering Techniques**: We could apply clustering techniques to the combined train and test sets, build models for each cluster, or use cluster information as input to the model or as auxiliary loss. This method might unintentionally create clusters based on location information. There’s a concern that participants might unknowingly use this cluster information. In data analysis, it’s common to perform cluster analysis on provided data and treat each cluster separately (e.g., normalizing each cluster) if mixed distributions are found. We haven’t tried this, but if this method generates clusters containing location information and improves R2 score, how would it be judged in light of the competition rules?\n\nThe current leakage issue is different from the typical target label leakage often seen on Kaggle. It relates to the inherent spatial and temporal nature of the data. Considering this, it might be difficult to completely exclude location and timestamp information from modeling, whether intentional or not. Furthermore, setting clear disqualification criteria might be challenging, though not impossible with extensive expertise and time for verification.\n\nOnce again, thank you for the guidance provided. We would appreciate your explanation on the above points to minimize the risk of participants being unintentionally disqualified.",
    "2917758": "@jerrylin96 \n\nHowever, I have already mentioned [here](url), I am concerned that it may get lost in the thread and the host might not find it, so I have decided to raise this discussion myself. Could you please answer the following questions?\n\n>Thank you for clarifying the measures to be taken when leaks are used. However, the criteria for what falls under a leak are still unclear, and I would like this to be clarified.\nFrom your intentions, we understand that using explicit leaks present in the test data (such as pbuf_ozone_2) to predict a single row from multiple rows is prohibited.\nHowever, for example:\n1. What about cases where location and timestamp information are inferred and restored using training data? This does not use the direct leak from pbuf_ozone_2. (The approach of inferring and restoring these information was known to be possible even before the leak involving pbuf_ozone_2 was discovered, and it was not prohibited.)\n2. Additionally, what happens if the restored data is used as features for a single row without using information from multiple rows? (e.g. month, day, hour, lat, lon, etc.)",
    "2918189": "Using multiple atmospheric columns corresponding to the same timestep to predict tendencies for an atmospheric column (e.g. 2D -> 1D regression or 2D -> 2D regression) will be disqualified. This is virtually impossible to do unintentionally.",
    "2918079": "IMO, both of your examples uses the leak, so both of them will lead to a disqualification. Honestly, if someone wants to use the leak, there are hundreds of implications. Just don’t use any of those information, then there is no any risk.",
    "2918094": "IMHO,\n1. Don't explicitly input month, day, hour, lat, lon, location_id, etc. [But I guess actually model can learn it well, just a guess. I didn't do a Ablation experiment so far]\n2. use 1d -> 1d model, do not use multi-rows inputs.",
    "2921024": "IMO, both of your examples uses the leak, so both of them will lead to a disqualification. Honestly, if someone wants to use the leak, there are hundreds of implications. Just don’t use any of those information, then there is no any risk.\n\n\nReply\n\nReact",
    "2918190": "",
    "2918017": "Since the organizers have repeatedly emphasized avoiding the use of data leaks and have even changed the test.csv .etc for this purpose, we shouldn't have to exploit the loopholes in the rules, and wouldn't it be nice to make this contest a little purer?\n"
  }
}