{
  "id": 519233,
  "title": "Does 'pseudo labeling' correspond to inference from multiple rows?",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/519233",
  "author_name": "deoxy",
  "post_date": "2024-07-10T09:01:55",
  "votes": 7,
  "comment_count": 11,
  "views": 0,
  "content": "<p>note:<br>\nI believe the host is referring to data linked to multiple \"sample_id\" when they mentions multiple-columns, but in order to avoid confusion with the terms Input Columns/Target Columns listed in the Data Tab, I will refer to it as multiple \"rows\" instead.</p>\n<p>I am under the impression that the host does not permit predictions based on multiple rows, regardless of whether or not that information is obtainable. However, there are many techniques that could potentially be construed as using predictions from multiple rows. From a competitive standpoint, I do not believe it is necessary to verify each and every one of those techniques. But I would like to request clear guidelines on whether or not pseudo labeling should be allowed. Pseudo labeling is a technique well known to many seasoned kagglers, and while it's easy to implement (at least for me), it is a technique that should be employed in order to improve scores.</p>\n<p>Pseudo labeling, in simple terms, is a method where a trained model makes predictions on a test dataset, then these predictions are added to the training data and used to retrain the model. Making predictions with a model trained using pseudo labeling may not seem like it constitutes predictions using multiple rows. However, if the score on public/private LB improves by using this method, it seems natural to assume that in order to reproduce equivalent performance in the real world, which is irrelevant to this competition, it would be necessary to rerun the three-step pipeline of inferring -&gt; retraining -&gt; inferring during the inference stage. One could argue that the score improvement was achieved due to an increase in the number of datasets. Therefore, it's not necessarily conclusive that this interpretation is correct. Therefore, I would like to request the host's opinion on whether or not utilizing pseudo labeling could be considered as inferring using multiple rows.</p>\n<p>Furthermore, I am hopeful that the host's opinion on this matter will guide other competitors in their selection of techniques and provide a fair judgment criterion after the competition ends.</p>",
  "messages": [
    {
      "id": 2914964,
      "postDate": "2024-07-10T09:01:55Z",
      "content": "<p>note:<br>\nI believe the host is referring to data linked to multiple \"sample_id\" when they mentions multiple-columns, but in order to avoid confusion with the terms Input Columns/Target Columns listed in the Data Tab, I will refer to it as multiple \"rows\" instead.</p>\n<p>I am under the impression that the host does not permit predictions based on multiple rows, regardless of whether or not that information is obtainable. However, there are many techniques that could potentially be construed as using predictions from multiple rows. From a competitive standpoint, I do not believe it is necessary to verify each and every one of those techniques. But I would like to request clear guidelines on whether or not pseudo labeling should be allowed. Pseudo labeling is a technique well known to many seasoned kagglers, and while it's easy to implement (at least for me), it is a technique that should be employed in order to improve scores.</p>\n<p>Pseudo labeling, in simple terms, is a method where a trained model makes predictions on a test dataset, then these predictions are added to the training data and used to retrain the model. Making predictions with a model trained using pseudo labeling may not seem like it constitutes predictions using multiple rows. However, if the score on public/private LB improves by using this method, it seems natural to assume that in order to reproduce equivalent performance in the real world, which is irrelevant to this competition, it would be necessary to rerun the three-step pipeline of inferring -&gt; retraining -&gt; inferring during the inference stage. One could argue that the score improvement was achieved due to an increase in the number of datasets. Therefore, it's not necessarily conclusive that this interpretation is correct. Therefore, I would like to request the host's opinion on whether or not utilizing pseudo labeling could be considered as inferring using multiple rows.</p>\n<p>Furthermore, I am hopeful that the host's opinion on this matter will guide other competitors in their selection of techniques and provide a fair judgment criterion after the competition ends.</p>",
      "rawMarkdown": "note:\nI believe the host is referring to data linked to multiple \"sample_id\" when they mentions multiple-columns, but in order to avoid confusion with the terms Input Columns/Target Columns listed in the Data Tab, I will refer to it as multiple \"rows\" instead.\n\n\n\nI am under the impression that the host does not permit predictions based on multiple rows, regardless of whether or not that information is obtainable. However, there are many techniques that could potentially be construed as using predictions from multiple rows. From a competitive standpoint, I do not believe it is necessary to verify each and every one of those techniques. But I would like to request clear guidelines on whether or not pseudo labeling should be allowed. Pseudo labeling is a technique well known to many seasoned kagglers, and while it's easy to implement (at least for me), it is a technique that should be employed in order to improve scores.\n\n\n\nPseudo labeling, in simple terms, is a method where a trained model makes predictions on a test dataset, then these predictions are added to the training data and used to retrain the model. Making predictions with a model trained using pseudo labeling may not seem like it constitutes predictions using multiple rows. However, if the score on public/private LB improves by using this method, it seems natural to assume that in order to reproduce equivalent performance in the real world, which is irrelevant to this competition, it would be necessary to rerun the three-step pipeline of inferring -> retraining -> inferring during the inference stage. One could argue that the score improvement was achieved due to an increase in the number of datasets. Therefore, it's not necessarily conclusive that this interpretation is correct. Therefore, I would like to request the host's opinion on whether or not utilizing pseudo labeling could be considered as inferring using multiple rows.\n\n\n\nFurthermore, I am hopeful that the host's opinion on this matter will guide other competitors in their selection of techniques and provide a fair judgment criterion after the competition ends.",
      "votes": 7
    },
    {
      "id": 2916340,
      "postDate": "2024-07-10T23:31:56.803Z",
      "content": "<p>We are primarily concerned about submissions that rely on using <strong>multiple atmospheric columns corresponding to the same timestep</strong> to predict the target column (i.e. 2D to 1D approaches). As far as I can tell, what you detailed seems fine.</p>",
      "rawMarkdown": "We are primarily concerned about submissions that rely on using **multiple atmospheric columns corresponding to the same timestep** to predict the target column (i.e. 2D to 1D approaches). As far as I can tell, what you detailed seems fine.",
      "votes": 3,
      "replies": [
        {
          "id": 2916430,
          "postDate": "2024-07-11T02:50:23.963Z",
          "content": "<p>Thank you for clearly outlining the host's competition design intentions. I am hoping that these standards will be appropriately applied!</p>",
          "rawMarkdown": "Thank you for clearly outlining the host's competition design intentions. I am hoping that these standards will be appropriately applied!"
        },
        {
          "id": 2916566,
          "postDate": "2024-07-11T05:16:59.253Z",
          "content": "<p>So, in this way, it seems entirely legitimate to include latitude, longitude, and timestamp as pseudo-labeling features extracted from pbuf_ozone_2, and then add some sort of layer normalization to 'unintentionally' exploit the leak.</p>",
          "rawMarkdown": "So, in this way, it seems entirely legitimate to include latitude, longitude, and timestamp as pseudo-labeling features extracted from pbuf_ozone_2, and then add some sort of layer normalization to 'unintentionally' exploit the leak.",
          "votes": 2
        },
        {
          "id": 2917005,
          "postDate": "2024-07-11T12:22:04.640Z",
          "content": "<p>How about this strategy?  Something along a two-person game mindset.</p>\n<p>(1) Create another track of this competition (call it Track 2).  Also, call the original competition (the regression problem) Track 1.</p>\n<p>(2) The goal of the competitors in Track 2 is this: Train an AI model so that given any test set solution S from Track 1 as an input, the model predicts the leak info from S <em>alone</em>.  Whoever can predict the leak info best will be winner(s) under Track 2.</p>\n<p>(3) A team is allowed to participate in either Track 1 or Track 2, but not both.</p>\n<p>(4) Decide which of the two Track's top winners are the very best, and reward them.  For the other Track, no rewards at all to any participant.</p>",
          "rawMarkdown": "How about this strategy?  Something along a two-person game mindset.\n\n(1) Create another track of this competition (call it Track 2).  Also, call the original competition (the regression problem) Track 1.\n\n(2) The goal of the competitors in Track 2 is this: Train an AI model so that given any test set solution S from Track 1 as an input, the model predicts the leak info from S *alone*.  Whoever can predict the leak info best will be winner(s) under Track 2.\n\n(3) A team is allowed to participate in either Track 1 or Track 2, but not both.\n\n(4) Decide which of the two Track's top winners are the very best, and reward them.  For the other Track, no rewards at all to any participant.\n",
          "replies": [
            {
              "id": 2917010,
              "postDate": "2024-07-11T12:23:59.690Z",
              "content": "<p>sounds tracky. </p>",
              "rawMarkdown": "sounds tracky. ",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2915122,
      "postDate": "2024-07-10T10:47:13.880Z",
      "content": "<p>Strictly saying, batch normalization is one of <em>multi-column</em> approach, since it uses information of other climate columns. So I believe it’s almost impossible to differentiate single column and multi column approach. Besides, probably all of participants’ models could implicitly learns location and timestamps through pbuf_ozone_2. The definition of what we can and what we cannot should become very subjective.</p>",
      "rawMarkdown": "Strictly saying, batch normalization is one of *multi-column* approach, since it uses information of other climate columns. So I believe it’s almost impossible to differentiate single column and multi column approach. Besides, probably all of participants’ models could implicitly learns location and timestamps through pbuf_ozone_2. The definition of what we can and what we cannot should become very subjective.",
      "votes": 3,
      "replies": [
        {
          "id": 2915126,
          "postDate": "2024-07-10T10:54:22.630Z",
          "content": "<p>oh damn, and i use layer norm everywhere! 😂</p>",
          "rawMarkdown": "oh damn, and i use layer norm everywhere! 😂"
        },
        {
          "id": 2915164,
          "postDate": "2024-07-10T11:20:08.863Z",
          "content": "<p>Batch normalization in eval mode is usually performed based on accumulated statistics, so it should not be affected by the variations between batches. There might indeed be a problem if you are inferring in train mode. It's simple to check this. Just verify that the behavior is the same when inferring with a batch size of 1 and when it's not. If the behavior differs, it might be beneficial to set the batch size to 1 for inference.<br>\ncf. <a href=\"https://pytorch.org/docs/stable/generated/torch.nn.BatchNorm1d.html\" target=\"_blank\">https://pytorch.org/docs/stable/generated/torch.nn.BatchNorm1d.html</a></p>\n<blockquote>\n  <p>Also by default, during training this layer keeps running estimates of its computed mean and variance, which are then used for normalization during evaluation. The running estimates are kept with a default momentum of 0.1.</p>\n</blockquote>\n<p>Also, this is my interpretation, but I believe that the host allows us to utilize the information even if timestamp/location information is implicitly embedded within a single row. What the host doesn’t allow is to output the prediction of a single row from multiple rows. However, using information that enables the use of multiple rows is interpreted as being permitted.</p>\n<p>In any case, I completely agree with the opinion that it is difficult, whether intentional or not, to differentiate completely between a single-row approach and multiple-row approach. Nonetheless, I believe that showing an example of the host’s guidelines before the end of the competition is beneficial for both the host and the participants.</p>",
          "rawMarkdown": "Batch normalization in eval mode is usually performed based on accumulated statistics, so it should not be affected by the variations between batches. There might indeed be a problem if you are inferring in train mode. It's simple to check this. Just verify that the behavior is the same when inferring with a batch size of 1 and when it's not. If the behavior differs, it might be beneficial to set the batch size to 1 for inference.\ncf. https://pytorch.org/docs/stable/generated/torch.nn.BatchNorm1d.html\n>Also by default, during training this layer keeps running estimates of its computed mean and variance, which are then used for normalization during evaluation. The running estimates are kept with a default momentum of 0.1.\n\nAlso, this is my interpretation, but I believe that the host allows us to utilize the information even if timestamp/location information is implicitly embedded within a single row. What the host doesn’t allow is to output the prediction of a single row from multiple rows. However, using information that enables the use of multiple rows is interpreted as being permitted.\n\nIn any case, I completely agree with the opinion that it is difficult, whether intentional or not, to differentiate completely between a single-row approach and multiple-row approach. Nonetheless, I believe that showing an example of the host’s guidelines before the end of the competition is beneficial for both the host and the participants.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2915002,
      "postDate": "2024-07-10T09:44:34.043Z",
      "content": "<p>Is Pseudo labeling working for this competition? I've not tried this.🤣</p>",
      "rawMarkdown": "Is Pseudo labeling working for this competition? I've not tried this.🤣"
    },
    {
      "id": 2915000,
      "postDate": "2024-07-10T09:43:50.480Z",
      "content": "<p>Hey thanks for this information, i am new in kaggle and i was not aware of this technique, do you mind explain me how that works? Do you use an ensemble of models to do the predictions and then you retrain the small models one by one? I dont get how a model can get better based on its own predictions since the targets should be a bit erroneous</p>",
      "rawMarkdown": "Hey thanks for this information, i am new in kaggle and i was not aware of this technique, do you mind explain me how that works? Do you use an ensemble of models to do the predictions and then you retrain the small models one by one? I dont get how a model can get better based on its own predictions since the targets should be a bit erroneous",
      "replies": [
        {
          "id": 2915194,
          "postDate": "2024-07-10T11:34:45.013Z",
          "content": "<p>Yes, it's exactly as you said. It's one of those deep learning 'magics' that you don't expect to work, but sometimes, it works well. Usually when you have a small training set and (comparatively) large test set.</p>",
          "rawMarkdown": "Yes, it's exactly as you said. It's one of those deep learning 'magics' that you don't expect to work, but sometimes, it works well. Usually when you have a small training set and (comparatively) large test set."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2916340,
      "author_name": "Jerry Lin",
      "author_url": "",
      "post_date": "2024-07-10T23:31:56.803000",
      "content": "<p>We are primarily concerned about submissions that rely on using <strong>multiple atmospheric columns corresponding to the same timestep</strong> to predict the target column (i.e. 2D to 1D approaches). As far as I can tell, what you detailed seems fine.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2916430,
          "author_name": "deoxy",
          "author_url": "",
          "post_date": "2024-07-11T02:50:23.963000",
          "content": "<p>Thank you for clearly outlining the host's competition design intentions. I am hoping that these standards will be appropriately applied!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2916566,
          "author_name": "steubk",
          "author_url": "",
          "post_date": "2024-07-11T05:16:59.253000",
          "content": "<p>So, in this way, it seems entirely legitimate to include latitude, longitude, and timestamp as pseudo-labeling features extracted from pbuf_ozone_2, and then add some sort of layer normalization to 'unintentionally' exploit the leak.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2917005,
          "author_name": "Truth Seeker",
          "author_url": "",
          "post_date": "2024-07-11T12:22:04.640000",
          "content": "<p>How about this strategy?  Something along a two-person game mindset.</p>\n<p>(1) Create another track of this competition (call it Track 2).  Also, call the original competition (the regression problem) Track 1.</p>\n<p>(2) The goal of the competitors in Track 2 is this: Train an AI model so that given any test set solution S from Track 1 as an input, the model predicts the leak info from S <em>alone</em>.  Whoever can predict the leak info best will be winner(s) under Track 2.</p>\n<p>(3) A team is allowed to participate in either Track 1 or Track 2, but not both.</p>\n<p>(4) Decide which of the two Track's top winners are the very best, and reward them.  For the other Track, no rewards at all to any participant.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2917010,
              "author_name": "Davut Polat",
              "author_url": "",
              "post_date": "2024-07-11T12:23:59.690000",
              "content": "<p>sounds tracky. </p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2915122,
      "author_name": "Bilzard",
      "author_url": "",
      "post_date": "2024-07-10T10:47:13.880000",
      "content": "<p>Strictly saying, batch normalization is one of <em>multi-column</em> approach, since it uses information of other climate columns. So I believe it’s almost impossible to differentiate single column and multi column approach. Besides, probably all of participants’ models could implicitly learns location and timestamps through pbuf_ozone_2. The definition of what we can and what we cannot should become very subjective.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2915126,
          "author_name": "Vasilis",
          "author_url": "",
          "post_date": "2024-07-10T10:54:22.630000",
          "content": "<p>oh damn, and i use layer norm everywhere! 😂</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2915164,
          "author_name": "deoxy",
          "author_url": "",
          "post_date": "2024-07-10T11:20:08.863000",
          "content": "<p>Batch normalization in eval mode is usually performed based on accumulated statistics, so it should not be affected by the variations between batches. There might indeed be a problem if you are inferring in train mode. It's simple to check this. Just verify that the behavior is the same when inferring with a batch size of 1 and when it's not. If the behavior differs, it might be beneficial to set the batch size to 1 for inference.<br>\ncf. <a href=\"https://pytorch.org/docs/stable/generated/torch.nn.BatchNorm1d.html\" target=\"_blank\">https://pytorch.org/docs/stable/generated/torch.nn.BatchNorm1d.html</a></p>\n<blockquote>\n  <p>Also by default, during training this layer keeps running estimates of its computed mean and variance, which are then used for normalization during evaluation. The running estimates are kept with a default momentum of 0.1.</p>\n</blockquote>\n<p>Also, this is my interpretation, but I believe that the host allows us to utilize the information even if timestamp/location information is implicitly embedded within a single row. What the host doesn’t allow is to output the prediction of a single row from multiple rows. However, using information that enables the use of multiple rows is interpreted as being permitted.</p>\n<p>In any case, I completely agree with the opinion that it is difficult, whether intentional or not, to differentiate completely between a single-row approach and multiple-row approach. Nonetheless, I believe that showing an example of the host’s guidelines before the end of the competition is beneficial for both the host and the participants.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2915002,
      "author_name": "ADAM.",
      "author_url": "",
      "post_date": "2024-07-10T09:44:34.043000",
      "content": "<p>Is Pseudo labeling working for this competition? I've not tried this.🤣</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2915000,
      "author_name": "Vasilis",
      "author_url": "",
      "post_date": "2024-07-10T09:43:50.480000",
      "content": "<p>Hey thanks for this information, i am new in kaggle and i was not aware of this technique, do you mind explain me how that works? Do you use an ensemble of models to do the predictions and then you retrain the small models one by one? I dont get how a model can get better based on its own predictions since the targets should be a bit erroneous</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2915194,
          "author_name": "greySnow",
          "author_url": "",
          "post_date": "2024-07-10T11:34:45.013000",
          "content": "<p>Yes, it's exactly as you said. It's one of those deep learning 'magics' that you don't expect to work, but sometimes, it works well. Usually when you have a small training set and (comparatively) large test set.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2914964": "note:\nI believe the host is referring to data linked to multiple \"sample_id\" when they mentions multiple-columns, but in order to avoid confusion with the terms Input Columns/Target Columns listed in the Data Tab, I will refer to it as multiple \"rows\" instead.\n\n\n\nI am under the impression that the host does not permit predictions based on multiple rows, regardless of whether or not that information is obtainable. However, there are many techniques that could potentially be construed as using predictions from multiple rows. From a competitive standpoint, I do not believe it is necessary to verify each and every one of those techniques. But I would like to request clear guidelines on whether or not pseudo labeling should be allowed. Pseudo labeling is a technique well known to many seasoned kagglers, and while it's easy to implement (at least for me), it is a technique that should be employed in order to improve scores.\n\n\n\nPseudo labeling, in simple terms, is a method where a trained model makes predictions on a test dataset, then these predictions are added to the training data and used to retrain the model. Making predictions with a model trained using pseudo labeling may not seem like it constitutes predictions using multiple rows. However, if the score on public/private LB improves by using this method, it seems natural to assume that in order to reproduce equivalent performance in the real world, which is irrelevant to this competition, it would be necessary to rerun the three-step pipeline of inferring -> retraining -> inferring during the inference stage. One could argue that the score improvement was achieved due to an increase in the number of datasets. Therefore, it's not necessarily conclusive that this interpretation is correct. Therefore, I would like to request the host's opinion on whether or not utilizing pseudo labeling could be considered as inferring using multiple rows.\n\n\n\nFurthermore, I am hopeful that the host's opinion on this matter will guide other competitors in their selection of techniques and provide a fair judgment criterion after the competition ends.",
    "2916340": "We are primarily concerned about submissions that rely on using **multiple atmospheric columns corresponding to the same timestep** to predict the target column (i.e. 2D to 1D approaches). As far as I can tell, what you detailed seems fine.",
    "2915122": "Strictly saying, batch normalization is one of *multi-column* approach, since it uses information of other climate columns. So I believe it’s almost impossible to differentiate single column and multi column approach. Besides, probably all of participants’ models could implicitly learns location and timestamps through pbuf_ozone_2. The definition of what we can and what we cannot should become very subjective.",
    "2915002": "Is Pseudo labeling working for this competition? I've not tried this.🤣",
    "2915000": "Hey thanks for this information, i am new in kaggle and i was not aware of this technique, do you mind explain me how that works? Do you use an ensemble of models to do the predictions and then you retrain the small models one by one? I dont get how a model can get better based on its own predictions since the targets should be a bit erroneous"
  }
}