{
  "id": 514626,
  "title": "Learner question: am I allowed to use test.csv along with train.csv to train the model?",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/514626",
  "author_name": "ibi",
  "post_date": "2024-06-25T02:09:47.386000",
  "votes": 0,
  "comment_count": 12,
  "views": 0,
  "content": "<p>This may sound obvious to experienced kagglers, but not clear to me. The test set was hidden and private in my other kaggle competitions.</p>\n<p>In the forum discussions, others imply \"they could exploit\". And to me, it sounds like this \"exploit\" is allowed in this competition as long as target variables are predicted from the same \"atmospheric column\". For example, the prediction of 368 columns of row test_0 in submission.csv must be made using only the 556 columns of row test_0 in test.csv. </p>\n<p>I appreciate any feedback on this. Thanks.</p>",
  "messages": [
    {
      "id": 2888651,
      "postDate": "2024-06-25T02:25:26.937Z",
      "content": "<blockquote>\n  <p>The test set was hidden and private in my other kaggle competitions.</p>\n</blockquote>\n<p>There are two types of completion in Kaggle. The type you mentioned is notebook competition. However, this competition is not. All test feature is public and you should expected to predict target columns. </p>\n<p>In the task description you wrote is mostly correct, except for the “exploit” part. Predicting row-by-row is allowed but multi-row-to-single-target approach seems not to be allowed according to hosts. The second approach possibly being considered as “exploit”. However they haven’t made clear definition of which is allowed and which is disallowed.</p>",
      "rawMarkdown": "> The test set was hidden and private in my other kaggle competitions.\n\nThere are two types of completion in Kaggle. The type you mentioned is notebook competition. However, this competition is not. All test feature is public and you should expected to predict target columns. \n\nIn the task description you wrote is mostly correct, except for the “exploit” part. Predicting row-by-row is allowed but multi-row-to-single-target approach seems not to be allowed according to hosts. The second approach possibly being considered as “exploit”. However they haven’t made clear definition of which is allowed and which is disallowed.",
      "votes": 1,
      "replies": [
        {
          "id": 2888661,
          "postDate": "2024-06-25T02:45:44.777Z",
          "content": "<p>In my personal feeling, both approach should be allowed because they both are normal, sane approach as long as not utilizing known/unknown leak in the test set.</p>\n<p>However, the definition of exploiting leak is really hard problem. As long as I know, this never be cleared in the past competition.</p>",
          "rawMarkdown": "In my personal feeling, both approach should be allowed because they both are normal, sane approach as long as not utilizing known/unknown leak in the test set.\n\nHowever, the definition of exploiting leak is really hard problem. As long as I know, this never be cleared in the past competition.",
          "votes": 2,
          "replies": [
            {
              "id": 2888693,
              "postDate": "2024-06-25T03:12:15.883Z",
              "content": "<p>Thanks for the kind explanation. </p>\n<p>In my understanding, the goal of this competition is to find a practical ml solution to replace the expensive simulation task. \"multi-row-to-single-target\" solution may not be useful for the desired task. </p>",
              "rawMarkdown": "Thanks for the kind explanation. \n\nIn my understanding, the goal of this competition is to find a practical ml solution to replace the expensive simulation task. \"multi-row-to-single-target\" solution may not be useful for the desired task. "
            },
            {
              "id": 2888724,
              "postDate": "2024-06-25T03:34:32.963Z",
              "content": "<p>Well, I partly agree the feeling of host that they don’t want solution of no use of their objectives, however, we are not host. We cannot know which solution is along their demand and which is not. So at least proper definition of acceptable models should be announced beforehand.</p>",
              "rawMarkdown": "Well, I partly agree the feeling of host that they don’t want solution of no use of their objectives, however, we are not host. We cannot know which solution is along their demand and which is not. So at least proper definition of acceptable models should be announced beforehand.",
              "votes": 1
            },
            {
              "id": 2889129,
              "postDate": "2024-06-25T09:19:06.583Z",
              "content": "<p>Kaggle is very simple, really. EVERYTHING is allowed…except what is written in the Rules tab.<br>\nDoes anything in said tab say, 'multi-row-to-single-target approach is forbidden'?<br>\nNo. Hence, it is allowed.<br>\nHowever, the host does not WANT this approach, so they scramble the data so we can't use it.<br>\nBut it does not make it forbidden. There is a difference there. It may look like a small one, but it's important.<br>\nIf I somehow manage to use the multi-row-to-single-target approach, for example, by predicting timespace on the test set (which can be done since we have spacetime data for train so we can build a model that predicts it for test), and then use this prediction to arrange test data for multi-row-to-single-target approach- well, my solution WOULD be eligible.<br>\nIf no one had revealed leak, the leak-based solutions WOULD have gotten the prices, even though they had gone against the intentions of the hosts.<br>\nIntentions do not make rules. Rules make rules. It's that simple.</p>",
              "rawMarkdown": "Kaggle is very simple, really. EVERYTHING is allowed...except what is written in the Rules tab.\nDoes anything in said tab say, 'multi-row-to-single-target approach is forbidden'?\nNo. Hence, it is allowed.\nHowever, the host does not WANT this approach, so they scramble the data so we can't use it.\nBut it does not make it forbidden. There is a difference there. It may look like a small one, but it's important.\nIf I somehow manage to use the multi-row-to-single-target approach, for example, by predicting timespace on the test set (which can be done since we have spacetime data for train so we can build a model that predicts it for test), and then use this prediction to arrange test data for multi-row-to-single-target approach- well, my solution WOULD be eligible.\nIf no one had revealed leak, the leak-based solutions WOULD have gotten the prices, even though they had gone against the intentions of the hosts.\nIntentions do not make rules. Rules make rules. It's that simple.",
              "votes": 2
            },
            {
              "id": 2889140,
              "postDate": "2024-06-25T09:28:31.507Z",
              "content": "<p><a href=\"https://www.kaggle.com/shlomoron\" target=\"_blank\">@shlomoron</a> Thanks. </p>\n<p>\"What host wants and what allowed in competition is different thing\".<br>\nYeah, its very simple criterion.</p>",
              "rawMarkdown": "@shlomoron Thanks. \n\n\"What host wants and what allowed in competition is different thing\".\nYeah, its very simple criterion.",
              "votes": 1
            },
            {
              "id": 2889195,
              "postDate": "2024-06-25T10:14:52.243Z",
              "content": "<blockquote>\n  <p>EVERYTHING is allowed…except what is written in the Rules tab.</p>\n</blockquote>\n<p>The Rules do not explicitly prohibit one to get into the host's computer, grab the test set ground truth, and use it to secure an ironclad first place.  Does that mean such atrocity is allowed?</p>\n<blockquote>\n  <p>Knowing this, it is fairly obvious that the intention of the competition designers was to not grant you access to this information. As Kaggle user <a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a> puts it, using this information indeed goes against the spirit of the competition.</p>\n</blockquote>\n<p>The quote above from <a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a> is a very natural reflection of the fact that it's simply impossible for any set of rules to foresee all possible cases and explicitly prohibit them.  To do that would require the ability to predict the future, which no one can.</p>\n<p>Rules are guideposts at best, and here is a pretty helpful one: If we want to do something and do not desire others to find out we're doing it, then chances are, this something shouldn't be allowed, whether the rules prohibit it or not.</p>",
              "rawMarkdown": ">EVERYTHING is allowed…except what is written in the Rules tab.\n\nThe Rules do not explicitly prohibit one to get into the host's computer, grab the test set ground truth, and use it to secure an ironclad first place.  Does that mean such atrocity is allowed?\n\n>Knowing this, it is fairly obvious that the intention of the competition designers was to not grant you access to this information. As Kaggle user @ryches puts it, using this information indeed goes against the spirit of the competition.\n\nThe quote above from @jerrylin96 is a very natural reflection of the fact that it's simply impossible for any set of rules to foresee all possible cases and explicitly prohibit them.  To do that would require the ability to predict the future, which no one can.\n\nRules are guideposts at best, and here is a pretty helpful one: If we want to do something and do not desire others to find out we're doing it, then chances are, this something shouldn't be allowed, whether the rules prohibit it or not."
            },
            {
              "id": 2889204,
              "postDate": "2024-06-25T10:29:26.927Z",
              "content": "<blockquote>\n  <p>The Rules do not explicitly prohibit one to get into the host's computer, grab the test set ground truth, and use it to secure an ironclad first place. Does that mean such atrocity is allowed?</p>\n</blockquote>\n<p>Well, this actually falls under external data ' However, you will ensure the External Data is publicly available and equally accessible to use by all participants of the Competition for purposes of the competition at no cost to the other participants.' <br>\nYour point may still be valid, but there is obvious cheating, and leak exploitations is not cheating. How do I know? Because leak exploitation was used before numerous times in Kaggle. It is known and accepted as part of the game.<br>\nSo the ROOT of your argument, i.e.:</p>\n<blockquote>\n  <p>If we want to do something and do not desire others to find out we're doing it, then chances are, this something shouldn't be allowed, whether the rules prohibit it or not.</p>\n</blockquote>\n<p>Is still invalid…</p>",
              "rawMarkdown": ">The Rules do not explicitly prohibit one to get into the host's computer, grab the test set ground truth, and use it to secure an ironclad first place. Does that mean such atrocity is allowed?\n\nWell, this actually falls under external data ' However, you will ensure the External Data is publicly available and equally accessible to use by all participants of the Competition for purposes of the competition at no cost to the other participants.' \nYour point may still be valid, but there is obvious cheating, and leak exploitations is not cheating. How do I know? Because leak exploitation was used before numerous times in Kaggle. It is known and accepted as part of the game.\nSo the ROOT of your argument, i.e.:\n>If we want to do something and do not desire others to find out we're doing it, then chances are, this something shouldn't be allowed, whether the rules prohibit it or not.\n\nIs still invalid...",
              "votes": 2
            },
            {
              "id": 2889416,
              "postDate": "2024-06-25T13:33:29.500Z",
              "content": "<blockquote>\n  <p>Well, this actually falls under external data ' However, you will ensure the External Data is publicly available and equally accessible to use by all participants of the Competition for purposes of the competition at no cost to the other participants.' </p>\n</blockquote>\n<p>So according to this reasoning, it's OK for one to grab the test set ground truth from the host's computer as long as one shares it publicly?</p>\n<p>I appreciate your faithfulness to Rules.  We have way too many Rule breakers in our world today.</p>\n<p>The root of my argument is very simple: No Rules are perfect, so when in doubt, maybe just ask.  Sure, one may win a Kaggle competition without violating the official Rules.  But if the solution doesn't exactly align with the way of the problem the host is trying to address (and in this case it's prediction not based on multi-row), then we're not advancing the state of the art as much as we could have.</p>",
              "rawMarkdown": ">Well, this actually falls under external data ' However, you will ensure the External Data is publicly available and equally accessible to use by all participants of the Competition for purposes of the competition at no cost to the other participants.' \n\nSo according to this reasoning, it's OK for one to grab the test set ground truth from the host's computer as long as one shares it publicly?\n\nI appreciate your faithfulness to Rules.  We have way too many Rule breakers in our world today.\n\nThe root of my argument is very simple: No Rules are perfect, so when in doubt, maybe just ask.  Sure, one may win a Kaggle competition without violating the official Rules.  But if the solution doesn't exactly align with the way of the problem the host is trying to address (and in this case it's prediction not based on multi-row), then we're not advancing the state of the art as much as we could have."
            },
            {
              "id": 2889441,
              "postDate": "2024-06-25T13:49:24.103Z",
              "content": "<p>You are missing the point. Even if your solution does not align with the host's aim, it wins the medal and price if it is within the rules. And THIS is what is important. Well, advancing sota is also important, but for most of us KAGGLERS medals and price come first :)<br>\nThen you go back to your 'grab solution from host computer' which is,  as I said is NOT THE SAME as leak exploitation which is a common and accepted practice on Kaggle and has already been used to win numerous medals and prices. You can love it, you can hate it, but this is how things works here.</p>",
              "rawMarkdown": "You are missing the point. Even if your solution does not align with the host's aim, it wins the medal and price if it is within the rules. And THIS is what is important. Well, advancing sota is also important, but for most of us KAGGLERS medals and price come first :)\nThen you go back to your 'grab solution from host computer' which is,  as I said is NOT THE SAME as leak exploitation which is a common and accepted practice on Kaggle and has already been used to win numerous medals and prices. You can love it, you can hate it, but this is how things works here.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2889453,
      "postDate": "2024-06-25T13:57:36.040Z",
      "content": "<blockquote>\n  <p>am I allowed to use test.csv along with train.csv to train the model</p>\n</blockquote>\n<p>If by \"to train\" you mean \"to carry out supervised learning\", I wonder how one could do supervised learning without ground truth.  But perhaps you've figured out a way of doing it.  I'd say, give it a shot!  I don't see how that would violate any sound machine learning practices.</p>\n<p>But every competition/application is different.  So of course if the host <a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a> says otherwise, his comment should override mine.</p>",
      "rawMarkdown": ">am I allowed to use test.csv along with train.csv to train the model\n\nIf by \"to train\" you mean \"to carry out supervised learning\", I wonder how one could do supervised learning without ground truth.  But perhaps you've figured out a way of doing it.  I'd say, give it a shot!  I don't see how that would violate any sound machine learning practices.\n\nBut every competition/application is different.  So of course if the host @jerrylin96 says otherwise, his comment should override mine.",
      "replies": [
        {
          "id": 2889945,
          "postDate": "2024-06-25T19:10:08.160Z",
          "content": "<p>Not directly by supervised learning, but by using the information from the test data. This is something that I never tried before.  In my other competitions, test data was simply not available at training time. But here test data is public and other kagglers seem to utilize it by \"exploring\" or \"exploiting\".  First thing that comes to my mind is to use for normalization, or even pseudo labeling. I will see if I can find a good use of it in some other ways.</p>",
          "rawMarkdown": "Not directly by supervised learning, but by using the information from the test data. This is something that I never tried before.  In my other competitions, test data was simply not available at training time. But here test data is public and other kagglers seem to utilize it by \"exploring\" or \"exploiting\".  First thing that comes to my mind is to use for normalization, or even pseudo labeling. I will see if I can find a good use of it in some other ways.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2888638,
      "postDate": "2024-06-25T02:09:47.387Z",
      "content": "<p>This may sound obvious to experienced kagglers, but not clear to me. The test set was hidden and private in my other kaggle competitions.</p>\n<p>In the forum discussions, others imply \"they could exploit\". And to me, it sounds like this \"exploit\" is allowed in this competition as long as target variables are predicted from the same \"atmospheric column\". For example, the prediction of 368 columns of row test_0 in submission.csv must be made using only the 556 columns of row test_0 in test.csv. </p>\n<p>I appreciate any feedback on this. Thanks.</p>",
      "rawMarkdown": "This may sound obvious to experienced kagglers, but not clear to me. The test set was hidden and private in my other kaggle competitions.\n\nIn the forum discussions, others imply \"they could exploit\". And to me, it sounds like this \"exploit\" is allowed in this competition as long as target variables are predicted from the same \"atmospheric column\". For example, the prediction of 368 columns of row test_0 in submission.csv must be made using only the 556 columns of row test_0 in test.csv. \n\nI appreciate any feedback on this. Thanks."
    }
  ],
  "comments": [
    {
      "id": 2888651,
      "author_name": "Bilzard",
      "author_url": "",
      "post_date": "2024-06-25T02:25:26.937000",
      "content": "<blockquote>\n  <p>The test set was hidden and private in my other kaggle competitions.</p>\n</blockquote>\n<p>There are two types of completion in Kaggle. The type you mentioned is notebook competition. However, this competition is not. All test feature is public and you should expected to predict target columns. </p>\n<p>In the task description you wrote is mostly correct, except for the “exploit” part. Predicting row-by-row is allowed but multi-row-to-single-target approach seems not to be allowed according to hosts. The second approach possibly being considered as “exploit”. However they haven’t made clear definition of which is allowed and which is disallowed.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2888661,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2024-06-25T02:45:44.777000",
          "content": "<p>In my personal feeling, both approach should be allowed because they both are normal, sane approach as long as not utilizing known/unknown leak in the test set.</p>\n<p>However, the definition of exploiting leak is really hard problem. As long as I know, this never be cleared in the past competition.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2888693,
              "author_name": "ibi",
              "author_url": "",
              "post_date": "2024-06-25T03:12:15.883000",
              "content": "<p>Thanks for the kind explanation. </p>\n<p>In my understanding, the goal of this competition is to find a practical ml solution to replace the expensive simulation task. \"multi-row-to-single-target\" solution may not be useful for the desired task. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2888724,
              "author_name": "Bilzard",
              "author_url": "",
              "post_date": "2024-06-25T03:34:32.963000",
              "content": "<p>Well, I partly agree the feeling of host that they don’t want solution of no use of their objectives, however, we are not host. We cannot know which solution is along their demand and which is not. So at least proper definition of acceptable models should be announced beforehand.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2889129,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2024-06-25T09:19:06.583000",
              "content": "<p>Kaggle is very simple, really. EVERYTHING is allowed…except what is written in the Rules tab.<br>\nDoes anything in said tab say, 'multi-row-to-single-target approach is forbidden'?<br>\nNo. Hence, it is allowed.<br>\nHowever, the host does not WANT this approach, so they scramble the data so we can't use it.<br>\nBut it does not make it forbidden. There is a difference there. It may look like a small one, but it's important.<br>\nIf I somehow manage to use the multi-row-to-single-target approach, for example, by predicting timespace on the test set (which can be done since we have spacetime data for train so we can build a model that predicts it for test), and then use this prediction to arrange test data for multi-row-to-single-target approach- well, my solution WOULD be eligible.<br>\nIf no one had revealed leak, the leak-based solutions WOULD have gotten the prices, even though they had gone against the intentions of the hosts.<br>\nIntentions do not make rules. Rules make rules. It's that simple.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2889140,
              "author_name": "Bilzard",
              "author_url": "",
              "post_date": "2024-06-25T09:28:31.507000",
              "content": "<p><a href=\"https://www.kaggle.com/shlomoron\" target=\"_blank\">@shlomoron</a> Thanks. </p>\n<p>\"What host wants and what allowed in competition is different thing\".<br>\nYeah, its very simple criterion.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2889195,
              "author_name": "Truth Seeker",
              "author_url": "",
              "post_date": "2024-06-25T10:14:52.243000",
              "content": "<blockquote>\n  <p>EVERYTHING is allowed…except what is written in the Rules tab.</p>\n</blockquote>\n<p>The Rules do not explicitly prohibit one to get into the host's computer, grab the test set ground truth, and use it to secure an ironclad first place.  Does that mean such atrocity is allowed?</p>\n<blockquote>\n  <p>Knowing this, it is fairly obvious that the intention of the competition designers was to not grant you access to this information. As Kaggle user <a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a> puts it, using this information indeed goes against the spirit of the competition.</p>\n</blockquote>\n<p>The quote above from <a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a> is a very natural reflection of the fact that it's simply impossible for any set of rules to foresee all possible cases and explicitly prohibit them.  To do that would require the ability to predict the future, which no one can.</p>\n<p>Rules are guideposts at best, and here is a pretty helpful one: If we want to do something and do not desire others to find out we're doing it, then chances are, this something shouldn't be allowed, whether the rules prohibit it or not.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2889204,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2024-06-25T10:29:26.927000",
              "content": "<blockquote>\n  <p>The Rules do not explicitly prohibit one to get into the host's computer, grab the test set ground truth, and use it to secure an ironclad first place. Does that mean such atrocity is allowed?</p>\n</blockquote>\n<p>Well, this actually falls under external data ' However, you will ensure the External Data is publicly available and equally accessible to use by all participants of the Competition for purposes of the competition at no cost to the other participants.' <br>\nYour point may still be valid, but there is obvious cheating, and leak exploitations is not cheating. How do I know? Because leak exploitation was used before numerous times in Kaggle. It is known and accepted as part of the game.<br>\nSo the ROOT of your argument, i.e.:</p>\n<blockquote>\n  <p>If we want to do something and do not desire others to find out we're doing it, then chances are, this something shouldn't be allowed, whether the rules prohibit it or not.</p>\n</blockquote>\n<p>Is still invalid…</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2889416,
              "author_name": "Truth Seeker",
              "author_url": "",
              "post_date": "2024-06-25T13:33:29.500000",
              "content": "<blockquote>\n  <p>Well, this actually falls under external data ' However, you will ensure the External Data is publicly available and equally accessible to use by all participants of the Competition for purposes of the competition at no cost to the other participants.' </p>\n</blockquote>\n<p>So according to this reasoning, it's OK for one to grab the test set ground truth from the host's computer as long as one shares it publicly?</p>\n<p>I appreciate your faithfulness to Rules.  We have way too many Rule breakers in our world today.</p>\n<p>The root of my argument is very simple: No Rules are perfect, so when in doubt, maybe just ask.  Sure, one may win a Kaggle competition without violating the official Rules.  But if the solution doesn't exactly align with the way of the problem the host is trying to address (and in this case it's prediction not based on multi-row), then we're not advancing the state of the art as much as we could have.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2889441,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2024-06-25T13:49:24.103000",
              "content": "<p>You are missing the point. Even if your solution does not align with the host's aim, it wins the medal and price if it is within the rules. And THIS is what is important. Well, advancing sota is also important, but for most of us KAGGLERS medals and price come first :)<br>\nThen you go back to your 'grab solution from host computer' which is,  as I said is NOT THE SAME as leak exploitation which is a common and accepted practice on Kaggle and has already been used to win numerous medals and prices. You can love it, you can hate it, but this is how things works here.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2889453,
      "author_name": "Truth Seeker",
      "author_url": "",
      "post_date": "2024-06-25T13:57:36.040000",
      "content": "<blockquote>\n  <p>am I allowed to use test.csv along with train.csv to train the model</p>\n</blockquote>\n<p>If by \"to train\" you mean \"to carry out supervised learning\", I wonder how one could do supervised learning without ground truth.  But perhaps you've figured out a way of doing it.  I'd say, give it a shot!  I don't see how that would violate any sound machine learning practices.</p>\n<p>But every competition/application is different.  So of course if the host <a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a> says otherwise, his comment should override mine.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2889945,
          "author_name": "ibi",
          "author_url": "",
          "post_date": "2024-06-25T19:10:08.160000",
          "content": "<p>Not directly by supervised learning, but by using the information from the test data. This is something that I never tried before.  In my other competitions, test data was simply not available at training time. But here test data is public and other kagglers seem to utilize it by \"exploring\" or \"exploiting\".  First thing that comes to my mind is to use for normalization, or even pseudo labeling. I will see if I can find a good use of it in some other ways.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2888651": "> The test set was hidden and private in my other kaggle competitions.\n\nThere are two types of completion in Kaggle. The type you mentioned is notebook competition. However, this competition is not. All test feature is public and you should expected to predict target columns. \n\nIn the task description you wrote is mostly correct, except for the “exploit” part. Predicting row-by-row is allowed but multi-row-to-single-target approach seems not to be allowed according to hosts. The second approach possibly being considered as “exploit”. However they haven’t made clear definition of which is allowed and which is disallowed.",
    "2889453": ">am I allowed to use test.csv along with train.csv to train the model\n\nIf by \"to train\" you mean \"to carry out supervised learning\", I wonder how one could do supervised learning without ground truth.  But perhaps you've figured out a way of doing it.  I'd say, give it a shot!  I don't see how that would violate any sound machine learning practices.\n\nBut every competition/application is different.  So of course if the host @jerrylin96 says otherwise, his comment should override mine.",
    "2888638": "This may sound obvious to experienced kagglers, but not clear to me. The test set was hidden and private in my other kaggle competitions.\n\nIn the forum discussions, others imply \"they could exploit\". And to me, it sounds like this \"exploit\" is allowed in this competition as long as target variables are predicted from the same \"atmospheric column\". For example, the prediction of 368 columns of row test_0 in submission.csv must be made using only the 556 columns of row test_0 in test.csv. \n\nI appreciate any feedback on this. Thanks."
  }
}