{
  "id": 116478,
  "title": "Is this commenting out of one line allowed in stage 2?",
  "url": "/competitions/rsna-intracranial-hemorrhage-detection/discussion/116478",
  "author_name": "Carlo Lepelaars",
  "post_date": "2019-11-09T13:03:31.683000",
  "votes": 7,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Our team has been experimenting with two versions of our model. One is where we upload our trained weights from stage 1 and fine-tune them. This is the format we handed in with the zip archive from stage 1.</p>\n\n<p>The second is a version where we train the same model without uploading the trained weights from stage 1. NO adjustments in the preprocessing, architecture or hyperparameters were made. This requires us to comment out one line of code (see image).</p>\n\n<p>@juliaelliott Is this adjustment allowed in stage 2?</p>",
  "messages": [
    {
      "id": 669081,
      "postDate": "2019-11-09T13:03:31.683Z",
      "content": "<p>Our team has been experimenting with two versions of our model. One is where we upload our trained weights from stage 1 and fine-tune them. This is the format we handed in with the zip archive from stage 1.</p>\n\n<p>The second is a version where we train the same model without uploading the trained weights from stage 1. NO adjustments in the preprocessing, architecture or hyperparameters were made. This requires us to comment out one line of code (see image).</p>\n\n<p>@juliaelliott Is this adjustment allowed in stage 2?</p>",
      "rawMarkdown": "Our team has been experimenting with two versions of our model. One is where we upload our trained weights from stage 1 and fine-tune them. This is the format we handed in with the zip archive from stage 1.\n\nThe second is a version where we train the same model without uploading the trained weights from stage 1. NO adjustments in the preprocessing, architecture or hyperparameters were made. This requires us to comment out one line of code (see image).\n\n@juliaelliott Is this adjustment allowed in stage 2?\n\n",
      "votes": 7
    },
    {
      "id": 669138,
      "postDate": "2019-11-09T14:55:18.320Z",
      "content": "<p>Training the same model without uploading == training from scratch and as far as I know, training from scratch is allowed during stage 2 if you've uploaded the training code. IMHO</p>",
      "rawMarkdown": "Training the same model without uploading == training from scratch and as far as I know, training from scratch is allowed during stage 2 if you've uploaded the training code. IMHO",
      "votes": 3,
      "replies": [
        {
          "id": 669167,
          "postDate": "2019-11-09T15:51:32.300Z",
          "content": "<p>Seems reasonable, thanks! Training from scratch may even be more fair and transparent than using the pre-trained weights. Still curious to hear what the Kaggle hosts will say.</p>",
          "rawMarkdown": "Seems reasonable, thanks! Training from scratch may even be more fair and transparent than using the pre-trained weights. Still curious to hear what the Kaggle hosts will say.",
          "votes": 1
        }
      ]
    },
    {
      "id": 669376,
      "postDate": "2019-11-10T00:38:14.600Z",
      "content": "<p>I am still yet to learn exactly how strict and explicit the model upload needs to be. I tried learning this before the model upload deadline but did not get any responses.  ie if your code is tested, is it expected that Kaggle can just run it and get your result without speaking to you? Or will they contact you and allow you to talk them through it?</p>\n\n<p>The FAQ suggests that you need to upload everything required to reproduce your results and if there is any change at all then you should provide instructions for how to do so. So technically speaking, you must know exactly what you are going to submit and how it is generated before stage 1 ends and explicit instructions for doing so should be included in the model upload. </p>\n\n<p>You say your team has been experimenting with two versions? Well technically that experimentation should have been completed by the model upload deadline and instructions for generating both models should have been included. Any kind of experimentation after the deadline is technically a scientific change.</p>\n\n<p>But, as I said, I have tried and failed to learn exactly how strictly this is enforced. I mean commenting out a single line in the way you describe is no scientific change at all if you plan to create these two submissions. However learning that you should retrain the model on all data (deciding to do two submissions) and you therefore need to comment out that line is technically a scientific change, even if it is a pretty darn obvious thing to do. </p>\n\n<p>So if you didn't already plan to do two submissions, then deciding to do two after the submission deadline would be a scientific change?? </p>\n\n<p>From <a href=\"https://www.kaggle.com/two-stage-frequently-asked-questions\">FAQ</a>\n&gt; Depending on the competition's rules, you may be permitted to select two submissions for final scoring. If so, don't forget to include the code/instructions for reproducing both! It can be totally different code, or it can be the same code with instructions about the modifications you would make to generate each.</p>",
      "rawMarkdown": "I am still yet to learn exactly how strict and explicit the model upload needs to be. I tried learning this before the model upload deadline but did not get any responses.  ie if your code is tested, is it expected that Kaggle can just run it and get your result without speaking to you? Or will they contact you and allow you to talk them through it?\n\nThe FAQ suggests that you need to upload everything required to reproduce your results and if there is any change at all then you should provide instructions for how to do so. So technically speaking, you must know exactly what you are going to submit and how it is generated before stage 1 ends and explicit instructions for doing so should be included in the model upload. \n\nYou say your team has been experimenting with two versions? Well technically that experimentation should have been completed by the model upload deadline and instructions for generating both models should have been included. Any kind of experimentation after the deadline is technically a scientific change.\n\nBut, as I said, I have tried and failed to learn exactly how strictly this is enforced. I mean commenting out a single line in the way you describe is no scientific change at all if you plan to create these two submissions. However learning that you should retrain the model on all data (deciding to do two submissions) and you therefore need to comment out that line is technically a scientific change, even if it is a pretty darn obvious thing to do. \n\nSo if you didn't already plan to do two submissions, then deciding to do two after the submission deadline would be a scientific change?? \n\nFrom [FAQ](https://www.kaggle.com/two-stage-frequently-asked-questions)\n&gt; Depending on the competition's rules, you may be permitted to select two submissions for final scoring. If so, don't forget to include the code/instructions for reproducing both! It can be totally different code, or it can be the same code with instructions about the modifications you would make to generate each.",
      "votes": 1,
      "replies": [
        {
          "id": 669378,
          "postDate": "2019-11-10T00:42:39.650Z",
          "content": "<p>Hmm, the reproducibility argument is a good point! We did not provide these instructions beforehand in stage 1. </p>",
          "rawMarkdown": "Hmm, the reproducibility argument is a good point! We did not provide these instructions beforehand in stage 1. ",
          "votes": 1
        },
        {
          "id": 669381,
          "postDate": "2019-11-10T00:49:45.410Z",
          "content": "<p>Yeah from a pure letter of the law perspective, it seems to be technically not allowed. However deciding to retrain is such a small and obvious 'scientific' change that surely there would be this much leniency. </p>",
          "rawMarkdown": "Yeah from a pure letter of the law perspective, it seems to be technically not allowed. However deciding to retrain is such a small and obvious 'scientific' change that surely there would be this much leniency. ",
          "votes": 1
        },
        {
          "id": 669382,
          "postDate": "2019-11-10T00:51:44.573Z",
          "content": "<p>👍 Agree! It has been a weird competition so far. Still excited to see the end results and what the top teams built!</p>",
          "rawMarkdown": "👍 Agree! It has been a weird competition so far. Still excited to see the end results and what the top teams built!",
          "votes": 1
        }
      ]
    },
    {
      "id": 669228,
      "postDate": "2019-11-09T18:13:43.897Z",
      "content": "<p>From the Two-Stage FAQ: </p>\n\n<blockquote>\n  <p>We expect you may need to make some \"non scientific\" alterations, such as changes to path names, in order to create your submissions for the second stage. You are allowed to re-train your model (including the stage one data), but your code should not change.</p>\n</blockquote>\n\n<p>Seems like this should fall under \"non scientific alterations.\" If you want to be extra, extra safe, maybe you can just change <code>weights_path</code> to point to randomly initialized weights? Haha.</p>",
      "rawMarkdown": "From the Two-Stage FAQ: \n&gt; We expect you may need to make some \"non scientific\" alterations, such as changes to path names, in order to create your submissions for the second stage. You are allowed to re-train your model (including the stage one data), but your code should not change.\n\nSeems like this should fall under \"non scientific alterations.\" If you want to be extra, extra safe, maybe you can just change `weights_path` to point to randomly initialized weights? Haha.",
      "votes": 1,
      "replies": [
        {
          "id": 669340,
          "postDate": "2019-11-09T22:33:15.040Z",
          "content": "<p>Thanks for the insight! I agree with you that it should fall under a non scientific alteration, but technically speaking it is also a slight code change. 😁 </p>",
          "rawMarkdown": "Thanks for the insight! I agree with you that it should fall under a non scientific alteration, but technically speaking it is also a slight code change. 😁 ",
          "votes": 2
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 669138,
      "author_name": "Oleg Yaroshevskiy",
      "author_url": "",
      "post_date": "2019-11-09T14:55:18.320000",
      "content": "<p>Training the same model without uploading == training from scratch and as far as I know, training from scratch is allowed during stage 2 if you've uploaded the training code. IMHO</p>",
      "votes": 3,
      "replies": [
        {
          "id": 669167,
          "author_name": "Carlo Lepelaars",
          "author_url": "",
          "post_date": "2019-11-09T15:51:32.300000",
          "content": "<p>Seems reasonable, thanks! Training from scratch may even be more fair and transparent than using the pre-trained weights. Still curious to hear what the Kaggle hosts will say.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 669376,
      "author_name": "cherring",
      "author_url": "",
      "post_date": "2019-11-10T00:38:14.600000",
      "content": "<p>I am still yet to learn exactly how strict and explicit the model upload needs to be. I tried learning this before the model upload deadline but did not get any responses.  ie if your code is tested, is it expected that Kaggle can just run it and get your result without speaking to you? Or will they contact you and allow you to talk them through it?</p>\n\n<p>The FAQ suggests that you need to upload everything required to reproduce your results and if there is any change at all then you should provide instructions for how to do so. So technically speaking, you must know exactly what you are going to submit and how it is generated before stage 1 ends and explicit instructions for doing so should be included in the model upload. </p>\n\n<p>You say your team has been experimenting with two versions? Well technically that experimentation should have been completed by the model upload deadline and instructions for generating both models should have been included. Any kind of experimentation after the deadline is technically a scientific change.</p>\n\n<p>But, as I said, I have tried and failed to learn exactly how strictly this is enforced. I mean commenting out a single line in the way you describe is no scientific change at all if you plan to create these two submissions. However learning that you should retrain the model on all data (deciding to do two submissions) and you therefore need to comment out that line is technically a scientific change, even if it is a pretty darn obvious thing to do. </p>\n\n<p>So if you didn't already plan to do two submissions, then deciding to do two after the submission deadline would be a scientific change?? </p>\n\n<p>From <a href=\"https://www.kaggle.com/two-stage-frequently-asked-questions\">FAQ</a>\n&gt; Depending on the competition's rules, you may be permitted to select two submissions for final scoring. If so, don't forget to include the code/instructions for reproducing both! It can be totally different code, or it can be the same code with instructions about the modifications you would make to generate each.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 669378,
          "author_name": "Carlo Lepelaars",
          "author_url": "",
          "post_date": "2019-11-10T00:42:39.650000",
          "content": "<p>Hmm, the reproducibility argument is a good point! We did not provide these instructions beforehand in stage 1. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 669381,
          "author_name": "cherring",
          "author_url": "",
          "post_date": "2019-11-10T00:49:45.410000",
          "content": "<p>Yeah from a pure letter of the law perspective, it seems to be technically not allowed. However deciding to retrain is such a small and obvious 'scientific' change that surely there would be this much leniency. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 669382,
          "author_name": "Carlo Lepelaars",
          "author_url": "",
          "post_date": "2019-11-10T00:51:44.573000",
          "content": "<p>👍 Agree! It has been a weird competition so far. Still excited to see the end results and what the top teams built!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 669228,
      "author_name": "Ryan Epp",
      "author_url": "",
      "post_date": "2019-11-09T18:13:43.897000",
      "content": "<p>From the Two-Stage FAQ: </p>\n\n<blockquote>\n  <p>We expect you may need to make some \"non scientific\" alterations, such as changes to path names, in order to create your submissions for the second stage. You are allowed to re-train your model (including the stage one data), but your code should not change.</p>\n</blockquote>\n\n<p>Seems like this should fall under \"non scientific alterations.\" If you want to be extra, extra safe, maybe you can just change <code>weights_path</code> to point to randomly initialized weights? Haha.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 669340,
          "author_name": "Carlo Lepelaars",
          "author_url": "",
          "post_date": "2019-11-09T22:33:15.040000",
          "content": "<p>Thanks for the insight! I agree with you that it should fall under a non scientific alteration, but technically speaking it is also a slight code change. 😁 </p>",
          "votes": 2,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "669081": "Our team has been experimenting with two versions of our model. One is where we upload our trained weights from stage 1 and fine-tune them. This is the format we handed in with the zip archive from stage 1.\n\nThe second is a version where we train the same model without uploading the trained weights from stage 1. NO adjustments in the preprocessing, architecture or hyperparameters were made. This requires us to comment out one line of code (see image).\n\n@juliaelliott Is this adjustment allowed in stage 2?\n\n",
    "669138": "Training the same model without uploading == training from scratch and as far as I know, training from scratch is allowed during stage 2 if you've uploaded the training code. IMHO",
    "669376": "I am still yet to learn exactly how strict and explicit the model upload needs to be. I tried learning this before the model upload deadline but did not get any responses.  ie if your code is tested, is it expected that Kaggle can just run it and get your result without speaking to you? Or will they contact you and allow you to talk them through it?\n\nThe FAQ suggests that you need to upload everything required to reproduce your results and if there is any change at all then you should provide instructions for how to do so. So technically speaking, you must know exactly what you are going to submit and how it is generated before stage 1 ends and explicit instructions for doing so should be included in the model upload. \n\nYou say your team has been experimenting with two versions? Well technically that experimentation should have been completed by the model upload deadline and instructions for generating both models should have been included. Any kind of experimentation after the deadline is technically a scientific change.\n\nBut, as I said, I have tried and failed to learn exactly how strictly this is enforced. I mean commenting out a single line in the way you describe is no scientific change at all if you plan to create these two submissions. However learning that you should retrain the model on all data (deciding to do two submissions) and you therefore need to comment out that line is technically a scientific change, even if it is a pretty darn obvious thing to do. \n\nSo if you didn't already plan to do two submissions, then deciding to do two after the submission deadline would be a scientific change?? \n\nFrom [FAQ](https://www.kaggle.com/two-stage-frequently-asked-questions)\n&gt; Depending on the competition's rules, you may be permitted to select two submissions for final scoring. If so, don't forget to include the code/instructions for reproducing both! It can be totally different code, or it can be the same code with instructions about the modifications you would make to generate each.",
    "669228": "From the Two-Stage FAQ: \n&gt; We expect you may need to make some \"non scientific\" alterations, such as changes to path names, in order to create your submissions for the second stage. You are allowed to re-train your model (including the stage one data), but your code should not change.\n\nSeems like this should fall under \"non scientific alterations.\" If you want to be extra, extra safe, maybe you can just change `weights_path` to point to randomly initialized weights? Haha."
  }
}