{
  "id": 115600,
  "title": "Does stage1 submissions matter?",
  "url": "/competitions/rsna-intracranial-hemorrhage-detection/discussion/115600",
  "author_name": "Appian",
  "post_date": "2019-11-04T01:59:44.537000",
  "votes": 3,
  "comment_count": 14,
  "views": 0,
  "content": "<p>@juliaelliott</p>\n\n<p>I'm still trying to understand the format of 2 stage competitions and would appreciate if you can clarify.</p>\n\n<p>I saw someone mentioned that the procedure must produce the csv files for both stage1 and stage2 but the official rules do not mention anything about stage 1 submissions.\n<a href=\"https://www.kaggle.com/two-stage-frequently-asked-questions\">https://www.kaggle.com/two-stage-frequently-asked-questions</a></p>\n\n<p>What if you choose two submissions files which predictions are all 0 for stage 1? Does this change anything?</p>",
  "messages": [
    {
      "id": 664715,
      "postDate": "2019-11-04T05:37:38.740Z",
      "content": "<p>From what I understand:</p>\n\n<ul>\n<li>You can choose whatever you want in stage1, it doesnt really matter.</li>\n<li>You have to upload your code and models (go to \"Team\") and upload there. </li>\n<li>Once stage 2 is released, you have to use the uploaded code and models to generate predictions on stage-2.</li>\n<li>Stage-1 submissions will be invalidated as soon as stage2 is released.</li>\n<li>In stage-2 you will have 5 submissions per day again.</li>\n<li>You have to choose the final submissions from Stage 2 submissions.</li>\n</ul>",
      "rawMarkdown": "From what I understand:\n\n- You can choose whatever you want in stage1, it doesnt really matter.\n- You have to upload your code and models (go to \"Team\") and upload there. \n- Once stage 2 is released, you have to use the uploaded code and models to generate predictions on stage-2.\n- Stage-1 submissions will be invalidated as soon as stage2 is released.\n- In stage-2 you will have 5 submissions per day again.\n- You have to choose the final submissions from Stage 2 submissions.\n",
      "votes": 4,
      "replies": [
        {
          "id": 664758,
          "postDate": "2019-11-04T07:18:05.600Z",
          "content": "<p>How will Kaggle know, whether we are using uploaded code? I think Kaggle checks only if we are in the money</p>",
          "rawMarkdown": "How will Kaggle know, whether we are using uploaded code? I think Kaggle checks only if we are in the money",
          "votes": 2
        },
        {
          "id": 665043,
          "postDate": "2019-11-04T15:11:32.203Z",
          "content": "<p>Abhishek’s response is correct.</p>",
          "rawMarkdown": "Abhishek’s response is correct.",
          "votes": 3
        },
        {
          "id": 665045,
          "postDate": "2019-11-04T15:14:14.997Z",
          "content": "<p>What is the use of having 5 submissions per day for stage 2? my understanding is that you should only need one - it should be exactly what your uploaded code produces? </p>",
          "rawMarkdown": "What is the use of having 5 submissions per day for stage 2? my understanding is that you should only need one - it should be exactly what your uploaded code produces? "
        },
        {
          "id": 665105,
          "postDate": "2019-11-04T16:23:32.507Z",
          "content": "<p><a href=\"/cherring\">@cherring</a>, because you can retrain your model with stage-1 train plus stage-2 test data. I do not know if they will release the stage-1 test labels in this competition but they have done that in the past. Note that the structure and parameters of your models for stage-1 should remain the same for stage-2 retraining.</p>",
          "rawMarkdown": "@cherring, because you can retrain your model with stage-1 train plus stage-2 test data. I do not know if they will release the stage-1 test labels in this competition but they have done that in the past. Note that the structure and parameters of your models for stage-1 should remain the same for stage-2 retraining.",
          "votes": 1
        },
        {
          "id": 665339,
          "postDate": "2019-11-04T22:27:57.040Z",
          "content": "<p>Thank you Julia and Abhishek for clarification.</p>",
          "rawMarkdown": "Thank you Julia and Abhishek for clarification.",
          "votes": 1
        },
        {
          "id": 665478,
          "postDate": "2019-11-05T03:33:30.453Z",
          "content": "<p>how does Kaggle exactly make sure we use the same structure and parameters as stage 1 ?</p>",
          "rawMarkdown": "how does Kaggle exactly make sure we use the same structure and parameters as stage 1 ?",
          "votes": 1
        },
        {
          "id": 665492,
          "postDate": "2019-11-05T03:52:21.233Z",
          "content": "<p><a href=\"/sheriytm\">@sheriytm</a> The rules explititely state that you CAN retrain the model, however you can not do any hyper parameter tuning unless it is automated;</p>\n\n<p><code>You are allowed to re-train your model (including the stage one data), but your code should not change. You should not be doing any hyper parameter tuning in the second stage. Parameter tuning is permitted as long as it is fully automated.</code></p>\n\n<p>The code can really only produce one model..</p>",
          "rawMarkdown": "@sheriytm The rules explititely state that you CAN retrain the model, however you can not do any hyper parameter tuning unless it is automated;\n\n`You are allowed to re-train your model (including the stage one data), but your code should not change. You should not be doing any hyper parameter tuning in the second stage. Parameter tuning is permitted as long as it is fully automated.`\n\nThe code can really only produce one model.."
        },
        {
          "id": 665546,
          "postDate": "2019-11-05T05:24:11.220Z",
          "content": "<p>Do we have to re-upload our models at the end of stage 2?</p>",
          "rawMarkdown": "Do we have to re-upload our models at the end of stage 2?"
        },
        {
          "id": 665550,
          "postDate": "2019-11-05T05:27:55.520Z",
          "content": "<p><a href=\"/ekan825\">@ekan825</a> I believe since no model change is allowed there is no need to reupload the models at the end of stage 2.</p>",
          "rawMarkdown": "@ekan825 I believe since no model change is allowed there is no need to reupload the models at the end of stage 2.",
          "votes": 1
        },
        {
          "id": 665555,
          "postDate": "2019-11-05T05:30:02.813Z",
          "content": "<p><a href=\"/chinhuic\">@chinhuic</a> that's what I thought haha, what's the point of re-uploading when all we can change are the filepaths.\nI was just confused how they're supposed to verify the fact that everybody is using the same code from stage 1.\nIf I'm understanding correctly, only the winners (top 10) are required to upload their final models?</p>",
          "rawMarkdown": "@chinhuic that's what I thought haha, what's the point of re-uploading when all we can change are the filepaths.\nI was just confused how they're supposed to verify the fact that everybody is using the same code from stage 1.\nIf I'm understanding correctly, only the winners (top 10) are required to upload their final models?"
        },
        {
          "id": 665634,
          "postDate": "2019-11-05T07:46:07.993Z",
          "content": "<p>Everyone is required to have a model uploaded to progress to stage 2, but yeah I think only people in the money will actually have their model looked at. Would be impossible to reproduce everyone's model.</p>",
          "rawMarkdown": "Everyone is required to have a model uploaded to progress to stage 2, but yeah I think only people in the money will actually have their model looked at. Would be impossible to reproduce everyone's model.",
          "votes": 1
        }
      ]
    },
    {
      "id": 664630,
      "postDate": "2019-11-04T01:59:44.537Z",
      "content": "<p>@juliaelliott</p>\n\n<p>I'm still trying to understand the format of 2 stage competitions and would appreciate if you can clarify.</p>\n\n<p>I saw someone mentioned that the procedure must produce the csv files for both stage1 and stage2 but the official rules do not mention anything about stage 1 submissions.\n<a href=\"https://www.kaggle.com/two-stage-frequently-asked-questions\">https://www.kaggle.com/two-stage-frequently-asked-questions</a></p>\n\n<p>What if you choose two submissions files which predictions are all 0 for stage 1? Does this change anything?</p>",
      "rawMarkdown": "@juliaelliott\n\nI'm still trying to understand the format of 2 stage competitions and would appreciate if you can clarify.\n\nI saw someone mentioned that the procedure must produce the csv files for both stage1 and stage2 but the official rules do not mention anything about stage 1 submissions.\nhttps://www.kaggle.com/two-stage-frequently-asked-questions\n\nWhat if you choose two submissions files which predictions are all 0 for stage 1? Does this change anything?",
      "votes": 3
    },
    {
      "id": 664754,
      "postDate": "2019-11-04T07:11:40.643Z",
      "content": "<p>you are suppose to choose your stage 1  solution before stage 2?\nyou can choose any one.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F4cc433589c41fae2d7da17adc98d72d2%2FSelection_100.png?generation=1572851498435923&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "you are suppose to choose your stage 1  solution before stage 2?\nyou can choose any one.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F4cc433589c41fae2d7da17adc98d72d2%2FSelection_100.png?generation=1572851498435923&amp;alt=media)\n\n\n\n",
      "replies": [
        {
          "id": 665023,
          "postDate": "2019-11-04T14:45:22.120Z",
          "content": "<p>i plan to upload 3 notebooks  which i think i would use for stage 2. Out of 3 i might not have used  1 notebook completely before stage1  end date and only 2 i might have used ,but i  do think third  can give good result. So can i use on new data set to </p>",
          "rawMarkdown": "i plan to upload 3 notebooks  which i think i would use for stage 2. Out of 3 i might not have used  1 notebook completely before stage1  end date and only 2 i might have used ,but i  do think third  can give good result. So can i use on new data set to "
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 664715,
      "author_name": "Abhishek Thakur",
      "author_url": "",
      "post_date": "2019-11-04T05:37:38.740000",
      "content": "<p>From what I understand:</p>\n\n<ul>\n<li>You can choose whatever you want in stage1, it doesnt really matter.</li>\n<li>You have to upload your code and models (go to \"Team\") and upload there. </li>\n<li>Once stage 2 is released, you have to use the uploaded code and models to generate predictions on stage-2.</li>\n<li>Stage-1 submissions will be invalidated as soon as stage2 is released.</li>\n<li>In stage-2 you will have 5 submissions per day again.</li>\n<li>You have to choose the final submissions from Stage 2 submissions.</li>\n</ul>",
      "votes": 4,
      "replies": [
        {
          "id": 664758,
          "author_name": "Suraj Soni",
          "author_url": "",
          "post_date": "2019-11-04T07:18:05.600000",
          "content": "<p>How will Kaggle know, whether we are using uploaded code? I think Kaggle checks only if we are in the money</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 665043,
          "author_name": "Julia Elliott",
          "author_url": "",
          "post_date": "2019-11-04T15:11:32.203000",
          "content": "<p>Abhishek’s response is correct.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 665045,
          "author_name": "cherring",
          "author_url": "",
          "post_date": "2019-11-04T15:14:14.997000",
          "content": "<p>What is the use of having 5 submissions per day for stage 2? my understanding is that you should only need one - it should be exactly what your uploaded code produces? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 665105,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2019-11-04T16:23:32.507000",
          "content": "<p><a href=\"/cherring\">@cherring</a>, because you can retrain your model with stage-1 train plus stage-2 test data. I do not know if they will release the stage-1 test labels in this competition but they have done that in the past. Note that the structure and parameters of your models for stage-1 should remain the same for stage-2 retraining.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 665339,
          "author_name": "Appian",
          "author_url": "",
          "post_date": "2019-11-04T22:27:57.040000",
          "content": "<p>Thank you Julia and Abhishek for clarification.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 665478,
          "author_name": "Kirayue",
          "author_url": "",
          "post_date": "2019-11-05T03:33:30.453000",
          "content": "<p>how does Kaggle exactly make sure we use the same structure and parameters as stage 1 ?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 665492,
          "author_name": "cherring",
          "author_url": "",
          "post_date": "2019-11-05T03:52:21.233000",
          "content": "<p><a href=\"/sheriytm\">@sheriytm</a> The rules explititely state that you CAN retrain the model, however you can not do any hyper parameter tuning unless it is automated;</p>\n\n<p><code>You are allowed to re-train your model (including the stage one data), but your code should not change. You should not be doing any hyper parameter tuning in the second stage. Parameter tuning is permitted as long as it is fully automated.</code></p>\n\n<p>The code can really only produce one model..</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 665546,
          "author_name": "CE Kan",
          "author_url": "",
          "post_date": "2019-11-05T05:24:11.220000",
          "content": "<p>Do we have to re-upload our models at the end of stage 2?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 665550,
          "author_name": "ChinHuiC",
          "author_url": "",
          "post_date": "2019-11-05T05:27:55.520000",
          "content": "<p><a href=\"/ekan825\">@ekan825</a> I believe since no model change is allowed there is no need to reupload the models at the end of stage 2.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 665555,
          "author_name": "CE Kan",
          "author_url": "",
          "post_date": "2019-11-05T05:30:02.813000",
          "content": "<p><a href=\"/chinhuic\">@chinhuic</a> that's what I thought haha, what's the point of re-uploading when all we can change are the filepaths.\nI was just confused how they're supposed to verify the fact that everybody is using the same code from stage 1.\nIf I'm understanding correctly, only the winners (top 10) are required to upload their final models?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 665634,
          "author_name": "cherring",
          "author_url": "",
          "post_date": "2019-11-05T07:46:07.993000",
          "content": "<p>Everyone is required to have a model uploaded to progress to stage 2, but yeah I think only people in the money will actually have their model looked at. Would be impossible to reproduce everyone's model.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 664754,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2019-11-04T07:11:40.643000",
      "content": "<p>you are suppose to choose your stage 1  solution before stage 2?\nyou can choose any one.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F4cc433589c41fae2d7da17adc98d72d2%2FSelection_100.png?generation=1572851498435923&amp;alt=media\" alt=\"\"></p>",
      "votes": 0,
      "replies": [
        {
          "id": 665023,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2019-11-04T14:45:22.120000",
          "content": "<p>i plan to upload 3 notebooks  which i think i would use for stage 2. Out of 3 i might not have used  1 notebook completely before stage1  end date and only 2 i might have used ,but i  do think third  can give good result. So can i use on new data set to </p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "664715": "From what I understand:\n\n- You can choose whatever you want in stage1, it doesnt really matter.\n- You have to upload your code and models (go to \"Team\") and upload there. \n- Once stage 2 is released, you have to use the uploaded code and models to generate predictions on stage-2.\n- Stage-1 submissions will be invalidated as soon as stage2 is released.\n- In stage-2 you will have 5 submissions per day again.\n- You have to choose the final submissions from Stage 2 submissions.\n",
    "664630": "@juliaelliott\n\nI'm still trying to understand the format of 2 stage competitions and would appreciate if you can clarify.\n\nI saw someone mentioned that the procedure must produce the csv files for both stage1 and stage2 but the official rules do not mention anything about stage 1 submissions.\nhttps://www.kaggle.com/two-stage-frequently-asked-questions\n\nWhat if you choose two submissions files which predictions are all 0 for stage 1? Does this change anything?",
    "664754": "you are suppose to choose your stage 1  solution before stage 2?\nyou can choose any one.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F4cc433589c41fae2d7da17adc98d72d2%2FSelection_100.png?generation=1572851498435923&amp;alt=media)\n\n\n\n"
  }
}