{
  "id": 113645,
  "title": "summary of Two-stage competitions",
  "url": "/competitions/rsna-intracranial-hemorrhage-detection/discussion/113645",
  "author_name": "Tian Bingyang",
  "post_date": "2019-10-21T03:49:10.646000",
  "votes": 4,
  "comment_count": 18,
  "views": 0,
  "content": "<p>in my opinion.\nWe should upload 3 things:\nmodel,\nthe ipynb/py file which produced the model\nsubmission.csv</p>\n\n<p>Am I right?\nThanks for your help~!</p>",
  "messages": [
    {
      "id": 655128,
      "postDate": "2019-10-22T17:46:05.633Z",
      "content": "<p>See <a href=\"https://www.kaggle.com/two-stage-frequently-asked-questions\">two-stage FAQ</a> for details on what to upload. </p>\n\n<p>Yes, you can upload the ipynb or whatever format your model is in. The output weights are not required. <code>submission.csv</code> is not required, since we already have that when you make your submission. You just need to ensure that the model(s) you are uploading match what you will be using to generate the final submissions that you ultimately select in stage 2.</p>",
      "rawMarkdown": "See [two-stage FAQ](https://www.kaggle.com/two-stage-frequently-asked-questions) for details on what to upload. \n\nYes, you can upload the ipynb or whatever format your model is in. The output weights are not required. `submission.csv` is not required, since we already have that when you make your submission. You just need to ensure that the model(s) you are uploading match what you will be using to generate the final submissions that you ultimately select in stage 2.",
      "votes": 3,
      "replies": [
        {
          "id": 655326,
          "postDate": "2019-10-22T23:22:45.333Z",
          "content": "<p>thank you !</p>",
          "rawMarkdown": "thank you !"
        },
        {
          "id": 655415,
          "postDate": "2019-10-23T03:03:42.683Z",
          "content": "<p>Thanks for your replies,\n<a href=\"/juliaelliott\">@juliaelliott</a>\nForgive my poor English please:\nHere some part from two-stage FAQ :\n`\nWhat happens to my pre-trained model?\n<strong>If you are using a pre-trained for which you don't have the source code, you must include the model as part of your upload.</strong></p>\n\n<p>What if my submission is too big?\nOur uploader will handle reasonably large files. If you still think your <strong>model will be too large</strong>, you can instead upload a checksum of your archive file (such as an md5 or sha hash). Note that if you do win, you will still have to upload your code. If you upload a checksum, finish in a prize spot, and are unable to subsequently provide an archive that matches the checksum, you will be removed from the competition.`</p>\n\n<hr>\n\n<p><strong>Almost all of us use pretrained model.</strong></p>\n\n<p>So what't the meaning of \"model\" in above document?\nit refers to \".py\" file or \".pth\" weights file?</p>\n\n<p>Does it mean uploading the** file which produces the pretrained model**?</p>\n\n<p>if if refers to \".py/ipynb\" file,there's no possibility that <strong>model will be too large</strong>.\nCorrect me where I am wrong,please,\nThanks for your help~!</p>",
          "rawMarkdown": "Thanks for your replies,\n@juliaelliott\nForgive my poor English please:\nHere some part from two-stage FAQ :\n`\nWhat happens to my pre-trained model?\n**If you are using a pre-trained for which you don't have the source code, you must include the model as part of your upload.**\n\nWhat if my submission is too big?\nOur uploader will handle reasonably large files. If you still think your **model will be too large**, you can instead upload a checksum of your archive file (such as an md5 or sha hash). Note that if you do win, you will still have to upload your code. If you upload a checksum, finish in a prize spot, and are unable to subsequently provide an archive that matches the checksum, you will be removed from the competition.`\n\n\n----------------------------------------------------------------\n\n\n**Almost all of us use pretrained model.**\n\nSo what't the meaning of \"model\" in above document?\nit refers to \".py\" file or \".pth\" weights file?\n\nDoes it mean uploading the** file which produces the pretrained model**?\n\nif if refers to \".py/ipynb\" file,there's no possibility that **model will be too large**.\nCorrect me where I am wrong,please,\nThanks for your help~!"
        },
        {
          "id": 655423,
          "postDate": "2019-10-23T03:15:42.493Z",
          "content": "<p><a href=\"/xinyan989\">@xinyan989</a> could you help me explain it?\nThe Kaggle team said only code is necessary.\nBut the document said:\n<strong>If you are using a pre-trained for which you don't have the source code, you must include the model as part of your upload.</strong></p>\n\n<p>There are two meanings of model in <a href=\"https://www.kaggle.com/two-stage-frequently-asked-questions\">Two-stage FAQ</a>\n1.py/ipynb file.\n2.weights file.</p>",
          "rawMarkdown": "@xinyan989 could you help me explain it?\nThe Kaggle team said only code is necessary.\nBut the document said:\n**If you are using a pre-trained for which you don't have the source code, you must include the model as part of your upload.**\n\n\nThere are two meanings of model in [Two-stage FAQ](https://www.kaggle.com/two-stage-frequently-asked-questions)\n1.py/ipynb file.\n2.weights file."
        },
        {
          "id": 656046,
          "postDate": "2019-10-23T20:44:25.137Z",
          "content": "<p>A README file indicating which pre-trained models you're using and the links to them will be sufficient to include within your zipped upload.</p>",
          "rawMarkdown": "A README file indicating which pre-trained models you're using and the links to them will be sufficient to include within your zipped upload.",
          "votes": 1
        },
        {
          "id": 656228,
          "postDate": "2019-10-24T02:58:41.347Z",
          "content": "<p>Very clear,Much Thanks</p>",
          "rawMarkdown": "Very clear,Much Thanks"
        },
        {
          "id": 659894,
          "postDate": "2019-10-28T11:56:27.593Z",
          "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a>\nSorry about two new question:\n1.\nIf I'm busy at stage1,\nbut I upload the model .py file and the sample_submission.csv(download from kaggle)\nand train my  weights file(.pth) in stage2,\nIs this allowed?\n2.\nIn stage2,I use new datasets,because the labels of test sets is released.\nThe Rule said:no code is allowed to change in stage2.\nSo ,Is it allowed to use code which merge the train and test(from stage1) in stage2?\nMuch Thanks for your help.</p>",
          "rawMarkdown": "@juliaelliott\nSorry about two new question:\n1.\nIf I'm busy at stage1,\nbut I upload the model .py file and the sample_submission.csv(download from kaggle)\nand train my  weights file(.pth) in stage2,\nIs this allowed?\n2.\nIn stage2,I use new datasets,because the labels of test sets is released.\nThe Rule said:no code is allowed to change in stage2.\nSo ,Is it allowed to use code which merge the train and test(from stage1) in stage2?\nMuch Thanks for your help."
        },
        {
          "id": 660536,
          "postDate": "2019-10-29T09:33:35.207Z",
          "content": "<p>Hello, could I ask if semi-supervised learning with test data is allowed?</p>",
          "rawMarkdown": "Hello, could I ask if semi-supervised learning with test data is allowed?"
        }
      ]
    },
    {
      "id": 653850,
      "postDate": "2019-10-21T03:49:10.647Z",
      "content": "<p>in my opinion.\nWe should upload 3 things:\nmodel,\nthe ipynb/py file which produced the model\nsubmission.csv</p>\n\n<p>Am I right?\nThanks for your help~!</p>",
      "rawMarkdown": "in my opinion.\nWe should upload 3 things:\nmodel,\nthe ipynb/py file which produced the model\nsubmission.csv\n\n \nAm I right?\nThanks for your help~!",
      "votes": 4
    },
    {
      "id": 661407,
      "postDate": "2019-10-30T08:56:32.597Z",
      "content": "<p>Just latching on to the questions here, regarding the possible retraining of the models:</p>\n\n<p>What happens if I want to change something in my code in the second stage?\nWe expect you may need to make some \"non scientific\" alterations, such as changes to path names, in order to create your submissions for the second stage. You are allowed to re-train your model (including the stage one data), but your code should not change. You should not be doing any hyper parameter tuning in the second stage. Parameter tuning is permitted as long as it is fully automated.</p>\n\n<p>Are Learning Rate and number of Epochs also considered hyper-parameters, which can only be tuned automatically? </p>",
      "rawMarkdown": "Just latching on to the questions here, regarding the possible retraining of the models:\n\nWhat happens if I want to change something in my code in the second stage?\nWe expect you may need to make some \"non scientific\" alterations, such as changes to path names, in order to create your submissions for the second stage. You are allowed to re-train your model (including the stage one data), but your code should not change. You should not be doing any hyper parameter tuning in the second stage. Parameter tuning is permitted as long as it is fully automated.\n\nAre Learning Rate and number of Epochs also considered hyper-parameters, which can only be tuned automatically? ",
      "replies": [
        {
          "id": 661409,
          "postDate": "2019-10-30T09:02:32.357Z",
          "content": "<p>\"You should not be doing any hyper parameter tuning in the second stage. Parameter tuning is permitted as long as it is fully automated.\"\nI think this sentence is contradictory</p>",
          "rawMarkdown": "\"You should not be doing any hyper parameter tuning in the second stage. Parameter tuning is permitted as long as it is fully automated.\"\nI think this sentence is contradictory"
        },
        {
          "id": 661447,
          "postDate": "2019-10-30T10:20:43.050Z",
          "content": "<p>In this context I believe it means that if you for example use a logistic regression and retrain your models etc. <em>you</em> should not decide manually on a new regularization (i.e. C). \nIf you, however, you sklearn's LogisticRegressionCV, which automatically chooses C based on cross-validation, that's fine.</p>\n\n<p>For Neural Networks,  you could use find_lr of the fastai learner to find a good learning rate, and also programatically use it for re-training. Or use Early Stopping with like a million training Epochs or something to decide on the number of Epochs for re-training.\nBut I think some examples on what constitutes as hyper-parameters in certain contexts would be nice. </p>",
          "rawMarkdown": "In this context I believe it means that if you for example use a logistic regression and retrain your models etc. *you* should not decide manually on a new regularization (i.e. C). \nIf you, however, you sklearn's LogisticRegressionCV, which automatically chooses C based on cross-validation, that's fine.\n\nFor Neural Networks,  you could use find_lr of the fastai learner to find a good learning rate, and also programatically use it for re-training. Or use Early Stopping with like a million training Epochs or something to decide on the number of Epochs for re-training.\nBut I think some examples on what constitutes as hyper-parameters in certain contexts would be nice. \n"
        },
        {
          "id": 661456,
          "postDate": "2019-10-30T10:27:18.517Z",
          "rawMarkdown": ""
        },
        {
          "id": 661749,
          "postDate": "2019-10-30T17:02:00.213Z",
          "content": "<p>You cannot change your code, except any changes required to handle reading in the new input files. Beyond that, ANY changes - to your model code, post-processing, etc. - potentially violate the rules.</p>\n\n<p>I wouldn't infer anything from the documentation about how far our checks go.</p>\n\n<p>Also, to <a href=\"/srsteinkamp\">@srsteinkamp</a> - good questions! Hopefully the above clears things up. Any code changes beyond those required to handle loading the new input files are potentially a rule violation.</p>",
          "rawMarkdown": "You cannot change your code, except any changes required to handle reading in the new input files. Beyond that, ANY changes - to your model code, post-processing, etc. - potentially violate the rules.\n\nI wouldn't infer anything from the documentation about how far our checks go.\n\nAlso, to @srsteinkamp - good questions! Hopefully the above clears things up. Any code changes beyond those required to handle loading the new input files are potentially a rule violation."
        },
        {
          "id": 662034,
          "postDate": "2019-10-31T02:27:54.973Z",
          "content": "<p><a href=\"/philculliton\">@philculliton</a> So Is it possible to retrain a couple of the models in an ensemble with full stage 1 train &amp; test and leave the remaining models in the ensemble trained only on stage 1 train? I'm asking this because it may not be possible to finish retraining all the models in an ensemble within the week time limit of stage 2. In other words is \"partial\" retraining allowed?</p>",
          "rawMarkdown": "@philculliton So Is it possible to retrain a couple of the models in an ensemble with full stage 1 train &amp; test and leave the remaining models in the ensemble trained only on stage 1 train? I'm asking this because it may not be possible to finish retraining all the models in an ensemble within the week time limit of stage 2. In other words is \"partial\" retraining allowed?"
        }
      ]
    },
    {
      "id": 654772,
      "postDate": "2019-10-22T09:46:55.760Z",
      "content": "<p>What does model mean? model weight?</p>",
      "rawMarkdown": "What does model mean? model weight?",
      "replies": [
        {
          "id": 654780,
          "postDate": "2019-10-22T09:59:32.853Z",
          "content": "<p>YES.such as pth file or others.</p>",
          "rawMarkdown": "YES.such as pth file or others."
        },
        {
          "id": 654825,
          "postDate": "2019-10-22T11:05:17.190Z",
          "content": "<p>I think no need to upload weight. just ipynb/py which produces weight is need to upload.\nI’m not confident. but weight file would be large....</p>",
          "rawMarkdown": "I think no need to upload weight. just ipynb/py which produces weight is need to upload.\nI’m not confident. but weight file would be large...."
        },
        {
          "id": 655017,
          "postDate": "2019-10-22T15:42:14.200Z",
          "content": "<p>I also don't want to upload model.\nif you are in prize region,you rank will be removed.</p>",
          "rawMarkdown": "I also don't want to upload model.\nif you are in prize region,you rank will be removed."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 655128,
      "author_name": "Julia Elliott",
      "author_url": "",
      "post_date": "2019-10-22T17:46:05.633000",
      "content": "<p>See <a href=\"https://www.kaggle.com/two-stage-frequently-asked-questions\">two-stage FAQ</a> for details on what to upload. </p>\n\n<p>Yes, you can upload the ipynb or whatever format your model is in. The output weights are not required. <code>submission.csv</code> is not required, since we already have that when you make your submission. You just need to ensure that the model(s) you are uploading match what you will be using to generate the final submissions that you ultimately select in stage 2.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 655326,
          "author_name": "XY",
          "author_url": "",
          "post_date": "2019-10-22T23:22:45.333000",
          "content": "<p>thank you !</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 655415,
          "author_name": "Tian Bingyang",
          "author_url": "",
          "post_date": "2019-10-23T03:03:42.683000",
          "content": "<p>Thanks for your replies,\n<a href=\"/juliaelliott\">@juliaelliott</a>\nForgive my poor English please:\nHere some part from two-stage FAQ :\n`\nWhat happens to my pre-trained model?\n<strong>If you are using a pre-trained for which you don't have the source code, you must include the model as part of your upload.</strong></p>\n\n<p>What if my submission is too big?\nOur uploader will handle reasonably large files. If you still think your <strong>model will be too large</strong>, you can instead upload a checksum of your archive file (such as an md5 or sha hash). Note that if you do win, you will still have to upload your code. If you upload a checksum, finish in a prize spot, and are unable to subsequently provide an archive that matches the checksum, you will be removed from the competition.`</p>\n\n<hr>\n\n<p><strong>Almost all of us use pretrained model.</strong></p>\n\n<p>So what't the meaning of \"model\" in above document?\nit refers to \".py\" file or \".pth\" weights file?</p>\n\n<p>Does it mean uploading the** file which produces the pretrained model**?</p>\n\n<p>if if refers to \".py/ipynb\" file,there's no possibility that <strong>model will be too large</strong>.\nCorrect me where I am wrong,please,\nThanks for your help~!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 655423,
          "author_name": "Tian Bingyang",
          "author_url": "",
          "post_date": "2019-10-23T03:15:42.493000",
          "content": "<p><a href=\"/xinyan989\">@xinyan989</a> could you help me explain it?\nThe Kaggle team said only code is necessary.\nBut the document said:\n<strong>If you are using a pre-trained for which you don't have the source code, you must include the model as part of your upload.</strong></p>\n\n<p>There are two meanings of model in <a href=\"https://www.kaggle.com/two-stage-frequently-asked-questions\">Two-stage FAQ</a>\n1.py/ipynb file.\n2.weights file.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 656046,
          "author_name": "Julia Elliott",
          "author_url": "",
          "post_date": "2019-10-23T20:44:25.137000",
          "content": "<p>A README file indicating which pre-trained models you're using and the links to them will be sufficient to include within your zipped upload.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 656228,
          "author_name": "Tian Bingyang",
          "author_url": "",
          "post_date": "2019-10-24T02:58:41.347000",
          "content": "<p>Very clear,Much Thanks</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 659894,
          "author_name": "Tian Bingyang",
          "author_url": "",
          "post_date": "2019-10-28T11:56:27.593000",
          "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a>\nSorry about two new question:\n1.\nIf I'm busy at stage1,\nbut I upload the model .py file and the sample_submission.csv(download from kaggle)\nand train my  weights file(.pth) in stage2,\nIs this allowed?\n2.\nIn stage2,I use new datasets,because the labels of test sets is released.\nThe Rule said:no code is allowed to change in stage2.\nSo ,Is it allowed to use code which merge the train and test(from stage1) in stage2?\nMuch Thanks for your help.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 660536,
          "author_name": "madamada",
          "author_url": "",
          "post_date": "2019-10-29T09:33:35.207000",
          "content": "<p>Hello, could I ask if semi-supervised learning with test data is allowed?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 661407,
      "author_name": "srs",
      "author_url": "",
      "post_date": "2019-10-30T08:56:32.597000",
      "content": "<p>Just latching on to the questions here, regarding the possible retraining of the models:</p>\n\n<p>What happens if I want to change something in my code in the second stage?\nWe expect you may need to make some \"non scientific\" alterations, such as changes to path names, in order to create your submissions for the second stage. You are allowed to re-train your model (including the stage one data), but your code should not change. You should not be doing any hyper parameter tuning in the second stage. Parameter tuning is permitted as long as it is fully automated.</p>\n\n<p>Are Learning Rate and number of Epochs also considered hyper-parameters, which can only be tuned automatically? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 661409,
          "author_name": "Tian Bingyang",
          "author_url": "",
          "post_date": "2019-10-30T09:02:32.357000",
          "content": "<p>\"You should not be doing any hyper parameter tuning in the second stage. Parameter tuning is permitted as long as it is fully automated.\"\nI think this sentence is contradictory</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 661447,
          "author_name": "srs",
          "author_url": "",
          "post_date": "2019-10-30T10:20:43.050000",
          "content": "<p>In this context I believe it means that if you for example use a logistic regression and retrain your models etc. <em>you</em> should not decide manually on a new regularization (i.e. C). \nIf you, however, you sklearn's LogisticRegressionCV, which automatically chooses C based on cross-validation, that's fine.</p>\n\n<p>For Neural Networks,  you could use find_lr of the fastai learner to find a good learning rate, and also programatically use it for re-training. Or use Early Stopping with like a million training Epochs or something to decide on the number of Epochs for re-training.\nBut I think some examples on what constitutes as hyper-parameters in certain contexts would be nice. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 661456,
          "author_name": "Tian Bingyang",
          "author_url": "",
          "post_date": "2019-10-30T10:27:18.517000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 661749,
          "author_name": "Phil Culliton",
          "author_url": "",
          "post_date": "2019-10-30T17:02:00.213000",
          "content": "<p>You cannot change your code, except any changes required to handle reading in the new input files. Beyond that, ANY changes - to your model code, post-processing, etc. - potentially violate the rules.</p>\n\n<p>I wouldn't infer anything from the documentation about how far our checks go.</p>\n\n<p>Also, to <a href=\"/srsteinkamp\">@srsteinkamp</a> - good questions! Hopefully the above clears things up. Any code changes beyond those required to handle loading the new input files are potentially a rule violation.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 662034,
          "author_name": "Tim Yee",
          "author_url": "",
          "post_date": "2019-10-31T02:27:54.973000",
          "content": "<p><a href=\"/philculliton\">@philculliton</a> So Is it possible to retrain a couple of the models in an ensemble with full stage 1 train &amp; test and leave the remaining models in the ensemble trained only on stage 1 train? I'm asking this because it may not be possible to finish retraining all the models in an ensemble within the week time limit of stage 2. In other words is \"partial\" retraining allowed?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 654772,
      "author_name": "XY",
      "author_url": "",
      "post_date": "2019-10-22T09:46:55.760000",
      "content": "<p>What does model mean? model weight?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 654780,
          "author_name": "Tian Bingyang",
          "author_url": "",
          "post_date": "2019-10-22T09:59:32.853000",
          "content": "<p>YES.such as pth file or others.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 654825,
          "author_name": "XY",
          "author_url": "",
          "post_date": "2019-10-22T11:05:17.190000",
          "content": "<p>I think no need to upload weight. just ipynb/py which produces weight is need to upload.\nI’m not confident. but weight file would be large....</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 655017,
          "author_name": "Tian Bingyang",
          "author_url": "",
          "post_date": "2019-10-22T15:42:14.200000",
          "content": "<p>I also don't want to upload model.\nif you are in prize region,you rank will be removed.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "655128": "See [two-stage FAQ](https://www.kaggle.com/two-stage-frequently-asked-questions) for details on what to upload. \n\nYes, you can upload the ipynb or whatever format your model is in. The output weights are not required. `submission.csv` is not required, since we already have that when you make your submission. You just need to ensure that the model(s) you are uploading match what you will be using to generate the final submissions that you ultimately select in stage 2.",
    "653850": "in my opinion.\nWe should upload 3 things:\nmodel,\nthe ipynb/py file which produced the model\nsubmission.csv\n\n \nAm I right?\nThanks for your help~!",
    "661407": "Just latching on to the questions here, regarding the possible retraining of the models:\n\nWhat happens if I want to change something in my code in the second stage?\nWe expect you may need to make some \"non scientific\" alterations, such as changes to path names, in order to create your submissions for the second stage. You are allowed to re-train your model (including the stage one data), but your code should not change. You should not be doing any hyper parameter tuning in the second stage. Parameter tuning is permitted as long as it is fully automated.\n\nAre Learning Rate and number of Epochs also considered hyper-parameters, which can only be tuned automatically? ",
    "654772": "What does model mean? model weight?"
  }
}