{
  "id": 115731,
  "title": "Stage 1 Ended. Stage 2 Update Underway!",
  "url": "/competitions/rsna-intracranial-hemorrhage-detection/discussion/115731",
  "author_name": "Julia Elliott",
  "post_date": "2019-11-05T01:12:58.801000",
  "votes": 15,
  "comment_count": 44,
  "views": 0,
  "content": "<p>Stage 1 has ended! We are updating the dataset and resetting the leaderboard. You can expect to see some shuffling and new data in this process. Stage 2 is not expected to begin until the evening of UTC time on 11/5. </p>\n\n<p>Submissions will be disabled until Stage 2 starts. We will announce on the forum when Stage 2 begins.</p>",
  "messages": [
    {
      "id": 665399,
      "postDate": "2019-11-05T01:12:58.800Z",
      "content": "<p>Stage 1 has ended! We are updating the dataset and resetting the leaderboard. You can expect to see some shuffling and new data in this process. Stage 2 is not expected to begin until the evening of UTC time on 11/5. </p>\n\n<p>Submissions will be disabled until Stage 2 starts. We will announce on the forum when Stage 2 begins.</p>",
      "rawMarkdown": "Stage 1 has ended! We are updating the dataset and resetting the leaderboard. You can expect to see some shuffling and new data in this process. Stage 2 is not expected to begin until the evening of UTC time on 11/5. \n\nSubmissions will be disabled until Stage 2 starts. We will announce on the forum when Stage 2 begins.",
      "votes": 15
    },
    {
      "id": 665517,
      "postDate": "2019-11-05T04:33:41.930Z",
      "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> can you please prepare test_stage_2_images as zip file?  The whole dataset is too big</p>",
      "rawMarkdown": "@juliaelliott can you please prepare test_stage_2_images as zip file?  The whole dataset is too big",
      "votes": 7
    },
    {
      "id": 665409,
      "postDate": "2019-11-05T01:21:18.243Z",
      "content": "<p>Thank you for your hard work!</p>",
      "rawMarkdown": "Thank you for your hard work!",
      "votes": 5,
      "replies": [
        {
          "id": 665411,
          "postDate": "2019-11-05T01:23:25.120Z",
          "content": "<p>Yes thank you for your hard work!</p>",
          "rawMarkdown": "Yes thank you for your hard work!"
        }
      ]
    },
    {
      "id": 665426,
      "postDate": "2019-11-05T02:07:12.107Z",
      "content": "<p>Thank you. Can we just download test data? The whole dataset is too big.</p>",
      "rawMarkdown": "Thank you. Can we just download test data? The whole dataset is too big.",
      "votes": 6,
      "replies": [
        {
          "id": 665443,
          "postDate": "2019-11-05T02:36:48.867Z",
          "content": "<p><a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/115726#latest-665432\">https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/115726#latest-665432</a></p>",
          "rawMarkdown": "https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/115726#latest-665432"
        },
        {
          "id": 665575,
          "postDate": "2019-11-05T06:12:03.807Z",
          "content": "<p>Downloading each image individually with its own API call is painfully slow, about 1 second per image. </p>\n\n<p>Anyone have any better ways?</p>",
          "rawMarkdown": "Downloading each image individually with its own API call is painfully slow, about 1 second per image. \n\nAnyone have any better ways?"
        },
        {
          "id": 665577,
          "postDate": "2019-11-05T06:14:59.923Z",
          "content": "<p>Havn't you updated your kaggle library? Kaggle 1.5.6 has solved this issue.</p>",
          "rawMarkdown": "Havn't you updated your kaggle library? Kaggle 1.5.6 has solved this issue."
        },
        {
          "id": 665647,
          "postDate": "2019-11-05T08:10:46.347Z",
          "content": "<p>what would be the command line to download just test_image folder ? <br>\nMy download failed twice today, but perhaps because the data is not ready yet. </p>",
          "rawMarkdown": "what would be the command line to download just test_image folder ?  \nMy download failed twice today, but perhaps because the data is not ready yet. ",
          "votes": 1
        },
        {
          "id": 666037,
          "postDate": "2019-11-05T16:54:03.083Z",
          "content": "<p>Thanks for your comments and questions - yep, we'll just be hosting the stage 2 test images, updated labels (train + stage 1 test), and stage 2 sample submission. You won't need to re-download stage 1 images.</p>",
          "rawMarkdown": "Thanks for your comments and questions - yep, we'll just be hosting the stage 2 test images, updated labels (train + stage 1 test), and stage 2 sample submission. You won't need to re-download stage 1 images."
        }
      ]
    },
    {
      "id": 666110,
      "postDate": "2019-11-05T18:26:53.973Z",
      "content": "<p>Thank you <a href=\"/juliaelliott\">@juliaelliott</a> and <a href=\"/philculliton\">@philculliton</a> for the updates a few hours ago on this forum.</p>\n\n<p>Some positive feedback that Kaggle might consider for the next  2 stage competition. Indeed split the dataset in 2 a train part and a test part. Offcourse it depends on the country but not all participants live in a country where a dataset of 180GB can be downloaded within one day. Let alone the unpacking and pre-processing.\nIn stage 1 there is enough time to lose a few days with it. In stage 2 not...\nBy offering the test-set as a seperate download this will be a lot easier. Also a lot of participants likely will not retrain there models. Anyway thanks for again a wonderfull competition...and I'am looking forward to submitting my first stage 2 submission :-)</p>",
      "rawMarkdown": "Thank you @juliaelliott and @philculliton for the updates a few hours ago on this forum.\n\nSome positive feedback that Kaggle might consider for the next  2 stage competition. Indeed split the dataset in 2 a train part and a test part. Offcourse it depends on the country but not all participants live in a country where a dataset of 180GB can be downloaded within one day. Let alone the unpacking and pre-processing.\nIn stage 1 there is enough time to lose a few days with it. In stage 2 not...\nBy offering the test-set as a seperate download this will be a lot easier. Also a lot of participants likely will not retrain there models. Anyway thanks for again a wonderfull competition...and I'am looking forward to submitting my first stage 2 submission :-)",
      "votes": 4
    },
    {
      "id": 665430,
      "postDate": "2019-11-05T02:23:09.603Z",
      "content": "<p>Excited for stage 2!</p>",
      "rawMarkdown": "Excited for stage 2!",
      "votes": 4
    },
    {
      "id": 666219,
      "postDate": "2019-11-05T22:16:44.370Z",
      "content": "<p>It's 11/5 UTC evening time. But, I don't see any updates. Anyone has any updae?</p>\n\n<p>10:13 PM\nTuesday, November 5, 2019\nCoordinated Universal Time (UTC)</p>",
      "rawMarkdown": "It's 11/5 UTC evening time. But, I don't see any updates. Anyone has any updae?\n\n10:13 PM\nTuesday, November 5, 2019\nCoordinated Universal Time (UTC)",
      "votes": 1
    },
    {
      "id": 665952,
      "postDate": "2019-11-05T15:15:37.237Z",
      "content": "<p>As shared yesterday, </p>\n\n<p>&gt; You can expect to see some shuffling and new data in this process.</p>\n\n<p>and</p>\n\n<p>&gt; Submissions will be disabled until Stage 2 starts. We will announce on the forum when Stage 2 begins.</p>\n\n<p>and</p>\n\n<p>&gt; Stage 2 is not expected to begin until the evening of UTC time on 11/5.</p>\n\n<p>We have not re-opened submissions yet and are still working on setting up stage 2. The leaderboard does <strong>not</strong> reflect who is eligible for stage 2, as it is still in the process of being reset. I ask for your patience in this process.</p>",
      "rawMarkdown": "As shared yesterday, \n\n&gt; You can expect to see some shuffling and new data in this process.\n\nand\n\n&gt; Submissions will be disabled until Stage 2 starts. We will announce on the forum when Stage 2 begins.\n\nand\n\n&gt; Stage 2 is not expected to begin until the evening of UTC time on 11/5.\n\nWe have not re-opened submissions yet and are still working on setting up stage 2. The leaderboard does **not** reflect who is eligible for stage 2, as it is still in the process of being reset. I ask for your patience in this process.",
      "votes": 1,
      "replies": [
        {
          "id": 665967,
          "postDate": "2019-11-05T15:38:37.077Z",
          "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> Could you please clarify, will it be possible to download stage2 test set separately? It may take quite a lot of time for many kagglers to download the whole data again. </p>",
          "rawMarkdown": "@juliaelliott Could you please clarify, will it be possible to download stage2 test set separately? It may take quite a lot of time for many kagglers to download the whole data again. "
        },
        {
          "id": 666038,
          "postDate": "2019-11-05T16:55:06.770Z",
          "content": "<p>Hi <a href=\"/cateek\">@cateek</a> - yep! We'll be providing just the stage 2 data for download, so you don't need to re-download stage 1 images again.</p>",
          "rawMarkdown": "Hi @cateek - yep! We'll be providing just the stage 2 data for download, so you don't need to re-download stage 1 images again.",
          "votes": 4
        },
        {
          "id": 666041,
          "postDate": "2019-11-05T16:56:53.813Z",
          "content": "<p>That's good news, thanks!</p>",
          "rawMarkdown": "That's good news, thanks!"
        }
      ]
    },
    {
      "id": 665764,
      "postDate": "2019-11-05T11:23:58.563Z",
      "content": "<p>Hi, \nMay you  explain if the file train_stage2 is different from train_stage1 ?\nMy understanding was that in the second stage was required to run inference on new test_stage2 only but now I  see also a new training file. \nDo we need to re-train our model from zero or we should train it starting from the  weights we calculated on stage 1 ?</p>",
      "rawMarkdown": "Hi, \nMay you  explain if the file train_stage2 is different from train_stage1 ?\nMy understanding was that in the second stage was required to run inference on new test_stage2 only but now I  see also a new training file. \nDo we need to re-train our model from zero or we should train it starting from the  weights we calculated on stage 1 ?",
      "votes": 1,
      "replies": [
        {
          "id": 667151,
          "postDate": "2019-11-06T21:44:18.857Z",
          "content": "<p>Stage 2's eligible training set includes Stage 1's test set. You may re-train your model on this additional training data, but it is not required.</p>",
          "rawMarkdown": "Stage 2's eligible training set includes Stage 1's test set. You may re-train your model on this additional training data, but it is not required."
        }
      ]
    },
    {
      "id": 665712,
      "postDate": "2019-11-05T10:06:37.320Z",
      "content": "<p>Why are there only 50 teams on the leaderboard? *what I’m seeing right now</p>",
      "rawMarkdown": "Why are there only 50 teams on the leaderboard? *what I’m seeing right now",
      "votes": 1,
      "replies": [
        {
          "id": 665726,
          "postDate": "2019-11-05T10:23:58.820Z",
          "content": "<p>Me too. Just a transient behavior, not the only bug that you can find now. I guess these small bugs are just not the top priority for Kaggle, and they prefer to notify that you can \"expect some shuffling\" each time, instead of fixing them. I am sure tomorrow all will settle down.</p>",
          "rawMarkdown": "Me too. Just a transient behavior, not the only bug that you can find now. I guess these small bugs are just not the top priority for Kaggle, and they prefer to notify that you can \"expect some shuffling\" each time, instead of fixing them. I am sure tomorrow all will settle down."
        },
        {
          "id": 665948,
          "postDate": "2019-11-05T15:08:27.443Z",
          "content": "<p>if you see your name on lb ,does it mean you are part of stage  2 ?</p>",
          "rawMarkdown": "if you see your name on lb ,does it mean you are part of stage  2 ?",
          "votes": 1
        }
      ]
    },
    {
      "id": 665445,
      "postDate": "2019-11-05T02:39:28.453Z",
      "content": "<p>Thanks, I will improve my model in the meantime.</p>",
      "rawMarkdown": "Thanks, I will improve my model in the meantime.",
      "votes": 1,
      "replies": [
        {
          "id": 665549,
          "postDate": "2019-11-05T05:25:54.703Z",
          "content": "<p>You are not supposed to make code change/model improvement after you submitted the model file in stage 1 :)</p>",
          "rawMarkdown": "You are not supposed to make code change/model improvement after you submitted the model file in stage 1 :)",
          "votes": 2
        },
        {
          "id": 665613,
          "postDate": "2019-11-05T07:16:44.993Z",
          "content": "<p><a href=\"/chinhuic\">@chinhuic</a> \npeople use their own code outside kaggle gpu. \nUnless you are in money rank ,is it possible to verify if  each n every submission made in stage 2 is made   from same uploaded model or there was change done</p>",
          "rawMarkdown": "@chinhuic \npeople use their own code outside kaggle gpu. \nUnless you are in money rank ,is it possible to verify if  each n every submission made in stage 2 is made   from same uploaded model or there was change done"
        },
        {
          "id": 666503,
          "postDate": "2019-11-06T06:37:56.010Z",
          "content": "<p><a href=\"/chinhuic\">@chinhuic</a> You can still retrain models with extended stage 1 train and test data even if it improves your model performance. But you shouldn't do any code or hyper parameter changes, e.g. architecture etc...</p>",
          "rawMarkdown": "@chinhuic You can still retrain models with extended stage 1 train and test data even if it improves your model performance. But you shouldn't do any code or hyper parameter changes, e.g. architecture etc..."
        }
      ]
    },
    {
      "id": 665661,
      "postDate": "2019-11-05T08:55:13.890Z",
      "content": "<p>Hi all, It's my first time with a 2-stage competition and it's not clear to me if one can re-train the model(s) on the stage2 training set. The rules say so, but the model must've been uploaded and in principle not changed from stage1. Can someone clarify? thanks :)</p>",
      "rawMarkdown": "Hi all, It's my first time with a 2-stage competition and it's not clear to me if one can re-train the model(s) on the stage2 training set. The rules say so, but the model must've been uploaded and in principle not changed from stage1. Can someone clarify? thanks :)",
      "votes": -1,
      "replies": [
        {
          "id": 665669,
          "postDate": "2019-11-05T09:12:20.063Z",
          "content": "<p>You can either upload model weights and inference code, or you can upload training and inference code. For the latter, you are allowed to change paths in order to re-train the model on stage-1 data. However, the only thing that you are allowed to change is the paths and \"non-scientific\" alterations. No hyper-parameter tuning is allowed unless it is fully automated (eg something like reduce lr on plateau)</p>",
          "rawMarkdown": "You can either upload model weights and inference code, or you can upload training and inference code. For the latter, you are allowed to change paths in order to re-train the model on stage-1 data. However, the only thing that you are allowed to change is the paths and \"non-scientific\" alterations. No hyper-parameter tuning is allowed unless it is fully automated (eg something like reduce lr on plateau)",
          "votes": 1
        },
        {
          "id": 665672,
          "postDate": "2019-11-05T09:16:48.353Z",
          "content": "<p>Thanks! Can I re-train if I uploaded the weights, training and inference code? Or the fact that I uploaded the weights prevents me from re-training? </p>",
          "rawMarkdown": "Thanks! Can I re-train if I uploaded the weights, training and inference code? Or the fact that I uploaded the weights prevents me from re-training? "
        },
        {
          "id": 665742,
          "postDate": "2019-11-05T10:49:23.063Z",
          "content": "<p>I think technically you are supposed to upload instructions to use what you uploaded to reproduce your results. However I am not sure how strict they are on this, if there are no instructions I wonder if you have the opportunity to explain how to use it? I have made a couple of posts to try and find out just how explicit instructions need to be, however had no response.</p>\n\n<p>Personally, I uploaded training code and weights, since we are allowed two submissions. I wrote some very brief instructions in the readme that basically said:\n- Submission 1: use these scripts to generate predictions for these models using these weights, then use this script to ensemble.\n- Submission 2: use these scripts to train new models on all data, then generate submission as per submission 1 instructions.</p>\n\n<hr>\n\n<p>So I have the opportunity to retrain models on all of the data, however I think it is very unlikely that all of the training code for my models will just work out of the box. This isn't a coding competition, so hopefully they are lenient on allowing people to fix some bugs. \nThat being said, I doubt it will be an issue for me as there is no chance that I will finish in the money. I expect those in the money will have really nice code.</p>\n\n<p>I would say just go ahead and re-train and use your retained model as one of your submissions. If you finish in the money, then it is easier to ask for forgiveness. And if you don't, then take the higher score regardless of if it is technically legal or not. Competitive integrity of Kaggle is dead anyway with people posting complete high scoring solutions half way through the comp. I am purely here to learn now, everything else is secondary. Unfortunately that makes it less fun but ohwell :(</p>",
          "rawMarkdown": "\nI think technically you are supposed to upload instructions to use what you uploaded to reproduce your results. However I am not sure how strict they are on this, if there are no instructions I wonder if you have the opportunity to explain how to use it? I have made a couple of posts to try and find out just how explicit instructions need to be, however had no response.\n\nPersonally, I uploaded training code and weights, since we are allowed two submissions. I wrote some very brief instructions in the readme that basically said:\n- Submission 1: use these scripts to generate predictions for these models using these weights, then use this script to ensemble.\n- Submission 2: use these scripts to train new models on all data, then generate submission as per submission 1 instructions.\n___\n\nSo I have the opportunity to retrain models on all of the data, however I think it is very unlikely that all of the training code for my models will just work out of the box. This isn't a coding competition, so hopefully they are lenient on allowing people to fix some bugs. \nThat being said, I doubt it will be an issue for me as there is no chance that I will finish in the money. I expect those in the money will have really nice code.\n\nI would say just go ahead and re-train and use your retained model as one of your submissions. If you finish in the money, then it is easier to ask for forgiveness. And if you don't, then take the higher score regardless of if it is technically legal or not. Competitive integrity of Kaggle is dead anyway with people posting complete high scoring solutions half way through the comp. I am purely here to learn now, everything else is secondary. Unfortunately that makes it less fun but ohwell :(",
          "votes": 1
        },
        {
          "id": 666496,
          "postDate": "2019-11-06T06:25:36.587Z",
          "content": "<p>You are allowed to retrain with the same code with the stage 2 train data (which is stage 1 train and test data)</p>",
          "rawMarkdown": "You are allowed to retrain with the same code with the stage 2 train data (which is stage 1 train and test data)",
          "votes": 1
        }
      ]
    },
    {
      "id": 667429,
      "postDate": "2019-11-07T08:14:48.970Z",
      "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> We use the dataset for a school project regardless of the leaderboard &amp;kaggle competition. Assuming it will stay online, we didn't actually download the dataset. Is there a way to still access the stage 1 data?</p>",
      "rawMarkdown": "@juliaelliott We use the dataset for a school project regardless of the leaderboard &amp;kaggle competition. Assuming it will stay online, we didn't actually download the dataset. Is there a way to still access the stage 1 data?",
      "replies": [
        {
          "id": 667953,
          "postDate": "2019-11-07T20:59:15.810Z",
          "content": "<p>We are working on it. I promise we'll update everyone when the datasets are confirmed fully available.</p>",
          "rawMarkdown": "We are working on it. I promise we'll update everyone when the datasets are confirmed fully available.",
          "votes": 2
        }
      ]
    },
    {
      "id": 666027,
      "postDate": "2019-11-05T16:46:57.837Z",
      "content": "<p>Can someone tell me what is happening with the leaderboard and why all scores all 0 ?</p>",
      "rawMarkdown": "Can someone tell me what is happening with the leaderboard and why all scores all 0 ?"
    },
    {
      "id": 665930,
      "postDate": "2019-11-05T14:49:38.687Z",
      "content": "<p>「We are updating the dataset and resetting the leaderboard. You can expect to see some shuffling and new data in this process. Stage 2 is not expected to begin until the evening of UTC time on 11/5.」</p>\n\n<p>which mean we cannot download the dataset until the stage2 annoucement?</p>",
      "rawMarkdown": "「We are updating the dataset and resetting the leaderboard. You can expect to see some shuffling and new data in this process. Stage 2 is not expected to begin until the evening of UTC time on 11/5.」\n\nwhich mean we cannot download the dataset until the stage2 annoucement?"
    },
    {
      "id": 665682,
      "postDate": "2019-11-05T09:29:54.943Z",
      "content": "<p>Thank you for the exciting dataset for stage 2 !</p>",
      "rawMarkdown": "Thank you for the exciting dataset for stage 2 !"
    },
    {
      "id": 665447,
      "postDate": "2019-11-05T02:46:34.413Z",
      "content": "<p>How many teams advanced to stage 2? Or anybody who uploaded their models automatically advance to stage 2?</p>",
      "rawMarkdown": "How many teams advanced to stage 2? Or anybody who uploaded their models automatically advance to stage 2?",
      "replies": [
        {
          "id": 665457,
          "postDate": "2019-11-05T02:59:15.240Z",
          "content": "<p>It's still in progress I think everyone will be able to submit for stage 2 but if you don't have a model uploaded by the end of stage there is a chance of you to be disqualified from leaderboard after the competition is over.</p>\n\n<p>You can check Timeline page :)\n<code>\nThis requirement is in place for the host team to verify the performance of the uploaded models matches the Stage 2 submission file. Compliance with the above will be verified by the host team. Submitters who fail to upload their model by the Stage 1 deadline, or are found not to be in compliance, may be disqualified from Stage 2 and removed from the final leaderboard.\n</code></p>",
          "rawMarkdown": "It's still in progress I think everyone will be able to submit for stage 2 but if you don't have a model uploaded by the end of stage there is a chance of you to be disqualified from leaderboard after the competition is over.\n\nYou can check Timeline page :)\n```\nThis requirement is in place for the host team to verify the performance of the uploaded models matches the Stage 2 submission file. Compliance with the above will be verified by the host team. Submitters who fail to upload their model by the Stage 1 deadline, or are found not to be in compliance, may be disqualified from Stage 2 and removed from the final leaderboard.\n```",
          "votes": 2
        },
        {
          "id": 665615,
          "postDate": "2019-11-05T07:19:57.497Z",
          "content": "<p>358 teams have qualified for stage 2 according to public leaderboard.</p>",
          "rawMarkdown": "358 teams have qualified for stage 2 according to public leaderboard."
        },
        {
          "id": 665825,
          "postDate": "2019-11-05T12:45:33.213Z",
          "content": "<p>how we know..\nso far i cant leader board beyond 576 ranks. In which i see there is my ranking still in. </p>",
          "rawMarkdown": "how we know..\nso far i cant leader board beyond 576 ranks. In which i see there is my ranking still in. \n"
        },
        {
          "id": 667154,
          "postDate": "2019-11-06T21:46:20.183Z",
          "content": "<p>Anyone who uploaded a model at the end of Stage 1 is eligible to make a submission to Stage 2. We wiped the leaderboard clean at the end of Stage 1, so you will not show up on the Stage 2 leaderboard until you make a Stage 2 submission. If you did not and you make a Stage 2 submission, you are at risk of being eliminated from the final leaderboard at the end of the competition.</p>",
          "rawMarkdown": "Anyone who uploaded a model at the end of Stage 1 is eligible to make a submission to Stage 2. We wiped the leaderboard clean at the end of Stage 1, so you will not show up on the Stage 2 leaderboard until you make a Stage 2 submission. If you did not and you make a Stage 2 submission, you are at risk of being eliminated from the final leaderboard at the end of the competition."
        }
      ]
    },
    {
      "id": 665416,
      "postDate": "2019-11-05T01:36:35.990Z",
      "content": "<p>Thank you for the timeliness of the data release, really appreciate it !</p>",
      "rawMarkdown": "Thank you for the timeliness of the data release, really appreciate it !"
    },
    {
      "id": 665628,
      "postDate": "2019-11-05T07:35:53.017Z",
      "content": "<p>So what if you didn't submit your model file at the end of stage 1?</p>",
      "rawMarkdown": "So what if you didn't submit your model file at the end of stage 1?",
      "isDeleted": true,
      "replies": [
        {
          "id": 665670,
          "postDate": "2019-11-05T09:13:59.143Z",
          "content": "<p>DQ</p>",
          "rawMarkdown": "DQ"
        }
      ]
    },
    {
      "id": 665406,
      "postDate": "2019-11-05T01:17:57.440Z",
      "content": "<p>Thank you for giving us a break :) </p>",
      "rawMarkdown": "Thank you for giving us a break :) ",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 665517,
      "author_name": "Alimbekov Renat [dsmlkz]",
      "author_url": "",
      "post_date": "2019-11-05T04:33:41.930000",
      "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> can you please prepare test_stage_2_images as zip file?  The whole dataset is too big</p>",
      "votes": 7,
      "replies": []
    },
    {
      "id": 665409,
      "author_name": "Guanshuo Xu",
      "author_url": "",
      "post_date": "2019-11-05T01:21:18.243000",
      "content": "<p>Thank you for your hard work!</p>",
      "votes": 5,
      "replies": [
        {
          "id": 665411,
          "author_name": "Hilal Shaath",
          "author_url": "",
          "post_date": "2019-11-05T01:23:25.120000",
          "content": "<p>Yes thank you for your hard work!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 665426,
      "author_name": "yelan",
      "author_url": "",
      "post_date": "2019-11-05T02:07:12.107000",
      "content": "<p>Thank you. Can we just download test data? The whole dataset is too big.</p>",
      "votes": 6,
      "replies": [
        {
          "id": 665443,
          "author_name": "Gabriel",
          "author_url": "",
          "post_date": "2019-11-05T02:36:48.867000",
          "content": "<p><a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/115726#latest-665432\">https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/115726#latest-665432</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 665575,
          "author_name": "cherring",
          "author_url": "",
          "post_date": "2019-11-05T06:12:03.807000",
          "content": "<p>Downloading each image individually with its own API call is painfully slow, about 1 second per image. </p>\n\n<p>Anyone have any better ways?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 665577,
          "author_name": "NitinKshatriya",
          "author_url": "",
          "post_date": "2019-11-05T06:14:59.923000",
          "content": "<p>Havn't you updated your kaggle library? Kaggle 1.5.6 has solved this issue.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 665647,
          "author_name": "yukiya",
          "author_url": "",
          "post_date": "2019-11-05T08:10:46.347000",
          "content": "<p>what would be the command line to download just test_image folder ? <br>\nMy download failed twice today, but perhaps because the data is not ready yet. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 666037,
          "author_name": "Phil Culliton",
          "author_url": "",
          "post_date": "2019-11-05T16:54:03.083000",
          "content": "<p>Thanks for your comments and questions - yep, we'll just be hosting the stage 2 test images, updated labels (train + stage 1 test), and stage 2 sample submission. You won't need to re-download stage 1 images.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 666110,
      "author_name": "Robin Smits",
      "author_url": "",
      "post_date": "2019-11-05T18:26:53.973000",
      "content": "<p>Thank you <a href=\"/juliaelliott\">@juliaelliott</a> and <a href=\"/philculliton\">@philculliton</a> for the updates a few hours ago on this forum.</p>\n\n<p>Some positive feedback that Kaggle might consider for the next  2 stage competition. Indeed split the dataset in 2 a train part and a test part. Offcourse it depends on the country but not all participants live in a country where a dataset of 180GB can be downloaded within one day. Let alone the unpacking and pre-processing.\nIn stage 1 there is enough time to lose a few days with it. In stage 2 not...\nBy offering the test-set as a seperate download this will be a lot easier. Also a lot of participants likely will not retrain there models. Anyway thanks for again a wonderfull competition...and I'am looking forward to submitting my first stage 2 submission :-)</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 665430,
      "author_name": "Carlo Lepelaars",
      "author_url": "",
      "post_date": "2019-11-05T02:23:09.603000",
      "content": "<p>Excited for stage 2!</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 666219,
      "author_name": "Hilal Shaath",
      "author_url": "",
      "post_date": "2019-11-05T22:16:44.370000",
      "content": "<p>It's 11/5 UTC evening time. But, I don't see any updates. Anyone has any updae?</p>\n\n<p>10:13 PM\nTuesday, November 5, 2019\nCoordinated Universal Time (UTC)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 665952,
      "author_name": "Julia Elliott",
      "author_url": "",
      "post_date": "2019-11-05T15:15:37.237000",
      "content": "<p>As shared yesterday, </p>\n\n<p>&gt; You can expect to see some shuffling and new data in this process.</p>\n\n<p>and</p>\n\n<p>&gt; Submissions will be disabled until Stage 2 starts. We will announce on the forum when Stage 2 begins.</p>\n\n<p>and</p>\n\n<p>&gt; Stage 2 is not expected to begin until the evening of UTC time on 11/5.</p>\n\n<p>We have not re-opened submissions yet and are still working on setting up stage 2. The leaderboard does <strong>not</strong> reflect who is eligible for stage 2, as it is still in the process of being reset. I ask for your patience in this process.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 665967,
          "author_name": "Eek The Cat",
          "author_url": "",
          "post_date": "2019-11-05T15:38:37.077000",
          "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> Could you please clarify, will it be possible to download stage2 test set separately? It may take quite a lot of time for many kagglers to download the whole data again. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 666038,
          "author_name": "Phil Culliton",
          "author_url": "",
          "post_date": "2019-11-05T16:55:06.770000",
          "content": "<p>Hi <a href=\"/cateek\">@cateek</a> - yep! We'll be providing just the stage 2 data for download, so you don't need to re-download stage 1 images again.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 666041,
          "author_name": "Eek The Cat",
          "author_url": "",
          "post_date": "2019-11-05T16:56:53.813000",
          "content": "<p>That's good news, thanks!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 665764,
      "author_name": "AlGiLa",
      "author_url": "",
      "post_date": "2019-11-05T11:23:58.563000",
      "content": "<p>Hi, \nMay you  explain if the file train_stage2 is different from train_stage1 ?\nMy understanding was that in the second stage was required to run inference on new test_stage2 only but now I  see also a new training file. \nDo we need to re-train our model from zero or we should train it starting from the  weights we calculated on stage 1 ?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 667151,
          "author_name": "Julia Elliott",
          "author_url": "",
          "post_date": "2019-11-06T21:44:18.857000",
          "content": "<p>Stage 2's eligible training set includes Stage 1's test set. You may re-train your model on this additional training data, but it is not required.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 665712,
      "author_name": "Nicholas Lyu",
      "author_url": "",
      "post_date": "2019-11-05T10:06:37.320000",
      "content": "<p>Why are there only 50 teams on the leaderboard? *what I’m seeing right now</p>",
      "votes": 1,
      "replies": [
        {
          "id": 665726,
          "author_name": "nosound",
          "author_url": "",
          "post_date": "2019-11-05T10:23:58.820000",
          "content": "<p>Me too. Just a transient behavior, not the only bug that you can find now. I guess these small bugs are just not the top priority for Kaggle, and they prefer to notify that you can \"expect some shuffling\" each time, instead of fixing them. I am sure tomorrow all will settle down.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 665948,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2019-11-05T15:08:27.443000",
          "content": "<p>if you see your name on lb ,does it mean you are part of stage  2 ?</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 665445,
      "author_name": "Claudio Moisés Valiense de Andrade",
      "author_url": "",
      "post_date": "2019-11-05T02:39:28.453000",
      "content": "<p>Thanks, I will improve my model in the meantime.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 665549,
          "author_name": "ChinHuiC",
          "author_url": "",
          "post_date": "2019-11-05T05:25:54.703000",
          "content": "<p>You are not supposed to make code change/model improvement after you submitted the model file in stage 1 :)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 665613,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2019-11-05T07:16:44.993000",
          "content": "<p><a href=\"/chinhuic\">@chinhuic</a> \npeople use their own code outside kaggle gpu. \nUnless you are in money rank ,is it possible to verify if  each n every submission made in stage 2 is made   from same uploaded model or there was change done</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 666503,
          "author_name": "Kerem Turgutlu",
          "author_url": "",
          "post_date": "2019-11-06T06:37:56.010000",
          "content": "<p><a href=\"/chinhuic\">@chinhuic</a> You can still retrain models with extended stage 1 train and test data even if it improves your model performance. But you shouldn't do any code or hyper parameter changes, e.g. architecture etc...</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 665661,
      "author_name": "Lorenzo Peppoloni",
      "author_url": "",
      "post_date": "2019-11-05T08:55:13.890000",
      "content": "<p>Hi all, It's my first time with a 2-stage competition and it's not clear to me if one can re-train the model(s) on the stage2 training set. The rules say so, but the model must've been uploaded and in principle not changed from stage1. Can someone clarify? thanks :)</p>",
      "votes": -1,
      "replies": [
        {
          "id": 665669,
          "author_name": "cherring",
          "author_url": "",
          "post_date": "2019-11-05T09:12:20.063000",
          "content": "<p>You can either upload model weights and inference code, or you can upload training and inference code. For the latter, you are allowed to change paths in order to re-train the model on stage-1 data. However, the only thing that you are allowed to change is the paths and \"non-scientific\" alterations. No hyper-parameter tuning is allowed unless it is fully automated (eg something like reduce lr on plateau)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 665672,
          "author_name": "Lorenzo Peppoloni",
          "author_url": "",
          "post_date": "2019-11-05T09:16:48.353000",
          "content": "<p>Thanks! Can I re-train if I uploaded the weights, training and inference code? Or the fact that I uploaded the weights prevents me from re-training? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 665742,
          "author_name": "cherring",
          "author_url": "",
          "post_date": "2019-11-05T10:49:23.063000",
          "content": "<p>I think technically you are supposed to upload instructions to use what you uploaded to reproduce your results. However I am not sure how strict they are on this, if there are no instructions I wonder if you have the opportunity to explain how to use it? I have made a couple of posts to try and find out just how explicit instructions need to be, however had no response.</p>\n\n<p>Personally, I uploaded training code and weights, since we are allowed two submissions. I wrote some very brief instructions in the readme that basically said:\n- Submission 1: use these scripts to generate predictions for these models using these weights, then use this script to ensemble.\n- Submission 2: use these scripts to train new models on all data, then generate submission as per submission 1 instructions.</p>\n\n<hr>\n\n<p>So I have the opportunity to retrain models on all of the data, however I think it is very unlikely that all of the training code for my models will just work out of the box. This isn't a coding competition, so hopefully they are lenient on allowing people to fix some bugs. \nThat being said, I doubt it will be an issue for me as there is no chance that I will finish in the money. I expect those in the money will have really nice code.</p>\n\n<p>I would say just go ahead and re-train and use your retained model as one of your submissions. If you finish in the money, then it is easier to ask for forgiveness. And if you don't, then take the higher score regardless of if it is technically legal or not. Competitive integrity of Kaggle is dead anyway with people posting complete high scoring solutions half way through the comp. I am purely here to learn now, everything else is secondary. Unfortunately that makes it less fun but ohwell :(</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 666496,
          "author_name": "ChinHuiC",
          "author_url": "",
          "post_date": "2019-11-06T06:25:36.587000",
          "content": "<p>You are allowed to retrain with the same code with the stage 2 train data (which is stage 1 train and test data)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 667429,
      "author_name": "Tomasz Lewicki",
      "author_url": "",
      "post_date": "2019-11-07T08:14:48.970000",
      "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> We use the dataset for a school project regardless of the leaderboard &amp;kaggle competition. Assuming it will stay online, we didn't actually download the dataset. Is there a way to still access the stage 1 data?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 667953,
          "author_name": "Julia Elliott",
          "author_url": "",
          "post_date": "2019-11-07T20:59:15.810000",
          "content": "<p>We are working on it. I promise we'll update everyone when the datasets are confirmed fully available.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 666027,
      "author_name": "Rrezon Pllana",
      "author_url": "",
      "post_date": "2019-11-05T16:46:57.837000",
      "content": "<p>Can someone tell me what is happening with the leaderboard and why all scores all 0 ?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 665930,
      "author_name": "liuzhangzhen",
      "author_url": "",
      "post_date": "2019-11-05T14:49:38.687000",
      "content": "<p>「We are updating the dataset and resetting the leaderboard. You can expect to see some shuffling and new data in this process. Stage 2 is not expected to begin until the evening of UTC time on 11/5.」</p>\n\n<p>which mean we cannot download the dataset until the stage2 annoucement?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 665682,
      "author_name": "Naruhiko Nakanishi",
      "author_url": "",
      "post_date": "2019-11-05T09:29:54.943000",
      "content": "<p>Thank you for the exciting dataset for stage 2 !</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 665447,
      "author_name": "CE Kan",
      "author_url": "",
      "post_date": "2019-11-05T02:46:34.413000",
      "content": "<p>How many teams advanced to stage 2? Or anybody who uploaded their models automatically advance to stage 2?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 665457,
          "author_name": "Kerem Turgutlu",
          "author_url": "",
          "post_date": "2019-11-05T02:59:15.240000",
          "content": "<p>It's still in progress I think everyone will be able to submit for stage 2 but if you don't have a model uploaded by the end of stage there is a chance of you to be disqualified from leaderboard after the competition is over.</p>\n\n<p>You can check Timeline page :)\n<code>\nThis requirement is in place for the host team to verify the performance of the uploaded models matches the Stage 2 submission file. Compliance with the above will be verified by the host team. Submitters who fail to upload their model by the Stage 1 deadline, or are found not to be in compliance, may be disqualified from Stage 2 and removed from the final leaderboard.\n</code></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 665615,
          "author_name": "Harveen Singh Chadha",
          "author_url": "",
          "post_date": "2019-11-05T07:19:57.497000",
          "content": "<p>358 teams have qualified for stage 2 according to public leaderboard.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 665825,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2019-11-05T12:45:33.213000",
          "content": "<p>how we know..\nso far i cant leader board beyond 576 ranks. In which i see there is my ranking still in. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 667154,
          "author_name": "Julia Elliott",
          "author_url": "",
          "post_date": "2019-11-06T21:46:20.183000",
          "content": "<p>Anyone who uploaded a model at the end of Stage 1 is eligible to make a submission to Stage 2. We wiped the leaderboard clean at the end of Stage 1, so you will not show up on the Stage 2 leaderboard until you make a Stage 2 submission. If you did not and you make a Stage 2 submission, you are at risk of being eliminated from the final leaderboard at the end of the competition.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 665416,
      "author_name": "yukiya",
      "author_url": "",
      "post_date": "2019-11-05T01:36:35.990000",
      "content": "<p>Thank you for the timeliness of the data release, really appreciate it !</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 665628,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-11-05T07:35:53.017000",
      "content": "<p>So what if you didn't submit your model file at the end of stage 1?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 665670,
          "author_name": "Yifeng (Ethan) Zou",
          "author_url": "",
          "post_date": "2019-11-05T09:13:59.143000",
          "content": "<p>DQ</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 665406,
      "author_name": "Hilal Shaath",
      "author_url": "",
      "post_date": "2019-11-05T01:17:57.440000",
      "content": "<p>Thank you for giving us a break :) </p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "665399": "Stage 1 has ended! We are updating the dataset and resetting the leaderboard. You can expect to see some shuffling and new data in this process. Stage 2 is not expected to begin until the evening of UTC time on 11/5. \n\nSubmissions will be disabled until Stage 2 starts. We will announce on the forum when Stage 2 begins.",
    "665517": "@juliaelliott can you please prepare test_stage_2_images as zip file?  The whole dataset is too big",
    "665409": "Thank you for your hard work!",
    "665426": "Thank you. Can we just download test data? The whole dataset is too big.",
    "666110": "Thank you @juliaelliott and @philculliton for the updates a few hours ago on this forum.\n\nSome positive feedback that Kaggle might consider for the next  2 stage competition. Indeed split the dataset in 2 a train part and a test part. Offcourse it depends on the country but not all participants live in a country where a dataset of 180GB can be downloaded within one day. Let alone the unpacking and pre-processing.\nIn stage 1 there is enough time to lose a few days with it. In stage 2 not...\nBy offering the test-set as a seperate download this will be a lot easier. Also a lot of participants likely will not retrain there models. Anyway thanks for again a wonderfull competition...and I'am looking forward to submitting my first stage 2 submission :-)",
    "665430": "Excited for stage 2!",
    "666219": "It's 11/5 UTC evening time. But, I don't see any updates. Anyone has any updae?\n\n10:13 PM\nTuesday, November 5, 2019\nCoordinated Universal Time (UTC)",
    "665952": "As shared yesterday, \n\n&gt; You can expect to see some shuffling and new data in this process.\n\nand\n\n&gt; Submissions will be disabled until Stage 2 starts. We will announce on the forum when Stage 2 begins.\n\nand\n\n&gt; Stage 2 is not expected to begin until the evening of UTC time on 11/5.\n\nWe have not re-opened submissions yet and are still working on setting up stage 2. The leaderboard does **not** reflect who is eligible for stage 2, as it is still in the process of being reset. I ask for your patience in this process.",
    "665764": "Hi, \nMay you  explain if the file train_stage2 is different from train_stage1 ?\nMy understanding was that in the second stage was required to run inference on new test_stage2 only but now I  see also a new training file. \nDo we need to re-train our model from zero or we should train it starting from the  weights we calculated on stage 1 ?",
    "665712": "Why are there only 50 teams on the leaderboard? *what I’m seeing right now",
    "665445": "Thanks, I will improve my model in the meantime.",
    "665661": "Hi all, It's my first time with a 2-stage competition and it's not clear to me if one can re-train the model(s) on the stage2 training set. The rules say so, but the model must've been uploaded and in principle not changed from stage1. Can someone clarify? thanks :)",
    "667429": "@juliaelliott We use the dataset for a school project regardless of the leaderboard &amp;kaggle competition. Assuming it will stay online, we didn't actually download the dataset. Is there a way to still access the stage 1 data?",
    "666027": "Can someone tell me what is happening with the leaderboard and why all scores all 0 ?",
    "665930": "「We are updating the dataset and resetting the leaderboard. You can expect to see some shuffling and new data in this process. Stage 2 is not expected to begin until the evening of UTC time on 11/5.」\n\nwhich mean we cannot download the dataset until the stage2 annoucement?",
    "665682": "Thank you for the exciting dataset for stage 2 !",
    "665447": "How many teams advanced to stage 2? Or anybody who uploaded their models automatically advance to stage 2?",
    "665416": "Thank you for the timeliness of the data release, really appreciate it !",
    "665628": "So what if you didn't submit your model file at the end of stage 1?",
    "665406": "Thank you for giving us a break :) "
  }
}