{
  "id": 115726,
  "title": "Downloading just test images for stage 2?",
  "url": "/competitions/rsna-intracranial-hemorrhage-detection/discussion/115726",
  "author_name": "Alexey Kotlik",
  "post_date": "2019-11-05T00:41:36.157000",
  "votes": 16,
  "comment_count": 21,
  "views": 0,
  "content": "<p>I understood that train set will remain the same from the stage 1, and the test set labels will become availabel so we could use it to extend the train set. \nIs there a way to download just 30Gb of new test images, instead of the whole pile of 180Gb of which we already supposedly have 150Gb on our HDDs? </p>",
  "messages": [
    {
      "id": 665384,
      "postDate": "2019-11-05T00:41:36.157Z",
      "content": "<p>I understood that train set will remain the same from the stage 1, and the test set labels will become availabel so we could use it to extend the train set. \nIs there a way to download just 30Gb of new test images, instead of the whole pile of 180Gb of which we already supposedly have 150Gb on our HDDs? </p>",
      "rawMarkdown": "I understood that train set will remain the same from the stage 1, and the test set labels will become availabel so we could use it to extend the train set. \nIs there a way to download just 30Gb of new test images, instead of the whole pile of 180Gb of which we already supposedly have 150Gb on our HDDs? ",
      "votes": 16
    },
    {
      "id": 666035,
      "postDate": "2019-11-05T16:51:48.113Z",
      "content": "<p>Hi all! Very sorry to hear about everyone's trouble so far.</p>\n\n<blockquote>\n  <p>Stage 1 has ended! We are updating the dataset and resetting the leaderboard. You can expect to see some shuffling and new data in this process. Stage 2 is not expected to begin until the evening of UTC time on 11/5.</p>\n</blockquote>\n\n<p>The above is what's happening now - we're still moving data around. We'll ultimately be hosting only the stage 2 images, updated labels, and new sample submission. You will not have to redownload stage 1 data. We're also working on longer-term design changes to our API and download sections so you can download selected archives again. Thanks for your patience, we'll let you know when the data is ready to go!</p>",
      "rawMarkdown": "Hi all! Very sorry to hear about everyone's trouble so far.\n\n&gt; Stage 1 has ended! We are updating the dataset and resetting the leaderboard. You can expect to see some shuffling and new data in this process. Stage 2 is not expected to begin until the evening of UTC time on 11/5.\n\nThe above is what's happening now - we're still moving data around. We'll ultimately be hosting only the stage 2 images, updated labels, and new sample submission. You will not have to redownload stage 1 data. We're also working on longer-term design changes to our API and download sections so you can download selected archives again. Thanks for your patience, we'll let you know when the data is ready to go!",
      "votes": 3
    },
    {
      "id": 665387,
      "postDate": "2019-11-05T00:46:22.760Z",
      "content": "<p>Yes. You can specify the file you want to download via the kaggle api using the <code>-f</code> flag. See below example for reference.</p>\n\n<p><code>kaggle competitions download -c rsna-intracranial-hemorrhage-detection -f stage_2_test_images</code></p>",
      "rawMarkdown": "Yes. You can specify the file you want to download via the kaggle api using the `-f` flag. See below example for reference.\n\n`kaggle competitions download -c rsna-intracranial-hemorrhage-detection -f stage_2_test_images`",
      "votes": 3,
      "replies": [
        {
          "id": 665408,
          "postDate": "2019-11-05T01:20:48.533Z",
          "content": "<p>Thanks for the quick reply James, however I couldn't figure it out - I get '404 not found' if using the folder name as in your example. Wildcard stage_2_test_images/*.dcm does not help either. \nA specific file from the folder downloads just fine. Do you know a workaround maybe? Or am I missing something entirely?</p>",
          "rawMarkdown": "Thanks for the quick reply James, however I couldn't figure it out - I get '404 not found' if using the folder name as in your example. Wildcard stage_2_test_images/*.dcm does not help either. \nA specific file from the folder downloads just fine. Do you know a workaround maybe? Or am I missing something entirely?",
          "votes": 1
        },
        {
          "id": 665415,
          "postDate": "2019-11-05T01:35:11.920Z",
          "content": "<p>Interesting...usually there is a zip file to download all of the files in a folder, but it looks like in this case they only have a folder with individual files to download. </p>\n\n<p>One workaround is to loop through all of the stage 2 ids and download them individually. The below script should do the trick, but when I just tested it some of the files are downloading but some of the files I am also getting 404 errors so perhaps the images aren't all available just yet. </p>\n\n<p>```\nimport numpy as np\nimport pandas as pd\nimport subprocess\nfrom tqdm import tqdm</p>\n\n<p>df = pd.read_csv('stage_2_sample_submission.csv')\ndicom_ids = ['ID_' + i.split('_')[1] + '.dcm' for i in df.ID.values]\ndicom_ids = np.unique(dicom_ids)</p>\n\n<p>for i in tqdm(range(len(dicom_ids))):\n    subprocess.call(['kaggle', 'competitions', 'download', '-c', 'rsna-intracranial-hemorrhage-detection', '-f', f'stage_2_test_images/{dicom_ids[i]}'])\n```</p>",
          "rawMarkdown": "Interesting...usually there is a zip file to download all of the files in a folder, but it looks like in this case they only have a folder with individual files to download. \n\nOne workaround is to loop through all of the stage 2 ids and download them individually. The below script should do the trick, but when I just tested it some of the files are downloading but some of the files I am also getting 404 errors so perhaps the images aren't all available just yet. \n\n```\nimport numpy as np\nimport pandas as pd\nimport subprocess\nfrom tqdm import tqdm\n\n\ndf = pd.read_csv('stage_2_sample_submission.csv')\ndicom_ids = ['ID_' + i.split('_')[1] + '.dcm' for i in df.ID.values]\ndicom_ids = np.unique(dicom_ids)\n\nfor i in tqdm(range(len(dicom_ids))):\n    subprocess.call(['kaggle', 'competitions', 'download', '-c', 'rsna-intracranial-hemorrhage-detection', '-f', f'stage_2_test_images/{dicom_ids[i]}'])\n```",
          "votes": 4
        },
        {
          "id": 665432,
          "postDate": "2019-11-05T02:26:35.503Z",
          "content": "<p>yes, I thought of doing it this way too. But ended up starting the download of the whole package. Will see how it will work - 67Gb and still 3 hours to go so far. But it's just an estimate of course. And as you noticed there maybe some problems with some files missing anyway I guess.</p>",
          "rawMarkdown": "yes, I thought of doing it this way too. But ended up starting the download of the whole package. Will see how it will work - 67Gb and still 3 hours to go so far. But it's just an estimate of course. And as you noticed there maybe some problems with some files missing anyway I guess."
        },
        {
          "id": 665573,
          "postDate": "2019-11-05T06:09:53.940Z",
          "content": "<p>Oof downloading individual images is painfully slow, like 1-2 seconds per image. I guess because it has to make a bunch of API calls. Anyone have a better way? I could multiprocess.. but even if it were to go 24 times as fast that is barely going to be 1MBps</p>",
          "rawMarkdown": "Oof downloading individual images is painfully slow, like 1-2 seconds per image. I guess because it has to make a bunch of API calls. Anyone have a better way? I could multiprocess.. but even if it were to go 24 times as fast that is barely going to be 1MBps"
        },
        {
          "id": 665623,
          "postDate": "2019-11-05T07:31:05.980Z",
          "content": "<p><a href=\"/jamesrequa\">@jamesrequa</a> \nI think i downloaded the test_stage1 images  during stage1  itself. \nThere shouldnt be any further need to redownload the set again as whole ?</p>",
          "rawMarkdown": "@jamesrequa \nI think i downloaded the test_stage1 images  during stage1  itself. \nThere shouldnt be any further need to redownload the set again as whole ?"
        },
        {
          "id": 665690,
          "postDate": "2019-11-05T09:38:03.600Z",
          "content": "<p>too many requests. Multiprocessing might make it worse. I start to fear that I won't get the dataset before the deadline end.  </p>",
          "rawMarkdown": "too many requests. Multiprocessing might make it worse. I start to fear that I won't get the dataset before the deadline end.  "
        }
      ]
    },
    {
      "id": 666312,
      "postDate": "2019-11-06T01:36:35.617Z",
      "content": "<p>Try this : <a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/115855\">https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/115855</a>\nStage 2 test set only !!</p>",
      "rawMarkdown": "Try this : https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/115855\nStage 2 test set only !!",
      "votes": 1
    },
    {
      "id": 666047,
      "postDate": "2019-11-05T17:05:26.777Z",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F753323%2F5c9f7de9254b0aac7cd109ad260d0a27%2F1.PNG?generation=1572973392088510&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F753323%2F30737a446e182db138c3547e36a42dc9%2F2.PNG?generation=1572973517413005&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F753323%2F5c9f7de9254b0aac7cd109ad260d0a27%2F1.PNG?generation=1572973392088510&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F753323%2F30737a446e182db138c3547e36a42dc9%2F2.PNG?generation=1572973517413005&amp;alt=media)\n",
      "votes": 1,
      "replies": [
        {
          "id": 666116,
          "postDate": "2019-11-05T18:39:37.050Z",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F753323%2Fc4cd33f432a47af3d6f812a9c97e834e%2F3.PNG?generation=1572979057625454&amp;alt=media\" alt=\"\">\nUsing API helps.\nWhy don't you just make a separate stage2test.zip? Just why? Why should I download the whole 181Gb dataset? </p>",
          "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F753323%2Fc4cd33f432a47af3d6f812a9c97e834e%2F3.PNG?generation=1572979057625454&amp;alt=media)\nUsing API helps.\nWhy don't you just make a separate stage2test.zip? Just why? Why should I download the whole 181Gb dataset? ",
          "votes": 1
        },
        {
          "id": 666139,
          "postDate": "2019-11-05T19:25:11.933Z",
          "content": "<p>Hi <a href=\"/vadiksadik\">@vadiksadik</a> - I'm sorry you're having trouble! Due to the way our API and download interfaces work at the moment, everything gets packaged together. We're working on a long-term fix for that.</p>\n\n<p>In the meantime, though, when the final stage 2 dataset is released, it will only include stage 2 test materials. It should be a significantly smaller download.</p>",
          "rawMarkdown": "Hi @vadiksadik - I'm sorry you're having trouble! Due to the way our API and download interfaces work at the moment, everything gets packaged together. We're working on a long-term fix for that.\n\nIn the meantime, though, when the final stage 2 dataset is released, it will only include stage 2 test materials. It should be a significantly smaller download.",
          "votes": 1
        },
        {
          "id": 666159,
          "postDate": "2019-11-05T19:59:36.893Z",
          "content": "<p><a href=\"/philculliton\">@philculliton</a> Thank you! So tomorrow, when stage2 will start, this line of code\n<code>kaggle competitions download -c rsna-intracranial-hemorrhage-detection</code>\nwill download only stage2 test data and not the whole 181Gb dataset?\nUPDATE: Yes, now in \"data\" page there is only stage2 test about 25Gb. Nice)</p>",
          "rawMarkdown": "@philculliton Thank you! So tomorrow, when stage2 will start, this line of code\n`kaggle competitions download -c rsna-intracranial-hemorrhage-detection`\nwill download only stage2 test data and not the whole 181Gb dataset?\nUPDATE: Yes, now in \"data\" page there is only stage2 test about 25Gb. Nice)"
        },
        {
          "id": 666168,
          "postDate": "2019-11-05T20:28:04.320Z",
          "content": "<p>It would be nice if Kaggle can give us one or two days extra, due to data download difficulties.</p>",
          "rawMarkdown": "It would be nice if Kaggle can give us one or two days extra, due to data download difficulties.",
          "votes": 1
        },
        {
          "id": 666179,
          "postDate": "2019-11-05T20:55:34.543Z",
          "content": "<p><a href=\"/philculliton\">@philculliton</a> Wait, are you saying that the existing stage_ two _ test_images folder in the 181GB dataset is NOT the final state 2 dataset? There will be another one?</p>\n\n<p>Edit: Never mind, I just saw your post below that states you will ultimately only be hosting Stage 2 Test. So presumably you will be removing the 181GB dataset and replacing it with the 30GB one. Cool :)</p>",
          "rawMarkdown": "@philculliton Wait, are you saying that the existing stage_ two _ test_images folder in the 181GB dataset is NOT the final state 2 dataset? There will be another one?\n\nEdit: Never mind, I just saw your post below that states you will ultimately only be hosting Stage 2 Test. So presumably you will be removing the 181GB dataset and replacing it with the 30GB one. Cool :)",
          "votes": 1
        }
      ]
    },
    {
      "id": 665931,
      "postDate": "2019-11-05T14:51:58.960Z",
      "content": "<p>「Stage 1 has ended! We are updating the dataset and resetting the leaderboard. You can expect to see some shuffling and new data in this process. Stage 2 is not expected to begin until the evening of UTC time on 11/5.」</p>\n\n<p>It seems that the dataset is under updating and it cannot be download until the stage 2 start announcement.</p>",
      "rawMarkdown": "「Stage 1 has ended! We are updating the dataset and resetting the leaderboard. You can expect to see some shuffling and new data in this process. Stage 2 is not expected to begin until the evening of UTC time on 11/5.」\n\nIt seems that the dataset is under updating and it cannot be download until the stage 2 start announcement.",
      "votes": 1
    },
    {
      "id": 665659,
      "postDate": "2019-11-05T08:47:44.147Z",
      "content": "<p>Really? But why to download full dataset two times?</p>",
      "rawMarkdown": "Really? But why to download full dataset two times?",
      "votes": 1
    },
    {
      "id": 666183,
      "postDate": "2019-11-05T21:13:06.763Z",
      "content": "<p>This was the first time I used API, and wondering if it is too difficult to implement the use of wildcards for file operations? It would be much easier, and faster to download just all files from the new test set: something like  '-f stage_2_test_images/*' . \nP.S.\nI guess I was lucky to start download earlier and get the whole dataset without hitches.\nSo far I've checked new test set and all files seems to be fine (121232 files, about 59Gb in total, unpacked).</p>",
      "rawMarkdown": "This was the first time I used API, and wondering if it is too difficult to implement the use of wildcards for file operations? It would be much easier, and faster to download just all files from the new test set: something like  '-f stage_2_test_images/*' . \nP.S.\nI guess I was lucky to start download earlier and get the whole dataset without hitches.\nSo far I've checked new test set and all files seems to be fine (121232 files, about 59Gb in total, unpacked)."
    },
    {
      "id": 665807,
      "postDate": "2019-11-05T12:19:40.477Z",
      "content": "<p>Damnit, downloaded 167/181GB and it died. </p>\n\n<p>Man that is frustrating</p>",
      "rawMarkdown": "Damnit, downloaded 167/181GB and it died. \n\nMan that is frustrating",
      "replies": [
        {
          "id": 665919,
          "postDate": "2019-11-05T14:30:26.017Z",
          "content": "<p>you are lucky. \nAfter 8 hours downloading, I only got 8 images, and endlessly \"too many requests\". </p>",
          "rawMarkdown": "you are lucky. \nAfter 8 hours downloading, I only got 8 images, and endlessly \"too many requests\". "
        }
      ]
    },
    {
      "id": 665877,
      "postDate": "2019-11-05T13:41:37.680Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 666035,
      "author_name": "Phil Culliton",
      "author_url": "",
      "post_date": "2019-11-05T16:51:48.113000",
      "content": "<p>Hi all! Very sorry to hear about everyone's trouble so far.</p>\n\n<blockquote>\n  <p>Stage 1 has ended! We are updating the dataset and resetting the leaderboard. You can expect to see some shuffling and new data in this process. Stage 2 is not expected to begin until the evening of UTC time on 11/5.</p>\n</blockquote>\n\n<p>The above is what's happening now - we're still moving data around. We'll ultimately be hosting only the stage 2 images, updated labels, and new sample submission. You will not have to redownload stage 1 data. We're also working on longer-term design changes to our API and download sections so you can download selected archives again. Thanks for your patience, we'll let you know when the data is ready to go!</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 665387,
      "author_name": "James Requa",
      "author_url": "",
      "post_date": "2019-11-05T00:46:22.760000",
      "content": "<p>Yes. You can specify the file you want to download via the kaggle api using the <code>-f</code> flag. See below example for reference.</p>\n\n<p><code>kaggle competitions download -c rsna-intracranial-hemorrhage-detection -f stage_2_test_images</code></p>",
      "votes": 3,
      "replies": [
        {
          "id": 665408,
          "author_name": "Alexey Kotlik",
          "author_url": "",
          "post_date": "2019-11-05T01:20:48.533000",
          "content": "<p>Thanks for the quick reply James, however I couldn't figure it out - I get '404 not found' if using the folder name as in your example. Wildcard stage_2_test_images/*.dcm does not help either. \nA specific file from the folder downloads just fine. Do you know a workaround maybe? Or am I missing something entirely?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 665415,
          "author_name": "James Requa",
          "author_url": "",
          "post_date": "2019-11-05T01:35:11.920000",
          "content": "<p>Interesting...usually there is a zip file to download all of the files in a folder, but it looks like in this case they only have a folder with individual files to download. </p>\n\n<p>One workaround is to loop through all of the stage 2 ids and download them individually. The below script should do the trick, but when I just tested it some of the files are downloading but some of the files I am also getting 404 errors so perhaps the images aren't all available just yet. </p>\n\n<p>```\nimport numpy as np\nimport pandas as pd\nimport subprocess\nfrom tqdm import tqdm</p>\n\n<p>df = pd.read_csv('stage_2_sample_submission.csv')\ndicom_ids = ['ID_' + i.split('_')[1] + '.dcm' for i in df.ID.values]\ndicom_ids = np.unique(dicom_ids)</p>\n\n<p>for i in tqdm(range(len(dicom_ids))):\n    subprocess.call(['kaggle', 'competitions', 'download', '-c', 'rsna-intracranial-hemorrhage-detection', '-f', f'stage_2_test_images/{dicom_ids[i]}'])\n```</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 665432,
          "author_name": "Alexey Kotlik",
          "author_url": "",
          "post_date": "2019-11-05T02:26:35.503000",
          "content": "<p>yes, I thought of doing it this way too. But ended up starting the download of the whole package. Will see how it will work - 67Gb and still 3 hours to go so far. But it's just an estimate of course. And as you noticed there maybe some problems with some files missing anyway I guess.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 665573,
          "author_name": "cherring",
          "author_url": "",
          "post_date": "2019-11-05T06:09:53.940000",
          "content": "<p>Oof downloading individual images is painfully slow, like 1-2 seconds per image. I guess because it has to make a bunch of API calls. Anyone have a better way? I could multiprocess.. but even if it were to go 24 times as fast that is barely going to be 1MBps</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 665623,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2019-11-05T07:31:05.980000",
          "content": "<p><a href=\"/jamesrequa\">@jamesrequa</a> \nI think i downloaded the test_stage1 images  during stage1  itself. \nThere shouldnt be any further need to redownload the set again as whole ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 665690,
          "author_name": "cjma",
          "author_url": "",
          "post_date": "2019-11-05T09:38:03.600000",
          "content": "<p>too many requests. Multiprocessing might make it worse. I start to fear that I won't get the dataset before the deadline end.  </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 666312,
      "author_name": "yukiya",
      "author_url": "",
      "post_date": "2019-11-06T01:36:35.617000",
      "content": "<p>Try this : <a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/115855\">https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/115855</a>\nStage 2 test set only !!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 666047,
      "author_name": "Vadik",
      "author_url": "",
      "post_date": "2019-11-05T17:05:26.777000",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F753323%2F5c9f7de9254b0aac7cd109ad260d0a27%2F1.PNG?generation=1572973392088510&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F753323%2F30737a446e182db138c3547e36a42dc9%2F2.PNG?generation=1572973517413005&amp;alt=media\" alt=\"\"></p>",
      "votes": 1,
      "replies": [
        {
          "id": 666116,
          "author_name": "Vadik",
          "author_url": "",
          "post_date": "2019-11-05T18:39:37.050000",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F753323%2Fc4cd33f432a47af3d6f812a9c97e834e%2F3.PNG?generation=1572979057625454&amp;alt=media\" alt=\"\">\nUsing API helps.\nWhy don't you just make a separate stage2test.zip? Just why? Why should I download the whole 181Gb dataset? </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 666139,
          "author_name": "Phil Culliton",
          "author_url": "",
          "post_date": "2019-11-05T19:25:11.933000",
          "content": "<p>Hi <a href=\"/vadiksadik\">@vadiksadik</a> - I'm sorry you're having trouble! Due to the way our API and download interfaces work at the moment, everything gets packaged together. We're working on a long-term fix for that.</p>\n\n<p>In the meantime, though, when the final stage 2 dataset is released, it will only include stage 2 test materials. It should be a significantly smaller download.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 666159,
          "author_name": "Vadik",
          "author_url": "",
          "post_date": "2019-11-05T19:59:36.893000",
          "content": "<p><a href=\"/philculliton\">@philculliton</a> Thank you! So tomorrow, when stage2 will start, this line of code\n<code>kaggle competitions download -c rsna-intracranial-hemorrhage-detection</code>\nwill download only stage2 test data and not the whole 181Gb dataset?\nUPDATE: Yes, now in \"data\" page there is only stage2 test about 25Gb. Nice)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 666168,
          "author_name": "Ren",
          "author_url": "",
          "post_date": "2019-11-05T20:28:04.320000",
          "content": "<p>It would be nice if Kaggle can give us one or two days extra, due to data download difficulties.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 666179,
          "author_name": "cherring",
          "author_url": "",
          "post_date": "2019-11-05T20:55:34.543000",
          "content": "<p><a href=\"/philculliton\">@philculliton</a> Wait, are you saying that the existing stage_ two _ test_images folder in the 181GB dataset is NOT the final state 2 dataset? There will be another one?</p>\n\n<p>Edit: Never mind, I just saw your post below that states you will ultimately only be hosting Stage 2 Test. So presumably you will be removing the 181GB dataset and replacing it with the 30GB one. Cool :)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 665931,
      "author_name": "liuzhangzhen",
      "author_url": "",
      "post_date": "2019-11-05T14:51:58.960000",
      "content": "<p>「Stage 1 has ended! We are updating the dataset and resetting the leaderboard. You can expect to see some shuffling and new data in this process. Stage 2 is not expected to begin until the evening of UTC time on 11/5.」</p>\n\n<p>It seems that the dataset is under updating and it cannot be download until the stage 2 start announcement.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 665659,
      "author_name": "Oleg Yaroshevskiy",
      "author_url": "",
      "post_date": "2019-11-05T08:47:44.147000",
      "content": "<p>Really? But why to download full dataset two times?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 666183,
      "author_name": "Alexey Kotlik",
      "author_url": "",
      "post_date": "2019-11-05T21:13:06.763000",
      "content": "<p>This was the first time I used API, and wondering if it is too difficult to implement the use of wildcards for file operations? It would be much easier, and faster to download just all files from the new test set: something like  '-f stage_2_test_images/*' . \nP.S.\nI guess I was lucky to start download earlier and get the whole dataset without hitches.\nSo far I've checked new test set and all files seems to be fine (121232 files, about 59Gb in total, unpacked).</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 665807,
      "author_name": "cherring",
      "author_url": "",
      "post_date": "2019-11-05T12:19:40.477000",
      "content": "<p>Damnit, downloaded 167/181GB and it died. </p>\n\n<p>Man that is frustrating</p>",
      "votes": 0,
      "replies": [
        {
          "id": 665919,
          "author_name": "cjma",
          "author_url": "",
          "post_date": "2019-11-05T14:30:26.017000",
          "content": "<p>you are lucky. \nAfter 8 hours downloading, I only got 8 images, and endlessly \"too many requests\". </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 665877,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-11-05T13:41:37.680000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "665384": "I understood that train set will remain the same from the stage 1, and the test set labels will become availabel so we could use it to extend the train set. \nIs there a way to download just 30Gb of new test images, instead of the whole pile of 180Gb of which we already supposedly have 150Gb on our HDDs? ",
    "666035": "Hi all! Very sorry to hear about everyone's trouble so far.\n\n&gt; Stage 1 has ended! We are updating the dataset and resetting the leaderboard. You can expect to see some shuffling and new data in this process. Stage 2 is not expected to begin until the evening of UTC time on 11/5.\n\nThe above is what's happening now - we're still moving data around. We'll ultimately be hosting only the stage 2 images, updated labels, and new sample submission. You will not have to redownload stage 1 data. We're also working on longer-term design changes to our API and download sections so you can download selected archives again. Thanks for your patience, we'll let you know when the data is ready to go!",
    "665387": "Yes. You can specify the file you want to download via the kaggle api using the `-f` flag. See below example for reference.\n\n`kaggle competitions download -c rsna-intracranial-hemorrhage-detection -f stage_2_test_images`",
    "666312": "Try this : https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/115855\nStage 2 test set only !!",
    "666047": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F753323%2F5c9f7de9254b0aac7cd109ad260d0a27%2F1.PNG?generation=1572973392088510&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F753323%2F30737a446e182db138c3547e36a42dc9%2F2.PNG?generation=1572973517413005&amp;alt=media)\n",
    "665931": "「Stage 1 has ended! We are updating the dataset and resetting the leaderboard. You can expect to see some shuffling and new data in this process. Stage 2 is not expected to begin until the evening of UTC time on 11/5.」\n\nIt seems that the dataset is under updating and it cannot be download until the stage 2 start announcement.",
    "665659": "Really? But why to download full dataset two times?",
    "666183": "This was the first time I used API, and wondering if it is too difficult to implement the use of wildcards for file operations? It would be much easier, and faster to download just all files from the new test set: something like  '-f stage_2_test_images/*' . \nP.S.\nI guess I was lucky to start download earlier and get the whole dataset without hitches.\nSo far I've checked new test set and all files seems to be fine (121232 files, about 59Gb in total, unpacked).",
    "665807": "Damnit, downloaded 167/181GB and it died. \n\nMan that is frustrating",
    "665877": ""
  }
}