{
  "id": 114986,
  "title": "Download of stage 2 images",
  "url": "/competitions/rsna-intracranial-hemorrhage-detection/discussion/114986",
  "author_name": "steelrose",
  "post_date": "2019-10-30T14:33:52.314000",
  "votes": 0,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi @philculliton and @juliaelliott, sorry to tag you specifically, but hope you'll understand :)</p>\n\n<p>As was already mentioned, the stage 2 will be available for download soon, but I have a question - how will it be downloadable?\nSpecifically questions to the kaggle API:\n- will we be able to download only the stage 2 test set? (if I remember correctly, only downloading the full dataset of a competition was possible lately)\n- some time ago there was a bug with the downloads - was this already fully fixed?\n- what will be the approx. size?</p>\n\n<p>And also a optional proposal for your consideration - wouldn't it be possible to:\n- publish a password protected data set before the planned release date (thus even people with slow internet will have no trouble downloading, also the load on kaggle servers would be probably split into multiple days); also publish a sha1 hash along the file so we can verify the file is not corrupted\n- and publish the password on the planned release date</p>\n\n<p>This was noone will have an advantage due to faster download speeds etc.</p>",
  "messages": [
    {
      "id": 661790,
      "postDate": "2019-10-30T17:53:59.777Z",
      "content": "<p>Hi! Thanks for your questions. :)</p>\n\n<ol>\n<li>Unfortunately, if you're using the API, you'll be downloading ALL files in one large zip. There's a longer-term design change planned, but it's still being worked on.</li>\n<li>Happy to let you know! Which download bug were you wondering about?</li>\n<li>The approximate (compressed) size will the current set - roughly 156 GB - plus an additional ~27 GB of stage 2 test images.</li>\n</ol>",
      "rawMarkdown": "Hi! Thanks for your questions. :)\n\n1. Unfortunately, if you're using the API, you'll be downloading ALL files in one large zip. There's a longer-term design change planned, but it's still being worked on.\n2. Happy to let you know! Which download bug were you wondering about?\n3. The approximate (compressed) size will the current set - roughly 156 GB - plus an additional ~27 GB of stage 2 test images.",
      "votes": 1,
      "replies": [
        {
          "id": 661899,
          "postDate": "2019-10-30T20:47:09.790Z",
          "content": "<p>thanks for the answers, will there be an easy way to download the stage 2 test images? \nthe train/test for stage 1 were either \"click on each image on the Data page\" which was unfeasible, so everyone most likely just used the kaggle api - which as you confirmed will download ALL data again :(</p>\n\n<p>regarding the bug - at the beginning there was the 429 error, then some people reported downloading only 20 images from each folder, not sure if that was all the bugs ... just wanted to make sure we won't have any issues at all</p>",
          "rawMarkdown": "thanks for the answers, will there be an easy way to download the stage 2 test images? \nthe train/test for stage 1 were either \"click on each image on the Data page\" which was unfeasible, so everyone most likely just used the kaggle api - which as you confirmed will download ALL data again :(\n\nregarding the bug - at the beginning there was the 429 error, then some people reported downloading only 20 images from each folder, not sure if that was all the bugs ... just wanted to make sure we won't have any issues at all"
        }
      ]
    },
    {
      "id": 661635,
      "postDate": "2019-10-30T14:33:52.313Z",
      "content": "<p>Hi @philculliton and @juliaelliott, sorry to tag you specifically, but hope you'll understand :)</p>\n\n<p>As was already mentioned, the stage 2 will be available for download soon, but I have a question - how will it be downloadable?\nSpecifically questions to the kaggle API:\n- will we be able to download only the stage 2 test set? (if I remember correctly, only downloading the full dataset of a competition was possible lately)\n- some time ago there was a bug with the downloads - was this already fully fixed?\n- what will be the approx. size?</p>\n\n<p>And also a optional proposal for your consideration - wouldn't it be possible to:\n- publish a password protected data set before the planned release date (thus even people with slow internet will have no trouble downloading, also the load on kaggle servers would be probably split into multiple days); also publish a sha1 hash along the file so we can verify the file is not corrupted\n- and publish the password on the planned release date</p>\n\n<p>This was noone will have an advantage due to faster download speeds etc.</p>",
      "rawMarkdown": "Hi @philculliton and @juliaelliott, sorry to tag you specifically, but hope you'll understand :)\n\nAs was already mentioned, the stage 2 will be available for download soon, but I have a question - how will it be downloadable?\nSpecifically questions to the kaggle API:\n- will we be able to download only the stage 2 test set? (if I remember correctly, only downloading the full dataset of a competition was possible lately)\n- some time ago there was a bug with the downloads - was this already fully fixed?\n- what will be the approx. size?\n\nAnd also a optional proposal for your consideration - wouldn't it be possible to:\n- publish a password protected data set before the planned release date (thus even people with slow internet will have no trouble downloading, also the load on kaggle servers would be probably split into multiple days); also publish a sha1 hash along the file so we can verify the file is not corrupted\n- and publish the password on the planned release date\n\nThis was noone will have an advantage due to faster download speeds etc."
    }
  ],
  "comments": [
    {
      "id": 661790,
      "author_name": "Phil Culliton",
      "author_url": "",
      "post_date": "2019-10-30T17:53:59.777000",
      "content": "<p>Hi! Thanks for your questions. :)</p>\n\n<ol>\n<li>Unfortunately, if you're using the API, you'll be downloading ALL files in one large zip. There's a longer-term design change planned, but it's still being worked on.</li>\n<li>Happy to let you know! Which download bug were you wondering about?</li>\n<li>The approximate (compressed) size will the current set - roughly 156 GB - plus an additional ~27 GB of stage 2 test images.</li>\n</ol>",
      "votes": 1,
      "replies": [
        {
          "id": 661899,
          "author_name": "steelrose",
          "author_url": "",
          "post_date": "2019-10-30T20:47:09.790000",
          "content": "<p>thanks for the answers, will there be an easy way to download the stage 2 test images? \nthe train/test for stage 1 were either \"click on each image on the Data page\" which was unfeasible, so everyone most likely just used the kaggle api - which as you confirmed will download ALL data again :(</p>\n\n<p>regarding the bug - at the beginning there was the 429 error, then some people reported downloading only 20 images from each folder, not sure if that was all the bugs ... just wanted to make sure we won't have any issues at all</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "661790": "Hi! Thanks for your questions. :)\n\n1. Unfortunately, if you're using the API, you'll be downloading ALL files in one large zip. There's a longer-term design change planned, but it's still being worked on.\n2. Happy to let you know! Which download bug were you wondering about?\n3. The approximate (compressed) size will the current set - roughly 156 GB - plus an additional ~27 GB of stage 2 test images.",
    "661635": "Hi @philculliton and @juliaelliott, sorry to tag you specifically, but hope you'll understand :)\n\nAs was already mentioned, the stage 2 will be available for download soon, but I have a question - how will it be downloadable?\nSpecifically questions to the kaggle API:\n- will we be able to download only the stage 2 test set? (if I remember correctly, only downloading the full dataset of a competition was possible lately)\n- some time ago there was a bug with the downloads - was this already fully fixed?\n- what will be the approx. size?\n\nAnd also a optional proposal for your consideration - wouldn't it be possible to:\n- publish a password protected data set before the planned release date (thus even people with slow internet will have no trouble downloading, also the load on kaggle servers would be probably split into multiple days); also publish a sha1 hash along the file so we can verify the file is not corrupted\n- and publish the password on the planned release date\n\nThis was noone will have an advantage due to faster download speeds etc."
  }
}