{
  "id": 186982,
  "title": "Saving time and GPU before submitting",
  "url": "/competitions/rsna-str-pulmonary-embolism-detection/discussion/186982",
  "author_name": "yuval reina",
  "post_date": "2020-09-26T20:09:32.381000",
  "votes": 74,
  "comment_count": 16,
  "views": 0,
  "content": "<p>It takes my inference  1-2h to run on the public test set, and then 4-8h on the full test set.<br>\nThis makes the submission procedure very long and uses a lot of GPU time for every submission.<br>\nI solved it using the following lines:</p>\n<pre><code>from os import path\nif path.exists('../input/rsna-str-pulmonary-embolism-detection/train') and not do_full:\n    df=pd.read_csv(params.path.data+'test.csv').head(2000)\nelse:\n    df=pd.read_csv(params.path.data+'test.csv')\n</code></pre>\n<p>If the notebook is running on the public test, then the training data exists and then it loads only the first 2000 lines and finish running in about 2 mins.<br>\nIf train folder does not exist, then it runs on the hidden test and need to load the full test data.<br>\n<code>do_full=True</code> when I want to run fully on the public test - to measure running time etc.</p>\n<p>**Be careful using this **- first test your pipeline  for a few submissions   </p>",
  "messages": [
    {
      "id": 1028378,
      "postDate": "2020-09-26T20:09:32.380Z",
      "content": "<p>It takes my inference  1-2h to run on the public test set, and then 4-8h on the full test set.<br>\nThis makes the submission procedure very long and uses a lot of GPU time for every submission.<br>\nI solved it using the following lines:</p>\n<pre><code>from os import path\nif path.exists('../input/rsna-str-pulmonary-embolism-detection/train') and not do_full:\n    df=pd.read_csv(params.path.data+'test.csv').head(2000)\nelse:\n    df=pd.read_csv(params.path.data+'test.csv')\n</code></pre>\n<p>If the notebook is running on the public test, then the training data exists and then it loads only the first 2000 lines and finish running in about 2 mins.<br>\nIf train folder does not exist, then it runs on the hidden test and need to load the full test data.<br>\n<code>do_full=True</code> when I want to run fully on the public test - to measure running time etc.</p>\n<p>**Be careful using this **- first test your pipeline  for a few submissions   </p>",
      "rawMarkdown": "It takes my inference  1-2h to run on the public test set, and then 4-8h on the full test set.\nThis makes the submission procedure very long and uses a lot of GPU time for every submission.\nI solved it using the following lines:\n```\nfrom os import path\nif path.exists('../input/rsna-str-pulmonary-embolism-detection/train') and not do_full:\n    df=pd.read_csv(params.path.data+'test.csv').head(2000)\nelse:\n    df=pd.read_csv(params.path.data+'test.csv')\n```\n\nIf the notebook is running on the public test, then the training data exists and then it loads only the first 2000 lines and finish running in about 2 mins.\nIf train folder does not exist, then it runs on the hidden test and need to load the full test data.\n`do_full=True` when I want to run fully on the public test - to measure running time etc.\n\n**Be careful using this **- first test your pipeline  for a few submissions   ",
      "votes": 73
    },
    {
      "id": 1029805,
      "postDate": "2020-09-28T06:35:37.647Z",
      "content": "<p>Nice solution to save some time! Thanks for sharing.</p>",
      "rawMarkdown": "Nice solution to save some time! Thanks for sharing.",
      "votes": 3
    },
    {
      "id": 1037610,
      "postDate": "2020-10-05T07:02:14.487Z",
      "content": "<p>Will there be a <code>train.csv</code> file during rerun? Also, does the public score depends on public test data or private test data. </p>",
      "rawMarkdown": "Will there be a `train.csv` file during rerun? Also, does the public score depends on public test data or private test data. ",
      "votes": 1,
      "replies": [
        {
          "id": 1037623,
          "postDate": "2020-10-05T07:10:47.707Z",
          "content": "<p>Yes, Julia indicated <a href=\"https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/182251\" target=\"_blank\">here</a> that <code>train.csv</code> will be available. </p>\n<p>The public score is based on the public test images, whereas the private score is computed on the hidden test images.</p>",
          "rawMarkdown": "Yes, Julia indicated [here](https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/182251) that `train.csv` will be available. \n\nThe public score is based on the public test images, whereas the private score is computed on the hidden test images.",
          "votes": 3
        }
      ]
    },
    {
      "id": 1047403,
      "postDate": "2020-10-12T15:06:54.613Z",
      "content": "<p>I might have misunderstood your post. If you limit reading the public test set to only 2000 lines, and the currently displayed public LB is dependent on the public test set, won't you get an inaccurate LB that you can't use it to gauge the performance of your model? </p>\n<p>So is this notebook just a test to make sure that your notebook run within the 9 hour limit for the private test?</p>",
      "rawMarkdown": "I might have misunderstood your post. If you limit reading the public test set to only 2000 lines, and the currently displayed public LB is dependent on the public test set, won't you get an inaccurate LB that you can't use it to gauge the performance of your model? \n\nSo is this notebook just a test to make sure that your notebook run within the 9 hour limit for the private test?",
      "replies": [
        {
          "id": 1047418,
          "postDate": "2020-10-12T15:31:32.657Z",
          "content": "<p>When you <strong>commit</strong> the notebook (\"Save and Run All\") it runs on the full public test dataset. After that you <strong>submit</strong> \"submission.csv\" and it runs again on the full test dataset (public + private), consuming your GPU quota. So, if you are sure your submitting pipeline has no errors, you can save some GPU-time and do <strong>commit</strong> only on 2000 samples (let's say ~2 minutes instead of ~2 hours). While you <strong>submit</strong> it anyway will run through the full testset.</p>\n<p>Hope it helps.</p>",
          "rawMarkdown": "When you **commit** the notebook (\"Save and Run All\") it runs on the full public test dataset. After that you **submit** \"submission.csv\" and it runs again on the full test dataset (public + private), consuming your GPU quota. So, if you are sure your submitting pipeline has no errors, you can save some GPU-time and do **commit** only on 2000 samples (let's say ~2 minutes instead of ~2 hours). While you **submit** it anyway will run through the full testset.\n\nHope it helps.",
          "votes": 5
        },
        {
          "id": 1047457,
          "postDate": "2020-10-12T16:13:43.497Z",
          "content": "<p>Got it. Thanks!</p>",
          "rawMarkdown": "Got it. Thanks!"
        }
      ]
    },
    {
      "id": 1032032,
      "postDate": "2020-09-29T21:48:13.167Z",
      "content": "<p>A really nice way to save our GPU quota. Thanks for sharing!</p>",
      "rawMarkdown": "A really nice way to save our GPU quota. Thanks for sharing!"
    },
    {
      "id": 1031888,
      "postDate": "2020-09-29T18:50:14.617Z",
      "content": "<p>Did you use TPU to get that quick time. I am using Pytorch lightning/TPU and it's quite slow</p>",
      "rawMarkdown": "Did you use TPU to get that quick time. I am using Pytorch lightning/TPU and it's quite slow\n",
      "replies": [
        {
          "id": 1031977,
          "postDate": "2020-09-29T20:21:00.343Z",
          "content": "<p>You can use TPU on the training stage, but using TPU in an inference notebook is prohibited. The submission is only possible with CPU/GPU.</p>",
          "rawMarkdown": "You can use TPU on the training stage, but using TPU in an inference notebook is prohibited. The submission is only possible with CPU/GPU.",
          "votes": 1
        },
        {
          "id": 1031987,
          "postDate": "2020-09-29T20:31:16.270Z",
          "content": "<p>Yep. I know that point. But config TPU in PyTorch needs more experience than TF. I tried but not success, now i am training with GPU</p>",
          "rawMarkdown": "Yep. I know that point. But config TPU in PyTorch needs more experience than TF. I tried but not success, now i am training with GPU"
        },
        {
          "id": 1032400,
          "postDate": "2020-09-30T07:18:26.017Z",
          "content": "<p>True. I am working on a PyTorch XLA pipeline and it does require a lot of effort to be set up efficiently. I will think about releasing a public notebook once the pipeline is up and running.</p>",
          "rawMarkdown": "True. I am working on a PyTorch XLA pipeline and it does require a lot of effort to be set up efficiently. I will think about releasing a public notebook once the pipeline is up and running."
        },
        {
          "id": 1032452,
          "postDate": "2020-09-30T08:01:27.370Z",
          "content": "<p>today I read about XLA 1.6. That's incredibly powerful and easy to use. Did you try it?</p>",
          "rawMarkdown": "today I read about XLA 1.6. That's incredibly powerful and easy to use. Did you try it?"
        },
        {
          "id": 1034108,
          "postDate": "2020-10-01T14:02:47.980Z",
          "content": "<p>I am working with XLA 1.6 but I am not sure I am utilizing its full potential yet :)</p>",
          "rawMarkdown": "I am working with XLA 1.6 but I am not sure I am utilizing its full potential yet :)"
        }
      ]
    },
    {
      "id": 1031785,
      "postDate": "2020-09-29T17:27:34.987Z",
      "content": "<p>Sorry to bother you, but I don't understand why you only loads the first 2000 lines rather than all?</p>",
      "rawMarkdown": "Sorry to bother you, but I don't understand why you only loads the first 2000 lines rather than all?",
      "replies": [
        {
          "id": 1031840,
          "postDate": "2020-09-29T17:57:07.773Z",
          "content": "<p><a href=\"https://www.kaggle.com/maxiuyu\" target=\"_blank\">@maxiuyu</a> If I'll load all the line it will take 1H when I commit, to process all the data and create the submission file. When I load only 2000 lines, it takes 2 min. When the commit is finished and I have a submission.csv file, I can submit, and then, when the notebook is rerun on the hidden test data, I load all files.</p>",
          "rawMarkdown": "@maxiuyu If I'll load all the line it will take 1H when I commit, to process all the data and create the submission file. When I load only 2000 lines, it takes 2 min. When the commit is finished and I have a submission.csv file, I can submit, and then, when the notebook is rerun on the hidden test data, I load all files.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1032468,
      "postDate": "2020-09-30T08:25:47.137Z",
      "content": "<p>Thanks for sharing!!</p>",
      "rawMarkdown": "Thanks for sharing!!"
    }
  ],
  "comments": [
    {
      "id": 1029805,
      "author_name": "Nikita Kozodoi",
      "author_url": "",
      "post_date": "2020-09-28T06:35:37.647000",
      "content": "<p>Nice solution to save some time! Thanks for sharing.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1037610,
      "author_name": "Jose",
      "author_url": "",
      "post_date": "2020-10-05T07:02:14.487000",
      "content": "<p>Will there be a <code>train.csv</code> file during rerun? Also, does the public score depends on public test data or private test data. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1037623,
          "author_name": "Nikita Kozodoi",
          "author_url": "",
          "post_date": "2020-10-05T07:10:47.707000",
          "content": "<p>Yes, Julia indicated <a href=\"https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/182251\" target=\"_blank\">here</a> that <code>train.csv</code> will be available. </p>\n<p>The public score is based on the public test images, whereas the private score is computed on the hidden test images.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1047403,
      "author_name": "Yee Ng",
      "author_url": "",
      "post_date": "2020-10-12T15:06:54.613000",
      "content": "<p>I might have misunderstood your post. If you limit reading the public test set to only 2000 lines, and the currently displayed public LB is dependent on the public test set, won't you get an inaccurate LB that you can't use it to gauge the performance of your model? </p>\n<p>So is this notebook just a test to make sure that your notebook run within the 9 hour limit for the private test?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1047418,
          "author_name": "A.Demyanchuk",
          "author_url": "",
          "post_date": "2020-10-12T15:31:32.657000",
          "content": "<p>When you <strong>commit</strong> the notebook (\"Save and Run All\") it runs on the full public test dataset. After that you <strong>submit</strong> \"submission.csv\" and it runs again on the full test dataset (public + private), consuming your GPU quota. So, if you are sure your submitting pipeline has no errors, you can save some GPU-time and do <strong>commit</strong> only on 2000 samples (let's say ~2 minutes instead of ~2 hours). While you <strong>submit</strong> it anyway will run through the full testset.</p>\n<p>Hope it helps.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1047457,
          "author_name": "Yee Ng",
          "author_url": "",
          "post_date": "2020-10-12T16:13:43.497000",
          "content": "<p>Got it. Thanks!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1032032,
      "author_name": "Gaofeng Huang",
      "author_url": "",
      "post_date": "2020-09-29T21:48:13.167000",
      "content": "<p>A really nice way to save our GPU quota. Thanks for sharing!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1031888,
      "author_name": "Manh Lab",
      "author_url": "",
      "post_date": "2020-09-29T18:50:14.617000",
      "content": "<p>Did you use TPU to get that quick time. I am using Pytorch lightning/TPU and it's quite slow</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1031977,
          "author_name": "Nikita Kozodoi",
          "author_url": "",
          "post_date": "2020-09-29T20:21:00.343000",
          "content": "<p>You can use TPU on the training stage, but using TPU in an inference notebook is prohibited. The submission is only possible with CPU/GPU.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1031987,
          "author_name": "Manh Lab",
          "author_url": "",
          "post_date": "2020-09-29T20:31:16.270000",
          "content": "<p>Yep. I know that point. But config TPU in PyTorch needs more experience than TF. I tried but not success, now i am training with GPU</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1032400,
          "author_name": "Nikita Kozodoi",
          "author_url": "",
          "post_date": "2020-09-30T07:18:26.017000",
          "content": "<p>True. I am working on a PyTorch XLA pipeline and it does require a lot of effort to be set up efficiently. I will think about releasing a public notebook once the pipeline is up and running.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1032452,
          "author_name": "Manh Lab",
          "author_url": "",
          "post_date": "2020-09-30T08:01:27.370000",
          "content": "<p>today I read about XLA 1.6. That's incredibly powerful and easy to use. Did you try it?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1034108,
          "author_name": "Nikita Kozodoi",
          "author_url": "",
          "post_date": "2020-10-01T14:02:47.980000",
          "content": "<p>I am working with XLA 1.6 but I am not sure I am utilizing its full potential yet :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1031785,
      "author_name": "MaXiuyu",
      "author_url": "",
      "post_date": "2020-09-29T17:27:34.987000",
      "content": "<p>Sorry to bother you, but I don't understand why you only loads the first 2000 lines rather than all?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1031840,
          "author_name": "yuval reina",
          "author_url": "",
          "post_date": "2020-09-29T17:57:07.773000",
          "content": "<p><a href=\"https://www.kaggle.com/maxiuyu\" target=\"_blank\">@maxiuyu</a> If I'll load all the line it will take 1H when I commit, to process all the data and create the submission file. When I load only 2000 lines, it takes 2 min. When the commit is finished and I have a submission.csv file, I can submit, and then, when the notebook is rerun on the hidden test data, I load all files.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1032468,
      "author_name": "Gryffindor",
      "author_url": "",
      "post_date": "2020-09-30T08:25:47.137000",
      "content": "<p>Thanks for sharing!!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1028378": "It takes my inference  1-2h to run on the public test set, and then 4-8h on the full test set.\nThis makes the submission procedure very long and uses a lot of GPU time for every submission.\nI solved it using the following lines:\n```\nfrom os import path\nif path.exists('../input/rsna-str-pulmonary-embolism-detection/train') and not do_full:\n    df=pd.read_csv(params.path.data+'test.csv').head(2000)\nelse:\n    df=pd.read_csv(params.path.data+'test.csv')\n```\n\nIf the notebook is running on the public test, then the training data exists and then it loads only the first 2000 lines and finish running in about 2 mins.\nIf train folder does not exist, then it runs on the hidden test and need to load the full test data.\n`do_full=True` when I want to run fully on the public test - to measure running time etc.\n\n**Be careful using this **- first test your pipeline  for a few submissions   ",
    "1029805": "Nice solution to save some time! Thanks for sharing.",
    "1037610": "Will there be a `train.csv` file during rerun? Also, does the public score depends on public test data or private test data. ",
    "1047403": "I might have misunderstood your post. If you limit reading the public test set to only 2000 lines, and the currently displayed public LB is dependent on the public test set, won't you get an inaccurate LB that you can't use it to gauge the performance of your model? \n\nSo is this notebook just a test to make sure that your notebook run within the 9 hour limit for the private test?",
    "1032032": "A really nice way to save our GPU quota. Thanks for sharing!",
    "1031888": "Did you use TPU to get that quick time. I am using Pytorch lightning/TPU and it's quite slow\n",
    "1031785": "Sorry to bother you, but I don't understand why you only loads the first 2000 lines rather than all?",
    "1032468": "Thanks for sharing!!"
  }
}