{
  "id": 349889,
  "title": "Submission Error",
  "url": "/competitions/rsna-2022-cervical-spine-fracture-detection/discussion/349889",
  "author_name": "Arka Bhowmick",
  "post_date": "2022-09-03T09:29:39.879000",
  "votes": 2,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hello, dear Kagglers,<br>\nI have a doubt regarding the submission procedure for this competition and any suggestions could be really helpful.<br>\nI have prepared the inference notebook, where I do not use the internet and I am loading a saved model which I trained on a different notebook. After submitting the notebook, every time I have an error of NOTEBOOK TIMEOUT and hence I believe I am not feeding in the test data in the most efficient manner. <br>\nI will be really grateful if someone could kindly help me with a sample notebook that has passed the submission part. I am really struggling here. <br>\nThanking you in anticipation.<br>\nBest,<br>\nArka </p>",
  "messages": [
    {
      "id": 1929160,
      "postDate": "2022-09-06T21:34:00.360Z",
      "content": "<p>one thing that tripped me up for a while is that the sample test.csv provided isn't valid…. In the \"real test.csv\" you likely have 8 rows per scan….  So depending on how you are generating your submission file you might be processing the \"test scans\" multiple times…  </p>\n<p>So I'm currently doing something like…</p>\n<p>test_scan_ids = pd.unique(df.StudyInstanceUID) </p>\n<p>to get a list of non-duped UIDs and looping thru that to drive the creation of the submission file… </p>",
      "rawMarkdown": "one thing that tripped me up for a while is that the sample test.csv provided isn't valid.... In the \"real test.csv\" you likely have 8 rows per scan....  So depending on how you are generating your submission file you might be processing the \"test scans\" multiple times...  \n\nSo I'm currently doing something like...\n\ntest_scan_ids = pd.unique(df.StudyInstanceUID) \n\nto get a list of non-duped UIDs and looping thru that to drive the creation of the submission file... ",
      "votes": 1
    },
    {
      "id": 1924672,
      "postDate": "2022-09-03T09:29:39.880Z",
      "content": "<p>Hello, dear Kagglers,<br>\nI have a doubt regarding the submission procedure for this competition and any suggestions could be really helpful.<br>\nI have prepared the inference notebook, where I do not use the internet and I am loading a saved model which I trained on a different notebook. After submitting the notebook, every time I have an error of NOTEBOOK TIMEOUT and hence I believe I am not feeding in the test data in the most efficient manner. <br>\nI will be really grateful if someone could kindly help me with a sample notebook that has passed the submission part. I am really struggling here. <br>\nThanking you in anticipation.<br>\nBest,<br>\nArka </p>",
      "rawMarkdown": "Hello, dear Kagglers,\nI have a doubt regarding the submission procedure for this competition and any suggestions could be really helpful.\nI have prepared the inference notebook, where I do not use the internet and I am loading a saved model which I trained on a different notebook. After submitting the notebook, every time I have an error of NOTEBOOK TIMEOUT and hence I believe I am not feeding in the test data in the most efficient manner. \nI will be really grateful if someone could kindly help me with a sample notebook that has passed the submission part. I am really struggling here. \nThanking you in anticipation.\nBest,\nArka ",
      "votes": 2
    },
    {
      "id": 1924815,
      "postDate": "2022-09-03T12:40:10.907Z",
      "content": "<p>NOTEBOOK TIMEOUT seems to indicate that your notebook ran more than the allowed 9 hours.  Is that consistent with what you see?  If so, realize that there are approximately 1500 patients in the real (hidden) test set.  You should convince yourself that your solution can handle that many patients in the allotted time.  You can test this out on some number of test/train patients.   If it can't, then you need to look at speeding things up.  Some ideas:</p>\n<ol>\n<li>Make use of the 4 CPUs you are allowed for some pre-processing and/or pipelining.  You want to make sure your GPU is utilized as much as possible</li>\n<li>You did enable the GPU, right?</li>\n<li>If you are doing K-Fold averaging, you might not be able to run all K-models.</li>\n<li>Similarly, if you are doing an ensemble, maybe you can't use that many models</li>\n<li>Use a smaller network model, or use a smaller input image size</li>\n</ol>",
      "rawMarkdown": "NOTEBOOK TIMEOUT seems to indicate that your notebook ran more than the allowed 9 hours.  Is that consistent with what you see?  If so, realize that there are approximately 1500 patients in the real (hidden) test set.  You should convince yourself that your solution can handle that many patients in the allotted time.  You can test this out on some number of test/train patients.   If it can't, then you need to look at speeding things up.  Some ideas:\n\n1.  Make use of the 4 CPUs you are allowed for some pre-processing and/or pipelining.  You want to make sure your GPU is utilized as much as possible\n2. You did enable the GPU, right?\n3. If you are doing K-Fold averaging, you might not be able to run all K-models.\n4. Similarly, if you are doing an ensemble, maybe you can't use that many models\n4. Use a smaller network model, or use a smaller input image size\n",
      "replies": [
        {
          "id": 1925220,
          "postDate": "2022-09-03T18:18:15.640Z",
          "content": "<p>Thank you so much for your detailed explanation.  And just so that you know,<br>\n-&gt; I have enabled the GPU<br>\n-&gt; Then there is no K-Fold averaging or ensemble involved. Everything is very simple, as this is my first trial.<br>\n-&gt; And I think as you have mentioned, a smaller input image might be helpful.</p>",
          "rawMarkdown": "Thank you so much for your detailed explanation.  And just so that you know,\n-> I have enabled the GPU\n-> Then there is no K-Fold averaging or ensemble involved. Everything is very simple, as this is my first trial.\n-> And I think as you have mentioned, a smaller input image might be helpful."
        },
        {
          "id": 1927885,
          "postDate": "2022-09-06T03:46:51.057Z",
          "content": "<p>It appears you only get 2 cpus if you activate the GPU</p>",
          "rawMarkdown": "It appears you only get 2 cpus if you activate the GPU"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1929160,
      "author_name": "John Robinson",
      "author_url": "",
      "post_date": "2022-09-06T21:34:00.360000",
      "content": "<p>one thing that tripped me up for a while is that the sample test.csv provided isn't valid…. In the \"real test.csv\" you likely have 8 rows per scan….  So depending on how you are generating your submission file you might be processing the \"test scans\" multiple times…  </p>\n<p>So I'm currently doing something like…</p>\n<p>test_scan_ids = pd.unique(df.StudyInstanceUID) </p>\n<p>to get a list of non-duped UIDs and looping thru that to drive the creation of the submission file… </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1924815,
      "author_name": "SolverWorld",
      "author_url": "",
      "post_date": "2022-09-03T12:40:10.907000",
      "content": "<p>NOTEBOOK TIMEOUT seems to indicate that your notebook ran more than the allowed 9 hours.  Is that consistent with what you see?  If so, realize that there are approximately 1500 patients in the real (hidden) test set.  You should convince yourself that your solution can handle that many patients in the allotted time.  You can test this out on some number of test/train patients.   If it can't, then you need to look at speeding things up.  Some ideas:</p>\n<ol>\n<li>Make use of the 4 CPUs you are allowed for some pre-processing and/or pipelining.  You want to make sure your GPU is utilized as much as possible</li>\n<li>You did enable the GPU, right?</li>\n<li>If you are doing K-Fold averaging, you might not be able to run all K-models.</li>\n<li>Similarly, if you are doing an ensemble, maybe you can't use that many models</li>\n<li>Use a smaller network model, or use a smaller input image size</li>\n</ol>",
      "votes": 0,
      "replies": [
        {
          "id": 1925220,
          "author_name": "Arka Bhowmick",
          "author_url": "",
          "post_date": "2022-09-03T18:18:15.640000",
          "content": "<p>Thank you so much for your detailed explanation.  And just so that you know,<br>\n-&gt; I have enabled the GPU<br>\n-&gt; Then there is no K-Fold averaging or ensemble involved. Everything is very simple, as this is my first trial.<br>\n-&gt; And I think as you have mentioned, a smaller input image might be helpful.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1927885,
          "author_name": "SolverWorld",
          "author_url": "",
          "post_date": "2022-09-06T03:46:51.057000",
          "content": "<p>It appears you only get 2 cpus if you activate the GPU</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1929160": "one thing that tripped me up for a while is that the sample test.csv provided isn't valid.... In the \"real test.csv\" you likely have 8 rows per scan....  So depending on how you are generating your submission file you might be processing the \"test scans\" multiple times...  \n\nSo I'm currently doing something like...\n\ntest_scan_ids = pd.unique(df.StudyInstanceUID) \n\nto get a list of non-duped UIDs and looping thru that to drive the creation of the submission file... ",
    "1924672": "Hello, dear Kagglers,\nI have a doubt regarding the submission procedure for this competition and any suggestions could be really helpful.\nI have prepared the inference notebook, where I do not use the internet and I am loading a saved model which I trained on a different notebook. After submitting the notebook, every time I have an error of NOTEBOOK TIMEOUT and hence I believe I am not feeding in the test data in the most efficient manner. \nI will be really grateful if someone could kindly help me with a sample notebook that has passed the submission part. I am really struggling here. \nThanking you in anticipation.\nBest,\nArka ",
    "1924815": "NOTEBOOK TIMEOUT seems to indicate that your notebook ran more than the allowed 9 hours.  Is that consistent with what you see?  If so, realize that there are approximately 1500 patients in the real (hidden) test set.  You should convince yourself that your solution can handle that many patients in the allotted time.  You can test this out on some number of test/train patients.   If it can't, then you need to look at speeding things up.  Some ideas:\n\n1.  Make use of the 4 CPUs you are allowed for some pre-processing and/or pipelining.  You want to make sure your GPU is utilized as much as possible\n2. You did enable the GPU, right?\n3. If you are doing K-Fold averaging, you might not be able to run all K-models.\n4. Similarly, if you are doing an ensemble, maybe you can't use that many models\n4. Use a smaller network model, or use a smaller input image size\n"
  }
}