{
  "id": 429050,
  "title": "Decent Test Set Please",
  "url": "/competitions/rsna-2023-abdominal-trauma-detection/discussion/429050",
  "author_name": "PC Jimmmy",
  "post_date": "2023-08-03T22:48:05.587000",
  "votes": 4,
  "comment_count": 14,
  "views": 0,
  "content": "<p>I really dislike kaggle competitions that require use of a kaggle notebook to run a private hidden test set.  The reason for this distaste is the lack of debug information available (almost none) when errors occur.  The discussion boards of all past completions of this type are filled with folks asking for help because of submission failures.</p>\n<p>In a past discussion from a long ago post a Kaggle staff member indicated that they wanted to force us to learn to error trap - clearly a good idea.  Surely that does not mean we need to start with such a bare bones public test set.</p>\n<p>The public test set provided for this competition is almost useless.  3 patients with only a single slice for each does very little to let us identify issues, understand timing, etc.  </p>\n<p>Kaggle staff - is it really that much of an issue to create a small but reasonable test set so I don't end up with a long list of submissions that fail.</p>",
  "messages": [
    {
      "id": 2372747,
      "postDate": "2023-08-03T22:48:05.587Z",
      "content": "<p>I really dislike kaggle competitions that require use of a kaggle notebook to run a private hidden test set.  The reason for this distaste is the lack of debug information available (almost none) when errors occur.  The discussion boards of all past completions of this type are filled with folks asking for help because of submission failures.</p>\n<p>In a past discussion from a long ago post a Kaggle staff member indicated that they wanted to force us to learn to error trap - clearly a good idea.  Surely that does not mean we need to start with such a bare bones public test set.</p>\n<p>The public test set provided for this competition is almost useless.  3 patients with only a single slice for each does very little to let us identify issues, understand timing, etc.  </p>\n<p>Kaggle staff - is it really that much of an issue to create a small but reasonable test set so I don't end up with a long list of submissions that fail.</p>",
      "rawMarkdown": "I really dislike kaggle competitions that require use of a kaggle notebook to run a private hidden test set.  The reason for this distaste is the lack of debug information available (almost none) when errors occur.  The discussion boards of all past completions of this type are filled with folks asking for help because of submission failures.\n\nIn a past discussion from a long ago post a Kaggle staff member indicated that they wanted to force us to learn to error trap - clearly a good idea.  Surely that does not mean we need to start with such a bare bones public test set.\n\nThe public test set provided for this competition is almost useless.  3 patients with only a single slice for each does very little to let us identify issues, understand timing, etc.  \n\nKaggle staff - is it really that much of an issue to create a small but reasonable test set so I don't end up with a long list of submissions that fail.",
      "votes": 4
    },
    {
      "id": 2424887,
      "postDate": "2023-09-05T14:33:18.343Z",
      "content": "<p>it would be better if the dummy test data are some data from the training set.</p>\n<p>now it is strange that there is <strong>only one dicom</strong> file in the dummy series folder.</p>\n<p>This create problem for my code. e.g. i need to take the first and last slice to compute distance of the scan, my model uses 3-channel input …</p>\n<p>i need to create additional dummy slice images just to let my notebook pass without error so that i can submit.</p>",
      "rawMarkdown": "it would be better if the dummy test data are some data from the training set.\n\nnow it is strange that there is **only one dicom** file in the dummy series folder.\n\nThis create problem for my code. e.g. i need to take the first and last slice to compute distance of the scan, my model uses 3-channel input ...\n\ni need to create additional dummy slice images just to let my notebook pass without error so that i can submit.\n\n\n",
      "votes": 4,
      "replies": [
        {
          "id": 2425144,
          "postDate": "2023-09-05T17:19:54.910Z",
          "content": "<p>That's why I always create two inference codes for commit and submission.</p>\n<pre><code>is_submission = df_test.shape[] != \n\n is_submission:\n    \n:\n    \n</code></pre>",
          "rawMarkdown": "That's why I always create two inference codes for commit and submission.\n```python\nis_submission = df_test.shape[0] != 3\n\nif is_submission:\n    # real inference code\nelse:\n    # dummy inference code\n```",
          "votes": 6,
          "replies": [
            {
              "id": 2425294,
              "postDate": "2023-09-05T19:03:47.093Z",
              "content": "<p>that is a good idea</p>",
              "rawMarkdown": "that is a good idea"
            },
            {
              "id": 2425493,
              "postDate": "2023-09-06T01:26:47.263Z",
              "content": "<p>Nice and simple - I like it!!  Thanks</p>",
              "rawMarkdown": "Nice and simple - I like it!!  Thanks"
            },
            {
              "id": 2425524,
              "postDate": "2023-09-06T02:08:29.607Z",
              "content": "<p>here is another alternative<br>\n<a href=\"https://www.kaggle.com/code/jamesmcguigan/kaggle-environment-variables-os-environ\" target=\"_blank\">https://www.kaggle.com/code/jamesmcguigan/kaggle-environment-variables-os-environ</a></p>\n<pre><code>i think you can use KAGGLE_KERNEL_RUN_TYPE toif it is running in submit test servercurrent notebook\n</code></pre>",
              "rawMarkdown": "here is another alternative\nhttps://www.kaggle.com/code/jamesmcguigan/kaggle-environment-variables-os-environ\n\n```\ni think you can use KAGGLE_KERNEL_RUN_TYPE to check if it is running in submit test server or current notebook\n```"
            },
            {
              "id": 2425930,
              "postDate": "2023-09-06T09:37:06.263Z",
              "content": "<p>for rerun aka hidden test -<br>\n<code>if os.getenv('KAGGLE_IS_COMPETITION_RERUN'):</code></p>\n<p>also posted here - <br>\n<a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/433658\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/433658</a><br>\n\"Subset of RSNA full dataset</p>\n<p>Dataset: <a href=\"https://www.kaggle.com/datasets/cheonresearch/rsna-submission-test-data\" target=\"_blank\">https://www.kaggle.com/datasets/cheonresearch/rsna-submission-test-data</a></p>\n<p>I made small subset dataset of RSNA full dataset. (It is too large to download!)</p>\n<p>It contains 6 patients from train. (10005, 10007, 10004, 10249, 11927, 12299)<br>\nSome patients have two series of CT and the others have one series.<br>\n(I used it to check submission on inference notebook)</p>\n<p>I hope that this small dataset helpful to you. :)\"</p>",
              "rawMarkdown": "for rerun aka hidden test -\n`if os.getenv('KAGGLE_IS_COMPETITION_RERUN'):`\n\nalso posted here - \nhttps://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/433658\n\"Subset of RSNA full dataset\n\nDataset: https://www.kaggle.com/datasets/cheonresearch/rsna-submission-test-data\n\nI made small subset dataset of RSNA full dataset. (It is too large to download!)\n\n\nIt contains 6 patients from train. (10005, 10007, 10004, 10249, 11927, 12299)\nSome patients have two series of CT and the others have one series.\n(I used it to check submission on inference notebook)\n\n\nI hope that this small dataset helpful to you. :)\""
            }
          ]
        }
      ]
    },
    {
      "id": 2374260,
      "postDate": "2023-08-04T19:49:19.873Z",
      "content": "<p>Just to be clear we do not use hidden test sets to force people to learn about error handling, though that may be an unintended consequence. We most commonly use hidden test sets to maintain the integrity of the leaderboard against attacks like hand labeling. In other cases we may only have permission to make a subset of the data available for download, and so on.</p>",
      "rawMarkdown": "Just to be clear we do not use hidden test sets to force people to learn about error handling, though that may be an unintended consequence. We most commonly use hidden test sets to maintain the integrity of the leaderboard against attacks like hand labeling. In other cases we may only have permission to make a subset of the data available for download, and so on.",
      "votes": 1,
      "replies": [
        {
          "id": 2424188,
          "postDate": "2023-09-05T05:34:27.547Z",
          "content": "<p>I agree to this, but some times it is impossible to know why my scoring is failing , as there are no errors. some times we spend days figuring out why the scoring fails , even though the submitted format is exactly as required ( including datatypes)  and example scoring methods work perfectly fine on our outputs.</p>\n<p>Its like we spend more time getting the submission to work , that building a solution.</p>",
          "rawMarkdown": "I agree to this, but some times it is impossible to know why my scoring is failing , as there are no errors. some times we spend days figuring out why the scoring fails , even though the submitted format is exactly as required ( including datatypes)  and example scoring methods work perfectly fine on our outputs.\n\nIts like we spend more time getting the submission to work , that building a solution.",
          "isDeleted": true,
          "replies": [
            {
              "id": 2424839,
              "postDate": "2023-09-05T14:16:59.967Z",
              "content": "<p>Maybe I did not say it well enough in my initial post - I am not asking that a huge set of the \"public\" data that is used for scoring is given to us - I am just pointing out that a 3 patient test with only a single slice per patient is USELESS for most debugging beyond the very simple level.  </p>\n<p>From the posts in this competition it seems that folks have had submission errors might have issues related to patients who have multiple series, scans that are upside down, timing issues when using some libraries to read dicom files, etc.</p>\n<p>I ended up creating my own test set using 7 patients that had these issues - which I than put on kaggle as a dataset, and than have to add to my notebook and than have to code in so the notebook knows when to use my test vs the private test set, when I submit.  A whole lot of wasted (IMO) time and cyber space when the issue could have been handled by kaggle making 7 of the training patients the test set.</p>\n<p>No loss of integrity and no additional data that needs sharing permission.</p>\n<p>For a discipline that has probability as one of its foundations, why not take a simple additional step in competition prep to INCREASE the probability that new and old users will have submission success.</p>",
              "rawMarkdown": "Maybe I did not say it well enough in my initial post - I am not asking that a huge set of the \"public\" data that is used for scoring is given to us - I am just pointing out that a 3 patient test with only a single slice per patient is USELESS for most debugging beyond the very simple level.  \n\nFrom the posts in this competition it seems that folks have had submission errors might have issues related to patients who have multiple series, scans that are upside down, timing issues when using some libraries to read dicom files, etc.\n\nI ended up creating my own test set using 7 patients that had these issues - which I than put on kaggle as a dataset, and than have to add to my notebook and than have to code in so the notebook knows when to use my test vs the private test set, when I submit.  A whole lot of wasted (IMO) time and cyber space when the issue could have been handled by kaggle making 7 of the training patients the test set.\n\nNo loss of integrity and no additional data that needs sharing permission.\n\nFor a discipline that has probability as one of its foundations, why not take a simple additional step in competition prep to INCREASE the probability that new and old users will have submission success.\n",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2373996,
      "postDate": "2023-08-04T15:15:57.910Z",
      "content": "<p>I know how you feel, but debugging is possible.</p>\n<p>Just use as many training images as test images and you infer the data.<br>\nIf you want to save GPU quota, you can generally debug at half or even quarter scale.</p>",
      "rawMarkdown": "I know how you feel, but debugging is possible.\n\nJust use as many training images as test images and you infer the data.\nIf you want to save GPU quota, you can generally debug at half or even quarter scale.",
      "votes": 2,
      "replies": [
        {
          "id": 2374180,
          "postDate": "2023-08-04T18:38:57.353Z",
          "content": "<p>Yes - for sure it's a simple task to use training data as test images.  </p>\n<p>My point would be that why do 6000 kagglers (a current competition I am in that's ending soon) have to add/complicate thier submission code to evaluate it using train converted to test when kaggle staff (one person) could build a simple but realistic test set for all 6000.</p>",
          "rawMarkdown": "Yes - for sure it's a simple task to use training data as test images.  \n\nMy point would be that why do 6000 kagglers (a current competition I am in that's ending soon) have to add/complicate thier submission code to evaluate it using train converted to test when kaggle staff (one person) could build a simple but realistic test set for all 6000.",
          "votes": 3,
          "replies": [
            {
              "id": 2374549,
              "postDate": "2023-08-05T04:58:11.203Z",
              "content": "<p>I understand what you are saying and partially agree.</p>\n<p>But if you're troubleshooting on a submission, it's not that difficult, and I think there are rather a lot of disadvantages to making the data public.<br>\nFor example, leaderboards would become useless and LB probing and pseudo labelling would be mandatory in all competitions.</p>\n<p>Even if the rules prohibit it, there are plenty of loopholes and I think it would be better to only publish a small number of test sets.<br>\nOf course, in some past competitions pseudo labelling is OK and test data has been published in the past, but not all of them should have to be the same.</p>",
              "rawMarkdown": "I understand what you are saying and partially agree.\n\nBut if you're troubleshooting on a submission, it's not that difficult, and I think there are rather a lot of disadvantages to making the data public.\nFor example, leaderboards would become useless and LB probing and pseudo labelling would be mandatory in all competitions.\n\nEven if the rules prohibit it, there are plenty of loopholes and I think it would be better to only publish a small number of test sets.\nOf course, in some past competitions pseudo labelling is OK and test data has been published in the past, but not all of them should have to be the same.",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2373721,
      "postDate": "2023-08-04T12:09:04.717Z",
      "content": "<p>What benefit would it be to have more public test images?</p>",
      "rawMarkdown": "What benefit would it be to have more public test images?",
      "replies": [
        {
          "id": 2374177,
          "postDate": "2023-08-04T18:35:21.620Z",
          "content": "<p>For Kaggle masters - probably little benefit.  My guess - you create a larger test set yourself when working on this type of project with a bad test set.</p>\n<p>For newbies - you probably have the skill set to tally up all the discussion posts over the past year - my bet - 95% of the posts regarding submission failures come from folks very new to AI/kaggle.  Lots of thier issues (for sure mine) are pretty simple mistakes that could be spotted with a larger test set.  Is kaggle a place for new folks to learn or is it a place for kaggle masters to build their resumes?  Lots of posts on this topic over the past year have responses from kaggle that zero feedback on submission failures is needed because lots of talented kagglers seem to want to probe the leader board.  My experience - 99% of the 'rules' in life are there because a handful of folks cheat and this always comes with an expense to honest folks.</p>\n<p>A pretty decent number of new folks 'give up' when they cannot get a successful submission.  Hard data to capture since almost all will fork a successful shared script or two.  Poor test data is responsible for some of these!</p>\n<p>When I recently ran my code with the standard 'test' set I found that I could sum 'kidney_healthy, kidney_low and kidney high predictions and they equaled very close to 1.0 - which is what they should be.  When I spent time adding a 7 patient \"test\" set to my system I found that those add up to 1.46.   Still have not found the source of this error but having a small decent 7 patient test set let me identify a logic/code issue.  I would NEVER see that error with the provided test set and the zero feedback that comes from a submission against the private test.</p>\n<p>The current test set executes it's predictions very fast.  Having to read only a single slice for each of the three patients.  When a submission is made unless you take manual notes of start time and watch closely for the end time you get no feedback on how long your script runs.  Let's pretend that you want to use tta to improve your predictions.  With the current test set it's extremely difficult to estimate submission run time for tta runs even as large as 100 (a serious overkill).  With a simple 7 patient test set I can clearly make estimates of run time with tta values from 1 to 1000 by doing a large number of runs within the notebook and perhaps as few as 3 actual submissions.</p>\n<p>Similar issues regarding submission timing for number of folds, etc.  all benefit from a small but realistic test set.  Saving kagglers on failed submissions and/or submissions needed to understand run time.</p>\n<p>Memory use with a test set that consists of a total of 3 slices can easily lead to code that fails in submission when the real test for a patient contains 400 or more slices.  A prediction batch size of 10000 will almost certainly work with the current test set, but pretty sure its a big failure with the real private test.  I can get a decent estimate of batch size (memory use) with my realistic 7 patient test set.</p>\n<p>I do understand that new coders do not appreciate the value of error trapping.  Given a small but realistic test set there will be lots of new coders who fix the errors associated but still not error trap.  But lets give everyone a fighting chance to find the simple stupid errors in their submission code.</p>\n<p>I think I could keep going on this list for some time - I am pretty sure that you also could build a nice list of reasons why realistic test data SHOULD be a part of every AI project in both the real world and kaggle.  </p>\n<p>I can see you now taking a project proposal to your boss - you tell him you have a 3,147 patient training set with dicom slices ranging from 1 to 400 for a total of 1,508,511 slices, and your going to run your prediction code against a test set with a total of 3 slices.  I think I can create an AI model that predicts your next actions - your reading the help wanted ads and touching up your CV for monster.com.</p>",
          "rawMarkdown": "For Kaggle masters - probably little benefit.  My guess - you create a larger test set yourself when working on this type of project with a bad test set.\n\nFor newbies - you probably have the skill set to tally up all the discussion posts over the past year - my bet - 95% of the posts regarding submission failures come from folks very new to AI/kaggle.  Lots of thier issues (for sure mine) are pretty simple mistakes that could be spotted with a larger test set.  Is kaggle a place for new folks to learn or is it a place for kaggle masters to build their resumes?  Lots of posts on this topic over the past year have responses from kaggle that zero feedback on submission failures is needed because lots of talented kagglers seem to want to probe the leader board.  My experience - 99% of the 'rules' in life are there because a handful of folks cheat and this always comes with an expense to honest folks.\n\nA pretty decent number of new folks 'give up' when they cannot get a successful submission.  Hard data to capture since almost all will fork a successful shared script or two.  Poor test data is responsible for some of these!\n\nWhen I recently ran my code with the standard 'test' set I found that I could sum 'kidney_healthy, kidney_low and kidney high predictions and they equaled very close to 1.0 - which is what they should be.  When I spent time adding a 7 patient \"test\" set to my system I found that those add up to 1.46.   Still have not found the source of this error but having a small decent 7 patient test set let me identify a logic/code issue.  I would NEVER see that error with the provided test set and the zero feedback that comes from a submission against the private test.\n\nThe current test set executes it's predictions very fast.  Having to read only a single slice for each of the three patients.  When a submission is made unless you take manual notes of start time and watch closely for the end time you get no feedback on how long your script runs.  Let's pretend that you want to use tta to improve your predictions.  With the current test set it's extremely difficult to estimate submission run time for tta runs even as large as 100 (a serious overkill).  With a simple 7 patient test set I can clearly make estimates of run time with tta values from 1 to 1000 by doing a large number of runs within the notebook and perhaps as few as 3 actual submissions.\n\nSimilar issues regarding submission timing for number of folds, etc.  all benefit from a small but realistic test set.  Saving kagglers on failed submissions and/or submissions needed to understand run time.\n\nMemory use with a test set that consists of a total of 3 slices can easily lead to code that fails in submission when the real test for a patient contains 400 or more slices.  A prediction batch size of 10000 will almost certainly work with the current test set, but pretty sure its a big failure with the real private test.  I can get a decent estimate of batch size (memory use) with my realistic 7 patient test set.\n\nI do understand that new coders do not appreciate the value of error trapping.  Given a small but realistic test set there will be lots of new coders who fix the errors associated but still not error trap.  But lets give everyone a fighting chance to find the simple stupid errors in their submission code.\n\nI think I could keep going on this list for some time - I am pretty sure that you also could build a nice list of reasons why realistic test data SHOULD be a part of every AI project in both the real world and kaggle.  \n\nI can see you now taking a project proposal to your boss - you tell him you have a 3,147 patient training set with dicom slices ranging from 1 to 400 for a total of 1,508,511 slices, and your going to run your prediction code against a test set with a total of 3 slices.  I think I can create an AI model that predicts your next actions - your reading the help wanted ads and touching up your CV for monster.com.",
          "votes": -3
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2424887,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-09-05T14:33:18.343000",
      "content": "<p>it would be better if the dummy test data are some data from the training set.</p>\n<p>now it is strange that there is <strong>only one dicom</strong> file in the dummy series folder.</p>\n<p>This create problem for my code. e.g. i need to take the first and last slice to compute distance of the scan, my model uses 3-channel input …</p>\n<p>i need to create additional dummy slice images just to let my notebook pass without error so that i can submit.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2425144,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2023-09-05T17:19:54.910000",
          "content": "<p>That's why I always create two inference codes for commit and submission.</p>\n<pre><code>is_submission = df_test.shape[] != \n\n is_submission:\n    \n:\n    \n</code></pre>",
          "votes": 6,
          "replies": [
            {
              "id": 2425294,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-09-05T19:03:47.093000",
              "content": "<p>that is a good idea</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2425493,
              "author_name": "PC Jimmmy",
              "author_url": "",
              "post_date": "2023-09-06T01:26:47.263000",
              "content": "<p>Nice and simple - I like it!!  Thanks</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2425524,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-09-06T02:08:29.607000",
              "content": "<p>here is another alternative<br>\n<a href=\"https://www.kaggle.com/code/jamesmcguigan/kaggle-environment-variables-os-environ\" target=\"_blank\">https://www.kaggle.com/code/jamesmcguigan/kaggle-environment-variables-os-environ</a></p>\n<pre><code>i think you can use KAGGLE_KERNEL_RUN_TYPE toif it is running in submit test servercurrent notebook\n</code></pre>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2425930,
              "author_name": "something4kag",
              "author_url": "",
              "post_date": "2023-09-06T09:37:06.263000",
              "content": "<p>for rerun aka hidden test -<br>\n<code>if os.getenv('KAGGLE_IS_COMPETITION_RERUN'):</code></p>\n<p>also posted here - <br>\n<a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/433658\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/433658</a><br>\n\"Subset of RSNA full dataset</p>\n<p>Dataset: <a href=\"https://www.kaggle.com/datasets/cheonresearch/rsna-submission-test-data\" target=\"_blank\">https://www.kaggle.com/datasets/cheonresearch/rsna-submission-test-data</a></p>\n<p>I made small subset dataset of RSNA full dataset. (It is too large to download!)</p>\n<p>It contains 6 patients from train. (10005, 10007, 10004, 10249, 11927, 12299)<br>\nSome patients have two series of CT and the others have one series.<br>\n(I used it to check submission on inference notebook)</p>\n<p>I hope that this small dataset helpful to you. :)\"</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2374260,
      "author_name": "Sohier Dane",
      "author_url": "",
      "post_date": "2023-08-04T19:49:19.873000",
      "content": "<p>Just to be clear we do not use hidden test sets to force people to learn about error handling, though that may be an unintended consequence. We most commonly use hidden test sets to maintain the integrity of the leaderboard against attacks like hand labeling. In other cases we may only have permission to make a subset of the data available for download, and so on.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2424188,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-09-05T05:34:27.547000",
          "content": "<p>I agree to this, but some times it is impossible to know why my scoring is failing , as there are no errors. some times we spend days figuring out why the scoring fails , even though the submitted format is exactly as required ( including datatypes)  and example scoring methods work perfectly fine on our outputs.</p>\n<p>Its like we spend more time getting the submission to work , that building a solution.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2424839,
              "author_name": "PC Jimmmy",
              "author_url": "",
              "post_date": "2023-09-05T14:16:59.967000",
              "content": "<p>Maybe I did not say it well enough in my initial post - I am not asking that a huge set of the \"public\" data that is used for scoring is given to us - I am just pointing out that a 3 patient test with only a single slice per patient is USELESS for most debugging beyond the very simple level.  </p>\n<p>From the posts in this competition it seems that folks have had submission errors might have issues related to patients who have multiple series, scans that are upside down, timing issues when using some libraries to read dicom files, etc.</p>\n<p>I ended up creating my own test set using 7 patients that had these issues - which I than put on kaggle as a dataset, and than have to add to my notebook and than have to code in so the notebook knows when to use my test vs the private test set, when I submit.  A whole lot of wasted (IMO) time and cyber space when the issue could have been handled by kaggle making 7 of the training patients the test set.</p>\n<p>No loss of integrity and no additional data that needs sharing permission.</p>\n<p>For a discipline that has probability as one of its foundations, why not take a simple additional step in competition prep to INCREASE the probability that new and old users will have submission success.</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2373996,
      "author_name": "YYama",
      "author_url": "",
      "post_date": "2023-08-04T15:15:57.910000",
      "content": "<p>I know how you feel, but debugging is possible.</p>\n<p>Just use as many training images as test images and you infer the data.<br>\nIf you want to save GPU quota, you can generally debug at half or even quarter scale.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2374180,
          "author_name": "PC Jimmmy",
          "author_url": "",
          "post_date": "2023-08-04T18:38:57.353000",
          "content": "<p>Yes - for sure it's a simple task to use training data as test images.  </p>\n<p>My point would be that why do 6000 kagglers (a current competition I am in that's ending soon) have to add/complicate thier submission code to evaluate it using train converted to test when kaggle staff (one person) could build a simple but realistic test set for all 6000.</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2374549,
              "author_name": "YYama",
              "author_url": "",
              "post_date": "2023-08-05T04:58:11.203000",
              "content": "<p>I understand what you are saying and partially agree.</p>\n<p>But if you're troubleshooting on a submission, it's not that difficult, and I think there are rather a lot of disadvantages to making the data public.<br>\nFor example, leaderboards would become useless and LB probing and pseudo labelling would be mandatory in all competitions.</p>\n<p>Even if the rules prohibit it, there are plenty of loopholes and I think it would be better to only publish a small number of test sets.<br>\nOf course, in some past competitions pseudo labelling is OK and test data has been published in the past, but not all of them should have to be the same.</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2373721,
      "author_name": "David Roberts",
      "author_url": "",
      "post_date": "2023-08-04T12:09:04.717000",
      "content": "<p>What benefit would it be to have more public test images?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2374177,
          "author_name": "PC Jimmmy",
          "author_url": "",
          "post_date": "2023-08-04T18:35:21.620000",
          "content": "<p>For Kaggle masters - probably little benefit.  My guess - you create a larger test set yourself when working on this type of project with a bad test set.</p>\n<p>For newbies - you probably have the skill set to tally up all the discussion posts over the past year - my bet - 95% of the posts regarding submission failures come from folks very new to AI/kaggle.  Lots of thier issues (for sure mine) are pretty simple mistakes that could be spotted with a larger test set.  Is kaggle a place for new folks to learn or is it a place for kaggle masters to build their resumes?  Lots of posts on this topic over the past year have responses from kaggle that zero feedback on submission failures is needed because lots of talented kagglers seem to want to probe the leader board.  My experience - 99% of the 'rules' in life are there because a handful of folks cheat and this always comes with an expense to honest folks.</p>\n<p>A pretty decent number of new folks 'give up' when they cannot get a successful submission.  Hard data to capture since almost all will fork a successful shared script or two.  Poor test data is responsible for some of these!</p>\n<p>When I recently ran my code with the standard 'test' set I found that I could sum 'kidney_healthy, kidney_low and kidney high predictions and they equaled very close to 1.0 - which is what they should be.  When I spent time adding a 7 patient \"test\" set to my system I found that those add up to 1.46.   Still have not found the source of this error but having a small decent 7 patient test set let me identify a logic/code issue.  I would NEVER see that error with the provided test set and the zero feedback that comes from a submission against the private test.</p>\n<p>The current test set executes it's predictions very fast.  Having to read only a single slice for each of the three patients.  When a submission is made unless you take manual notes of start time and watch closely for the end time you get no feedback on how long your script runs.  Let's pretend that you want to use tta to improve your predictions.  With the current test set it's extremely difficult to estimate submission run time for tta runs even as large as 100 (a serious overkill).  With a simple 7 patient test set I can clearly make estimates of run time with tta values from 1 to 1000 by doing a large number of runs within the notebook and perhaps as few as 3 actual submissions.</p>\n<p>Similar issues regarding submission timing for number of folds, etc.  all benefit from a small but realistic test set.  Saving kagglers on failed submissions and/or submissions needed to understand run time.</p>\n<p>Memory use with a test set that consists of a total of 3 slices can easily lead to code that fails in submission when the real test for a patient contains 400 or more slices.  A prediction batch size of 10000 will almost certainly work with the current test set, but pretty sure its a big failure with the real private test.  I can get a decent estimate of batch size (memory use) with my realistic 7 patient test set.</p>\n<p>I do understand that new coders do not appreciate the value of error trapping.  Given a small but realistic test set there will be lots of new coders who fix the errors associated but still not error trap.  But lets give everyone a fighting chance to find the simple stupid errors in their submission code.</p>\n<p>I think I could keep going on this list for some time - I am pretty sure that you also could build a nice list of reasons why realistic test data SHOULD be a part of every AI project in both the real world and kaggle.  </p>\n<p>I can see you now taking a project proposal to your boss - you tell him you have a 3,147 patient training set with dicom slices ranging from 1 to 400 for a total of 1,508,511 slices, and your going to run your prediction code against a test set with a total of 3 slices.  I think I can create an AI model that predicts your next actions - your reading the help wanted ads and touching up your CV for monster.com.</p>",
          "votes": -3,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2372747": "I really dislike kaggle competitions that require use of a kaggle notebook to run a private hidden test set.  The reason for this distaste is the lack of debug information available (almost none) when errors occur.  The discussion boards of all past completions of this type are filled with folks asking for help because of submission failures.\n\nIn a past discussion from a long ago post a Kaggle staff member indicated that they wanted to force us to learn to error trap - clearly a good idea.  Surely that does not mean we need to start with such a bare bones public test set.\n\nThe public test set provided for this competition is almost useless.  3 patients with only a single slice for each does very little to let us identify issues, understand timing, etc.  \n\nKaggle staff - is it really that much of an issue to create a small but reasonable test set so I don't end up with a long list of submissions that fail.",
    "2424887": "it would be better if the dummy test data are some data from the training set.\n\nnow it is strange that there is **only one dicom** file in the dummy series folder.\n\nThis create problem for my code. e.g. i need to take the first and last slice to compute distance of the scan, my model uses 3-channel input ...\n\ni need to create additional dummy slice images just to let my notebook pass without error so that i can submit.\n\n\n",
    "2374260": "Just to be clear we do not use hidden test sets to force people to learn about error handling, though that may be an unintended consequence. We most commonly use hidden test sets to maintain the integrity of the leaderboard against attacks like hand labeling. In other cases we may only have permission to make a subset of the data available for download, and so on.",
    "2373996": "I know how you feel, but debugging is possible.\n\nJust use as many training images as test images and you infer the data.\nIf you want to save GPU quota, you can generally debug at half or even quarter scale.",
    "2373721": "What benefit would it be to have more public test images?"
  }
}