{
  "id": 388859,
  "title": "Hidden test dataset/images not read.",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/388859",
  "author_name": "Fabio Souza",
  "post_date": "2023-02-19T20:36:44.733000",
  "votes": 0,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Hi. I am submitting my code through the \"Submit\" button in the \"Submit to competition\" section in the right sidebar of the code. But, during the processing, the code reads just four samples, not the complete hidden test set. So, please, what am I doing wrong? I appreciate any help.</p>",
  "messages": [
    {
      "id": 2155415,
      "postDate": "2023-02-22T16:11:20.717Z",
      "content": "<pre><code>csv_path = '/kaggle/input/rsna-breast-cancer-detection/test.csv'\ndata_path  = '/kaggle/input/rsna-breast-cancer-detection/test_images'\n\nsave_imgs_path = '/kaggle/tmp/~png'\nos.makedirs(save_imgs_path, exist_ok=True)\n</code></pre>\n<p>save your imgs under there and load when needed.</p>",
      "rawMarkdown": "```\ncsv_path = '/kaggle/input/rsna-breast-cancer-detection/test.csv'\ndata_path  = '/kaggle/input/rsna-breast-cancer-detection/test_images'\n\nsave_imgs_path = '/kaggle/tmp/~png'\nos.makedirs(save_imgs_path, exist_ok=True)\n```\n\nsave your imgs under there and load when needed."
    },
    {
      "id": 2151177,
      "postDate": "2023-02-19T22:16:31.910Z",
      "content": "<blockquote>\n  <p>But, during the processing, the code reads just four samples, not the complete hidden test set.</p>\n</blockquote>\n<p>How do you know that? When editing your notebook, you are only able to access 4 samples of the hidden test set. Did this solve your problem?</p>",
      "rawMarkdown": ">But, during the processing, the code reads just four samples, not the complete hidden test set.\n\nHow do you know that? When editing your notebook, you are only able to access 4 samples of the hidden test set. Did this solve your problem?",
      "replies": [
        {
          "id": 2151254,
          "postDate": "2023-02-20T00:45:00.213Z",
          "content": "<p>My notebook fails when running for the submission. I was not expecting to be able to see the log, but I was able. After reading, a print about the test.csv shape shows just (4, 9). I wonder if I am submitting in the right way, as I am being able to see the log.</p>",
          "rawMarkdown": "My notebook fails when running for the submission. I was not expecting to be able to see the log, but I was able. After reading, a print about the test.csv shape shows just (4, 9). I wonder if I am submitting in the right way, as I am being able to see the log.",
          "replies": [
            {
              "id": 2151265,
              "postDate": "2023-02-20T00:56:18.387Z",
              "content": "<p>Perhaps you are sending too many rows? See here <a href=\"https://www.kaggle.com/code/hengck23/zero-submission-for-debug\" target=\"_blank\">https://www.kaggle.com/code/hengck23/zero-submission-for-debug</a> how to combine the predictions for MLO and CC and format your submission.</p>",
              "rawMarkdown": "Perhaps you are sending too many rows? See here https://www.kaggle.com/code/hengck23/zero-submission-for-debug how to combine the predictions for MLO and CC and format your submission."
            },
            {
              "id": 2151314,
              "postDate": "2023-02-20T02:49:38.487Z",
              "content": "<p>A good point to consider before saving the submission file. Thanks. But my problem occurs when reading the test file before predictions. We expect to have data about around 8k patients, but my code reads just few lines in the test.csv when submitting to the competition.</p>",
              "rawMarkdown": "A good point to consider before saving the submission file. Thanks. But my problem occurs when reading the test file before predictions. We expect to have data about around 8k patients, but my code reads just few lines in the test.csv when submitting to the competition."
            },
            {
              "id": 2152152,
              "postDate": "2023-02-20T15:49:01.673Z",
              "content": "<p>But it reads more from the train.csv? If so, it should be fine. In this competition, the provided test samples are available only to see that your scripts run.</p>",
              "rawMarkdown": "But it reads more from the train.csv? If so, it should be fine. In this competition, the provided test samples are available only to see that your scripts run."
            },
            {
              "id": 2152397,
              "postDate": "2023-02-20T19:20:17.940Z",
              "content": "<p>That is the point. It reads four lines, the same amount when I run not for submission.</p>",
              "rawMarkdown": "That is the point. It reads four lines, the same amount when I run not for submission."
            },
            {
              "id": 2152434,
              "postDate": "2023-02-20T19:49:43.273Z",
              "content": "<p>Getting only four lines at least rules out an accidental df.head() command, which returns 5 lines. 🤔</p>",
              "rawMarkdown": "Getting only four lines at least rules out an accidental df.head() command, which returns 5 lines. 🤔"
            },
            {
              "id": 2155055,
              "postDate": "2023-02-22T12:08:03.800Z",
              "content": "<p>Can you try these:</p>\n<pre><code>train_csv_file = Path()\ntrain_meta = pd.read_csv(train_csv_file)\n((train_meta))\n</code></pre>\n<pre><code>test_csv_file = Path()\ntest_meta = pd.read_csv(test_csv_file)\n((test_meta))\n</code></pre>",
              "rawMarkdown": "Can you try these:\n\n```python\ntrain_csv_file = Path('/kaggle/input/rsna-breast-cancer-detection/train.csv')\ntrain_meta = pd.read_csv(train_csv_file)\nprint(len(train_meta))\n```\n\n```python\ntest_csv_file = Path('/kaggle/input/rsna-breast-cancer-detection/test.csv')\ntest_meta = pd.read_csv(test_csv_file)\nprint(len(test_meta))\n```"
            },
            {
              "id": 2155331,
              "postDate": "2023-02-22T15:16:58.303Z",
              "content": "<p>Yes, I am working with two paths, one for the train and the other for the test. Also, I am making predictions on a holdout set when not running for submission for accuracy evaluation purposes. Still, my notebook can't read the full hidden test set in the submission environment, just the four samples. <br>\nIt seems to be, for me at this moment, something about the environment itself or the way I am submitting, although I can see nothing wrong with the way I am doing that 🙁<br>\nBut it is not the end of the world 🙂 It was a good exercise anyway—many thanks for being a valuable contributor giving me tips on my issue. Of course, I am always open to discussing that.</p>",
              "rawMarkdown": "Yes, I am working with two paths, one for the train and the other for the test. Also, I am making predictions on a holdout set when not running for submission for accuracy evaluation purposes. Still, my notebook can't read the full hidden test set in the submission environment, just the four samples. \nIt seems to be, for me at this moment, something about the environment itself or the way I am submitting, although I can see nothing wrong with the way I am doing that 🙁\nBut it is not the end of the world 🙂 It was a good exercise anyway—many thanks for being a valuable contributor giving me tips on my issue. Of course, I am always open to discussing that."
            },
            {
              "id": 2155370,
              "postDate": "2023-02-22T15:49:59.440Z",
              "content": "<p>So, you have made a Kaggle datasets of all the data and the libraries you need to run your data offline? And tested this with the training data with internet off?</p>\n<p>And your training script can process each and every training image (reading/decoding different kind of DICOM etc.)? Check this out: <a href=\"https://www.kaggle.com/code/anttiisosalo/16-bit-post-processing-rsna-bc-detection\" target=\"_blank\">https://www.kaggle.com/code/anttiisosalo/16-bit-post-processing-rsna-bc-detection</a> .</p>\n<p>And you tried the zero submission script from above and that does not work either?</p>\n<p>It's true that it is not the end of the world, but it would be good to have a positive experience and get a successful submission. I am also struggling to get something other than random numbers from my model and would like to do that before the competition ends. 😀</p>",
              "rawMarkdown": "So, you have made a Kaggle datasets of all the data and the libraries you need to run your data offline? And tested this with the training data with internet off?\n\nAnd your training script can process each and every training image (reading/decoding different kind of DICOM etc.)? Check this out: https://www.kaggle.com/code/anttiisosalo/16-bit-post-processing-rsna-bc-detection .\n\nAnd you tried the zero submission script from above and that does not work either?\n\nIt's true that it is not the end of the world, but it would be good to have a positive experience and get a successful submission. I am also struggling to get something other than random numbers from my model and would like to do that before the competition ends. 😀"
            },
            {
              "id": 2155651,
              "postDate": "2023-02-22T19:29:20.283Z",
              "content": "<p>\"So, you have made Kaggle datasets of all the libraries you need to run your data offline?\"<br>\n<strong>&gt;&gt;&gt;</strong> I am simply running in the Kaggle environment (but not in the submission mode). In this scenario, I separated a subset of the train set to be used to evaluate the accuracy (holdout). Then, when submitting, I change the code to predict on the test set instead of the holdout set.</p>\n<p>\"And your training script can process each and every training image (reading/decoding different kind of DICOM etc.)?\"<br>\n<strong>&gt;&gt;&gt;</strong> Yes, in the train set and holdout set.</p>\n<p>\"And you tried the zero submission script from above and that does not work either?\"<br>\n<strong>&gt;&gt;&gt;</strong> Yes, more specifically, a version of that.</p>\n<p>\"I am also struggling to get something other than random numbers from my model. I would like to do that before the competition ends\"<br>\n<strong>&gt;&gt;&gt;</strong> The way I worked on that: I Reduced the dimensionality of the images so that they could be trained in the available memory, trying to loss as less as possible important information (checked throughout the accuracy measurement with the holdout set). Another point concerns the architecture of your neural network. In a CNN architecture, tweaking the parameters and quantity of the convolutional layers, pooling layers, and fully connected layers can be done through experiments (varying parameters) to find a reasonable configuration.</p>",
              "rawMarkdown": "\"So, you have made Kaggle datasets of all the libraries you need to run your data offline?\"\n**>>>** I am simply running in the Kaggle environment (but not in the submission mode). In this scenario, I separated a subset of the train set to be used to evaluate the accuracy (holdout). Then, when submitting, I change the code to predict on the test set instead of the holdout set.\n\n\"And your training script can process each and every training image (reading/decoding different kind of DICOM etc.)?\"\n**>>>** Yes, in the train set and holdout set.\n\n\"And you tried the zero submission script from above and that does not work either?\"\n**>>>** Yes, more specifically, a version of that.\n\n\"I am also struggling to get something other than random numbers from my model. I would like to do that before the competition ends\"\n**>>>** The way I worked on that: I Reduced the dimensionality of the images so that they could be trained in the available memory, trying to loss as less as possible important information (checked throughout the accuracy measurement with the holdout set). Another point concerns the architecture of your neural network. In a CNN architecture, tweaking the parameters and quantity of the convolutional layers, pooling layers, and fully connected layers can be done through experiments (varying parameters) to find a reasonable configuration.",
              "votes": 1
            },
            {
              "id": 2156876,
              "postDate": "2023-02-23T15:35:29.207Z",
              "content": "<p>Perhaps it would be good to try also submitting through the competition page. First save new version using Save and Run all (Commit). And then from the competition page Submit Predictions (by choosing the particular notebook you want to try out). Good luck! 👍</p>",
              "rawMarkdown": "Perhaps it would be good to try also submitting through the competition page. First save new version using Save and Run all (Commit). And then from the competition page Submit Predictions (by choosing the particular notebook you want to try out). Good luck! 👍"
            },
            {
              "id": 2160634,
              "postDate": "2023-02-26T20:12:23.657Z",
              "content": "<p>Yes. Good tip. I am trying that. I received an error saying that the submission file had an incorrect format; in other words, the notebook ran completely, but there was some problem with the output file. I already checked the fields being generated and the type of fields in the code, which are okay. Now, I am reviewing the possibility of some predictions not being included in the output, which is a little bit hard to do without precisely seeing the point in a log. Unfortunately, I have been busy lately to dedicate enough to this project, and now the time is finishing, but I will try yet.</p>",
              "rawMarkdown": "Yes. Good tip. I am trying that. I received an error saying that the submission file had an incorrect format; in other words, the notebook ran completely, but there was some problem with the output file. I already checked the fields being generated and the type of fields in the code, which are okay. Now, I am reviewing the possibility of some predictions not being included in the output, which is a little bit hard to do without precisely seeing the point in a log. Unfortunately, I have been busy lately to dedicate enough to this project, and now the time is finishing, but I will try yet."
            }
          ]
        }
      ]
    },
    {
      "id": 2151109,
      "postDate": "2023-02-19T20:36:44.733Z",
      "content": "<p>Hi. I am submitting my code through the \"Submit\" button in the \"Submit to competition\" section in the right sidebar of the code. But, during the processing, the code reads just four samples, not the complete hidden test set. So, please, what am I doing wrong? I appreciate any help.</p>",
      "rawMarkdown": "Hi. I am submitting my code through the \"Submit\" button in the \"Submit to competition\" section in the right sidebar of the code. But, during the processing, the code reads just four samples, not the complete hidden test set. So, please, what am I doing wrong? I appreciate any help."
    }
  ],
  "comments": [
    {
      "id": 2155415,
      "author_name": "Eleftherios Fanioudakis",
      "author_url": "",
      "post_date": "2023-02-22T16:11:20.717000",
      "content": "<pre><code>csv_path = '/kaggle/input/rsna-breast-cancer-detection/test.csv'\ndata_path  = '/kaggle/input/rsna-breast-cancer-detection/test_images'\n\nsave_imgs_path = '/kaggle/tmp/~png'\nos.makedirs(save_imgs_path, exist_ok=True)\n</code></pre>\n<p>save your imgs under there and load when needed.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2151177,
      "author_name": "Antti Isosalo",
      "author_url": "",
      "post_date": "2023-02-19T22:16:31.910000",
      "content": "<blockquote>\n  <p>But, during the processing, the code reads just four samples, not the complete hidden test set.</p>\n</blockquote>\n<p>How do you know that? When editing your notebook, you are only able to access 4 samples of the hidden test set. Did this solve your problem?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2151254,
          "author_name": "Fabio Souza",
          "author_url": "",
          "post_date": "2023-02-20T00:45:00.213000",
          "content": "<p>My notebook fails when running for the submission. I was not expecting to be able to see the log, but I was able. After reading, a print about the test.csv shape shows just (4, 9). I wonder if I am submitting in the right way, as I am being able to see the log.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2151265,
              "author_name": "Antti Isosalo",
              "author_url": "",
              "post_date": "2023-02-20T00:56:18.387000",
              "content": "<p>Perhaps you are sending too many rows? See here <a href=\"https://www.kaggle.com/code/hengck23/zero-submission-for-debug\" target=\"_blank\">https://www.kaggle.com/code/hengck23/zero-submission-for-debug</a> how to combine the predictions for MLO and CC and format your submission.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2151314,
              "author_name": "Fabio Souza",
              "author_url": "",
              "post_date": "2023-02-20T02:49:38.487000",
              "content": "<p>A good point to consider before saving the submission file. Thanks. But my problem occurs when reading the test file before predictions. We expect to have data about around 8k patients, but my code reads just few lines in the test.csv when submitting to the competition.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2152152,
              "author_name": "Antti Isosalo",
              "author_url": "",
              "post_date": "2023-02-20T15:49:01.673000",
              "content": "<p>But it reads more from the train.csv? If so, it should be fine. In this competition, the provided test samples are available only to see that your scripts run.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2152397,
              "author_name": "Fabio Souza",
              "author_url": "",
              "post_date": "2023-02-20T19:20:17.940000",
              "content": "<p>That is the point. It reads four lines, the same amount when I run not for submission.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2152434,
              "author_name": "Antti Isosalo",
              "author_url": "",
              "post_date": "2023-02-20T19:49:43.273000",
              "content": "<p>Getting only four lines at least rules out an accidental df.head() command, which returns 5 lines. 🤔</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2155055,
              "author_name": "Antti Isosalo",
              "author_url": "",
              "post_date": "2023-02-22T12:08:03.800000",
              "content": "<p>Can you try these:</p>\n<pre><code>train_csv_file = Path()\ntrain_meta = pd.read_csv(train_csv_file)\n((train_meta))\n</code></pre>\n<pre><code>test_csv_file = Path()\ntest_meta = pd.read_csv(test_csv_file)\n((test_meta))\n</code></pre>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2155331,
              "author_name": "Fabio Souza",
              "author_url": "",
              "post_date": "2023-02-22T15:16:58.303000",
              "content": "<p>Yes, I am working with two paths, one for the train and the other for the test. Also, I am making predictions on a holdout set when not running for submission for accuracy evaluation purposes. Still, my notebook can't read the full hidden test set in the submission environment, just the four samples. <br>\nIt seems to be, for me at this moment, something about the environment itself or the way I am submitting, although I can see nothing wrong with the way I am doing that 🙁<br>\nBut it is not the end of the world 🙂 It was a good exercise anyway—many thanks for being a valuable contributor giving me tips on my issue. Of course, I am always open to discussing that.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2155370,
              "author_name": "Antti Isosalo",
              "author_url": "",
              "post_date": "2023-02-22T15:49:59.440000",
              "content": "<p>So, you have made a Kaggle datasets of all the data and the libraries you need to run your data offline? And tested this with the training data with internet off?</p>\n<p>And your training script can process each and every training image (reading/decoding different kind of DICOM etc.)? Check this out: <a href=\"https://www.kaggle.com/code/anttiisosalo/16-bit-post-processing-rsna-bc-detection\" target=\"_blank\">https://www.kaggle.com/code/anttiisosalo/16-bit-post-processing-rsna-bc-detection</a> .</p>\n<p>And you tried the zero submission script from above and that does not work either?</p>\n<p>It's true that it is not the end of the world, but it would be good to have a positive experience and get a successful submission. I am also struggling to get something other than random numbers from my model and would like to do that before the competition ends. 😀</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2155651,
              "author_name": "Fabio Souza",
              "author_url": "",
              "post_date": "2023-02-22T19:29:20.283000",
              "content": "<p>\"So, you have made Kaggle datasets of all the libraries you need to run your data offline?\"<br>\n<strong>&gt;&gt;&gt;</strong> I am simply running in the Kaggle environment (but not in the submission mode). In this scenario, I separated a subset of the train set to be used to evaluate the accuracy (holdout). Then, when submitting, I change the code to predict on the test set instead of the holdout set.</p>\n<p>\"And your training script can process each and every training image (reading/decoding different kind of DICOM etc.)?\"<br>\n<strong>&gt;&gt;&gt;</strong> Yes, in the train set and holdout set.</p>\n<p>\"And you tried the zero submission script from above and that does not work either?\"<br>\n<strong>&gt;&gt;&gt;</strong> Yes, more specifically, a version of that.</p>\n<p>\"I am also struggling to get something other than random numbers from my model. I would like to do that before the competition ends\"<br>\n<strong>&gt;&gt;&gt;</strong> The way I worked on that: I Reduced the dimensionality of the images so that they could be trained in the available memory, trying to loss as less as possible important information (checked throughout the accuracy measurement with the holdout set). Another point concerns the architecture of your neural network. In a CNN architecture, tweaking the parameters and quantity of the convolutional layers, pooling layers, and fully connected layers can be done through experiments (varying parameters) to find a reasonable configuration.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2156876,
              "author_name": "Antti Isosalo",
              "author_url": "",
              "post_date": "2023-02-23T15:35:29.207000",
              "content": "<p>Perhaps it would be good to try also submitting through the competition page. First save new version using Save and Run all (Commit). And then from the competition page Submit Predictions (by choosing the particular notebook you want to try out). Good luck! 👍</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2160634,
              "author_name": "Fabio Souza",
              "author_url": "",
              "post_date": "2023-02-26T20:12:23.657000",
              "content": "<p>Yes. Good tip. I am trying that. I received an error saying that the submission file had an incorrect format; in other words, the notebook ran completely, but there was some problem with the output file. I already checked the fields being generated and the type of fields in the code, which are okay. Now, I am reviewing the possibility of some predictions not being included in the output, which is a little bit hard to do without precisely seeing the point in a log. Unfortunately, I have been busy lately to dedicate enough to this project, and now the time is finishing, but I will try yet.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2155415": "```\ncsv_path = '/kaggle/input/rsna-breast-cancer-detection/test.csv'\ndata_path  = '/kaggle/input/rsna-breast-cancer-detection/test_images'\n\nsave_imgs_path = '/kaggle/tmp/~png'\nos.makedirs(save_imgs_path, exist_ok=True)\n```\n\nsave your imgs under there and load when needed.",
    "2151177": ">But, during the processing, the code reads just four samples, not the complete hidden test set.\n\nHow do you know that? When editing your notebook, you are only able to access 4 samples of the hidden test set. Did this solve your problem?",
    "2151109": "Hi. I am submitting my code through the \"Submit\" button in the \"Submit to competition\" section in the right sidebar of the code. But, during the processing, the code reads just four samples, not the complete hidden test set. So, please, what am I doing wrong? I appreciate any help."
  }
}