{
  "id": 155959,
  "title": "How do we reach the test data?",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/155959",
  "author_name": "Tolga",
  "post_date": "2020-06-03T18:30:50.025000",
  "votes": 2,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Seriously, how do we read the test data? I submitted a submission.csv file with random numbers but still don't know how to see the test data. Is it visible at all? Any help is appreciated.</p>\n\n<p>P.S. My model is dying to make some predictions on the test data.</p>",
  "messages": [
    {
      "id": 873735,
      "postDate": "2020-06-04T11:46:37.993Z",
      "content": "<p>I'm having troubles as well and came upon the same answer but the folder <code>/kaggle/input/prostate-cancer-grade-assessment/test_images</code> is somehow not accessible when I compile my notebooks. Here is a minimal example:\n<a href=\"https://www.kaggle.com/lvulliard/test-panda-data\">https://www.kaggle.com/lvulliard/test-panda-data</a>\nAny idea what is going wrong here? Cheers.</p>",
      "rawMarkdown": "I'm having troubles as well and came upon the same answer but the folder `/kaggle/input/prostate-cancer-grade-assessment/test_images` is somehow not accessible when I compile my notebooks. Here is a minimal example:\nhttps://www.kaggle.com/lvulliard/test-panda-data\nAny idea what is going wrong here? Cheers.",
      "votes": 1,
      "replies": [
        {
          "id": 873787,
          "postDate": "2020-06-04T12:29:39.440Z",
          "content": "<p>As far as I understood, there is no way to see the test_images folder and its content.\nSo the best strategy is the following:\n1. Train your model in a notebook and save the model.\n2. For the inference, use a different notebook. In this case you will need to add your saved model to this new notebook by adding data. This time, your code will be reading the images from the test_images, doing some preprocessing etc, and sending the data to the model for inference. Try this with the train_images. Once you see this is working perfectly. You change your code to get the images from the test_images folder and then click on Save Version and select Save &amp; Run All (Commit). Your code must produce a file named submission.csv in the requested format. And voila! That's it.</p>",
          "rawMarkdown": "As far as I understood, there is no way to see the test_images folder and its content.\nSo the best strategy is the following:\n1. Train your model in a notebook and save the model.\n2. For the inference, use a different notebook. In this case you will need to add your saved model to this new notebook by adding data. This time, your code will be reading the images from the test_images, doing some preprocessing etc, and sending the data to the model for inference. Try this with the train_images. Once you see this is working perfectly. You change your code to get the images from the test_images folder and then click on Save Version and select Save &amp; Run All (Commit). Your code must produce a file named submission.csv in the requested format. And voila! That's it."
        },
        {
          "id": 873811,
          "postDate": "2020-06-04T12:52:47.787Z",
          "content": "<p>Thanks for the suggestion. I indeed train my model separately, and so far I tried separating the preprocessing and the prediction in different notebooks. This would look like:</p>\n\n<ol>\n<li>Preprocess images (both train and test sets), save transformed images.</li>\n<li>Load train images from Notebook 1, train and export model.</li>\n<li>Load test images from Notebook 1, load model from Notebook 2, predict labels and save in a submission file.</li>\n</ol>\n\n<p>I don't mind not \"seeing\" the test images in my live notebooks as I can test everything with training images, but so far when I compile my first notebook (with <em>Save &amp; Run All</em>), I only get preprocessed training images in the output. Should I process the testing images in the same notebook where the inference is done? How does Kaggle know that and \"unlock\" the test dataset?</p>\n\n<p>Edit: Okay, now I thing I get it. What I was missing was that your notebook basically run twice for submissions: \n1. When you press <em>Save &amp; Run All</em>, in which case the test set is not available (hence the output).\n2. When you press <em>Submit</em> next to your <em>submission.csv</em> output, in which case the code is ran but you don't see its output at all. The test set is unlocked and the result is directly evaluated for the leaderboard.\nThis means that I indeed need to do the preprocessing and the infer labels in a single notebook (which makes sense if you consider that the organizers want the full prediction from scratch to end-results to be constrained in time and resources). Cheers.</p>",
          "rawMarkdown": "Thanks for the suggestion. I indeed train my model separately, and so far I tried separating the preprocessing and the prediction in different notebooks. This would look like:\n\n1. Preprocess images (both train and test sets), save transformed images.\n2. Load train images from Notebook 1, train and export model.\n3. Load test images from Notebook 1, load model from Notebook 2, predict labels and save in a submission file.\n\nI don't mind not \"seeing\" the test images in my live notebooks as I can test everything with training images, but so far when I compile my first notebook (with *Save &amp; Run All*), I only get preprocessed training images in the output. Should I process the testing images in the same notebook where the inference is done? How does Kaggle know that and \"unlock\" the test dataset?\n\nEdit: Okay, now I thing I get it. What I was missing was that your notebook basically run twice for submissions: \n1. When you press *Save &amp; Run All*, in which case the test set is not available (hence the output).\n2. When you press *Submit* next to your *submission.csv* output, in which case the code is ran but you don't see its output at all. The test set is unlocked and the result is directly evaluated for the leaderboard.\nThis means that I indeed need to do the preprocessing and the infer labels in a single notebook (which makes sense if you consider that the organizers want the full prediction from scratch to end-results to be constrained in time and resources). Cheers.",
          "votes": 2
        },
        {
          "id": 873839,
          "postDate": "2020-06-04T13:22:03.263Z",
          "content": "<p>In your first step, you only process the train set.\nIn your second step, you only load the model and its weights to the second notebook.\nIn your third step, you preprocess the images in the test_images folder without seeing them. This means that you loop over the columns in the test.csv - you get the file names and join them with the test_images directory. Then you read each test file, preprocess them, send them to your model, get the results, and save them in a submission.csv file. When you commit your notebook(save &amp; run), your code will find the images and produce your submission.csv file correctly. It won't show you the test images ever. In short, do the inference like you are doing it with the train_images folder and the train.csv file, and then change your variables to use the test_images and test.csv. And commit.</p>\n\n<p>Yes, you should preprocess the test images in the second notebook where the inference is done. Unlocking of the test data happens when the code is run when you commit.</p>",
          "rawMarkdown": "In your first step, you only process the train set.\nIn your second step, you only load the model and its weights to the second notebook.\nIn your third step, you preprocess the images in the test_images folder without seeing them. This means that you loop over the columns in the test.csv - you get the file names and join them with the test_images directory. Then you read each test file, preprocess them, send them to your model, get the results, and save them in a submission.csv file. When you commit your notebook(save &amp; run), your code will find the images and produce your submission.csv file correctly. It won't show you the test images ever. In short, do the inference like you are doing it with the train_images folder and the train.csv file, and then change your variables to use the test_images and test.csv. And commit.\n\nYes, you should preprocess the test images in the second notebook where the inference is done. Unlocking of the test data happens when the code is run when you commit.\n",
          "votes": 1
        },
        {
          "id": 874202,
          "postDate": "2020-06-04T17:48:05.197Z",
          "content": "<p>I managed to send my result. There is a final step as Yovin mentioned below. While comitting the inference, there should be an if statement. If the test_image exists, your submission.csv file will process the test_image files, else it has to produce a submission.csv preferably with the values from sample_submission.csv.\nAfter the comit ends, you will see a submission.csv file as an output. Then you submit the submission.csv file in the submission page. When you do that kaggle will rerun your notebook and this time it will see the test_image folder.</p>",
          "rawMarkdown": "I managed to send my result. There is a final step as Yovin mentioned below. While comitting the inference, there should be an if statement. If the test_image exists, your submission.csv file will process the test_image files, else it has to produce a submission.csv preferably with the values from sample_submission.csv.\nAfter the comit ends, you will see a submission.csv file as an output. Then you submit the submission.csv file in the submission page. When you do that kaggle will rerun your notebook and this time it will see the test_image folder."
        }
      ]
    },
    {
      "id": 873388,
      "postDate": "2020-06-04T05:22:59.457Z",
      "content": "<p>You will need to do something like this to predict on unseen test data</p>\n\n<p><code>DATA = '../input/prostate-cancer-grade-assessment/test_images'</code></p>\n\n<p><code>\nsub_df = pd.read_csv(SAMPLE)\n  if os.path.exists(DATA):\n    test_dataset = ProstateDataset(TEST['image_id'], labels=TEST, mode='test', \n                     transform=get_transforms(data='valid'))\n    test_loader = DataLoader(test_dataset, batch_size=batch_size, shuffle=False, num_workers=5)\n    test()\n else:\n    sub_df.to_csv(\"submission.csv\", index=False)\n</code></p>",
      "rawMarkdown": "You will need to do something like this to predict on unseen test data\n\n`DATA = '../input/prostate-cancer-grade-assessment/test_images'`\n\n```\nsub_df = pd.read_csv(SAMPLE)\n  if os.path.exists(DATA):\n    test_dataset = ProstateDataset(TEST['image_id'], labels=TEST, mode='test', \n                     transform=get_transforms(data='valid'))\n    test_loader = DataLoader(test_dataset, batch_size=batch_size, shuffle=False, num_workers=5)\n    test()\n else:\n    sub_df.to_csv(\"submission.csv\", index=False)\n```",
      "votes": 1,
      "replies": [
        {
          "id": 873722,
          "postDate": "2020-06-04T11:38:48.880Z",
          "content": "<p>Thank you! I will give it a try.</p>",
          "rawMarkdown": "Thank you! I will give it a try."
        }
      ]
    },
    {
      "id": 873061,
      "postDate": "2020-06-03T18:30:50.027Z",
      "content": "<p>Seriously, how do we read the test data? I submitted a submission.csv file with random numbers but still don't know how to see the test data. Is it visible at all? Any help is appreciated.</p>\n\n<p>P.S. My model is dying to make some predictions on the test data.</p>",
      "rawMarkdown": "Seriously, how do we read the test data? I submitted a submission.csv file with random numbers but still don't know how to see the test data. Is it visible at all? Any help is appreciated.\n\nP.S. My model is dying to make some predictions on the test data.",
      "votes": 2
    },
    {
      "id": 873385,
      "postDate": "2020-06-04T05:15:50.317Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 873735,
      "author_name": "Loan Vulliard",
      "author_url": "",
      "post_date": "2020-06-04T11:46:37.993000",
      "content": "<p>I'm having troubles as well and came upon the same answer but the folder <code>/kaggle/input/prostate-cancer-grade-assessment/test_images</code> is somehow not accessible when I compile my notebooks. Here is a minimal example:\n<a href=\"https://www.kaggle.com/lvulliard/test-panda-data\">https://www.kaggle.com/lvulliard/test-panda-data</a>\nAny idea what is going wrong here? Cheers.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 873787,
          "author_name": "Tolga",
          "author_url": "",
          "post_date": "2020-06-04T12:29:39.440000",
          "content": "<p>As far as I understood, there is no way to see the test_images folder and its content.\nSo the best strategy is the following:\n1. Train your model in a notebook and save the model.\n2. For the inference, use a different notebook. In this case you will need to add your saved model to this new notebook by adding data. This time, your code will be reading the images from the test_images, doing some preprocessing etc, and sending the data to the model for inference. Try this with the train_images. Once you see this is working perfectly. You change your code to get the images from the test_images folder and then click on Save Version and select Save &amp; Run All (Commit). Your code must produce a file named submission.csv in the requested format. And voila! That's it.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 873811,
          "author_name": "Loan Vulliard",
          "author_url": "",
          "post_date": "2020-06-04T12:52:47.787000",
          "content": "<p>Thanks for the suggestion. I indeed train my model separately, and so far I tried separating the preprocessing and the prediction in different notebooks. This would look like:</p>\n\n<ol>\n<li>Preprocess images (both train and test sets), save transformed images.</li>\n<li>Load train images from Notebook 1, train and export model.</li>\n<li>Load test images from Notebook 1, load model from Notebook 2, predict labels and save in a submission file.</li>\n</ol>\n\n<p>I don't mind not \"seeing\" the test images in my live notebooks as I can test everything with training images, but so far when I compile my first notebook (with <em>Save &amp; Run All</em>), I only get preprocessed training images in the output. Should I process the testing images in the same notebook where the inference is done? How does Kaggle know that and \"unlock\" the test dataset?</p>\n\n<p>Edit: Okay, now I thing I get it. What I was missing was that your notebook basically run twice for submissions: \n1. When you press <em>Save &amp; Run All</em>, in which case the test set is not available (hence the output).\n2. When you press <em>Submit</em> next to your <em>submission.csv</em> output, in which case the code is ran but you don't see its output at all. The test set is unlocked and the result is directly evaluated for the leaderboard.\nThis means that I indeed need to do the preprocessing and the infer labels in a single notebook (which makes sense if you consider that the organizers want the full prediction from scratch to end-results to be constrained in time and resources). Cheers.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 873839,
          "author_name": "Tolga",
          "author_url": "",
          "post_date": "2020-06-04T13:22:03.263000",
          "content": "<p>In your first step, you only process the train set.\nIn your second step, you only load the model and its weights to the second notebook.\nIn your third step, you preprocess the images in the test_images folder without seeing them. This means that you loop over the columns in the test.csv - you get the file names and join them with the test_images directory. Then you read each test file, preprocess them, send them to your model, get the results, and save them in a submission.csv file. When you commit your notebook(save &amp; run), your code will find the images and produce your submission.csv file correctly. It won't show you the test images ever. In short, do the inference like you are doing it with the train_images folder and the train.csv file, and then change your variables to use the test_images and test.csv. And commit.</p>\n\n<p>Yes, you should preprocess the test images in the second notebook where the inference is done. Unlocking of the test data happens when the code is run when you commit.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 874202,
          "author_name": "Tolga",
          "author_url": "",
          "post_date": "2020-06-04T17:48:05.197000",
          "content": "<p>I managed to send my result. There is a final step as Yovin mentioned below. While comitting the inference, there should be an if statement. If the test_image exists, your submission.csv file will process the test_image files, else it has to produce a submission.csv preferably with the values from sample_submission.csv.\nAfter the comit ends, you will see a submission.csv file as an output. Then you submit the submission.csv file in the submission page. When you do that kaggle will rerun your notebook and this time it will see the test_image folder.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 873388,
      "author_name": "Yovin Yahathugoda",
      "author_url": "",
      "post_date": "2020-06-04T05:22:59.457000",
      "content": "<p>You will need to do something like this to predict on unseen test data</p>\n\n<p><code>DATA = '../input/prostate-cancer-grade-assessment/test_images'</code></p>\n\n<p><code>\nsub_df = pd.read_csv(SAMPLE)\n  if os.path.exists(DATA):\n    test_dataset = ProstateDataset(TEST['image_id'], labels=TEST, mode='test', \n                     transform=get_transforms(data='valid'))\n    test_loader = DataLoader(test_dataset, batch_size=batch_size, shuffle=False, num_workers=5)\n    test()\n else:\n    sub_df.to_csv(\"submission.csv\", index=False)\n</code></p>",
      "votes": 1,
      "replies": [
        {
          "id": 873722,
          "author_name": "Tolga",
          "author_url": "",
          "post_date": "2020-06-04T11:38:48.880000",
          "content": "<p>Thank you! I will give it a try.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 873385,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-04T05:15:50.317000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "873735": "I'm having troubles as well and came upon the same answer but the folder `/kaggle/input/prostate-cancer-grade-assessment/test_images` is somehow not accessible when I compile my notebooks. Here is a minimal example:\nhttps://www.kaggle.com/lvulliard/test-panda-data\nAny idea what is going wrong here? Cheers.",
    "873388": "You will need to do something like this to predict on unseen test data\n\n`DATA = '../input/prostate-cancer-grade-assessment/test_images'`\n\n```\nsub_df = pd.read_csv(SAMPLE)\n  if os.path.exists(DATA):\n    test_dataset = ProstateDataset(TEST['image_id'], labels=TEST, mode='test', \n                     transform=get_transforms(data='valid'))\n    test_loader = DataLoader(test_dataset, batch_size=batch_size, shuffle=False, num_workers=5)\n    test()\n else:\n    sub_df.to_csv(\"submission.csv\", index=False)\n```",
    "873061": "Seriously, how do we read the test data? I submitted a submission.csv file with random numbers but still don't know how to see the test data. Is it visible at all? Any help is appreciated.\n\nP.S. My model is dying to make some predictions on the test data.",
    "873385": ""
  }
}