{
  "id": 165065,
  "title": "Submission Score of -0.0!!",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/165065",
  "author_name": "Ankit Gupta",
  "post_date": "2020-07-08T12:51:33.155000",
  "votes": 5,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi guys,</p>\n\n<p>I'm new to this Competition and when I first got this score I looked at discussions and realise the different code for submission and testing. So I modified my code accordingly to this.\n```\nbase_dir = Path('/kaggle/input/prostate-cancer-grade-assessment')\ntest_dir = base_dir / 'test_images'\ntrain_dir = base_dir / 'train_images'\nif not test_dir.exists():\n    print(\"WARNING: Not in competition mode -&gt; test images not available!\")\n    eval_image_names = pd.read_csv(base_dir/'train.csv').head()\n    eval_dir = base_dir / 'train_images'</p>\n\n<p>else:\n    eval_image_names = pd.read_csv(base_dir/'test.csv')\n    eval_dir = base_dir / 'test_images'\n<code>\nNow I just iterate through `eval_image_names` and generate inference like this:\n</code>\neval_images = []\neval_isup = []\nfor (idx, entry) in eval_image_names.iterrows():\n    image_id = entry.image_id\n    with torch.no_grad():\n    ##inference\n    isup = model(input)\n    isup = int(op.argmax().item())\n    eval_images.append(image_id)\n    eval_isup.append(isup)\n<code>\nAnd then save the csv as:\n</code>\nsubmission = pd.DataFrame({\"image_id\":eval_images, \"isup_grade\": eval_isup})\nsubmission['isup_grade'] = submission['isup_grade'].astype(int)\nsubmission.to_csv('submission.csv', index=False)\n```</p>\n\n<p>I've tried the submission 2 times now with different notebooks but the result is still the same. Both the times, the submission ran successfully and took 6+ hours on CPU so I know it's processing all the test images. I've also tested the kappa score on 100 images on train data and it was around 0.5 (not -0.0) and also tried different models which are working on CPU.</p>\n\n<p>At this point, I'm not sure what I'm doing wrong. Any help would be greatly appreciated. :)</p>",
  "messages": [
    {
      "id": 920217,
      "postDate": "2020-07-08T12:51:33.157Z",
      "content": "<p>Hi guys,</p>\n\n<p>I'm new to this Competition and when I first got this score I looked at discussions and realise the different code for submission and testing. So I modified my code accordingly to this.\n```\nbase_dir = Path('/kaggle/input/prostate-cancer-grade-assessment')\ntest_dir = base_dir / 'test_images'\ntrain_dir = base_dir / 'train_images'\nif not test_dir.exists():\n    print(\"WARNING: Not in competition mode -&gt; test images not available!\")\n    eval_image_names = pd.read_csv(base_dir/'train.csv').head()\n    eval_dir = base_dir / 'train_images'</p>\n\n<p>else:\n    eval_image_names = pd.read_csv(base_dir/'test.csv')\n    eval_dir = base_dir / 'test_images'\n<code>\nNow I just iterate through `eval_image_names` and generate inference like this:\n</code>\neval_images = []\neval_isup = []\nfor (idx, entry) in eval_image_names.iterrows():\n    image_id = entry.image_id\n    with torch.no_grad():\n    ##inference\n    isup = model(input)\n    isup = int(op.argmax().item())\n    eval_images.append(image_id)\n    eval_isup.append(isup)\n<code>\nAnd then save the csv as:\n</code>\nsubmission = pd.DataFrame({\"image_id\":eval_images, \"isup_grade\": eval_isup})\nsubmission['isup_grade'] = submission['isup_grade'].astype(int)\nsubmission.to_csv('submission.csv', index=False)\n```</p>\n\n<p>I've tried the submission 2 times now with different notebooks but the result is still the same. Both the times, the submission ran successfully and took 6+ hours on CPU so I know it's processing all the test images. I've also tested the kappa score on 100 images on train data and it was around 0.5 (not -0.0) and also tried different models which are working on CPU.</p>\n\n<p>At this point, I'm not sure what I'm doing wrong. Any help would be greatly appreciated. :)</p>",
      "rawMarkdown": "Hi guys,\n\nI'm new to this Competition and when I first got this score I looked at discussions and realise the different code for submission and testing. So I modified my code accordingly to this.\n```\nbase_dir = Path('/kaggle/input/prostate-cancer-grade-assessment')\ntest_dir = base_dir / 'test_images'\ntrain_dir = base_dir / 'train_images'\nif not test_dir.exists():\n    print(\"WARNING: Not in competition mode -&gt; test images not available!\")\n    eval_image_names = pd.read_csv(base_dir/'train.csv').head()\n    eval_dir = base_dir / 'train_images'\n    \nelse:\n    eval_image_names = pd.read_csv(base_dir/'test.csv')\n    eval_dir = base_dir / 'test_images'\n```\nNow I just iterate through `eval_image_names` and generate inference like this:\n```\neval_images = []\neval_isup = []\nfor (idx, entry) in eval_image_names.iterrows():\n    image_id = entry.image_id\n    with torch.no_grad():\n    ##inference\n    isup = model(input)\n    isup = int(op.argmax().item())\n    eval_images.append(image_id)\n    eval_isup.append(isup)\n```\nAnd then save the csv as:\n```\nsubmission = pd.DataFrame({\"image_id\":eval_images, \"isup_grade\": eval_isup})\nsubmission['isup_grade'] = submission['isup_grade'].astype(int)\nsubmission.to_csv('submission.csv', index=False)\n```\n\nI've tried the submission 2 times now with different notebooks but the result is still the same. Both the times, the submission ran successfully and took 6+ hours on CPU so I know it's processing all the test images. I've also tested the kappa score on 100 images on train data and it was around 0.5 (not -0.0) and also tried different models which are working on CPU.\n\nAt this point, I'm not sure what I'm doing wrong. Any help would be greatly appreciated. :)",
      "votes": 5
    },
    {
      "id": 921062,
      "postDate": "2020-07-09T03:59:15.097Z",
      "content": "<p>Try doing prediction on the train images and look at the file generated. Most likely there is something major wrong but you can't see the prediction on the actual test data. </p>",
      "rawMarkdown": "Try doing prediction on the train images and look at the file generated. Most likely there is something major wrong but you can't see the prediction on the actual test data. ",
      "replies": [
        {
          "id": 922517,
          "postDate": "2020-07-10T07:02:19.773Z",
          "content": "<p>Thanks for the reply. It turns out, I was doing a couple of things that were causing wrong predictions. Fixed it now. :)</p>",
          "rawMarkdown": "Thanks for the reply. It turns out, I was doing a couple of things that were causing wrong predictions. Fixed it now. :)"
        },
        {
          "id": 925185,
          "postDate": "2020-07-11T21:44:45.563Z",
          "content": "<p>Hi,</p>\n\n<p>What was the issue?</p>\n\n<p>I experience the same. submission.csv looks fine, when I generate it for the part of train images.\nBut when I submit the code, that works for test images I get exact 0.0 score.</p>",
          "rawMarkdown": "Hi,\n\nWhat was the issue?\n\nI experience the same. submission.csv looks fine, when I generate it for the part of train images.\nBut when I submit the code, that works for test images I get exact 0.0 score."
        },
        {
          "id": 925738,
          "postDate": "2020-07-12T09:24:29.390Z",
          "content": "<p>Hey, in my case, I was loading data wrongly which was causing the model to give the same output for any input. I would suggest trying to run the network on the training images and see if it's working and then run it on test images. :) </p>",
          "rawMarkdown": "Hey, in my case, I was loading data wrongly which was causing the model to give the same output for any input. I would suggest trying to run the network on the training images and see if it's working and then run it on test images. :) ",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 921062,
      "author_name": "andy jennings",
      "author_url": "",
      "post_date": "2020-07-09T03:59:15.097000",
      "content": "<p>Try doing prediction on the train images and look at the file generated. Most likely there is something major wrong but you can't see the prediction on the actual test data. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 922517,
          "author_name": "Ankit Gupta",
          "author_url": "",
          "post_date": "2020-07-10T07:02:19.773000",
          "content": "<p>Thanks for the reply. It turns out, I was doing a couple of things that were causing wrong predictions. Fixed it now. :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 925185,
          "author_name": "Dmitry A. Grechka",
          "author_url": "",
          "post_date": "2020-07-11T21:44:45.563000",
          "content": "<p>Hi,</p>\n\n<p>What was the issue?</p>\n\n<p>I experience the same. submission.csv looks fine, when I generate it for the part of train images.\nBut when I submit the code, that works for test images I get exact 0.0 score.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 925738,
          "author_name": "Ankit Gupta",
          "author_url": "",
          "post_date": "2020-07-12T09:24:29.390000",
          "content": "<p>Hey, in my case, I was loading data wrongly which was causing the model to give the same output for any input. I would suggest trying to run the network on the training images and see if it's working and then run it on test images. :) </p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "920217": "Hi guys,\n\nI'm new to this Competition and when I first got this score I looked at discussions and realise the different code for submission and testing. So I modified my code accordingly to this.\n```\nbase_dir = Path('/kaggle/input/prostate-cancer-grade-assessment')\ntest_dir = base_dir / 'test_images'\ntrain_dir = base_dir / 'train_images'\nif not test_dir.exists():\n    print(\"WARNING: Not in competition mode -&gt; test images not available!\")\n    eval_image_names = pd.read_csv(base_dir/'train.csv').head()\n    eval_dir = base_dir / 'train_images'\n    \nelse:\n    eval_image_names = pd.read_csv(base_dir/'test.csv')\n    eval_dir = base_dir / 'test_images'\n```\nNow I just iterate through `eval_image_names` and generate inference like this:\n```\neval_images = []\neval_isup = []\nfor (idx, entry) in eval_image_names.iterrows():\n    image_id = entry.image_id\n    with torch.no_grad():\n    ##inference\n    isup = model(input)\n    isup = int(op.argmax().item())\n    eval_images.append(image_id)\n    eval_isup.append(isup)\n```\nAnd then save the csv as:\n```\nsubmission = pd.DataFrame({\"image_id\":eval_images, \"isup_grade\": eval_isup})\nsubmission['isup_grade'] = submission['isup_grade'].astype(int)\nsubmission.to_csv('submission.csv', index=False)\n```\n\nI've tried the submission 2 times now with different notebooks but the result is still the same. Both the times, the submission ran successfully and took 6+ hours on CPU so I know it's processing all the test images. I've also tested the kappa score on 100 images on train data and it was around 0.5 (not -0.0) and also tried different models which are working on CPU.\n\nAt this point, I'm not sure what I'm doing wrong. Any help would be greatly appreciated. :)",
    "921062": "Try doing prediction on the train images and look at the file generated. Most likely there is something major wrong but you can't see the prediction on the actual test data. "
  }
}