{
  "id": 149182,
  "title": "Score 0.0 when submitting",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/149182",
  "author_name": "TheStoneMX",
  "post_date": "2020-05-07T07:02:13.536000",
  "votes": 2,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Hi there all, every time I submit I get 0.0 this is driving me crazy as it is the same code to make my submissions, I have only change the trained models that I do at home, 5 different with five different submissions and all give me 0.0 and there is no error on the kernel that will give me a hint......</p>\n\n<p>Can someone help me?</p>\n\n<p>The only thing I am doing differently is I am using:</p>\n\n<p>model = torch.nn.DataParallel(learn.model) \nas I train with multiple GPUs</p>\n\n<p>Thanks.</p>",
  "messages": [
    {
      "id": 836810,
      "postDate": "2020-05-07T08:57:55.027Z",
      "content": "<p><a href=\"/oscarrangel\">@oscarrangel</a> Since you trained on multiple GPUs your trained weights file may not have been correctly loaded when running on a single GPU available on Kaggle. Therefore, only the <strong>sample_submission.csv</strong> file with only 3 entries have been scored.</p>\n\n<p>I came across the following code in a forum a while back for loading a multiple GPU weights file to a singe GPU and has worked for me so far.</p>\n\n<p><code>class WrappedModel(nn.Module):</code>\n    <code>def __init__(self, module):</code>\n        <code>super(WrappedModel, self).__init__()</code>\n        <code>self.module = module</code>\n    <code>def forward(self, x):</code>\n        <code>return self.module(x)</code></p>\n\n<p>call it as follows afterwards:</p>\n\n<p><code>predictor = PretrainedCNN(in_channels=3, out_dim=n_total, model_name=model_name, pretrained=None)</code>\n  <code>predictor = WrappedModel(predictor)</code>\n  <code>predictor.load_state_dict(torch.load(load_model_path))</code></p>\n\n<p>I hope this solves the problem. Its more likely your model weights did not load and no predictions were actually made in your case.</p>",
      "rawMarkdown": "@oscarrangel Since you trained on multiple GPUs your trained weights file may not have been correctly loaded when running on a single GPU available on Kaggle. Therefore, only the **sample_submission.csv** file with only 3 entries have been scored.\n\nI came across the following code in a forum a while back for loading a multiple GPU weights file to a singe GPU and has worked for me so far.\n\n`class WrappedModel(nn.Module):`\n\t`  def __init__(self, module):`\n\t\t`super(WrappedModel, self).__init__()`\n\t\t`self.module = module`\n\t`def forward(self, x):`\n\t\t`return self.module(x)`\n\ncall it as follows afterwards:\n\n`predictor = PretrainedCNN(in_channels=3, out_dim=n_total, model_name=model_name, pretrained=None)`\n  `predictor = WrappedModel(predictor)`\n  `predictor.load_state_dict(torch.load(load_model_path))`\n\nI hope this solves the problem. Its more likely your model weights did not load and no predictions were actually made in your case.",
      "votes": 1,
      "replies": [
        {
          "id": 837026,
          "postDate": "2020-05-07T13:07:47.277Z",
          "content": "<p><a href=\"/yovinyahathugoda\">@yovinyahathugoda</a> thanks a lot for your post! let me try it and let you know, I am still learning so anything ith kaggle.</p>",
          "rawMarkdown": "@yovinyahathugoda thanks a lot for your post! let me try it and let you know, I am still learning so anything ith kaggle."
        },
        {
          "id": 837042,
          "postDate": "2020-05-07T13:28:57.037Z",
          "content": "<p><a href=\"/yovinyahathugoda\">@yovinyahathugoda</a> thanks but did not work.... even though it did no complain and submitted ok.</p>\n\n<p>PANDA concat tile pooling starter [inference](version 12/12)</p>\n\n<p>From Kernel [PANDA Concat tile pooling starter [inference]] Succeeded 0.00</p>\n\n<p>I am going to create some pth with single processor. it may take longer to create them and see what happen</p>",
          "rawMarkdown": "@yovinyahathugoda thanks but did not work.... even though it did no complain and submitted ok.\n\nPANDA concat tile pooling starter [inference](version 12/12)\n\nFrom Kernel [PANDA Concat tile pooling starter [inference]] Succeeded 0.00\n\nI am going to create some pth with single processor. it may take longer to create them and see what happen\n\n"
        },
        {
          "id": 837107,
          "postDate": "2020-05-07T14:21:08.517Z",
          "content": "<p><a href=\"/oscarrangel\">@oscarrangel</a> could you try running the submission kernel in draft mode with the training images and the train.csv file to see where an error could be occurring? </p>",
          "rawMarkdown": "@oscarrangel could you try running the submission kernel in draft mode with the training images and the train.csv file to see where an error could be occurring? "
        },
        {
          "id": 837396,
          "postDate": "2020-05-07T18:27:53.163Z",
          "content": "<p>hi <a href=\"/yovinyahathugoda\">@yovinyahathugoda</a> , trying.... </p>",
          "rawMarkdown": "hi @yovinyahathugoda , trying.... "
        },
        {
          "id": 837401,
          "postDate": "2020-05-07T18:31:48.043Z",
          "content": "<p>ahhhh!!! nice!!! I finally got to see the problem!!!</p>\n\n<p>RuntimeError: CUDA out of memory. Tried to allocate 6.12 GiB (GPU 0; 15.90 GiB total capacity; 7.46 GiB already allocated; 1.52 GiB free; 13.72 GiB reserved in total by PyTorch)</p>\n\n<p>Then let me lower the batch size, and see</p>",
          "rawMarkdown": "ahhhh!!! nice!!! I finally got to see the problem!!!\n\nRuntimeError: CUDA out of memory. Tried to allocate 6.12 GiB (GPU 0; 15.90 GiB total capacity; 7.46 GiB already allocated; 1.52 GiB free; 13.72 GiB reserved in total by PyTorch)\n\nThen let me lower the batch size, and see\n"
        },
        {
          "id": 837407,
          "postDate": "2020-05-07T18:41:05.217Z",
          "content": "<p>I am using \nsz = 224\nbs = 16</p>\n\n<p>I had to change it to 224, 16.....</p>\n\n<p>But then this means that I will never be able to use 512x512images? very limited resources, what do you suggest?</p>\n\n<p>Thanks a lot for your help <a href=\"/yovinyahathugoda\">@yovinyahathugoda</a> it was driving me crazy!!! </p>",
          "rawMarkdown": "I am using \nsz = 224\nbs = 16\n\n\nI had to change it to 224, 16.....\n\nBut then this means that I will never be able to use 512x512images? very limited resources, what do you suggest?\n\nThanks a lot for your help @yovinyahathugoda it was driving me crazy!!! "
        },
        {
          "id": 837439,
          "postDate": "2020-05-07T19:15:50.270Z",
          "content": "<p>PANDA concat tile pooling starter [inference](version 7/7)\n31 minutes ago by TheStoneMX</p>\n\n<p>From Kernel [PANDA concat tile pooling starter [inference]] Succeeded 0.00</p>\n\n<p>🤕 😭 </p>",
          "rawMarkdown": "PANDA concat tile pooling starter [inference](version 7/7)\n31 minutes ago by TheStoneMX\n\nFrom Kernel [PANDA concat tile pooling starter [inference]] Succeeded 0.00\n\n🤕 😭 "
        },
        {
          "id": 837728,
          "postDate": "2020-05-08T02:18:36.413Z",
          "content": "<p>If you want to use 512x512 you will need to drop the batch size even further. But using 512 image size sometimes gives only a small improvement so it is not worth the time for me personally. </p>\n\n<p>However, Iafoss made a great kernel where the entire image is broken down into smaller tiles. In this method you can reduce the individual tile image size and set the number of tiles accordingly to keep the original image size and you would not be down-sampling. </p>\n\n<p><a href=\"https://www.kaggle.com/iafoss/panda-16x128x128-tiles\">https://www.kaggle.com/iafoss/panda-16x128x128-tiles</a></p>",
          "rawMarkdown": "If you want to use 512x512 you will need to drop the batch size even further. But using 512 image size sometimes gives only a small improvement so it is not worth the time for me personally. \n\nHowever, Iafoss made a great kernel where the entire image is broken down into smaller tiles. In this method you can reduce the individual tile image size and set the number of tiles accordingly to keep the original image size and you would not be down-sampling. \n\n[https://www.kaggle.com/iafoss/panda-16x128x128-tiles](https://www.kaggle.com/iafoss/panda-16x128x128-tiles)",
          "votes": 1
        },
        {
          "id": 838246,
          "postDate": "2020-05-08T12:11:02.067Z",
          "content": "<p>That is the kernel I am using as a base. from lafoss.</p>",
          "rawMarkdown": "That is the kernel I am using as a base. from lafoss."
        }
      ]
    },
    {
      "id": 837259,
      "postDate": "2020-05-07T16:39:59.817Z",
      "content": "<p><a href=\"/oscarrangel\">@oscarrangel</a> Make sure you have an if/else statement in your kernel like the below, since the hidden test set is only loaded when committing the kernel, otherwise you may be just submitting the sample csv file which might explain a 0 score. </p>\n\n<p>```\nTEST_DATA = f'{DATA_DIR}/test_images'</p>\n\n<p>if os.path.exists(TEST_DATA):\n    # code for submission df with preds\nelse:\n    sub_df = pd.read_csv(f'{DATA_DIR}/sample_submission.csv’)</p>\n\n<p>sub_df.to_csv(\"submission.csv\", index=False)\n```</p>",
      "rawMarkdown": "@oscarrangel Make sure you have an if/else statement in your kernel like the below, since the hidden test set is only loaded when committing the kernel, otherwise you may be just submitting the sample csv file which might explain a 0 score. \n\n```\nTEST_DATA = f'{DATA_DIR}/test_images'\n\nif os.path.exists(TEST_DATA):\n    # code for submission df with preds\nelse:\n    sub_df = pd.read_csv(f'{DATA_DIR}/sample_submission.csv’)\n\nsub_df.to_csv(\"submission.csv\", index=False)\n```",
      "votes": 2
    },
    {
      "id": 836706,
      "postDate": "2020-05-07T07:02:13.537Z",
      "content": "<p>Hi there all, every time I submit I get 0.0 this is driving me crazy as it is the same code to make my submissions, I have only change the trained models that I do at home, 5 different with five different submissions and all give me 0.0 and there is no error on the kernel that will give me a hint......</p>\n\n<p>Can someone help me?</p>\n\n<p>The only thing I am doing differently is I am using:</p>\n\n<p>model = torch.nn.DataParallel(learn.model) \nas I train with multiple GPUs</p>\n\n<p>Thanks.</p>",
      "rawMarkdown": "Hi there all, every time I submit I get 0.0 this is driving me crazy as it is the same code to make my submissions, I have only change the trained models that I do at home, 5 different with five different submissions and all give me 0.0 and there is no error on the kernel that will give me a hint......\n\nCan someone help me?\n\nThe only thing I am doing differently is I am using:\n\nmodel = torch.nn.DataParallel(learn.model) \nas I train with multiple GPUs\n\nThanks.",
      "votes": 2
    },
    {
      "id": 846324,
      "postDate": "2020-05-13T17:28:47.017Z",
      "content": "<p>There is a problem with tqmd on the kaggle servers.... follow this link.</p>\n\n<p><a href=\"https://www.kaggle.com/product-feedback/150817\">https://www.kaggle.com/product-feedback/150817</a></p>",
      "rawMarkdown": "There is a problem with tqmd on the kaggle servers.... follow this link.\n\nhttps://www.kaggle.com/product-feedback/150817"
    },
    {
      "id": 838507,
      "postDate": "2020-05-08T15:58:36.693Z",
      "content": "<p>for anyone else that has the same problem you cant train on multiple GPUs and infer on a single one, it will not work on kaggle you need to do something like this. taken from here.</p>\n\n<p><a href=\"https://discuss.pytorch.org/t/how-could-i-train-on-multi-gpu-and-infer-with-single-gpu/22838/3\">https://discuss.pytorch.org/t/how-could-i-train-on-multi-gpu-and-infer-with-single-gpu/22838/3</a></p>\n\n<p>Your Problem lies within your saving/loading code:</p>\n\n<p>You save your model via torch.save(net, save_path) where net is an instance of nn.DataParallel</p>\n\n<p>I would recommend to change a few things:</p>\n\n<p>You should only save the state dict instead of your model. See this 97 post for further details</p>\n\n<p>You should not save the stat dict of your DataParallel instance but your model’s state dict since the data parallel is simply a wrapper and also contains things like the used GPUs and copies of your model. You can access the model by calling net.module just as @JuanFMontesinos mentioned.</p>\n\n<p>Taking these steps into account you would save your model with code like this:</p>\n\n<p>net = nn.DataParallel(net)\n    .....\n    torch.save(net.module.state_dict(), save_path)\nand load your model like this:</p>\n\n<p>model = YourNetworkClass() # create an instance of your network\nmodel.load_state_dict(torch.load(save_path))\nIf you want to infer on multiple GPUs or continue training on multiple GPUs you would have to wrap your model again with nn.DataParallel.</p>\n\n<p>Also a good practice would be to move the model to cpu before saving it’s state_dict and move it back to GPU afterwards. This way the state dict will also be loaded to CPU instead of being loaded to GPU directly (which happens if you save the state_dict of a GPU-Model since you would save CUDA-Tensors). Loading a model on CPU is better practice since you could also deploy it to machines wich aren’t CUDA-capable</p>",
      "rawMarkdown": "for anyone else that has the same problem you cant train on multiple GPUs and infer on a single one, it will not work on kaggle you need to do something like this. taken from here.\n\nhttps://discuss.pytorch.org/t/how-could-i-train-on-multi-gpu-and-infer-with-single-gpu/22838/3\n\nYour Problem lies within your saving/loading code:\n\nYou save your model via torch.save(net, save_path) where net is an instance of nn.DataParallel\n\nI would recommend to change a few things:\n\nYou should only save the state dict instead of your model. See this 97 post for further details\n\nYou should not save the stat dict of your DataParallel instance but your model’s state dict since the data parallel is simply a wrapper and also contains things like the used GPUs and copies of your model. You can access the model by calling net.module just as @JuanFMontesinos mentioned.\n\nTaking these steps into account you would save your model with code like this:\n\nnet = nn.DataParallel(net)\n    .....\n    torch.save(net.module.state_dict(), save_path)\nand load your model like this:\n\nmodel = YourNetworkClass() # create an instance of your network\nmodel.load_state_dict(torch.load(save_path))\nIf you want to infer on multiple GPUs or continue training on multiple GPUs you would have to wrap your model again with nn.DataParallel.\n\nAlso a good practice would be to move the model to cpu before saving it’s state_dict and move it back to GPU afterwards. This way the state dict will also be loaded to CPU instead of being loaded to GPU directly (which happens if you save the state_dict of a GPU-Model since you would save CUDA-Tensors). Loading a model on CPU is better practice since you could also deploy it to machines wich aren’t CUDA-capable"
    }
  ],
  "comments": [
    {
      "id": 836810,
      "author_name": "Yovin Yahathugoda",
      "author_url": "",
      "post_date": "2020-05-07T08:57:55.027000",
      "content": "<p><a href=\"/oscarrangel\">@oscarrangel</a> Since you trained on multiple GPUs your trained weights file may not have been correctly loaded when running on a single GPU available on Kaggle. Therefore, only the <strong>sample_submission.csv</strong> file with only 3 entries have been scored.</p>\n\n<p>I came across the following code in a forum a while back for loading a multiple GPU weights file to a singe GPU and has worked for me so far.</p>\n\n<p><code>class WrappedModel(nn.Module):</code>\n    <code>def __init__(self, module):</code>\n        <code>super(WrappedModel, self).__init__()</code>\n        <code>self.module = module</code>\n    <code>def forward(self, x):</code>\n        <code>return self.module(x)</code></p>\n\n<p>call it as follows afterwards:</p>\n\n<p><code>predictor = PretrainedCNN(in_channels=3, out_dim=n_total, model_name=model_name, pretrained=None)</code>\n  <code>predictor = WrappedModel(predictor)</code>\n  <code>predictor.load_state_dict(torch.load(load_model_path))</code></p>\n\n<p>I hope this solves the problem. Its more likely your model weights did not load and no predictions were actually made in your case.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 837026,
          "author_name": "TheStoneMX",
          "author_url": "",
          "post_date": "2020-05-07T13:07:47.277000",
          "content": "<p><a href=\"/yovinyahathugoda\">@yovinyahathugoda</a> thanks a lot for your post! let me try it and let you know, I am still learning so anything ith kaggle.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 837042,
          "author_name": "TheStoneMX",
          "author_url": "",
          "post_date": "2020-05-07T13:28:57.037000",
          "content": "<p><a href=\"/yovinyahathugoda\">@yovinyahathugoda</a> thanks but did not work.... even though it did no complain and submitted ok.</p>\n\n<p>PANDA concat tile pooling starter [inference](version 12/12)</p>\n\n<p>From Kernel [PANDA Concat tile pooling starter [inference]] Succeeded 0.00</p>\n\n<p>I am going to create some pth with single processor. it may take longer to create them and see what happen</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 837107,
          "author_name": "Yovin Yahathugoda",
          "author_url": "",
          "post_date": "2020-05-07T14:21:08.517000",
          "content": "<p><a href=\"/oscarrangel\">@oscarrangel</a> could you try running the submission kernel in draft mode with the training images and the train.csv file to see where an error could be occurring? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 837396,
          "author_name": "TheStoneMX",
          "author_url": "",
          "post_date": "2020-05-07T18:27:53.163000",
          "content": "<p>hi <a href=\"/yovinyahathugoda\">@yovinyahathugoda</a> , trying.... </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 837401,
          "author_name": "TheStoneMX",
          "author_url": "",
          "post_date": "2020-05-07T18:31:48.043000",
          "content": "<p>ahhhh!!! nice!!! I finally got to see the problem!!!</p>\n\n<p>RuntimeError: CUDA out of memory. Tried to allocate 6.12 GiB (GPU 0; 15.90 GiB total capacity; 7.46 GiB already allocated; 1.52 GiB free; 13.72 GiB reserved in total by PyTorch)</p>\n\n<p>Then let me lower the batch size, and see</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 837407,
          "author_name": "TheStoneMX",
          "author_url": "",
          "post_date": "2020-05-07T18:41:05.217000",
          "content": "<p>I am using \nsz = 224\nbs = 16</p>\n\n<p>I had to change it to 224, 16.....</p>\n\n<p>But then this means that I will never be able to use 512x512images? very limited resources, what do you suggest?</p>\n\n<p>Thanks a lot for your help <a href=\"/yovinyahathugoda\">@yovinyahathugoda</a> it was driving me crazy!!! </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 837439,
          "author_name": "TheStoneMX",
          "author_url": "",
          "post_date": "2020-05-07T19:15:50.270000",
          "content": "<p>PANDA concat tile pooling starter [inference](version 7/7)\n31 minutes ago by TheStoneMX</p>\n\n<p>From Kernel [PANDA concat tile pooling starter [inference]] Succeeded 0.00</p>\n\n<p>🤕 😭 </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 837728,
          "author_name": "Yovin Yahathugoda",
          "author_url": "",
          "post_date": "2020-05-08T02:18:36.413000",
          "content": "<p>If you want to use 512x512 you will need to drop the batch size even further. But using 512 image size sometimes gives only a small improvement so it is not worth the time for me personally. </p>\n\n<p>However, Iafoss made a great kernel where the entire image is broken down into smaller tiles. In this method you can reduce the individual tile image size and set the number of tiles accordingly to keep the original image size and you would not be down-sampling. </p>\n\n<p><a href=\"https://www.kaggle.com/iafoss/panda-16x128x128-tiles\">https://www.kaggle.com/iafoss/panda-16x128x128-tiles</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 838246,
          "author_name": "TheStoneMX",
          "author_url": "",
          "post_date": "2020-05-08T12:11:02.067000",
          "content": "<p>That is the kernel I am using as a base. from lafoss.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 837259,
      "author_name": "James Requa",
      "author_url": "",
      "post_date": "2020-05-07T16:39:59.817000",
      "content": "<p><a href=\"/oscarrangel\">@oscarrangel</a> Make sure you have an if/else statement in your kernel like the below, since the hidden test set is only loaded when committing the kernel, otherwise you may be just submitting the sample csv file which might explain a 0 score. </p>\n\n<p>```\nTEST_DATA = f'{DATA_DIR}/test_images'</p>\n\n<p>if os.path.exists(TEST_DATA):\n    # code for submission df with preds\nelse:\n    sub_df = pd.read_csv(f'{DATA_DIR}/sample_submission.csv’)</p>\n\n<p>sub_df.to_csv(\"submission.csv\", index=False)\n```</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 846324,
      "author_name": "TheStoneMX",
      "author_url": "",
      "post_date": "2020-05-13T17:28:47.017000",
      "content": "<p>There is a problem with tqmd on the kaggle servers.... follow this link.</p>\n\n<p><a href=\"https://www.kaggle.com/product-feedback/150817\">https://www.kaggle.com/product-feedback/150817</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 838507,
      "author_name": "TheStoneCa",
      "author_url": "",
      "post_date": "2020-05-08T15:58:36.693000",
      "content": "<p>for anyone else that has the same problem you cant train on multiple GPUs and infer on a single one, it will not work on kaggle you need to do something like this. taken from here.</p>\n\n<p><a href=\"https://discuss.pytorch.org/t/how-could-i-train-on-multi-gpu-and-infer-with-single-gpu/22838/3\">https://discuss.pytorch.org/t/how-could-i-train-on-multi-gpu-and-infer-with-single-gpu/22838/3</a></p>\n\n<p>Your Problem lies within your saving/loading code:</p>\n\n<p>You save your model via torch.save(net, save_path) where net is an instance of nn.DataParallel</p>\n\n<p>I would recommend to change a few things:</p>\n\n<p>You should only save the state dict instead of your model. See this 97 post for further details</p>\n\n<p>You should not save the stat dict of your DataParallel instance but your model’s state dict since the data parallel is simply a wrapper and also contains things like the used GPUs and copies of your model. You can access the model by calling net.module just as @JuanFMontesinos mentioned.</p>\n\n<p>Taking these steps into account you would save your model with code like this:</p>\n\n<p>net = nn.DataParallel(net)\n    .....\n    torch.save(net.module.state_dict(), save_path)\nand load your model like this:</p>\n\n<p>model = YourNetworkClass() # create an instance of your network\nmodel.load_state_dict(torch.load(save_path))\nIf you want to infer on multiple GPUs or continue training on multiple GPUs you would have to wrap your model again with nn.DataParallel.</p>\n\n<p>Also a good practice would be to move the model to cpu before saving it’s state_dict and move it back to GPU afterwards. This way the state dict will also be loaded to CPU instead of being loaded to GPU directly (which happens if you save the state_dict of a GPU-Model since you would save CUDA-Tensors). Loading a model on CPU is better practice since you could also deploy it to machines wich aren’t CUDA-capable</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "836810": "@oscarrangel Since you trained on multiple GPUs your trained weights file may not have been correctly loaded when running on a single GPU available on Kaggle. Therefore, only the **sample_submission.csv** file with only 3 entries have been scored.\n\nI came across the following code in a forum a while back for loading a multiple GPU weights file to a singe GPU and has worked for me so far.\n\n`class WrappedModel(nn.Module):`\n\t`  def __init__(self, module):`\n\t\t`super(WrappedModel, self).__init__()`\n\t\t`self.module = module`\n\t`def forward(self, x):`\n\t\t`return self.module(x)`\n\ncall it as follows afterwards:\n\n`predictor = PretrainedCNN(in_channels=3, out_dim=n_total, model_name=model_name, pretrained=None)`\n  `predictor = WrappedModel(predictor)`\n  `predictor.load_state_dict(torch.load(load_model_path))`\n\nI hope this solves the problem. Its more likely your model weights did not load and no predictions were actually made in your case.",
    "837259": "@oscarrangel Make sure you have an if/else statement in your kernel like the below, since the hidden test set is only loaded when committing the kernel, otherwise you may be just submitting the sample csv file which might explain a 0 score. \n\n```\nTEST_DATA = f'{DATA_DIR}/test_images'\n\nif os.path.exists(TEST_DATA):\n    # code for submission df with preds\nelse:\n    sub_df = pd.read_csv(f'{DATA_DIR}/sample_submission.csv’)\n\nsub_df.to_csv(\"submission.csv\", index=False)\n```",
    "836706": "Hi there all, every time I submit I get 0.0 this is driving me crazy as it is the same code to make my submissions, I have only change the trained models that I do at home, 5 different with five different submissions and all give me 0.0 and there is no error on the kernel that will give me a hint......\n\nCan someone help me?\n\nThe only thing I am doing differently is I am using:\n\nmodel = torch.nn.DataParallel(learn.model) \nas I train with multiple GPUs\n\nThanks.",
    "846324": "There is a problem with tqmd on the kaggle servers.... follow this link.\n\nhttps://www.kaggle.com/product-feedback/150817",
    "838507": "for anyone else that has the same problem you cant train on multiple GPUs and infer on a single one, it will not work on kaggle you need to do something like this. taken from here.\n\nhttps://discuss.pytorch.org/t/how-could-i-train-on-multi-gpu-and-infer-with-single-gpu/22838/3\n\nYour Problem lies within your saving/loading code:\n\nYou save your model via torch.save(net, save_path) where net is an instance of nn.DataParallel\n\nI would recommend to change a few things:\n\nYou should only save the state dict instead of your model. See this 97 post for further details\n\nYou should not save the stat dict of your DataParallel instance but your model’s state dict since the data parallel is simply a wrapper and also contains things like the used GPUs and copies of your model. You can access the model by calling net.module just as @JuanFMontesinos mentioned.\n\nTaking these steps into account you would save your model with code like this:\n\nnet = nn.DataParallel(net)\n    .....\n    torch.save(net.module.state_dict(), save_path)\nand load your model like this:\n\nmodel = YourNetworkClass() # create an instance of your network\nmodel.load_state_dict(torch.load(save_path))\nIf you want to infer on multiple GPUs or continue training on multiple GPUs you would have to wrap your model again with nn.DataParallel.\n\nAlso a good practice would be to move the model to cpu before saving it’s state_dict and move it back to GPU afterwards. This way the state dict will also be loaded to CPU instead of being loaded to GPU directly (which happens if you save the state_dict of a GPU-Model since you would save CUDA-Tensors). Loading a model on CPU is better practice since you could also deploy it to machines wich aren’t CUDA-capable"
  }
}