{
  "id": 151570,
  "title": "model.eval() leads to Random predictions",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/151570",
  "author_name": "Jaideep",
  "post_date": "2020-05-16T03:41:35.102000",
  "votes": 0,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I find strange issue as never before  .Adding model.eval( ) makes predictions worst compared to without it . I have been doing all this earlier with no issue ,as it is way as well when u have batch norm in place .\nI dont know what unknown mistake i could be making</p>\n\n<p><code>\ngpu = torch.device(\"cuda:0\")\nstate_dict = torch.load('../work/xyz.pth',map_location=gpu)\nmodel=Model()\nmodel.load_state_dict(state_dict   )\nmodel.eval()\np=model.cuda()(x.half())\n</code></p>",
  "messages": [
    {
      "id": 849822,
      "postDate": "2020-05-16T05:19:17.750Z",
      "content": "<p>With <code>model.eval()</code> , batch normalization layers use fixed mean/std values calculated during training phase. Due to this, sometimes performance gets worse when there is difference between train and test batches. As far as I remember, the same occurred to me in <a href=\"https://www.kaggle.com/c/recursion-cellular-image-classification\">Recursion competition</a>. </p>",
      "rawMarkdown": "With `model.eval()` , batch normalization layers use fixed mean/std values calculated during training phase. Due to this, sometimes performance gets worse when there is difference between train and test batches. As far as I remember, the same occurred to me in [Recursion competition](https://www.kaggle.com/c/recursion-cellular-image-classification). ",
      "votes": 1,
      "replies": [
        {
          "id": 849838,
          "postDate": "2020-05-16T05:37:57.360Z",
          "content": "<p>m doing this for a train data</p>",
          "rawMarkdown": "m doing this for a train data"
        }
      ]
    },
    {
      "id": 850077,
      "postDate": "2020-05-16T10:08:21.073Z",
      "content": "<p>I had this the other day when using an efficientnet with reasonable batch sizes. Testing accuracy was diabolical on a randomly-split train/test dataset but was good if I didn't switch to eval() model. I gave up trying to figure it out, but I assumed I was doing something stupid. Absolutely bizarre. It's was the nightly pytorch build and efficientnet-pytorch, any chance you're using the same?</p>",
      "rawMarkdown": "I had this the other day when using an efficientnet with reasonable batch sizes. Testing accuracy was diabolical on a randomly-split train/test dataset but was good if I didn't switch to eval() model. I gave up trying to figure it out, but I assumed I was doing something stupid. Absolutely bizarre. It's was the nightly pytorch build and efficientnet-pytorch, any chance you're using the same?",
      "replies": [
        {
          "id": 850122,
          "postDate": "2020-05-16T11:00:36.960Z",
          "content": "<p>i use resnext...50\n1 ) Are you using  split methodology ,if yes then are you keeping constant no of split for all tiff  images . I use split method what i did in pre processing was  to generate the tiles for each image but each had different no of tiles ,in Test i used same pipeline except i maintained same no of tiles,not sure if this change is causing it </p>",
          "rawMarkdown": "i use resnext...50\n1 ) Are you using  split methodology ,if yes then are you keeping constant no of split for all tiff  images . I use split method what i did in pre processing was  to generate the tiles for each image but each had different no of tiles ,in Test i used same pipeline except i maintained same no of tiles,not sure if this change is causing it \n"
        },
        {
          "id": 850443,
          "postDate": "2020-05-16T16:23:06.503Z",
          "content": "<p>finally resolved all the issues on my way to next level :)...\nIt was silly mistake which made a vast difference between Train and Test preprocessing</p>",
          "rawMarkdown": "finally resolved all the issues on my way to next level :)...\nIt was silly mistake which made a vast difference between Train and Test preprocessing"
        }
      ]
    },
    {
      "id": 849761,
      "postDate": "2020-05-16T03:41:35.103Z",
      "content": "<p>I find strange issue as never before  .Adding model.eval( ) makes predictions worst compared to without it . I have been doing all this earlier with no issue ,as it is way as well when u have batch norm in place .\nI dont know what unknown mistake i could be making</p>\n\n<p><code>\ngpu = torch.device(\"cuda:0\")\nstate_dict = torch.load('../work/xyz.pth',map_location=gpu)\nmodel=Model()\nmodel.load_state_dict(state_dict   )\nmodel.eval()\np=model.cuda()(x.half())\n</code></p>",
      "rawMarkdown": "I find strange issue as never before  .Adding model.eval( ) makes predictions worst compared to without it . I have been doing all this earlier with no issue ,as it is way as well when u have batch norm in place .\nI dont know what unknown mistake i could be making\n\n```\ngpu = torch.device(\"cuda:0\")\nstate_dict = torch.load('../work/xyz.pth',map_location=gpu)\nmodel=Model()\nmodel.load_state_dict(state_dict   )\nmodel.eval()\np=model.cuda()(x.half())\n```"
    }
  ],
  "comments": [
    {
      "id": 849822,
      "author_name": "RabotniKuma",
      "author_url": "",
      "post_date": "2020-05-16T05:19:17.750000",
      "content": "<p>With <code>model.eval()</code> , batch normalization layers use fixed mean/std values calculated during training phase. Due to this, sometimes performance gets worse when there is difference between train and test batches. As far as I remember, the same occurred to me in <a href=\"https://www.kaggle.com/c/recursion-cellular-image-classification\">Recursion competition</a>. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 849838,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-05-16T05:37:57.360000",
          "content": "<p>m doing this for a train data</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 850077,
      "author_name": "James Howard",
      "author_url": "",
      "post_date": "2020-05-16T10:08:21.073000",
      "content": "<p>I had this the other day when using an efficientnet with reasonable batch sizes. Testing accuracy was diabolical on a randomly-split train/test dataset but was good if I didn't switch to eval() model. I gave up trying to figure it out, but I assumed I was doing something stupid. Absolutely bizarre. It's was the nightly pytorch build and efficientnet-pytorch, any chance you're using the same?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 850122,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-05-16T11:00:36.960000",
          "content": "<p>i use resnext...50\n1 ) Are you using  split methodology ,if yes then are you keeping constant no of split for all tiff  images . I use split method what i did in pre processing was  to generate the tiles for each image but each had different no of tiles ,in Test i used same pipeline except i maintained same no of tiles,not sure if this change is causing it </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 850443,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-05-16T16:23:06.503000",
          "content": "<p>finally resolved all the issues on my way to next level :)...\nIt was silly mistake which made a vast difference between Train and Test preprocessing</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "849822": "With `model.eval()` , batch normalization layers use fixed mean/std values calculated during training phase. Due to this, sometimes performance gets worse when there is difference between train and test batches. As far as I remember, the same occurred to me in [Recursion competition](https://www.kaggle.com/c/recursion-cellular-image-classification). ",
    "850077": "I had this the other day when using an efficientnet with reasonable batch sizes. Testing accuracy was diabolical on a randomly-split train/test dataset but was good if I didn't switch to eval() model. I gave up trying to figure it out, but I assumed I was doing something stupid. Absolutely bizarre. It's was the nightly pytorch build and efficientnet-pytorch, any chance you're using the same?",
    "849761": "I find strange issue as never before  .Adding model.eval( ) makes predictions worst compared to without it . I have been doing all this earlier with no issue ,as it is way as well when u have batch norm in place .\nI dont know what unknown mistake i could be making\n\n```\ngpu = torch.device(\"cuda:0\")\nstate_dict = torch.load('../work/xyz.pth',map_location=gpu)\nmodel=Model()\nmodel.load_state_dict(state_dict   )\nmodel.eval()\np=model.cuda()(x.half())\n```"
  }
}