{
  "id": 110857,
  "title": "Reproducibility of fastai",
  "url": "/competitions/rsna-intracranial-hemorrhage-detection/discussion/110857",
  "author_name": "kambarakun",
  "post_date": "2019-10-01T14:52:11.288000",
  "votes": 0,
  "comment_count": 8,
  "views": 0,
  "content": "<p><strong>I am not used to fastai so sorry for any rudimentary question. It may not be the content to be discussed in this discussion.</strong>  </p>\n\n<p>If you use the pre-trained models of ImageNet, I think we use <code>DataBunch.normalize()</code>, but \n reproducibility cannot be ensured. I thik this is caused by the effect applied to each batch. <br>\nI fix numpy random seed.  Is it the influence of the seed related to GPU? Or is it a problem caused by multiple cpu_workers?</p>\n\n<p>```</p>\n\n<h1>Pseudo code block</h1>\n\n<h1>I run these lines, everytime preds are different.</h1>\n\n<h1>I try delete normalize(), so preds are same.</h1>\n\n<p>np.random.seed(0)\nimg_list = ImageList.from_df(df_train, path_train)\ndata     = (img_list.split_by_idxs(index_train, index_valid)\n                    .label_from_df(label_delim=' ')\n                    .add_test(ImageList.from_df(df_test, path=path_test))\n                    .databunch(bs=256, num_workers=os.cpu_count())\n                    .normalize())\n...\nreturn learn.get_preds(DatasetType.Valid)\n```</p>",
  "messages": [
    {
      "id": 638413,
      "postDate": "2019-10-01T21:46:03.160Z",
      "content": "<p>this should help</p>\n\n<p><code>\ndef seed_everything(seed):\n    random.seed(seed)\n    os.environ['PYTHONHASHSEED'] = str(seed)\n    np.random.seed(seed)\n    torch.manual_seed(seed)\n    torch.cuda.manual_seed(seed)\n    torch.backends.cudnn.deterministic = True\n    torch.backends.cudnn.benchmark = False\nSEED = 1\nseed_everything(SEED)\n</code></p>",
      "rawMarkdown": "this should help\n\n\n```\ndef seed_everything(seed):\n    random.seed(seed)\n    os.environ['PYTHONHASHSEED'] = str(seed)\n    np.random.seed(seed)\n    torch.manual_seed(seed)\n    torch.cuda.manual_seed(seed)\n    torch.backends.cudnn.deterministic = True\n    torch.backends.cudnn.benchmark = False\nSEED = 1\nseed_everything(SEED)\n```",
      "votes": 5,
      "replies": [
        {
          "id": 638978,
          "postDate": "2019-10-02T16:32:19.453Z",
          "content": "<p>Thanks <a href=\"/valanm\">@valanm</a> for your reply! I tried in my local machine with 1080ti, but I got different results...\nI think this caused by multiprocessing or other processes or my GPU-seed trouble.</p>",
          "rawMarkdown": "Thanks @valanm for your reply! I tried in my local machine with 1080ti, but I got different results...\nI think this caused by multiprocessing or other processes or my GPU-seed trouble."
        }
      ]
    },
    {
      "id": 639963,
      "postDate": "2019-10-03T18:13:30.473Z",
      "content": "<p>See Val's post for setting appropriate seeds, though he missed <code>torch.backends.cudnn.benchmark = False</code> which is in the <a href=\"https://pytorch.org/docs/stable/notes/randomness.html\">PyTorch guide</a> on this.</p>\n\n<p>Passing no options to <code>normalize</code> will cause it to use the mean/std of a couple of batches and so cause non-determinism. Passing stats to use will avoid this.\nHowever the imagenet stats are not really appropriate. They are the mean and std of imagenet images and so will normalize them appropriately to create values of the appropriate range for pretrained models (mean of 0, std of 1, i.e. Z-scores). You want to use the mean and std of your images so they are scaled to the same range the saved model was trained on.\nSo you want to calculate the actual mean and std of your training data, or in this case probably a subset given the large size, and then pass that to <code>normalize</code>. This <a href=\"https://forums.fast.ai/t/calculating-our-own-image-stats-imagenet-stats-cifar-stats-etc/40355/5\">post</a> on the fastai forums contains a link to code I wrote for collecting stats (to do a subset just do something like <code>collect_stats(data.train_ds.x[:1000])</code>).</p>",
      "rawMarkdown": "See Val's post for setting appropriate seeds, though he missed `torch.backends.cudnn.benchmark = False` which is in the [PyTorch guide](https://pytorch.org/docs/stable/notes/randomness.html) on this.\n\nPassing no options to `normalize` will cause it to use the mean/std of a couple of batches and so cause non-determinism. Passing stats to use will avoid this.\nHowever the imagenet stats are not really appropriate. They are the mean and std of imagenet images and so will normalize them appropriately to create values of the appropriate range for pretrained models (mean of 0, std of 1, i.e. Z-scores). You want to use the mean and std of your images so they are scaled to the same range the saved model was trained on.\nSo you want to calculate the actual mean and std of your training data, or in this case probably a subset given the large size, and then pass that to `normalize`. This [post](https://forums.fast.ai/t/calculating-our-own-image-stats-imagenet-stats-cifar-stats-etc/40355/5) on the fastai forums contains a link to code I wrote for collecting stats (to do a subset just do something like `collect_stats(data.train_ds.x[:1000])`).",
      "votes": 1,
      "replies": [
        {
          "id": 640026,
          "postDate": "2019-10-03T18:56:34.527Z",
          "content": "<p>Thx. Fixed as suggested </p>",
          "rawMarkdown": "Thx. Fixed as suggested "
        },
        {
          "id": 642323,
          "postDate": "2019-10-05T22:08:44.223Z",
          "content": "<p>It works fine, thanks a lot!!</p>",
          "rawMarkdown": "It works fine, thanks a lot!!"
        }
      ]
    },
    {
      "id": 638123,
      "postDate": "2019-10-01T14:52:11.290Z",
      "content": "<p><strong>I am not used to fastai so sorry for any rudimentary question. It may not be the content to be discussed in this discussion.</strong>  </p>\n\n<p>If you use the pre-trained models of ImageNet, I think we use <code>DataBunch.normalize()</code>, but \n reproducibility cannot be ensured. I thik this is caused by the effect applied to each batch. <br>\nI fix numpy random seed.  Is it the influence of the seed related to GPU? Or is it a problem caused by multiple cpu_workers?</p>\n\n<p>```</p>\n\n<h1>Pseudo code block</h1>\n\n<h1>I run these lines, everytime preds are different.</h1>\n\n<h1>I try delete normalize(), so preds are same.</h1>\n\n<p>np.random.seed(0)\nimg_list = ImageList.from_df(df_train, path_train)\ndata     = (img_list.split_by_idxs(index_train, index_valid)\n                    .label_from_df(label_delim=' ')\n                    .add_test(ImageList.from_df(df_test, path=path_test))\n                    .databunch(bs=256, num_workers=os.cpu_count())\n                    .normalize())\n...\nreturn learn.get_preds(DatasetType.Valid)\n```</p>",
      "rawMarkdown": "**I am not used to fastai so sorry for any rudimentary question. It may not be the content to be discussed in this discussion.**  \n\nIf you use the pre-trained models of ImageNet, I think we use `DataBunch.normalize()`, but \n reproducibility cannot be ensured. I thik this is caused by the effect applied to each batch.  \nI fix numpy random seed.  Is it the influence of the seed related to GPU? Or is it a problem caused by multiple cpu_workers?\n\n```\n# Pseudo code block\n# I run these lines, everytime preds are different.\n# I try delete normalize(), so preds are same.\n\nnp.random.seed(0)\nimg_list = ImageList.from_df(df_train, path_train)\ndata     = (img_list.split_by_idxs(index_train, index_valid)\n                    .label_from_df(label_delim=' ')\n                    .add_test(ImageList.from_df(df_test, path=path_test))\n                    .databunch(bs=256, num_workers=os.cpu_count())\n                    .normalize())\n...\nreturn learn.get_preds(DatasetType.Valid)\n```"
    },
    {
      "id": 638279,
      "postDate": "2019-10-01T17:21:47.553Z",
      "content": "<p>Please, have a look here: <a href=\"https://docs.fast.ai/vision.data.html#ImageDataBunch.normalize\">https://docs.fast.ai/vision.data.html#ImageDataBunch.normalize</a>\n <code>\nAdd normalize transform using stats (defaults to DataBunch.batch_stats).\nIn the fast.ai library we have imagenet_stats, ... so we can add normalization easily with any of these datasets.\n</code>\nSo <code>.normalize(imagenet_stats)</code> should do the trick!</p>",
      "rawMarkdown": "Please, have a look here: https://docs.fast.ai/vision.data.html#ImageDataBunch.normalize\n ```\nAdd normalize transform using stats (defaults to DataBunch.batch_stats).\nIn the fast.ai library we have imagenet_stats, ... so we can add normalization easily with any of these datasets.\n```\nSo `.normalize(imagenet_stats)` should do the trick!",
      "votes": 1,
      "replies": [
        {
          "id": 638981,
          "postDate": "2019-10-02T16:35:12.290Z",
          "content": "<p>Thank you <a href=\"/micpie\">@micpie</a>! <br>\nAre there loss differences between no-normalize and <code>normalize()</code> and <code>normalize(imagenet_stats)</code> in your model?</p>",
          "rawMarkdown": "Thank you @micpie!  \nAre there loss differences between no-normalize and `normalize()` and `normalize(imagenet_stats)` in your model?"
        }
      ]
    },
    {
      "id": 638180,
      "postDate": "2019-10-01T15:46:35.153Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 638413,
      "author_name": "Miroslav Valan",
      "author_url": "",
      "post_date": "2019-10-01T21:46:03.160000",
      "content": "<p>this should help</p>\n\n<p><code>\ndef seed_everything(seed):\n    random.seed(seed)\n    os.environ['PYTHONHASHSEED'] = str(seed)\n    np.random.seed(seed)\n    torch.manual_seed(seed)\n    torch.cuda.manual_seed(seed)\n    torch.backends.cudnn.deterministic = True\n    torch.backends.cudnn.benchmark = False\nSEED = 1\nseed_everything(SEED)\n</code></p>",
      "votes": 5,
      "replies": [
        {
          "id": 638978,
          "author_name": "kambarakun",
          "author_url": "",
          "post_date": "2019-10-02T16:32:19.453000",
          "content": "<p>Thanks <a href=\"/valanm\">@valanm</a> for your reply! I tried in my local machine with 1080ti, but I got different results...\nI think this caused by multiprocessing or other processes or my GPU-seed trouble.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 639963,
      "author_name": "Thomas Brandon",
      "author_url": "",
      "post_date": "2019-10-03T18:13:30.473000",
      "content": "<p>See Val's post for setting appropriate seeds, though he missed <code>torch.backends.cudnn.benchmark = False</code> which is in the <a href=\"https://pytorch.org/docs/stable/notes/randomness.html\">PyTorch guide</a> on this.</p>\n\n<p>Passing no options to <code>normalize</code> will cause it to use the mean/std of a couple of batches and so cause non-determinism. Passing stats to use will avoid this.\nHowever the imagenet stats are not really appropriate. They are the mean and std of imagenet images and so will normalize them appropriately to create values of the appropriate range for pretrained models (mean of 0, std of 1, i.e. Z-scores). You want to use the mean and std of your images so they are scaled to the same range the saved model was trained on.\nSo you want to calculate the actual mean and std of your training data, or in this case probably a subset given the large size, and then pass that to <code>normalize</code>. This <a href=\"https://forums.fast.ai/t/calculating-our-own-image-stats-imagenet-stats-cifar-stats-etc/40355/5\">post</a> on the fastai forums contains a link to code I wrote for collecting stats (to do a subset just do something like <code>collect_stats(data.train_ds.x[:1000])</code>).</p>",
      "votes": 1,
      "replies": [
        {
          "id": 640026,
          "author_name": "Miroslav Valan",
          "author_url": "",
          "post_date": "2019-10-03T18:56:34.527000",
          "content": "<p>Thx. Fixed as suggested </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 642323,
          "author_name": "kambarakun",
          "author_url": "",
          "post_date": "2019-10-05T22:08:44.223000",
          "content": "<p>It works fine, thanks a lot!!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 638279,
      "author_name": "Michael Pieler",
      "author_url": "",
      "post_date": "2019-10-01T17:21:47.553000",
      "content": "<p>Please, have a look here: <a href=\"https://docs.fast.ai/vision.data.html#ImageDataBunch.normalize\">https://docs.fast.ai/vision.data.html#ImageDataBunch.normalize</a>\n <code>\nAdd normalize transform using stats (defaults to DataBunch.batch_stats).\nIn the fast.ai library we have imagenet_stats, ... so we can add normalization easily with any of these datasets.\n</code>\nSo <code>.normalize(imagenet_stats)</code> should do the trick!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 638981,
          "author_name": "kambarakun",
          "author_url": "",
          "post_date": "2019-10-02T16:35:12.290000",
          "content": "<p>Thank you <a href=\"/micpie\">@micpie</a>! <br>\nAre there loss differences between no-normalize and <code>normalize()</code> and <code>normalize(imagenet_stats)</code> in your model?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 638180,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-10-01T15:46:35.153000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "638413": "this should help\n\n\n```\ndef seed_everything(seed):\n    random.seed(seed)\n    os.environ['PYTHONHASHSEED'] = str(seed)\n    np.random.seed(seed)\n    torch.manual_seed(seed)\n    torch.cuda.manual_seed(seed)\n    torch.backends.cudnn.deterministic = True\n    torch.backends.cudnn.benchmark = False\nSEED = 1\nseed_everything(SEED)\n```",
    "639963": "See Val's post for setting appropriate seeds, though he missed `torch.backends.cudnn.benchmark = False` which is in the [PyTorch guide](https://pytorch.org/docs/stable/notes/randomness.html) on this.\n\nPassing no options to `normalize` will cause it to use the mean/std of a couple of batches and so cause non-determinism. Passing stats to use will avoid this.\nHowever the imagenet stats are not really appropriate. They are the mean and std of imagenet images and so will normalize them appropriately to create values of the appropriate range for pretrained models (mean of 0, std of 1, i.e. Z-scores). You want to use the mean and std of your images so they are scaled to the same range the saved model was trained on.\nSo you want to calculate the actual mean and std of your training data, or in this case probably a subset given the large size, and then pass that to `normalize`. This [post](https://forums.fast.ai/t/calculating-our-own-image-stats-imagenet-stats-cifar-stats-etc/40355/5) on the fastai forums contains a link to code I wrote for collecting stats (to do a subset just do something like `collect_stats(data.train_ds.x[:1000])`).",
    "638123": "**I am not used to fastai so sorry for any rudimentary question. It may not be the content to be discussed in this discussion.**  \n\nIf you use the pre-trained models of ImageNet, I think we use `DataBunch.normalize()`, but \n reproducibility cannot be ensured. I thik this is caused by the effect applied to each batch.  \nI fix numpy random seed.  Is it the influence of the seed related to GPU? Or is it a problem caused by multiple cpu_workers?\n\n```\n# Pseudo code block\n# I run these lines, everytime preds are different.\n# I try delete normalize(), so preds are same.\n\nnp.random.seed(0)\nimg_list = ImageList.from_df(df_train, path_train)\ndata     = (img_list.split_by_idxs(index_train, index_valid)\n                    .label_from_df(label_delim=' ')\n                    .add_test(ImageList.from_df(df_test, path=path_test))\n                    .databunch(bs=256, num_workers=os.cpu_count())\n                    .normalize())\n...\nreturn learn.get_preds(DatasetType.Valid)\n```",
    "638279": "Please, have a look here: https://docs.fast.ai/vision.data.html#ImageDataBunch.normalize\n ```\nAdd normalize transform using stats (defaults to DataBunch.batch_stats).\nIn the fast.ai library we have imagenet_stats, ... so we can add normalization easily with any of these datasets.\n```\nSo `.normalize(imagenet_stats)` should do the trick!",
    "638180": ""
  }
}