{
  "id": 34514,
  "title": "MXNet resnet-152 fine-tune baseline (pub LB 0.117, no ensemble)",
  "url": "/competitions/inaturalist-challenge-at-fgvc-2017/discussion/34514",
  "author_name": "phunter",
  "post_date": "2017-06-10T19:21:26.527000",
  "votes": 9,
  "comment_count": 12,
  "views": 0,
  "content": "<p>update: added <code>.rec</code> format generator script, see the updated README.md at github repo. using <code>.rec</code> may help speed up IO and data augmentation.</p>\n\n<p>I have prepared for a simple baseline model by fine-tuning a pretrained ResNet 152 layers (Imagenet 11k) from MXNet's model zoo. Github link is at <a href=\"https://github.com/phunterlau/iNaturalist\">https://github.com/phunterlau/iNaturalist</a> with detailed instruction. Its public LB score is 0.117 from a single epoch without ensemble on models or n-crop on inference.</p>\n\n<p>To run this script, one may need a kind of decent machine, e.g. 4 GPUs. I can reach about 70-80 images per second on a 4 GTX 1080 machine when all data augmentation options are on. Please tune up the parameters in <code>run.sh</code> for your machine.</p>\n\n<p>As always, MXNet is a great deep learning framework, please star and fork : <a href=\"https://github.com/dmlc/mxnet\">https://github.com/dmlc/mxnet</a> </p>",
  "messages": [
    {
      "id": 191541,
      "postDate": "2017-06-10T19:21:26.527Z",
      "content": "<p>update: added <code>.rec</code> format generator script, see the updated README.md at github repo. using <code>.rec</code> may help speed up IO and data augmentation.</p>\n\n<p>I have prepared for a simple baseline model by fine-tuning a pretrained ResNet 152 layers (Imagenet 11k) from MXNet's model zoo. Github link is at <a href=\"https://github.com/phunterlau/iNaturalist\">https://github.com/phunterlau/iNaturalist</a> with detailed instruction. Its public LB score is 0.117 from a single epoch without ensemble on models or n-crop on inference.</p>\n\n<p>To run this script, one may need a kind of decent machine, e.g. 4 GPUs. I can reach about 70-80 images per second on a 4 GTX 1080 machine when all data augmentation options are on. Please tune up the parameters in <code>run.sh</code> for your machine.</p>\n\n<p>As always, MXNet is a great deep learning framework, please star and fork : <a href=\"https://github.com/dmlc/mxnet\">https://github.com/dmlc/mxnet</a> </p>",
      "rawMarkdown": "update: added `.rec` format generator script, see the updated README.md at github repo. using `.rec` may help speed up IO and data augmentation.\n\nI have prepared for a simple baseline model by fine-tuning a pretrained ResNet 152 layers (Imagenet 11k) from MXNet's model zoo. Github link is at [https://github.com/phunterlau/iNaturalist][1] with detailed instruction. Its public LB score is 0.117 from a single epoch without ensemble on models or n-crop on inference.\n\nTo run this script, one may need a kind of decent machine, e.g. 4 GPUs. I can reach about 70-80 images per second on a 4 GTX 1080 machine when all data augmentation options are on. Please tune up the parameters in `run.sh` for your machine.\n\nAs always, MXNet is a great deep learning framework, please star and fork : [https://github.com/dmlc/mxnet][2] \n\n\n  [1]: https://github.com/phunterlau/iNaturalist\n  [2]: https://github.com/dmlc/mxnet",
      "votes": 9
    },
    {
      "id": 198318,
      "postDate": "2017-07-01T22:00:56.903Z",
      "content": "<p>Just found this fact: this baseline output can produce a top 10 result on the current public LB without any modification. Just run it and you can get to top 10, let's do it!</p>",
      "rawMarkdown": "Just found this fact: this baseline output can produce a top 10 result on the current public LB without any modification. Just run it and you can get to top 10, let's do it!",
      "votes": 1,
      "replies": [
        {
          "id": 200445,
          "postDate": "2017-07-07T23:08:27.850Z",
          "content": "<p>Thank you for providing this baseline, actually the best single epoch(pub lb 0.087) I have is from your code. </p>",
          "rawMarkdown": "Thank you for providing this baseline, actually the best single epoch(pub lb 0.087) I have is from your code. ",
          "votes": 1
        },
        {
          "id": 200446,
          "postDate": "2017-07-07T23:24:11.767Z",
          "content": "<p>Thanks for using MXNet...looking forward to your report about how to get to 0.087, honestly I was surprised since I was not able to get that but switched to an alternative method :-(\n<img src=\"https://s-media-cache-ak0.pinimg.com/736x/32/ba/19/32ba192e9eb99068a91fd2e8a1e21189--vet-office-funny-ideas.jpg\" alt=\"em?\" title=\"\"></p>",
          "rawMarkdown": "Thanks for using MXNet...looking forward to your report about how to get to 0.087, honestly I was surprised since I was not able to get that but switched to an alternative method :-(\n![em?][1]\n\n  [1]: https://s-media-cache-ak0.pinimg.com/736x/32/ba/19/32ba192e9eb99068a91fd2e8a1e21189--vet-office-funny-ideas.jpg",
          "votes": 1
        },
        {
          "id": 200586,
          "postDate": "2017-07-08T15:29:37.687Z",
          "content": "<p>As a fact, I found a single epoch with single pub lb 0.8384 when I tried to reproduce this result.\nI don't have enough resource to do comparison test of which changes/changes are effective.</p>\n\n<p><a href=\"https://drive.google.com/open?id=0B7-6AXcnzhEFX1NINWNWaGZyTm8\">These coefficients</a> are trained after 38 + 22 epochs,  First 38 epochs are trained on given split of training data. Then the resulting net work without last fully connected layer was used as initial values and then trained for 22 epochs(which was just finished yesterday).\nThe crop value is changes to 333 and batch size is set to 36.\nThe interesting part seems to be in the sub.py, if img_sz and  crop_sz are set to 360 then the result got 0.87 public lb. When I tried to reproduce today I set them to 499 and got 0.8384 public  lb(0.84 private lb),  but I haven't found an explanation yet. </p>",
          "rawMarkdown": "As a fact, I found a single epoch with single pub lb 0.8384 when I tried to reproduce this result.\nI don't have enough resource to do comparison test of which changes/changes are effective.\n\n[These coefficients][1] are trained after 38 + 22 epochs,  First 38 epochs are trained on given split of training data. Then the resulting net work without last fully connected layer was used as initial values and then trained for 22 epochs(which was just finished yesterday).\nThe crop value is changes to 333 and batch size is set to 36.\nThe interesting part seems to be in the sub.py, if img_sz and  crop_sz are set to 360 then the result got 0.87 public lb. When I tried to reproduce today I set them to 499 and got 0.8384 public  lb(0.84 private lb),  but I haven't found an explanation yet. \n\n\n  [1]: https://drive.google.com/open?id=0B7-6AXcnzhEFX1NINWNWaGZyTm8",
          "votes": 1
        },
        {
          "id": 200689,
          "postDate": "2017-07-08T22:30:05.757Z",
          "content": "<p>Nice finding of going further than 30 epochs: I thought 30 was quite enough so I just stopped at 30th epoch and picked 30th + 23rd</p>",
          "rawMarkdown": "Nice finding of going further than 30 epochs: I thought 30 was quite enough so I just stopped at 30th epoch and picked 30th + 23rd"
        }
      ]
    },
    {
      "id": 194279,
      "postDate": "2017-06-20T00:05:58.847Z",
      "content": "<p>Hi, </p>\n\n<p>Have you tried to compare with the same network but pretrained on Imagenet 1k ? Does it make any difference?</p>",
      "rawMarkdown": "Hi, \n\nHave you tried to compare with the same network but pretrained on Imagenet 1k ? Does it make any difference?",
      "replies": [
        {
          "id": 194286,
          "postDate": "2017-06-20T00:23:06.613Z",
          "content": "<p>Imagenet 1k pretrain has some less prediction power, e.g. 0.14x vs 0.11x on LB score. But it might be a good back up for ensemble.</p>",
          "rawMarkdown": "Imagenet 1k pretrain has some less prediction power, e.g. 0.14x vs 0.11x on LB score. But it might be a good back up for ensemble."
        }
      ]
    },
    {
      "id": 193559,
      "postDate": "2017-06-16T22:15:51.767Z",
      "content": "<p>Hi phunter.</p>\n\n<p>Thanks for opening up the code.</p>\n\n<p>I actually tried to use the .rec format before your update, but didn't get any performance enhancements, since the bottleneck was already in the GPU. My guess is that for people with multiple GPUs, or slower IO than a normal HDD, using recordIO could be a benefit. What are your thoughts?</p>",
      "rawMarkdown": "Hi phunter.\n\nThanks for opening up the code.\n\nI actually tried to use the .rec format before your update, but didn't get any performance enhancements, since the bottleneck was already in the GPU. My guess is that for people with multiple GPUs, or slower IO than a normal HDD, using recordIO could be a benefit. What are your thoughts?",
      "replies": [
        {
          "id": 193562,
          "postDate": "2017-06-16T22:24:53.433Z",
          "content": "<p>That is true. My SSD controller may have some problem of loading data to CPU: CPUs can't have full load and GPUs spent much time waiting for data.</p>\n\n<p>P.S. I have tried resnet-50 pretrained model from <a href=\"http://data.mxnet.io/models/imagenet-11k-place365-ch/\">http://data.mxnet.io/models/imagenet-11k-place365-ch/</a> it has about double the speed of resnet 152, costs much less memory, and a single crop on 21 epoch can have 0.144, not bad at all. You can have a try with ResNet 50 for a quick result.</p>",
          "rawMarkdown": "That is true. My SSD controller may have some problem of loading data to CPU: CPUs can't have full load and GPUs spent much time waiting for data.\n\nP.S. I have tried resnet-50 pretrained model from http://data.mxnet.io/models/imagenet-11k-place365-ch/ it has about double the speed of resnet 152, costs much less memory, and a single crop on 21 epoch can have 0.144, not bad at all. You can have a try with ResNet 50 for a quick result."
        }
      ]
    },
    {
      "id": 193298,
      "postDate": "2017-06-16T01:38:36.970Z",
      "content": "<p>update: added .rec format generator script, see the updated README.md at github repo. using .rec may help speed up IO and data augmentation.</p>",
      "rawMarkdown": "update: added .rec format generator script, see the updated README.md at github repo. using .rec may help speed up IO and data augmentation."
    },
    {
      "id": 191557,
      "postDate": "2017-06-10T20:38:40.473Z",
      "content": "<p>Nice one  !!!\nOne thing I need to know that how much time this code take while running Resnet 152  on single GTX 1080 with the given training dataset.</p>",
      "rawMarkdown": "Nice one  !!!\nOne thing I need to know that how much time this code take while running Resnet 152  on single GTX 1080 with the given training dataset.",
      "replies": [
        {
          "id": 191587,
          "postDate": "2017-06-10T23:49:43.490Z",
          "content": "<p>A fully used GTX 1080 can have 20-25 images per second if disk IO is good, considering the sample size is 579184, it needs about 6.5 hours per epoch :-) I have Couple of tips to speed up on a single card:</p>\n\n<ol>\n<li>the current crop size is 320x320. reducing to 256x256 may hurt a little bit precision but can speed up much.</li>\n<li>I turned on all data augmentation by default. Choosing a good set of data aug operations may help precision while saving run time. Ref <a href=\"http://mxnet.io/api/python/io.html\">http://mxnet.io/api/python/io.html</a></li>\n<li>Try other models at <a href=\"http://data.mxnet.io/models/\">http://data.mxnet.io/models/</a> : ResNet 152 on 11k Imagenet is a very deep model. Other models like inception, ResNet 101 or 50 can be faster.</li>\n</ol>",
          "rawMarkdown": "A fully used GTX 1080 can have 20-25 images per second if disk IO is good, considering the sample size is 579184, it needs about 6.5 hours per epoch :-) I have Couple of tips to speed up on a single card:\n\n1. the current crop size is 320x320. reducing to 256x256 may hurt a little bit precision but can speed up much.\n2. I turned on all data augmentation by default. Choosing a good set of data aug operations may help precision while saving run time. Ref http://mxnet.io/api/python/io.html\n3. Try other models at http://data.mxnet.io/models/ : ResNet 152 on 11k Imagenet is a very deep model. Other models like inception, ResNet 101 or 50 can be faster.",
          "votes": 2
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 198318,
      "author_name": "phunter",
      "author_url": "",
      "post_date": "2017-07-01T22:00:56.903000",
      "content": "<p>Just found this fact: this baseline output can produce a top 10 result on the current public LB without any modification. Just run it and you can get to top 10, let's do it!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 200445,
          "author_name": "zz",
          "author_url": "",
          "post_date": "2017-07-07T23:08:27.850000",
          "content": "<p>Thank you for providing this baseline, actually the best single epoch(pub lb 0.087) I have is from your code. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 200446,
          "author_name": "phunter",
          "author_url": "",
          "post_date": "2017-07-07T23:24:11.767000",
          "content": "<p>Thanks for using MXNet...looking forward to your report about how to get to 0.087, honestly I was surprised since I was not able to get that but switched to an alternative method :-(\n<img src=\"https://s-media-cache-ak0.pinimg.com/736x/32/ba/19/32ba192e9eb99068a91fd2e8a1e21189--vet-office-funny-ideas.jpg\" alt=\"em?\" title=\"\"></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 200586,
          "author_name": "zz",
          "author_url": "",
          "post_date": "2017-07-08T15:29:37.687000",
          "content": "<p>As a fact, I found a single epoch with single pub lb 0.8384 when I tried to reproduce this result.\nI don't have enough resource to do comparison test of which changes/changes are effective.</p>\n\n<p><a href=\"https://drive.google.com/open?id=0B7-6AXcnzhEFX1NINWNWaGZyTm8\">These coefficients</a> are trained after 38 + 22 epochs,  First 38 epochs are trained on given split of training data. Then the resulting net work without last fully connected layer was used as initial values and then trained for 22 epochs(which was just finished yesterday).\nThe crop value is changes to 333 and batch size is set to 36.\nThe interesting part seems to be in the sub.py, if img_sz and  crop_sz are set to 360 then the result got 0.87 public lb. When I tried to reproduce today I set them to 499 and got 0.8384 public  lb(0.84 private lb),  but I haven't found an explanation yet. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 200689,
          "author_name": "phunter",
          "author_url": "",
          "post_date": "2017-07-08T22:30:05.757000",
          "content": "<p>Nice finding of going further than 30 epochs: I thought 30 was quite enough so I just stopped at 30th epoch and picked 30th + 23rd</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 194279,
      "author_name": "Artem.Sanakoev",
      "author_url": "",
      "post_date": "2017-06-20T00:05:58.847000",
      "content": "<p>Hi, </p>\n\n<p>Have you tried to compare with the same network but pretrained on Imagenet 1k ? Does it make any difference?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 194286,
          "author_name": "phunter",
          "author_url": "",
          "post_date": "2017-06-20T00:23:06.613000",
          "content": "<p>Imagenet 1k pretrain has some less prediction power, e.g. 0.14x vs 0.11x on LB score. But it might be a good back up for ensemble.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 193559,
      "author_name": "Gyuri Im",
      "author_url": "",
      "post_date": "2017-06-16T22:15:51.767000",
      "content": "<p>Hi phunter.</p>\n\n<p>Thanks for opening up the code.</p>\n\n<p>I actually tried to use the .rec format before your update, but didn't get any performance enhancements, since the bottleneck was already in the GPU. My guess is that for people with multiple GPUs, or slower IO than a normal HDD, using recordIO could be a benefit. What are your thoughts?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 193562,
          "author_name": "phunter",
          "author_url": "",
          "post_date": "2017-06-16T22:24:53.433000",
          "content": "<p>That is true. My SSD controller may have some problem of loading data to CPU: CPUs can't have full load and GPUs spent much time waiting for data.</p>\n\n<p>P.S. I have tried resnet-50 pretrained model from <a href=\"http://data.mxnet.io/models/imagenet-11k-place365-ch/\">http://data.mxnet.io/models/imagenet-11k-place365-ch/</a> it has about double the speed of resnet 152, costs much less memory, and a single crop on 21 epoch can have 0.144, not bad at all. You can have a try with ResNet 50 for a quick result.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 193298,
      "author_name": "phunter",
      "author_url": "",
      "post_date": "2017-06-16T01:38:36.970000",
      "content": "<p>update: added .rec format generator script, see the updated README.md at github repo. using .rec may help speed up IO and data augmentation.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 191557,
      "author_name": "Rishikesh",
      "author_url": "",
      "post_date": "2017-06-10T20:38:40.473000",
      "content": "<p>Nice one  !!!\nOne thing I need to know that how much time this code take while running Resnet 152  on single GTX 1080 with the given training dataset.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 191587,
          "author_name": "phunter",
          "author_url": "",
          "post_date": "2017-06-10T23:49:43.490000",
          "content": "<p>A fully used GTX 1080 can have 20-25 images per second if disk IO is good, considering the sample size is 579184, it needs about 6.5 hours per epoch :-) I have Couple of tips to speed up on a single card:</p>\n\n<ol>\n<li>the current crop size is 320x320. reducing to 256x256 may hurt a little bit precision but can speed up much.</li>\n<li>I turned on all data augmentation by default. Choosing a good set of data aug operations may help precision while saving run time. Ref <a href=\"http://mxnet.io/api/python/io.html\">http://mxnet.io/api/python/io.html</a></li>\n<li>Try other models at <a href=\"http://data.mxnet.io/models/\">http://data.mxnet.io/models/</a> : ResNet 152 on 11k Imagenet is a very deep model. Other models like inception, ResNet 101 or 50 can be faster.</li>\n</ol>",
          "votes": 2,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "191541": "update: added `.rec` format generator script, see the updated README.md at github repo. using `.rec` may help speed up IO and data augmentation.\n\nI have prepared for a simple baseline model by fine-tuning a pretrained ResNet 152 layers (Imagenet 11k) from MXNet's model zoo. Github link is at [https://github.com/phunterlau/iNaturalist][1] with detailed instruction. Its public LB score is 0.117 from a single epoch without ensemble on models or n-crop on inference.\n\nTo run this script, one may need a kind of decent machine, e.g. 4 GPUs. I can reach about 70-80 images per second on a 4 GTX 1080 machine when all data augmentation options are on. Please tune up the parameters in `run.sh` for your machine.\n\nAs always, MXNet is a great deep learning framework, please star and fork : [https://github.com/dmlc/mxnet][2] \n\n\n  [1]: https://github.com/phunterlau/iNaturalist\n  [2]: https://github.com/dmlc/mxnet",
    "198318": "Just found this fact: this baseline output can produce a top 10 result on the current public LB without any modification. Just run it and you can get to top 10, let's do it!",
    "194279": "Hi, \n\nHave you tried to compare with the same network but pretrained on Imagenet 1k ? Does it make any difference?",
    "193559": "Hi phunter.\n\nThanks for opening up the code.\n\nI actually tried to use the .rec format before your update, but didn't get any performance enhancements, since the bottleneck was already in the GPU. My guess is that for people with multiple GPUs, or slower IO than a normal HDD, using recordIO could be a benefit. What are your thoughts?",
    "193298": "update: added .rec format generator script, see the updated README.md at github repo. using .rec may help speed up IO and data augmentation.",
    "191557": "Nice one  !!!\nOne thing I need to know that how much time this code take while running Resnet 152  on single GTX 1080 with the given training dataset."
  }
}