{
  "id": 35960,
  "title": "Competition Wrap-Up",
  "url": "/competitions/inaturalist-challenge-at-fgvc-2017/discussion/35960",
  "author_name": "macaodha",
  "post_date": "2017-07-08T00:24:06.401000",
  "votes": 4,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Congratulations to all the competitors! We are very impressed by the quality of the results. We are going to take some time to go though the submissions and will soon post the final leader-board.</p>\n\n<p>In the mean time we would ask competitors to reply to this thread with a description of their method, specify if they used any additional annotations during training, and mention if the team consists of industry, university, or individual competitors.</p>\n\n<p>Thanks and well done!</p>",
  "messages": [
    {
      "id": 201561,
      "postDate": "2017-07-11T03:51:38.157Z",
      "content": "<p>Sadly, I did not spend that much time on this problem as I would love. My final submission was based on the mxnet code provided by @phunter and was just an average of fine-tuned Resnet 152 and Resnext 101. I also used train and test time augmentations.</p>\n\n<p>It was the first time when the size of the input images was that important. I would believe that the reason is that in this competition the difference between some classes was very subtle and downsampling original images to something like standard 224x224 would make many important details almost invisible. I would believe (although I did not check) that more careful fine tuning on something like 480x480 input images will lead in the 0.5-0.6 top5 error range.   </p>\n\n<p>I am really grateful to admins that organized this problem. The dataset is amazing.</p>\n\n<p>Sadly there was no monetary prize or at least Kaggle points for this problem and this pushed nearly all competitors toward other problems. Plus 1 month for a competition with the dataset of this size is a bit too short. I would prefer if this problem was going for 3 months, had standard $100,000 prize, and would award Kaggle points. My guess that in this case, we would have hundreds of competitors with all creativity that Kagglers are famous for bringing new insights algorithms and approaches to this task. </p>",
      "rawMarkdown": "Sadly, I did not spend that much time on this problem as I would love. My final submission was based on the mxnet code provided by @phunter and was just an average of fine-tuned Resnet 152 and Resnext 101. I also used train and test time augmentations.\n\nIt was the first time when the size of the input images was that important. I would believe that the reason is that in this competition the difference between some classes was very subtle and downsampling original images to something like standard 224x224 would make many important details almost invisible. I would believe (although I did not check) that more careful fine tuning on something like 480x480 input images will lead in the 0.5-0.6 top5 error range.   \n\nI am really grateful to admins that organized this problem. The dataset is amazing.\n\nSadly there was no monetary prize or at least Kaggle points for this problem and this pushed nearly all competitors toward other problems. Plus 1 month for a competition with the dataset of this size is a bit too short. I would prefer if this problem was going for 3 months, had standard $100,000 prize, and would award Kaggle points. My guess that in this case, we would have hundreds of competitors with all creativity that Kagglers are famous for bringing new insights algorithms and approaches to this task. ",
      "votes": 5
    },
    {
      "id": 201114,
      "postDate": "2017-07-10T13:26:14.733Z",
      "content": "<p>Team \"MunMM\" tried a number of 'elaborate' techniques employing the implicit hierarchy and attention but unfortunately none of these were ever successful. We tried mapping the problem to an \"image captioning\" approach where the output would be two-word sequences consisting of first the super category and then the specific category (basically \"show, attend, and tell\" where the labels would be \"[SUPERCAT] [SPECIFIC CAT] [END SEQ]\"). The idea was that the attention model would learn to recognize the super category (putting emphasis on those parts of the image) and then the RNN would help to constrain the problem to the smaller set of classes within that super category for predicting the specific category (and of course the attention would then direct the model to the parts of the image that were discriminative for differentiating more specific species within that super category). This never gave us any meaningfully good results and in retrospect it was probably because it didn't have a strong enough base model. We only used VGGNet (like in the original paper) and ResNet-18 with fine-tuning which clearly wasn't expressive enough for the problem. It might be interesting to first have a high quality InceptionV3 or Resnet152 model fine-tuned (or trained from scratch -- see below) as the encoder and then perhaps due some very light fine-tuning (or none at all) for the main model but then again the whole RNN might just be overkill.\nWe also tried several variants of \"MixDCNN\" with different numbers of \"sub\"-convnets and freezing different amounts/layers of the model to be able to both share weights optimally and fit everything into memory. Here we had slightly more success than the above captioning-based approaches but still nothing that would get us anywhere close to the basic approach of fine-tuning powerful ImageNet-type convnets.\nIn the end, our score of ~0.011 came from an Ensemble of 3 InceptionV3 models that were all trained slightly differently. Two were done using fine-tuning (both in different toolkits -- CNTK and PyTorch, just for fun :-)) and the other was InceptionV3 trained from scratch. This latter model was our best performing stand-alone model but we just ran out of time to make it even better. The model is still training now and the validation error is continuing to improve with top-5 now under 10% error. With ~600,000 training images, we are close to the size of ImageNet and it looks like training-from-scratch -- while slow (especially with only 1 GPU) -- can work quite nicely for this problem. We were interested to read about the success of fine-tuning with the ImageNet\"11k\" ResNet152 model in mxnet, but didn't have the resources to try it out ourselves.</p>\n\n<p>We are in industry but competed just for fun. We did not use any additional annotations. We are very interested/excited to hear about the techniques that led to the very impressive scores that others in the competition were able to achieve. Thanks for organizing this competition!</p>",
      "rawMarkdown": "Team \"MunMM\" tried a number of 'elaborate' techniques employing the implicit hierarchy and attention but unfortunately none of these were ever successful. We tried mapping the problem to an \"image captioning\" approach where the output would be two-word sequences consisting of first the super category and then the specific category (basically \"show, attend, and tell\" where the labels would be \"[SUPERCAT] [SPECIFIC CAT] [END SEQ]\"). The idea was that the attention model would learn to recognize the super category (putting emphasis on those parts of the image) and then the RNN would help to constrain the problem to the smaller set of classes within that super category for predicting the specific category (and of course the attention would then direct the model to the parts of the image that were discriminative for differentiating more specific species within that super category). This never gave us any meaningfully good results and in retrospect it was probably because it didn't have a strong enough base model. We only used VGGNet (like in the original paper) and ResNet-18 with fine-tuning which clearly wasn't expressive enough for the problem. It might be interesting to first have a high quality InceptionV3 or Resnet152 model fine-tuned (or trained from scratch -- see below) as the encoder and then perhaps due some very light fine-tuning (or none at all) for the main model but then again the whole RNN might just be overkill.\nWe also tried several variants of \"MixDCNN\" with different numbers of \"sub\"-convnets and freezing different amounts/layers of the model to be able to both share weights optimally and fit everything into memory. Here we had slightly more success than the above captioning-based approaches but still nothing that would get us anywhere close to the basic approach of fine-tuning powerful ImageNet-type convnets.\nIn the end, our score of ~0.011 came from an Ensemble of 3 InceptionV3 models that were all trained slightly differently. Two were done using fine-tuning (both in different toolkits -- CNTK and PyTorch, just for fun :-)) and the other was InceptionV3 trained from scratch. This latter model was our best performing stand-alone model but we just ran out of time to make it even better. The model is still training now and the validation error is continuing to improve with top-5 now under 10% error. With ~600,000 training images, we are close to the size of ImageNet and it looks like training-from-scratch -- while slow (especially with only 1 GPU) -- can work quite nicely for this problem. We were interested to read about the success of fine-tuning with the ImageNet\"11k\" ResNet152 model in mxnet, but didn't have the resources to try it out ourselves.\n\nWe are in industry but competed just for fun. We did not use any additional annotations. We are very interested/excited to hear about the techniques that led to the very impressive scores that others in the competition were able to achieve. Thanks for organizing this competition!",
      "votes": 4
    },
    {
      "id": 200453,
      "postDate": "2017-07-08T00:24:06.400Z",
      "content": "<p>Congratulations to all the competitors! We are very impressed by the quality of the results. We are going to take some time to go though the submissions and will soon post the final leader-board.</p>\n\n<p>In the mean time we would ask competitors to reply to this thread with a description of their method, specify if they used any additional annotations during training, and mention if the team consists of industry, university, or individual competitors.</p>\n\n<p>Thanks and well done!</p>",
      "rawMarkdown": "Congratulations to all the competitors! We are very impressed by the quality of the results. We are going to take some time to go though the submissions and will soon post the final leader-board.\n\nIn the mean time we would ask competitors to reply to this thread with a description of their method, specify if they used any additional annotations during training, and mention if the team consists of industry, university, or individual competitors.\n\nThanks and well done!",
      "votes": 4
    },
    {
      "id": 202253,
      "postDate": "2017-07-12T05:49:25.153Z",
      "rawMarkdown": ""
    },
    {
      "id": 202252,
      "postDate": "2017-07-12T05:49:16.803Z",
      "content": "<p>First, thank you to the people who organized this competition along with the (huge) dataset!</p>\n\n<p>My method was quite simple, finetuning Inception-v3 that was pretrained in ImageNet.</p>\n\n<p>I wanted to explore various data augmentation strategies, and found the <a href=\"https://github.com/aleju/imgaug\">imgaug library</a> helpful in getting a few more percentages in the leaderboard.</p>\n\n<p>I also noticed that some species, such as Eacles Imperialis, contain both the larvae state and the full grown state, which have completely different characteristics. Therefore I separated the images into different lables, and combined them before submission.</p>\n\n<p>While most of the image sizes were either 600x800 or 800x600, there were some images that had disproportionate width/height ratios. In extreme cases I tried to preserve the image ratio, both in train and test time, but am unsure if this was helpful.</p>",
      "rawMarkdown": "First, thank you to the people who organized this competition along with the (huge) dataset!\n\nMy method was quite simple, finetuning Inception-v3 that was pretrained in ImageNet.\n\nI wanted to explore various data augmentation strategies, and found the [imgaug library][1] helpful in getting a few more percentages in the leaderboard.\n\nI also noticed that some species, such as Eacles Imperialis, contain both the larvae state and the full grown state, which have completely different characteristics. Therefore I separated the images into different lables, and combined them before submission.\n\nWhile most of the image sizes were either 600x800 or 800x600, there were some images that had disproportionate width/height ratios. In extreme cases I tried to preserve the image ratio, both in train and test time, but am unsure if this was helpful.\n\n\n  [1]: https://github.com/aleju/imgaug",
      "replies": [
        {
          "id": 202496,
          "postDate": "2017-07-12T16:28:41.457Z",
          "content": "<p>Out of interest, what percentage performance increase did you get when including the augmentations?</p>",
          "rawMarkdown": "Out of interest, what percentage performance increase did you get when including the augmentations?"
        },
        {
          "id": 202733,
          "postDate": "2017-07-13T06:23:42.587Z",
          "content": "<p>It's difficult to be exact, but around 1.5% ~ 2%.</p>",
          "rawMarkdown": "It's difficult to be exact, but around 1.5% ~ 2%.",
          "votes": 1
        },
        {
          "id": 207737,
          "postDate": "2017-07-27T10:25:53.227Z",
          "content": "<p>Thank you Gyuri! If I may ask, what is your conclusion about the best augmentation methods? And what is your validation/test score? </p>",
          "rawMarkdown": "Thank you Gyuri! If I may ask, what is your conclusion about the best augmentation methods? And what is your validation/test score? "
        }
      ]
    },
    {
      "id": 200468,
      "postDate": "2017-07-08T02:38:57.250Z",
      "content": "<p>I was attempting to train a faster-rcnn network (<a href=\"https://arxiv.org/abs/1506.01497\">https://arxiv.org/abs/1506.01497</a>) on the problem. I started by generating region proposals using a sliding window/pyramid and a well trained (80%+ validation accuracy) inception network. Spot checking the proposed RoIs showed that they were not <em>perfect</em>, but in tests on a small subset of the full 5k problem, the faster-rcnn approach worked quite well using these proposals to train on. Unfortunately I didn't have enough time to adapt the existing frcnn implementations to run on multiple GPUs and it simply did not progress fast enough on a single GPU (even with weeks of training) to converge on a satisfactory solution for the full 5k problem. I suspect that the existing architecture needs some tweaking to better perform on a problem with so many classes.</p>",
      "rawMarkdown": "I was attempting to train a faster-rcnn network (https://arxiv.org/abs/1506.01497) on the problem. I started by generating region proposals using a sliding window/pyramid and a well trained (80%+ validation accuracy) inception network. Spot checking the proposed RoIs showed that they were not *perfect*, but in tests on a small subset of the full 5k problem, the faster-rcnn approach worked quite well using these proposals to train on. Unfortunately I didn't have enough time to adapt the existing frcnn implementations to run on multiple GPUs and it simply did not progress fast enough on a single GPU (even with weeks of training) to converge on a satisfactory solution for the full 5k problem. I suspect that the existing architecture needs some tweaking to better perform on a problem with so many classes."
    }
  ],
  "comments": [
    {
      "id": 201561,
      "author_name": "Vladimir Iglovikov",
      "author_url": "",
      "post_date": "2017-07-11T03:51:38.157000",
      "content": "<p>Sadly, I did not spend that much time on this problem as I would love. My final submission was based on the mxnet code provided by @phunter and was just an average of fine-tuned Resnet 152 and Resnext 101. I also used train and test time augmentations.</p>\n\n<p>It was the first time when the size of the input images was that important. I would believe that the reason is that in this competition the difference between some classes was very subtle and downsampling original images to something like standard 224x224 would make many important details almost invisible. I would believe (although I did not check) that more careful fine tuning on something like 480x480 input images will lead in the 0.5-0.6 top5 error range.   </p>\n\n<p>I am really grateful to admins that organized this problem. The dataset is amazing.</p>\n\n<p>Sadly there was no monetary prize or at least Kaggle points for this problem and this pushed nearly all competitors toward other problems. Plus 1 month for a competition with the dataset of this size is a bit too short. I would prefer if this problem was going for 3 months, had standard $100,000 prize, and would award Kaggle points. My guess that in this case, we would have hundreds of competitors with all creativity that Kagglers are famous for bringing new insights algorithms and approaches to this task. </p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 201114,
      "author_name": "William Darling",
      "author_url": "",
      "post_date": "2017-07-10T13:26:14.733000",
      "content": "<p>Team \"MunMM\" tried a number of 'elaborate' techniques employing the implicit hierarchy and attention but unfortunately none of these were ever successful. We tried mapping the problem to an \"image captioning\" approach where the output would be two-word sequences consisting of first the super category and then the specific category (basically \"show, attend, and tell\" where the labels would be \"[SUPERCAT] [SPECIFIC CAT] [END SEQ]\"). The idea was that the attention model would learn to recognize the super category (putting emphasis on those parts of the image) and then the RNN would help to constrain the problem to the smaller set of classes within that super category for predicting the specific category (and of course the attention would then direct the model to the parts of the image that were discriminative for differentiating more specific species within that super category). This never gave us any meaningfully good results and in retrospect it was probably because it didn't have a strong enough base model. We only used VGGNet (like in the original paper) and ResNet-18 with fine-tuning which clearly wasn't expressive enough for the problem. It might be interesting to first have a high quality InceptionV3 or Resnet152 model fine-tuned (or trained from scratch -- see below) as the encoder and then perhaps due some very light fine-tuning (or none at all) for the main model but then again the whole RNN might just be overkill.\nWe also tried several variants of \"MixDCNN\" with different numbers of \"sub\"-convnets and freezing different amounts/layers of the model to be able to both share weights optimally and fit everything into memory. Here we had slightly more success than the above captioning-based approaches but still nothing that would get us anywhere close to the basic approach of fine-tuning powerful ImageNet-type convnets.\nIn the end, our score of ~0.011 came from an Ensemble of 3 InceptionV3 models that were all trained slightly differently. Two were done using fine-tuning (both in different toolkits -- CNTK and PyTorch, just for fun :-)) and the other was InceptionV3 trained from scratch. This latter model was our best performing stand-alone model but we just ran out of time to make it even better. The model is still training now and the validation error is continuing to improve with top-5 now under 10% error. With ~600,000 training images, we are close to the size of ImageNet and it looks like training-from-scratch -- while slow (especially with only 1 GPU) -- can work quite nicely for this problem. We were interested to read about the success of fine-tuning with the ImageNet\"11k\" ResNet152 model in mxnet, but didn't have the resources to try it out ourselves.</p>\n\n<p>We are in industry but competed just for fun. We did not use any additional annotations. We are very interested/excited to hear about the techniques that led to the very impressive scores that others in the competition were able to achieve. Thanks for organizing this competition!</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 202253,
      "author_name": "Gyuri Im",
      "author_url": "",
      "post_date": "2017-07-12T05:49:25.153000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 202252,
      "author_name": "Gyuri Im",
      "author_url": "",
      "post_date": "2017-07-12T05:49:16.803000",
      "content": "<p>First, thank you to the people who organized this competition along with the (huge) dataset!</p>\n\n<p>My method was quite simple, finetuning Inception-v3 that was pretrained in ImageNet.</p>\n\n<p>I wanted to explore various data augmentation strategies, and found the <a href=\"https://github.com/aleju/imgaug\">imgaug library</a> helpful in getting a few more percentages in the leaderboard.</p>\n\n<p>I also noticed that some species, such as Eacles Imperialis, contain both the larvae state and the full grown state, which have completely different characteristics. Therefore I separated the images into different lables, and combined them before submission.</p>\n\n<p>While most of the image sizes were either 600x800 or 800x600, there were some images that had disproportionate width/height ratios. In extreme cases I tried to preserve the image ratio, both in train and test time, but am unsure if this was helpful.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 202496,
          "author_name": "macaodha",
          "author_url": "",
          "post_date": "2017-07-12T16:28:41.457000",
          "content": "<p>Out of interest, what percentage performance increase did you get when including the augmentations?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 202733,
          "author_name": "Gyuri Im",
          "author_url": "",
          "post_date": "2017-07-13T06:23:42.587000",
          "content": "<p>It's difficult to be exact, but around 1.5% ~ 2%.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 207737,
          "author_name": "Ravid Cohen",
          "author_url": "",
          "post_date": "2017-07-27T10:25:53.227000",
          "content": "<p>Thank you Gyuri! If I may ask, what is your conclusion about the best augmentation methods? And what is your validation/test score? </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 200468,
      "author_name": "dapper",
      "author_url": "",
      "post_date": "2017-07-08T02:38:57.250000",
      "content": "<p>I was attempting to train a faster-rcnn network (<a href=\"https://arxiv.org/abs/1506.01497\">https://arxiv.org/abs/1506.01497</a>) on the problem. I started by generating region proposals using a sliding window/pyramid and a well trained (80%+ validation accuracy) inception network. Spot checking the proposed RoIs showed that they were not <em>perfect</em>, but in tests on a small subset of the full 5k problem, the faster-rcnn approach worked quite well using these proposals to train on. Unfortunately I didn't have enough time to adapt the existing frcnn implementations to run on multiple GPUs and it simply did not progress fast enough on a single GPU (even with weeks of training) to converge on a satisfactory solution for the full 5k problem. I suspect that the existing architecture needs some tweaking to better perform on a problem with so many classes.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "201561": "Sadly, I did not spend that much time on this problem as I would love. My final submission was based on the mxnet code provided by @phunter and was just an average of fine-tuned Resnet 152 and Resnext 101. I also used train and test time augmentations.\n\nIt was the first time when the size of the input images was that important. I would believe that the reason is that in this competition the difference between some classes was very subtle and downsampling original images to something like standard 224x224 would make many important details almost invisible. I would believe (although I did not check) that more careful fine tuning on something like 480x480 input images will lead in the 0.5-0.6 top5 error range.   \n\nI am really grateful to admins that organized this problem. The dataset is amazing.\n\nSadly there was no monetary prize or at least Kaggle points for this problem and this pushed nearly all competitors toward other problems. Plus 1 month for a competition with the dataset of this size is a bit too short. I would prefer if this problem was going for 3 months, had standard $100,000 prize, and would award Kaggle points. My guess that in this case, we would have hundreds of competitors with all creativity that Kagglers are famous for bringing new insights algorithms and approaches to this task. ",
    "201114": "Team \"MunMM\" tried a number of 'elaborate' techniques employing the implicit hierarchy and attention but unfortunately none of these were ever successful. We tried mapping the problem to an \"image captioning\" approach where the output would be two-word sequences consisting of first the super category and then the specific category (basically \"show, attend, and tell\" where the labels would be \"[SUPERCAT] [SPECIFIC CAT] [END SEQ]\"). The idea was that the attention model would learn to recognize the super category (putting emphasis on those parts of the image) and then the RNN would help to constrain the problem to the smaller set of classes within that super category for predicting the specific category (and of course the attention would then direct the model to the parts of the image that were discriminative for differentiating more specific species within that super category). This never gave us any meaningfully good results and in retrospect it was probably because it didn't have a strong enough base model. We only used VGGNet (like in the original paper) and ResNet-18 with fine-tuning which clearly wasn't expressive enough for the problem. It might be interesting to first have a high quality InceptionV3 or Resnet152 model fine-tuned (or trained from scratch -- see below) as the encoder and then perhaps due some very light fine-tuning (or none at all) for the main model but then again the whole RNN might just be overkill.\nWe also tried several variants of \"MixDCNN\" with different numbers of \"sub\"-convnets and freezing different amounts/layers of the model to be able to both share weights optimally and fit everything into memory. Here we had slightly more success than the above captioning-based approaches but still nothing that would get us anywhere close to the basic approach of fine-tuning powerful ImageNet-type convnets.\nIn the end, our score of ~0.011 came from an Ensemble of 3 InceptionV3 models that were all trained slightly differently. Two were done using fine-tuning (both in different toolkits -- CNTK and PyTorch, just for fun :-)) and the other was InceptionV3 trained from scratch. This latter model was our best performing stand-alone model but we just ran out of time to make it even better. The model is still training now and the validation error is continuing to improve with top-5 now under 10% error. With ~600,000 training images, we are close to the size of ImageNet and it looks like training-from-scratch -- while slow (especially with only 1 GPU) -- can work quite nicely for this problem. We were interested to read about the success of fine-tuning with the ImageNet\"11k\" ResNet152 model in mxnet, but didn't have the resources to try it out ourselves.\n\nWe are in industry but competed just for fun. We did not use any additional annotations. We are very interested/excited to hear about the techniques that led to the very impressive scores that others in the competition were able to achieve. Thanks for organizing this competition!",
    "200453": "Congratulations to all the competitors! We are very impressed by the quality of the results. We are going to take some time to go though the submissions and will soon post the final leader-board.\n\nIn the mean time we would ask competitors to reply to this thread with a description of their method, specify if they used any additional annotations during training, and mention if the team consists of industry, university, or individual competitors.\n\nThanks and well done!",
    "202253": "",
    "202252": "First, thank you to the people who organized this competition along with the (huge) dataset!\n\nMy method was quite simple, finetuning Inception-v3 that was pretrained in ImageNet.\n\nI wanted to explore various data augmentation strategies, and found the [imgaug library][1] helpful in getting a few more percentages in the leaderboard.\n\nI also noticed that some species, such as Eacles Imperialis, contain both the larvae state and the full grown state, which have completely different characteristics. Therefore I separated the images into different lables, and combined them before submission.\n\nWhile most of the image sizes were either 600x800 or 800x600, there were some images that had disproportionate width/height ratios. In extreme cases I tried to preserve the image ratio, both in train and test time, but am unsure if this was helpful.\n\n\n  [1]: https://github.com/aleju/imgaug",
    "200468": "I was attempting to train a faster-rcnn network (https://arxiv.org/abs/1506.01497) on the problem. I started by generating region proposals using a sliding window/pyramid and a well trained (80%+ validation accuracy) inception network. Spot checking the proposed RoIs showed that they were not *perfect*, but in tests on a small subset of the full 5k problem, the faster-rcnn approach worked quite well using these proposals to train on. Unfortunately I didn't have enough time to adapt the existing frcnn implementations to run on multiple GPUs and it simply did not progress fast enough on a single GPU (even with weeks of training) to converge on a satisfactory solution for the full 5k problem. I suspect that the existing architecture needs some tweaking to better perform on a problem with so many classes."
  }
}