{
  "id": 167426,
  "title": "Noisy students vs Imagenet weights in the competition context",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/167426",
  "author_name": "Vlad Vaduva",
  "post_date": "2020-07-16T12:19:47.169000",
  "votes": 12,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I have tested the noisy students weights applied in the context of this application.\nFor people unfamiliar with this weights this is a quick description of what are they and how they are obtained Aakash Nain explains it very good in <a href=\"https://medium.com/@nainaakash012/self-training-with-noisy-student-f33640edbab2\">https://medium.com/@nainaakash012/self-training-with-noisy-student-f33640edbab2</a>.</p>\n\n<p>\"Self-training first uses the labeled data to train a good teacher model, then use the teacher model to label the unlabeled data. As we know that all the predictions of the teacher model on the unlabeled data can’t be good, hence in classical self-training, we select a subset of this unlabeled data by filtering out the predictions(aka pseudo labels) using a threshold for the score. This subset is now combined with the original labeled data and a new model, student model, is jointly trained on this combined data. This whole procedure can be repeated for n number of times until the convergence is reached.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4005865%2F2ab58f57352bce4a5259b9a788d47d66%2F1%20auzrU0rW7A6pn0dX1oBVOw.png?generation=1594900925004921&amp;alt=media\" alt=\"\"></p>\n\n<p>Nice approach to make use of the unlabeled data. But the paper emphasizes on something else, a noisy student. What’s the deal with that? Is it different from the classical approach?</p>\n\n<p>Yes, you are correct. The authors found out that for this method to work at scale, the student model should be noised during its training while the teacher model should not be noised during the generation of pseudo-labels. Because noise is an important piece (which we will be looking into detail in a minute) for this whole idea to work, that’s why they called it a noisy student.\nThe Algorithm</p>\n\n<p>The algorithm is similar to classical self-training with some minor differences.</p>\n\n<p>Noisy Student algorithm\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4005865%2Fa776df4725ea7f88c55d252545041ac4%2F1%20D-exTIZ9ekT9Zdq97AmGIQ.png?generation=1594900890365816&amp;alt=media\" alt=\"\"></p>\n\n<p>The main difference is the addition of noise to the student using different techniques like dropout, stochastic depth, and augmentation. It should be noted that the teacher is not noised when it generates the pseudo labels.\"</p>\n\n<p>So, my tests for this competition were made on EffNets B0, B1 and B2 and compared with the results in the identical conditions with classical Imagenet weighs.\nThe results were interesting, while on B0 and B2 the accuracy did not improve, on B1 I got a good improvement, so it's no guarantee of functionality improvement but I suggest to try them in this last competition days.</p>\n\n<p>For you guys which preffe pytorch you can use:\n- <a href=\"https://github.com/tensorflow/tpu/tree/master/models/official/efficientnet\">https://github.com/tensorflow/tpu/tree/master/models/official/efficientnet</a> for downloading TensorFlow NoisyStudent weights and then \n- <a href=\"https://github.com/lukemelas/EfficientNet-PyTorch/tree/master/tf_to_pytorch\">https://github.com/lukemelas/EfficientNet-PyTorch/tree/master/tf_to_pytorch</a> to convert the tensorflow weight file to pytorch format.</p>",
  "messages": [
    {
      "id": 931740,
      "postDate": "2020-07-16T12:19:47.170Z",
      "content": "<p>I have tested the noisy students weights applied in the context of this application.\nFor people unfamiliar with this weights this is a quick description of what are they and how they are obtained Aakash Nain explains it very good in <a href=\"https://medium.com/@nainaakash012/self-training-with-noisy-student-f33640edbab2\">https://medium.com/@nainaakash012/self-training-with-noisy-student-f33640edbab2</a>.</p>\n\n<p>\"Self-training first uses the labeled data to train a good teacher model, then use the teacher model to label the unlabeled data. As we know that all the predictions of the teacher model on the unlabeled data can’t be good, hence in classical self-training, we select a subset of this unlabeled data by filtering out the predictions(aka pseudo labels) using a threshold for the score. This subset is now combined with the original labeled data and a new model, student model, is jointly trained on this combined data. This whole procedure can be repeated for n number of times until the convergence is reached.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4005865%2F2ab58f57352bce4a5259b9a788d47d66%2F1%20auzrU0rW7A6pn0dX1oBVOw.png?generation=1594900925004921&amp;alt=media\" alt=\"\"></p>\n\n<p>Nice approach to make use of the unlabeled data. But the paper emphasizes on something else, a noisy student. What’s the deal with that? Is it different from the classical approach?</p>\n\n<p>Yes, you are correct. The authors found out that for this method to work at scale, the student model should be noised during its training while the teacher model should not be noised during the generation of pseudo-labels. Because noise is an important piece (which we will be looking into detail in a minute) for this whole idea to work, that’s why they called it a noisy student.\nThe Algorithm</p>\n\n<p>The algorithm is similar to classical self-training with some minor differences.</p>\n\n<p>Noisy Student algorithm\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4005865%2Fa776df4725ea7f88c55d252545041ac4%2F1%20D-exTIZ9ekT9Zdq97AmGIQ.png?generation=1594900890365816&amp;alt=media\" alt=\"\"></p>\n\n<p>The main difference is the addition of noise to the student using different techniques like dropout, stochastic depth, and augmentation. It should be noted that the teacher is not noised when it generates the pseudo labels.\"</p>\n\n<p>So, my tests for this competition were made on EffNets B0, B1 and B2 and compared with the results in the identical conditions with classical Imagenet weighs.\nThe results were interesting, while on B0 and B2 the accuracy did not improve, on B1 I got a good improvement, so it's no guarantee of functionality improvement but I suggest to try them in this last competition days.</p>\n\n<p>For you guys which preffe pytorch you can use:\n- <a href=\"https://github.com/tensorflow/tpu/tree/master/models/official/efficientnet\">https://github.com/tensorflow/tpu/tree/master/models/official/efficientnet</a> for downloading TensorFlow NoisyStudent weights and then \n- <a href=\"https://github.com/lukemelas/EfficientNet-PyTorch/tree/master/tf_to_pytorch\">https://github.com/lukemelas/EfficientNet-PyTorch/tree/master/tf_to_pytorch</a> to convert the tensorflow weight file to pytorch format.</p>",
      "rawMarkdown": "I have tested the noisy students weights applied in the context of this application.\nFor people unfamiliar with this weights this is a quick description of what are they and how they are obtained Aakash Nain explains it very good in https://medium.com/@nainaakash012/self-training-with-noisy-student-f33640edbab2.\n\n\"Self-training first uses the labeled data to train a good teacher model, then use the teacher model to label the unlabeled data. As we know that all the predictions of the teacher model on the unlabeled data can’t be good, hence in classical self-training, we select a subset of this unlabeled data by filtering out the predictions(aka pseudo labels) using a threshold for the score. This subset is now combined with the original labeled data and a new model, student model, is jointly trained on this combined data. This whole procedure can be repeated for n number of times until the convergence is reached.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4005865%2F2ab58f57352bce4a5259b9a788d47d66%2F1%20auzrU0rW7A6pn0dX1oBVOw.png?generation=1594900925004921&amp;alt=media)\n\nNice approach to make use of the unlabeled data. But the paper emphasizes on something else, a noisy student. What’s the deal with that? Is it different from the classical approach?\n\nYes, you are correct. The authors found out that for this method to work at scale, the student model should be noised during its training while the teacher model should not be noised during the generation of pseudo-labels. Because noise is an important piece (which we will be looking into detail in a minute) for this whole idea to work, that’s why they called it a noisy student.\nThe Algorithm\n\nThe algorithm is similar to classical self-training with some minor differences.\n\nNoisy Student algorithm\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4005865%2Fa776df4725ea7f88c55d252545041ac4%2F1%20D-exTIZ9ekT9Zdq97AmGIQ.png?generation=1594900890365816&amp;alt=media)\n\nThe main difference is the addition of noise to the student using different techniques like dropout, stochastic depth, and augmentation. It should be noted that the teacher is not noised when it generates the pseudo labels.\"\n\nSo, my tests for this competition were made on EffNets B0, B1 and B2 and compared with the results in the identical conditions with classical Imagenet weighs.\nThe results were interesting, while on B0 and B2 the accuracy did not improve, on B1 I got a good improvement, so it's no guarantee of functionality improvement but I suggest to try them in this last competition days.\n\nFor you guys which preffe pytorch you can use:\n- https://github.com/tensorflow/tpu/tree/master/models/official/efficientnet for downloading TensorFlow NoisyStudent weights and then \n- https://github.com/lukemelas/EfficientNet-PyTorch/tree/master/tf_to_pytorch to convert the tensorflow weight file to pytorch format.\n\n",
      "votes": 12
    },
    {
      "id": 931896,
      "postDate": "2020-07-16T14:37:18.550Z",
      "content": "<p>I tried to use noisy student pre-training for b0, but couldn't get any improvement. By the way, the weights are already converted from tf to pytorch in the timm package - very easy to use</p>",
      "rawMarkdown": "I tried to use noisy student pre-training for b0, but couldn't get any improvement. By the way, the weights are already converted from tf to pytorch in the timm package - very easy to use",
      "votes": 3,
      "replies": [
        {
          "id": 931906,
          "postDate": "2020-07-16T14:48:42.140Z",
          "content": "<p><a href=\"/abzaliev\">@abzaliev</a> Good to know. Thanks for the info</p>",
          "rawMarkdown": "@abzaliev Good to know. Thanks for the info",
          "votes": 1
        }
      ]
    },
    {
      "id": 932419,
      "postDate": "2020-07-17T03:53:08.417Z",
      "content": "<p><a href=\"https://www.kaggle.com/vladvdv\" target=\"_blank\">@vladvdv</a>  how do we define convergence here..<br>\nThanks for posting.. I was doing it bit incorrectly i think..</p>",
      "rawMarkdown": "@vladvdv  how do we define convergence here..\nThanks for posting.. I was doing it bit incorrectly i think..",
      "votes": 1
    },
    {
      "id": 931814,
      "postDate": "2020-07-16T13:20:47.240Z",
      "content": "<p>You just loads the efficientNet weights,<code>load_state_dict(torch.load(pretrained_model[backbone]))</code>\nWhere pretrained is a dict with different arhitecture types and backbone is the arhitecture where you want to use\n<code>\npretrained_model = {\n    'efficientnet-b2': 'efficientnet-b2.pth'\n}\n</code></p>",
      "rawMarkdown": "You just loads the efficientNet weights,` load_state_dict(torch.load(pretrained_model[backbone]))`\nWhere pretrained is a dict with different arhitecture types and backbone is the arhitecture where you want to use\n```\npretrained_model = {\n    'efficientnet-b2': 'efficientnet-b2.pth'\n}\n```",
      "votes": 1
    },
    {
      "id": 934804,
      "postDate": "2020-07-18T18:36:04.923Z",
      "content": "<p>Interesting post, thanks for sharing! Do you think we should consider that labels in the dataset are already noisy, given that this is medical annotation and that labels may be incorrect?</p>\n<p>Please also find <a href=\"https://www.kaggle.com/mika30/efficientnet-weights-for-keras\" target=\"_blank\">here</a> all EfficientNet weights for Keras, I hope this dataset will make easier to compare weights for TF/Keras users!</p>",
      "rawMarkdown": "Interesting post, thanks for sharing! Do you think we should consider that labels in the dataset are already noisy, given that this is medical annotation and that labels may be incorrect?\n\nPlease also find [here](https://www.kaggle.com/mika30/efficientnet-weights-for-keras) all EfficientNet weights for Keras, I hope this dataset will make easier to compare weights for TF/Keras users!"
    },
    {
      "id": 934056,
      "postDate": "2020-07-18T07:55:14.893Z",
      "content": "<p><a href=\"/jaideepvalani\">@jaideepvalani</a> From what I read they seem to use this auto-loop until the improvement per teacher- student paradigm iteration is smaller than a minimum convergence step</p>",
      "rawMarkdown": "@jaideepvalani From what I read they seem to use this auto-loop until the improvement per teacher- student paradigm iteration is smaller than a minimum convergence step"
    },
    {
      "id": 931795,
      "postDate": "2020-07-16T13:05:50.330Z",
      "content": "<p>Thanks for share, the repo doesnt' share an example to implement the pretrained model with \"NoisyStudent\". Could be that it's implemented like a normal EfficientNet model?</p>",
      "rawMarkdown": "Thanks for share, the repo doesnt' share an example to implement the pretrained model with \"NoisyStudent\". Could be that it's implemented like a normal EfficientNet model?"
    },
    {
      "id": 931813,
      "postDate": "2020-07-16T13:20:18.960Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 931896,
      "author_name": "abzaliev",
      "author_url": "",
      "post_date": "2020-07-16T14:37:18.550000",
      "content": "<p>I tried to use noisy student pre-training for b0, but couldn't get any improvement. By the way, the weights are already converted from tf to pytorch in the timm package - very easy to use</p>",
      "votes": 3,
      "replies": [
        {
          "id": 931906,
          "author_name": "Vlad Vaduva",
          "author_url": "",
          "post_date": "2020-07-16T14:48:42.140000",
          "content": "<p><a href=\"/abzaliev\">@abzaliev</a> Good to know. Thanks for the info</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 932419,
      "author_name": "Jaideep",
      "author_url": "",
      "post_date": "2020-07-17T03:53:08.417000",
      "content": "<p><a href=\"https://www.kaggle.com/vladvdv\" target=\"_blank\">@vladvdv</a>  how do we define convergence here..<br>\nThanks for posting.. I was doing it bit incorrectly i think..</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 931814,
      "author_name": "Vlad Vaduva",
      "author_url": "",
      "post_date": "2020-07-16T13:20:47.240000",
      "content": "<p>You just loads the efficientNet weights,<code>load_state_dict(torch.load(pretrained_model[backbone]))</code>\nWhere pretrained is a dict with different arhitecture types and backbone is the arhitecture where you want to use\n<code>\npretrained_model = {\n    'efficientnet-b2': 'efficientnet-b2.pth'\n}\n</code></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 934804,
      "author_name": "Michaël Karpe",
      "author_url": "",
      "post_date": "2020-07-18T18:36:04.923000",
      "content": "<p>Interesting post, thanks for sharing! Do you think we should consider that labels in the dataset are already noisy, given that this is medical annotation and that labels may be incorrect?</p>\n<p>Please also find <a href=\"https://www.kaggle.com/mika30/efficientnet-weights-for-keras\" target=\"_blank\">here</a> all EfficientNet weights for Keras, I hope this dataset will make easier to compare weights for TF/Keras users!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 934056,
      "author_name": "Vlad Vaduva",
      "author_url": "",
      "post_date": "2020-07-18T07:55:14.893000",
      "content": "<p><a href=\"/jaideepvalani\">@jaideepvalani</a> From what I read they seem to use this auto-loop until the improvement per teacher- student paradigm iteration is smaller than a minimum convergence step</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 931795,
      "author_name": "Hiram Coria 🧬",
      "author_url": "",
      "post_date": "2020-07-16T13:05:50.330000",
      "content": "<p>Thanks for share, the repo doesnt' share an example to implement the pretrained model with \"NoisyStudent\". Could be that it's implemented like a normal EfficientNet model?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 931813,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-16T13:20:18.960000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "931740": "I have tested the noisy students weights applied in the context of this application.\nFor people unfamiliar with this weights this is a quick description of what are they and how they are obtained Aakash Nain explains it very good in https://medium.com/@nainaakash012/self-training-with-noisy-student-f33640edbab2.\n\n\"Self-training first uses the labeled data to train a good teacher model, then use the teacher model to label the unlabeled data. As we know that all the predictions of the teacher model on the unlabeled data can’t be good, hence in classical self-training, we select a subset of this unlabeled data by filtering out the predictions(aka pseudo labels) using a threshold for the score. This subset is now combined with the original labeled data and a new model, student model, is jointly trained on this combined data. This whole procedure can be repeated for n number of times until the convergence is reached.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4005865%2F2ab58f57352bce4a5259b9a788d47d66%2F1%20auzrU0rW7A6pn0dX1oBVOw.png?generation=1594900925004921&amp;alt=media)\n\nNice approach to make use of the unlabeled data. But the paper emphasizes on something else, a noisy student. What’s the deal with that? Is it different from the classical approach?\n\nYes, you are correct. The authors found out that for this method to work at scale, the student model should be noised during its training while the teacher model should not be noised during the generation of pseudo-labels. Because noise is an important piece (which we will be looking into detail in a minute) for this whole idea to work, that’s why they called it a noisy student.\nThe Algorithm\n\nThe algorithm is similar to classical self-training with some minor differences.\n\nNoisy Student algorithm\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4005865%2Fa776df4725ea7f88c55d252545041ac4%2F1%20D-exTIZ9ekT9Zdq97AmGIQ.png?generation=1594900890365816&amp;alt=media)\n\nThe main difference is the addition of noise to the student using different techniques like dropout, stochastic depth, and augmentation. It should be noted that the teacher is not noised when it generates the pseudo labels.\"\n\nSo, my tests for this competition were made on EffNets B0, B1 and B2 and compared with the results in the identical conditions with classical Imagenet weighs.\nThe results were interesting, while on B0 and B2 the accuracy did not improve, on B1 I got a good improvement, so it's no guarantee of functionality improvement but I suggest to try them in this last competition days.\n\nFor you guys which preffe pytorch you can use:\n- https://github.com/tensorflow/tpu/tree/master/models/official/efficientnet for downloading TensorFlow NoisyStudent weights and then \n- https://github.com/lukemelas/EfficientNet-PyTorch/tree/master/tf_to_pytorch to convert the tensorflow weight file to pytorch format.\n\n",
    "931896": "I tried to use noisy student pre-training for b0, but couldn't get any improvement. By the way, the weights are already converted from tf to pytorch in the timm package - very easy to use",
    "932419": "@vladvdv  how do we define convergence here..\nThanks for posting.. I was doing it bit incorrectly i think..",
    "931814": "You just loads the efficientNet weights,` load_state_dict(torch.load(pretrained_model[backbone]))`\nWhere pretrained is a dict with different arhitecture types and backbone is the arhitecture where you want to use\n```\npretrained_model = {\n    'efficientnet-b2': 'efficientnet-b2.pth'\n}\n```",
    "934804": "Interesting post, thanks for sharing! Do you think we should consider that labels in the dataset are already noisy, given that this is medical annotation and that labels may be incorrect?\n\nPlease also find [here](https://www.kaggle.com/mika30/efficientnet-weights-for-keras) all EfficientNet weights for Keras, I hope this dataset will make easier to compare weights for TF/Keras users!",
    "934056": "@jaideepvalani From what I read they seem to use this auto-loop until the improvement per teacher- student paradigm iteration is smaller than a minimum convergence step",
    "931795": "Thanks for share, the repo doesnt' share an example to implement the pretrained model with \"NoisyStudent\". Could be that it's implemented like a normal EfficientNet model?",
    "931813": ""
  }
}