{
  "id": 159084,
  "title": "EfficientNets far superior?",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/159084",
  "author_name": "Alex",
  "post_date": "2020-06-16T11:40:58.746000",
  "votes": 16,
  "comment_count": 19,
  "views": 0,
  "content": "<p>From my experiments, it seems that EfficientNets are working way better than any other architecture that I've tried (on CV). Those architectures include ResNet50, InceptionV3, Xception and DenseNet121 (see <a href=\"https://www.tensorflow.org/api_docs/python/tf/keras/applications\">here</a>). I've also tried multiple different <code>preprocess_input()</code> modes, they all seem to give similar results, which is also interesting. I would like to ask you guys, do you have similar experience?</p>",
  "messages": [
    {
      "id": 888477,
      "postDate": "2020-06-16T11:40:58.747Z",
      "content": "<p>From my experiments, it seems that EfficientNets are working way better than any other architecture that I've tried (on CV). Those architectures include ResNet50, InceptionV3, Xception and DenseNet121 (see <a href=\"https://www.tensorflow.org/api_docs/python/tf/keras/applications\">here</a>). I've also tried multiple different <code>preprocess_input()</code> modes, they all seem to give similar results, which is also interesting. I would like to ask you guys, do you have similar experience?</p>",
      "rawMarkdown": "From my experiments, it seems that EfficientNets are working way better than any other architecture that I've tried (on CV). Those architectures include ResNet50, InceptionV3, Xception and DenseNet121 (see [here](https://www.tensorflow.org/api_docs/python/tf/keras/applications)). I've also tried multiple different `preprocess_input()` modes, they all seem to give similar results, which is also interesting. I would like to ask you guys, do you have similar experience?",
      "votes": 16
    },
    {
      "id": 888901,
      "postDate": "2020-06-16T16:27:27.303Z",
      "content": "<p>For me MixNets work better, especially with small batch size</p>",
      "rawMarkdown": "For me MixNets work better, especially with small batch size",
      "votes": 7,
      "replies": [
        {
          "id": 890278,
          "postDate": "2020-06-17T12:26:36.803Z",
          "content": "<p>Did you implemented the MixNet from scratch?</p>",
          "rawMarkdown": "Did you implemented the MixNet from scratch?"
        },
        {
          "id": 890489,
          "postDate": "2020-06-17T14:28:33.230Z",
          "content": "<p>No, took it from this repo \n<a href=\"https://github.com/rwightman/pytorch-image-models\">https://github.com/rwightman/pytorch-image-models</a></p>",
          "rawMarkdown": "No, took it from this repo \nhttps://github.com/rwightman/pytorch-image-models\n",
          "votes": 5
        }
      ]
    },
    {
      "id": 911612,
      "postDate": "2020-07-01T21:47:57.447Z",
      "content": "<p>From what I have found with my experiments, it seems that Efficientnets are more prone to achieve a higher score on Radboud but a lower score on Karolinska.</p>",
      "rawMarkdown": "From what I have found with my experiments, it seems that Efficientnets are more prone to achieve a higher score on Radboud but a lower score on Karolinska.",
      "votes": 1
    },
    {
      "id": 888564,
      "postDate": "2020-06-16T12:45:05.727Z",
      "content": "<p>That's funny because I observed the contrary when I tested EfficientNets against resnext. \nFor example:\nResnext =&gt; VAL : 0.87 =&gt; LB : 0.88 with TTA (my best single model so far)\nEfficientNets=&gt; VAL : 0.87 =&gt; LB : 0.85 with TTA\nI'm stucked at 0.88 LB, I thought that mooving to EfficientNets would boost my score but this is not the case :( </p>",
      "rawMarkdown": "That's funny because I observed the contrary when I tested EfficientNets against resnext. \nFor example:\nResnext =&gt; VAL : 0.87 =&gt; LB : 0.88 with TTA (my best single model so far)\nEfficientNets=&gt; VAL : 0.87 =&gt; LB : 0.85 with TTA\nI'm stucked at 0.88 LB, I thought that mooving to EfficientNets would boost my score but this is not the case :( ",
      "votes": 1
    },
    {
      "id": 898520,
      "postDate": "2020-06-23T15:09:26.820Z",
      "content": "<p>Im trying out efficientnet but it seems like im doing something wrong, with renext i get 0.87 CV and with the same set up it get 0.83 Cv with effnetb0. Im not sure what i need to change to make it work, learning rate?</p>",
      "rawMarkdown": "Im trying out efficientnet but it seems like im doing something wrong, with renext i get 0.87 CV and with the same set up it get 0.83 Cv with effnetb0. Im not sure what i need to change to make it work, learning rate?"
    },
    {
      "id": 888559,
      "postDate": "2020-06-16T12:41:39.873Z",
      "content": "<p>I used EfficientNet but i don't have de criterion to chose what version, so i used B6 and the model toke more than 5 hours to train. Any advice? and, how much time did you spend in the training of the EfficientNet model?</p>",
      "rawMarkdown": "I used EfficientNet but i don't have de criterion to chose what version, so i used B6 and the model toke more than 5 hours to train. Any advice? and, how much time did you spend in the training of the EfficientNet model?",
      "replies": [
        {
          "id": 888588,
          "postDate": "2020-06-16T12:57:41.873Z",
          "content": "<p>I am not entirely sure that you need <code>EfficientNet</code>. Currently my best model is <code>resnet34</code> and single fold can reach <code>0.90</code>.</p>",
          "rawMarkdown": "I am not entirely sure that you need `EfficientNet`. Currently my best model is `resnet34` and single fold can reach `0.90`.",
          "votes": 14
        },
        {
          "id": 888614,
          "postDate": "2020-06-16T13:16:54.240Z",
          "content": "<p><a href=\"/drhabib\">@drhabib</a>  Is it more important for preprocessing, and using the correct data for this competition, rather than madly changing from one architecture to another?</p>",
          "rawMarkdown": "@drhabib  Is it more important for preprocessing, and using the correct data for this competition, rather than madly changing from one architecture to another?",
          "votes": 2
        },
        {
          "id": 888638,
          "postDate": "2020-06-16T13:30:41.647Z",
          "content": "<p><code>madly changing from one architecture to another</code> never a solution =) . Start from a simple network, train data, check validation score and see how you can improve. If you stuck go to forum and try to read what people did and see if it increases score or not =) Good luck! </p>",
          "rawMarkdown": "`madly changing from one architecture to another` never a solution =) . Start from a simple network, train data, check validation score and see how you can improve. If you stuck go to forum and try to read what people did and see if it increases score or not =) Good luck! ",
          "votes": 12
        },
        {
          "id": 888642,
          "postDate": "2020-06-16T13:33:23.883Z",
          "content": "<p>Btw, I'm using 'Big images' as input (similar to <a href=\"https://www.kaggle.com/haqishen/train-efficientnet-b0-w-36-tiles-256-lb0-87\">this kernel</a> but in TensorFlow). <a href=\"/hiramcho\">@hiramcho</a> I'm working with EfficientNetB0-B2 (My current LB score is a B1). I would focus on B0 first, then work my way up if needed. <a href=\"/drhabib\">@drhabib</a> for me Resnet just doesn't work (Perhaps I should go lower than ResNet50?) --- I'm likely doing something wrong, but haven't been able to figure out the problem.</p>",
          "rawMarkdown": "Btw, I'm using 'Big images' as input (similar to [this kernel](https://www.kaggle.com/haqishen/train-efficientnet-b0-w-36-tiles-256-lb0-87) but in TensorFlow). @hiramcho I'm working with EfficientNetB0-B2 (My current LB score is a B1). I would focus on B0 first, then work my way up if needed. @drhabib for me Resnet just doesn't work (Perhaps I should go lower than ResNet50?) --- I'm likely doing something wrong, but haven't been able to figure out the problem.",
          "votes": 3
        },
        {
          "id": 888673,
          "postDate": "2020-06-16T13:48:52.630Z",
          "content": "<p><code>(Perhaps I should go lower than ResNet50?)</code>  To be perfectly honest I havent  spend much time on other networks yet. I usually do it at the end... The reason I choose small network because you can experiment much quickly ... </p>\n\n<p>I wish you all the best with experiments =) </p>",
          "rawMarkdown": "`(Perhaps I should go lower than ResNet50?)`  To be perfectly honest I havent  spend much time on other networks yet. I usually do it at the end... The reason I choose small network because you can experiment much quickly ... \n\n I wish you all the best with experiments =) ",
          "votes": 5
        },
        {
          "id": 888675,
          "postDate": "2020-06-16T13:50:35.500Z",
          "content": "<p><a href=\"/drhabib\">@drhabib</a> Thanks, likewise!</p>",
          "rawMarkdown": "@drhabib Thanks, likewise!",
          "votes": 2
        },
        {
          "id": 888877,
          "postDate": "2020-06-16T16:06:34.320Z",
          "content": "<p><a href=\"/drhabib\">@drhabib</a>  good to see u are able to work with basic network arch so well.\n1) I think it would be impossible to reach to cv close to 0.9 based on just tile extraction methods used in public kernels .  More than net arch &amp; tile extraction what should contribute further ? ,\n2) Are you using simple loss  like bce /mse  or you implemented some ordinality </p>",
          "rawMarkdown": "@drhabib  good to see u are able to work with basic network arch so well.\n1) I think it would be impossible to reach to cv close to 0.9 based on just tile extraction methods used in public kernels .  More than net arch &amp; tile extraction what should contribute further ? ,\n2) Are you using simple loss  like bce /mse  or you implemented some ordinality ",
          "votes": 2
        },
        {
          "id": 889641,
          "postDate": "2020-06-17T03:44:58.350Z",
          "content": "<p>1) Yes, its combination =)\n2) Yes, simple loss =)</p>",
          "rawMarkdown": "1) Yes, its combination =)\n2) Yes, simple loss =)\n"
        },
        {
          "id": 891111,
          "postDate": "2020-06-17T22:50:22.390Z",
          "content": "<p>What is net arch extraction?</p>",
          "rawMarkdown": "What is net arch extraction?"
        },
        {
          "id": 897094,
          "postDate": "2020-06-22T16:02:48.167Z",
          "content": "<p><a href=\"/thomasx\">@thomasx</a> I think he meant \"net architecture\" ... \"and tile extraction\", two different things, (although your way of reading it is perfectly valid english as well ;) )</p>",
          "rawMarkdown": "@thomasx I think he meant \"net architecture\" ... \"and tile extraction\", two different things, (although your way of reading it is perfectly valid english as well ;) )"
        },
        {
          "id": 904259,
          "postDate": "2020-06-27T13:33:53.157Z",
          "content": "<p>Indeed, thanks for the clarification 👍🏼</p>",
          "rawMarkdown": "Indeed, thanks for the clarification 👍🏼"
        },
        {
          "id": 904301,
          "postDate": "2020-06-27T14:10:26.653Z",
          "content": "<p>Could you tell how many epochs did you train your model?</p>",
          "rawMarkdown": "Could you tell how many epochs did you train your model?"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 888901,
      "author_name": "Eek The Cat",
      "author_url": "",
      "post_date": "2020-06-16T16:27:27.303000",
      "content": "<p>For me MixNets work better, especially with small batch size</p>",
      "votes": 7,
      "replies": [
        {
          "id": 890278,
          "author_name": "Hiram Coria 🧬",
          "author_url": "",
          "post_date": "2020-06-17T12:26:36.803000",
          "content": "<p>Did you implemented the MixNet from scratch?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 890489,
          "author_name": "Eek The Cat",
          "author_url": "",
          "post_date": "2020-06-17T14:28:33.230000",
          "content": "<p>No, took it from this repo \n<a href=\"https://github.com/rwightman/pytorch-image-models\">https://github.com/rwightman/pytorch-image-models</a></p>",
          "votes": 5,
          "replies": []
        }
      ]
    },
    {
      "id": 911612,
      "author_name": "Richard Xiao",
      "author_url": "",
      "post_date": "2020-07-01T21:47:57.447000",
      "content": "<p>From what I have found with my experiments, it seems that Efficientnets are more prone to achieve a higher score on Radboud but a lower score on Karolinska.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 888564,
      "author_name": "Alex",
      "author_url": "",
      "post_date": "2020-06-16T12:45:05.727000",
      "content": "<p>That's funny because I observed the contrary when I tested EfficientNets against resnext. \nFor example:\nResnext =&gt; VAL : 0.87 =&gt; LB : 0.88 with TTA (my best single model so far)\nEfficientNets=&gt; VAL : 0.87 =&gt; LB : 0.85 with TTA\nI'm stucked at 0.88 LB, I thought that mooving to EfficientNets would boost my score but this is not the case :( </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 898520,
      "author_name": "Yann Majewski",
      "author_url": "",
      "post_date": "2020-06-23T15:09:26.820000",
      "content": "<p>Im trying out efficientnet but it seems like im doing something wrong, with renext i get 0.87 CV and with the same set up it get 0.83 Cv with effnetb0. Im not sure what i need to change to make it work, learning rate?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 888559,
      "author_name": "Hiram Coria 🧬",
      "author_url": "",
      "post_date": "2020-06-16T12:41:39.873000",
      "content": "<p>I used EfficientNet but i don't have de criterion to chose what version, so i used B6 and the model toke more than 5 hours to train. Any advice? and, how much time did you spend in the training of the EfficientNet model?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 888588,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2020-06-16T12:57:41.873000",
          "content": "<p>I am not entirely sure that you need <code>EfficientNet</code>. Currently my best model is <code>resnet34</code> and single fold can reach <code>0.90</code>.</p>",
          "votes": 14,
          "replies": []
        },
        {
          "id": 888614,
          "author_name": "Kurian Benoy",
          "author_url": "",
          "post_date": "2020-06-16T13:16:54.240000",
          "content": "<p><a href=\"/drhabib\">@drhabib</a>  Is it more important for preprocessing, and using the correct data for this competition, rather than madly changing from one architecture to another?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 888638,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2020-06-16T13:30:41.647000",
          "content": "<p><code>madly changing from one architecture to another</code> never a solution =) . Start from a simple network, train data, check validation score and see how you can improve. If you stuck go to forum and try to read what people did and see if it increases score or not =) Good luck! </p>",
          "votes": 12,
          "replies": []
        },
        {
          "id": 888642,
          "author_name": "Alex",
          "author_url": "",
          "post_date": "2020-06-16T13:33:23.883000",
          "content": "<p>Btw, I'm using 'Big images' as input (similar to <a href=\"https://www.kaggle.com/haqishen/train-efficientnet-b0-w-36-tiles-256-lb0-87\">this kernel</a> but in TensorFlow). <a href=\"/hiramcho\">@hiramcho</a> I'm working with EfficientNetB0-B2 (My current LB score is a B1). I would focus on B0 first, then work my way up if needed. <a href=\"/drhabib\">@drhabib</a> for me Resnet just doesn't work (Perhaps I should go lower than ResNet50?) --- I'm likely doing something wrong, but haven't been able to figure out the problem.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 888673,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2020-06-16T13:48:52.630000",
          "content": "<p><code>(Perhaps I should go lower than ResNet50?)</code>  To be perfectly honest I havent  spend much time on other networks yet. I usually do it at the end... The reason I choose small network because you can experiment much quickly ... </p>\n\n<p>I wish you all the best with experiments =) </p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 888675,
          "author_name": "Alex",
          "author_url": "",
          "post_date": "2020-06-16T13:50:35.500000",
          "content": "<p><a href=\"/drhabib\">@drhabib</a> Thanks, likewise!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 888877,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-06-16T16:06:34.320000",
          "content": "<p><a href=\"/drhabib\">@drhabib</a>  good to see u are able to work with basic network arch so well.\n1) I think it would be impossible to reach to cv close to 0.9 based on just tile extraction methods used in public kernels .  More than net arch &amp; tile extraction what should contribute further ? ,\n2) Are you using simple loss  like bce /mse  or you implemented some ordinality </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 889641,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2020-06-17T03:44:58.350000",
          "content": "<p>1) Yes, its combination =)\n2) Yes, simple loss =)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 891111,
          "author_name": "None",
          "author_url": "",
          "post_date": "2020-06-17T22:50:22.390000",
          "content": "<p>What is net arch extraction?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 897094,
          "author_name": "Benjamin Dubreu",
          "author_url": "",
          "post_date": "2020-06-22T16:02:48.167000",
          "content": "<p><a href=\"/thomasx\">@thomasx</a> I think he meant \"net architecture\" ... \"and tile extraction\", two different things, (although your way of reading it is perfectly valid english as well ;) )</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 904259,
          "author_name": "None",
          "author_url": "",
          "post_date": "2020-06-27T13:33:53.157000",
          "content": "<p>Indeed, thanks for the clarification 👍🏼</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 904301,
          "author_name": "Son of Anton v3.0",
          "author_url": "",
          "post_date": "2020-06-27T14:10:26.653000",
          "content": "<p>Could you tell how many epochs did you train your model?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "888477": "From my experiments, it seems that EfficientNets are working way better than any other architecture that I've tried (on CV). Those architectures include ResNet50, InceptionV3, Xception and DenseNet121 (see [here](https://www.tensorflow.org/api_docs/python/tf/keras/applications)). I've also tried multiple different `preprocess_input()` modes, they all seem to give similar results, which is also interesting. I would like to ask you guys, do you have similar experience?",
    "888901": "For me MixNets work better, especially with small batch size",
    "911612": "From what I have found with my experiments, it seems that Efficientnets are more prone to achieve a higher score on Radboud but a lower score on Karolinska.",
    "888564": "That's funny because I observed the contrary when I tested EfficientNets against resnext. \nFor example:\nResnext =&gt; VAL : 0.87 =&gt; LB : 0.88 with TTA (my best single model so far)\nEfficientNets=&gt; VAL : 0.87 =&gt; LB : 0.85 with TTA\nI'm stucked at 0.88 LB, I thought that mooving to EfficientNets would boost my score but this is not the case :( ",
    "898520": "Im trying out efficientnet but it seems like im doing something wrong, with renext i get 0.87 CV and with the same set up it get 0.83 Cv with effnetb0. Im not sure what i need to change to make it work, learning rate?",
    "888559": "I used EfficientNet but i don't have de criterion to chose what version, so i used B6 and the model toke more than 5 hours to train. Any advice? and, how much time did you spend in the training of the EfficientNet model?"
  }
}