{
  "id": 373957,
  "title": "There's something about efficientnets",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/373957",
  "author_name": "James Howard",
  "post_date": "2022-12-24T11:58:50.900000",
  "votes": 30,
  "comment_count": 22,
  "views": 0,
  "content": "<p>In the HuBMAP competition 3 months ago, most people found transformers performed excellently for histology segmentation. However, there was an exception; for the slices of lung tissue, nothing could match efficientnets, and I couldn't find a great explanation for this.</p>\n<p>Lung tissue was more sparse than the other organs, and often involved very fine tissue planes and thin features. I assumed it was that. However, there is a dataset called ImageNet-Sketch, which is like ImageNet for hand drawn images. You'd think if being able to identify fine planes rather than texture was relevant, EfficientNet would do well at this? Well it doesn't - Ross Wightman shows <a href=\"https://github.com/rwightman/pytorch-image-models/blob/main/results/results-sketch.csv\" target=\"_blank\">here</a> that actually ResNext based models and transformers outperform EfficientNet. This is paralleled by the results of the Bengali challenge, where sketched characters had to be identified. Here, the winners had more success with SE-ResNeXT models rather than EfficientNets.</p>\n<p>In this challenge, where mammograms are made up of spindly ethereal structures in relatively sparse imagines, EfficientNets again are proving unreasonably effective for me. SE-ResNeXT and other novel models just consistently underperform for me, whereas a modest EfficientNet-B2 works fine.</p>\n<p>Is this other people's experience? Can they explain it?</p>\n<p>Last night I adapted an EfficientNet-B3-pruned model (brown in the image below) from Ross' timm repo and trained it for the challenge with identical hyperparameters (including augmentations and batch size) to my standard EfficientNet-B2 training, and  could not match the standard B2's results (blue line in image below). Bizarre. I also show the performance of a similarly trained seresnext50_32x4d model in green. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3108186%2F27e28e56e34e678505404cbdaeaadfb9%2FScreenshot%202022-12-24%20at%2012.02.21.png?generation=1671883395618303&amp;alt=media\" alt=\"\"></p>\n<p>(Apologies for not logging / plotting the competition metric, pF1, but it's so noisy and I strongly suspect AUC is the validation metric to track).</p>\n<p>Note, a STANDARD EfficientNet-B3 model (non pruned) does much better than the pruned model, and similarly to EfficientNet-B2 - there seems to be something subtle here that is of great significance, and the architecture choices almost certainly are going to matter a lot in this competition.</p>\n<p>Maybe there is something about my training pipeline where over the years I have subconsciously tweaked it for EfficientNet (LRs, optimizers, schedulers etc.) - but currently I am struggling to find anything that rivals these 3-4 year old models</p>",
  "messages": [
    {
      "id": 2074611,
      "postDate": "2022-12-24T11:58:50.900Z",
      "content": "<p>In the HuBMAP competition 3 months ago, most people found transformers performed excellently for histology segmentation. However, there was an exception; for the slices of lung tissue, nothing could match efficientnets, and I couldn't find a great explanation for this.</p>\n<p>Lung tissue was more sparse than the other organs, and often involved very fine tissue planes and thin features. I assumed it was that. However, there is a dataset called ImageNet-Sketch, which is like ImageNet for hand drawn images. You'd think if being able to identify fine planes rather than texture was relevant, EfficientNet would do well at this? Well it doesn't - Ross Wightman shows <a href=\"https://github.com/rwightman/pytorch-image-models/blob/main/results/results-sketch.csv\" target=\"_blank\">here</a> that actually ResNext based models and transformers outperform EfficientNet. This is paralleled by the results of the Bengali challenge, where sketched characters had to be identified. Here, the winners had more success with SE-ResNeXT models rather than EfficientNets.</p>\n<p>In this challenge, where mammograms are made up of spindly ethereal structures in relatively sparse imagines, EfficientNets again are proving unreasonably effective for me. SE-ResNeXT and other novel models just consistently underperform for me, whereas a modest EfficientNet-B2 works fine.</p>\n<p>Is this other people's experience? Can they explain it?</p>\n<p>Last night I adapted an EfficientNet-B3-pruned model (brown in the image below) from Ross' timm repo and trained it for the challenge with identical hyperparameters (including augmentations and batch size) to my standard EfficientNet-B2 training, and  could not match the standard B2's results (blue line in image below). Bizarre. I also show the performance of a similarly trained seresnext50_32x4d model in green. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3108186%2F27e28e56e34e678505404cbdaeaadfb9%2FScreenshot%202022-12-24%20at%2012.02.21.png?generation=1671883395618303&amp;alt=media\" alt=\"\"></p>\n<p>(Apologies for not logging / plotting the competition metric, pF1, but it's so noisy and I strongly suspect AUC is the validation metric to track).</p>\n<p>Note, a STANDARD EfficientNet-B3 model (non pruned) does much better than the pruned model, and similarly to EfficientNet-B2 - there seems to be something subtle here that is of great significance, and the architecture choices almost certainly are going to matter a lot in this competition.</p>\n<p>Maybe there is something about my training pipeline where over the years I have subconsciously tweaked it for EfficientNet (LRs, optimizers, schedulers etc.) - but currently I am struggling to find anything that rivals these 3-4 year old models</p>",
      "rawMarkdown": "In the HuBMAP competition 3 months ago, most people found transformers performed excellently for histology segmentation. However, there was an exception; for the slices of lung tissue, nothing could match efficientnets, and I couldn't find a great explanation for this.\n\nLung tissue was more sparse than the other organs, and often involved very fine tissue planes and thin features. I assumed it was that. However, there is a dataset called ImageNet-Sketch, which is like ImageNet for hand drawn images. You'd think if being able to identify fine planes rather than texture was relevant, EfficientNet would do well at this? Well it doesn't - Ross Wightman shows [here](https://github.com/rwightman/pytorch-image-models/blob/main/results/results-sketch.csv) that actually ResNext based models and transformers outperform EfficientNet. This is paralleled by the results of the Bengali challenge, where sketched characters had to be identified. Here, the winners had more success with SE-ResNeXT models rather than EfficientNets.\n\nIn this challenge, where mammograms are made up of spindly ethereal structures in relatively sparse imagines, EfficientNets again are proving unreasonably effective for me. SE-ResNeXT and other novel models just consistently underperform for me, whereas a modest EfficientNet-B2 works fine.\n\nIs this other people's experience? Can they explain it?\n\nLast night I adapted an EfficientNet-B3-pruned model (brown in the image below) from Ross' timm repo and trained it for the challenge with identical hyperparameters (including augmentations and batch size) to my standard EfficientNet-B2 training, and  could not match the standard B2's results (blue line in image below). Bizarre. I also show the performance of a similarly trained seresnext50_32x4d model in green. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3108186%2F27e28e56e34e678505404cbdaeaadfb9%2FScreenshot%202022-12-24%20at%2012.02.21.png?generation=1671883395618303&alt=media)\n\n(Apologies for not logging / plotting the competition metric, pF1, but it's so noisy and I strongly suspect AUC is the validation metric to track).\n\nNote, a STANDARD EfficientNet-B3 model (non pruned) does much better than the pruned model, and similarly to EfficientNet-B2 - there seems to be something subtle here that is of great significance, and the architecture choices almost certainly are going to matter a lot in this competition.\n\nMaybe there is something about my training pipeline where over the years I have subconsciously tweaked it for EfficientNet (LRs, optimizers, schedulers etc.) - but currently I am struggling to find anything that rivals these 3-4 year old models",
      "votes": 30
    },
    {
      "id": 2074832,
      "postDate": "2022-12-24T16:37:23.727Z",
      "content": "<p>Maybe the key is that effnets were designed to benefit directly from image size, which is not true for all other architectures (? if you find any please post a link to the paper), and in this competition we have to use large image sizes (e.g. &gt;=1024px) to get good performance, maybe it's somehow related to the superiority of effnets in this challenge, by the way, I'm also using a model from effnet family.</p>",
      "rawMarkdown": "Maybe the key is that effnets were designed to benefit directly from image size, which is not true for all other architectures (? if you find any please post a link to the paper), and in this competition we have to use large image sizes (e.g. >=1024px) to get good performance, maybe it's somehow related to the superiority of effnets in this challenge, by the way, I'm also using a model from effnet family.",
      "votes": 5
    },
    {
      "id": 2079070,
      "postDate": "2022-12-28T23:59:02.470Z",
      "content": "<p>anyone want to try this:<br>\n<a href=\"https://github.com/BMEII-AI/RadImageNet\" target=\"_blank\">https://github.com/BMEII-AI/RadImageNet</a></p>\n<p>radImageNet pretrained models </p>",
      "rawMarkdown": "anyone want to try this:\nhttps://github.com/BMEII-AI/RadImageNet\n\nradImageNet pretrained models ",
      "votes": 3
    },
    {
      "id": 2077897,
      "postDate": "2022-12-27T23:56:04.040Z",
      "content": "<p>i read at least 2 papers that claims inception Resnet-v2 is the best. you may want to try that<br>\n(i think it is because of high size variability of size and shape in the abnormality. There can be no \"clear shape\" in some cases if you check BIRADS images from the web)</p>\n<p><img src=\"https://i.ibb.co/FDWp911/Selection-323.png\" alt=\"https://i.ibb.co/FDWp911/Selection-323.png\"></p>\n<p>[1]MULTI-VIEW DEEP EVIDENTIAL FUSION NEURAL NETWORK FOR ASSESSMENT OF SCREENING MAMMOGRAMS<br>\nUnder review as a conference paper at ICLR 202<br>\n<a href=\"https://openreview.net/pdf?id=snjmwYRuqh\" target=\"_blank\">https://openreview.net/pdf?id=snjmwYRuqh</a></p>",
      "rawMarkdown": "i read at least 2 papers that claims inception Resnet-v2 is the best. you may want to try that\n(i think it is because of high size variability of size and shape in the abnormality. There can be no \"clear shape\" in some cases if you check BIRADS images from the web)\n\n\n![https://i.ibb.co/FDWp911/Selection-323.png](https://i.ibb.co/FDWp911/Selection-323.png)\n\n[1]MULTI-VIEW DEEP EVIDENTIAL FUSION NEURAL NETWORK FOR ASSESSMENT OF SCREENING MAMMOGRAMS\nUnder review as a conference paper at ICLR 202\nhttps://openreview.net/pdf?id=snjmwYRuqh",
      "votes": 3,
      "replies": [
        {
          "id": 2078380,
          "postDate": "2022-12-28T09:38:17.447Z",
          "content": "<p>I tested both - Inception_Resnet_v2 and Densenet169. So far Effnet is better for me (but I am not happy with my current solution and somebody will be able to train inception and densenet better then me).</p>",
          "rawMarkdown": "I tested both - Inception_Resnet_v2 and Densenet169. So far Effnet is better for me (but I am not happy with my current solution and somebody will be able to train inception and densenet better then me).",
          "votes": 1
        }
      ]
    },
    {
      "id": 2074647,
      "postDate": "2022-12-24T13:14:57.113Z",
      "content": "<p>I participated in the same competition and recently had a similar experience. I'm doing research in medical center to classify the conditions of certain cells, find that Transformer model(including ConvNeXT) learns nothing. But SE-ResNeXT, EfficientNet model classifies quite accurately. Interestingly, in Grad-CAM, EfficientNet showed much more detail than SE-ResNeXT. </p>",
      "rawMarkdown": "I participated in the same competition and recently had a similar experience. I'm doing research in medical center to classify the conditions of certain cells, find that Transformer model(including ConvNeXT) learns nothing. But SE-ResNeXT, EfficientNet model classifies quite accurately. Interestingly, in Grad-CAM, EfficientNet showed much more detail than SE-ResNeXT. ",
      "votes": 3
    },
    {
      "id": 2079566,
      "postDate": "2022-12-29T12:31:09.013Z",
      "content": "<p>Same experience here. Usually in most competitions EfficientNets work best for me. I guess their easily adaptable architecture works well for most simple classification tasks also considering the paper increasing image size benefits accuracy:</p>\n<p><a href=\"https://arxiv.org/pdf/1905.11946.pdf\" target=\"_blank\">https://arxiv.org/pdf/1905.11946.pdf</a> </p>\n<p>What is also great about them is their low parameter size and efficient throughput and latency.</p>",
      "rawMarkdown": "Same experience here. Usually in most competitions EfficientNets work best for me. I guess their easily adaptable architecture works well for most simple classification tasks also considering the paper increasing image size benefits accuracy:\n \nhttps://arxiv.org/pdf/1905.11946.pdf \n\nWhat is also great about them is their low parameter size and efficient throughput and latency.",
      "votes": 1
    },
    {
      "id": 2077974,
      "postDate": "2022-12-28T02:07:19.130Z",
      "content": "<p>Thank you for sharing, very helpful 🙏<br>\nIn the validation set oof my model kept assigning very low probabilities to the two classes.<br>\nMay I ask what I'm doing wrong ? </p>\n<ul>\n<li>Training only on the original dataset.</li>\n<li>Image size 1024*1024</li>\n<li>Batch size 32</li>\n<li>stratifying folds with non-overlapping groups(patient_id)</li>\n<li>Used a BalanceSampler in the dataloader to ensure that every batch has a positive case(s). </li>\n<li>resnet18 with drop_rate = 0.3 and drop_path_rate = 0.2 to prevent overfitting.</li>\n<li>Used BCEWithLogitsLoss (I applied weights also pos_weight)</li>\n<li>fixed lr=3e-4<br>\nI'm working only with kaggle GPU.</li>\n</ul>",
      "rawMarkdown": "Thank you for sharing, very helpful 🙏\nIn the validation set oof my model kept assigning very low probabilities to the two classes.\nMay I ask what I'm doing wrong ? \n* Training only on the original dataset.\n* Image size 1024*1024\n* Batch size 32\n* stratifying folds with non-overlapping groups(patient_id)\n* Used a BalanceSampler in the dataloader to ensure that every batch has a positive case(s). \n* resnet18 with drop_rate = 0.3 and drop_path_rate = 0.2 to prevent overfitting.\n* Used BCEWithLogitsLoss (I applied weights also pos_weight)\n* fixed lr=3e-4\nI'm working only with kaggle GPU.",
      "votes": 1,
      "replies": [
        {
          "id": 2077977,
          "postDate": "2022-12-28T02:11:23.410Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 2075074,
      "postDate": "2022-12-25T03:12:29.007Z",
      "content": "<p>\"Is this other people's experience? Can they explain it?\"</p>\n<p>you can check the CAM map of Effnet-B2 and another model.<br>\nyou can try other AI explainability lib or paper.</p>\n<p>it is how to conclude without looking at interpretability results.</p>\n<hr>\n<p>in theory, final performance is limited by data. </p>\n<p>when you consider only the model, its performance is limited on how you can prevent overfitting (assuming model capacity &gt; data capacity). </p>\n<p>the classic example is efficient-net. It out performs the initial resnet paper.<br>\nbut in recent resnet paper, they apply modern training techniques (like new augmentation, optimizers, initialisation,  and those found in the efficient net paper), resnet out performs efficient net.</p>",
      "rawMarkdown": "\"Is this other people's experience? Can they explain it?\"\n\nyou can check the CAM map of Effnet-B2 and another model.\nyou can try other AI explainability lib or paper.\n\nit is how to conclude without looking at interpretability results.\n\n---\n\nin theory, final performance is limited by data. \n\nwhen you consider only the model, its performance is limited on how you can prevent overfitting (assuming model capacity > data capacity). \n\nthe classic example is efficient-net. It out performs the initial resnet paper.\nbut in recent resnet paper, they apply modern training techniques (like new augmentation, optimizers, initialisation,  and those found in the efficient net paper), resnet out performs efficient net.\n\n \n",
      "votes": 1,
      "replies": [
        {
          "id": 2075321,
          "postDate": "2022-12-25T08:52:12.170Z",
          "content": "<p>Yes, except I am finding SE-ResNeXT struggled to overfit as well as fit appropriately (see the training accuracy and AUC plots on the left)…</p>",
          "rawMarkdown": "Yes, except I am finding SE-ResNeXT struggled to overfit as well as fit appropriately (see the training accuracy and AUC plots on the left)..."
        }
      ]
    },
    {
      "id": 2075070,
      "postDate": "2022-12-25T03:01:26.703Z",
      "content": "<p>\"In the HuBMAP competition 3 months ago, most people found transformers performed excellently for histology segmentation.\"</p>\n<p>i am now doing experiments on transforms (like HuBMAP) to predict 1/32 scaled heatmap on external data (aux loss) + image label on kaggle. The preliminary results are pretty good. (very fast convergence and good CAM maps/attention maps)</p>\n<p>transformers are very good for this type of problem. the attention captures global interation ( (unlike convolution which is restricted to the receptive field) </p>\n<p>i suspect transformers will be the winner. … more results to come soon in my post later.</p>",
      "rawMarkdown": "\"In the HuBMAP competition 3 months ago, most people found transformers performed excellently for histology segmentation.\"\n\ni am now doing experiments on transforms (like HuBMAP) to predict 1/32 scaled heatmap on external data (aux loss) + image label on kaggle. The preliminary results are pretty good. (very fast convergence and good CAM maps/attention maps)\n\ntransformers are very good for this type of problem. the attention captures global interation ( (unlike convolution which is restricted to the receptive field) \n\ni suspect transformers will be the winner. ... more results to come soon in my post later.",
      "votes": 1
    },
    {
      "id": 2074826,
      "postDate": "2022-12-24T16:33:39.070Z",
      "content": "<p>The pruned model usually is poor for transfer learning.<br>\nThe pruning is based on Image Net, i.e. the features specialised/biased to Image Net are kept and the rest (important for other domains) are pruned/removed</p>",
      "rawMarkdown": "The pruned model usually is poor for transfer learning.\nThe pruning is based on Image Net, i.e. the features specialised/biased to Image Net are kept and the rest (important for other domains) are pruned/removed",
      "votes": 1,
      "replies": [
        {
          "id": 2074834,
          "postDate": "2022-12-24T16:37:56.403Z",
          "content": "<p>Whilst I sort of accept that, transfer learning from ImageNet is generally overhyped as it is. Google have shown you lose almost no performance resetting your pretrained model weights to random, if you have as big a domain shift as going into medical imaging:</p>\n<p><a href=\"https://ai.googleblog.com/2019/12/understanding-transfer-learning-for.html\" target=\"_blank\">https://ai.googleblog.com/2019/12/understanding-transfer-learning-for.html</a></p>\n<p>So whilst this might contribute a bit, I don't think it can be the predominant explanation?</p>",
          "rawMarkdown": "Whilst I sort of accept that, transfer learning from ImageNet is generally overhyped as it is. Google have shown you lose almost no performance resetting your pretrained model weights to random, if you have as big a domain shift as going into medical imaging:\n\nhttps://ai.googleblog.com/2019/12/understanding-transfer-learning-for.html\n\nSo whilst this might contribute a bit, I don't think it can be the predominant explanation?",
          "votes": 3
        }
      ]
    },
    {
      "id": 2075478,
      "postDate": "2022-12-25T13:29:35.090Z",
      "content": "<p>My resnext50 model gave me 0.5 LB. It's much faster than effnets.</p>",
      "rawMarkdown": "My resnext50 model gave me 0.5 LB. It's much faster than effnets.",
      "votes": 2,
      "replies": [
        {
          "id": 2075486,
          "postDate": "2022-12-25T13:46:33.993Z",
          "content": "<p>That's very good - is that a single fold or 4/5fold?</p>",
          "rawMarkdown": "That's very good - is that a single fold or 4/5fold?",
          "replies": [
            {
              "id": 2075491,
              "postDate": "2022-12-25T13:55:20.323Z",
              "content": "<blockquote>\n  <p>is that a single fold or 4/5fold?</p>\n</blockquote>\n<p>5fold.</p>",
              "rawMarkdown": ">is that a single fold or 4/5fold?\n\n5fold."
            },
            {
              "id": 2075575,
              "postDate": "2022-12-25T15:48:20.360Z",
              "content": "<p>single efficientnet-b2 at 2048 for one fold is LB 0.51.<br>\nmaybe the split for this fold is special….</p>",
              "rawMarkdown": "single efficientnet-b2 at 2048 for one fold is LB 0.51.\nmaybe the split for this fold is special....",
              "votes": 3
            },
            {
              "id": 2078739,
              "postDate": "2022-12-28T15:43:59.870Z",
              "content": "<p>2048!!! Is this single laterality single view image or have you combined CC and MLO for L and R?</p>",
              "rawMarkdown": "2048!!! Is this single laterality single view image or have you combined CC and MLO for L and R?\n"
            }
          ]
        },
        {
          "id": 2075792,
          "postDate": "2022-12-25T22:04:16.787Z",
          "content": "<p>Do you think it was just lucky weights, or was the CV good too?</p>",
          "rawMarkdown": "Do you think it was just lucky weights, or was the CV good too?",
          "replies": [
            {
              "id": 2076590,
              "postDate": "2022-12-26T16:11:28.623Z",
              "content": "<p>\"Do you think it was just lucky weights,\"</p>\n<p>my suggestion is not to spend time on producing a stable CV-LB. </p>\n<p>I check the CAM of train images. some of them are wrong, i.e.  even the train are wrong so you can forget about the validation. the fact that the train are wrong shows that there isn't enough train data and i would expect LB/CV score to be unstable (which is proven by my observations and others too )</p>\n<p>Rather, try to use external data with mask annotation to build a patch classifier or heat map (coarse segmentation) predictor. Then think of a way to use these results. The prediction will be more stable.<br>\n(e.g. pesudo mask label the kaggle train image, etc …)</p>\n<hr>\n<p>i also check the kaggle pos train images. i feel that some of the mass, calcification are very small. i don't think 1024 is enough. (in fact when i check the literature, people are using image in the near 2400~2600) </p>\n<hr>\n<p>i am expecting score of near 0.65 (using high resolution +  multi-view input + heatmap) in the final kaggle ranking after 2 months</p>",
              "rawMarkdown": "\"Do you think it was just lucky weights,\"\n\nmy suggestion is not to spend time on producing a stable CV-LB. \n\nI check the CAM of train images. some of them are wrong, i.e.  even the train are wrong so you can forget about the validation. the fact that the train are wrong shows that there isn't enough train data and i would expect LB/CV score to be unstable (which is proven by my observations and others too )\n\nRather, try to use external data with mask annotation to build a patch classifier or heat map (coarse segmentation) predictor. Then think of a way to use these results. The prediction will be more stable.\n(e.g. pesudo mask label the kaggle train image, etc ...)\n\n---\n\ni also check the kaggle pos train images. i feel that some of the mass, calcification are very small. i don't think 1024 is enough. (in fact when i check the literature, people are using image in the near 2400~2600) \n\n---\n\ni am expecting score of near 0.65 (using high resolution +  multi-view input + heatmap) in the final kaggle ranking after 2 months",
              "votes": 3
            }
          ]
        },
        {
          "id": 2076544,
          "postDate": "2022-12-26T15:36:49.117Z",
          "content": "<p>ya thought same.</p>",
          "rawMarkdown": "ya thought same."
        }
      ]
    },
    {
      "id": 2077111,
      "postDate": "2022-12-27T08:40:38.093Z",
      "content": "<p>Could you share your LB score for efficientnetb2 and b3?</p>",
      "rawMarkdown": "Could you share your LB score for efficientnetb2 and b3?"
    }
  ],
  "comments": [
    {
      "id": 2074832,
      "author_name": "slime",
      "author_url": "",
      "post_date": "2022-12-24T16:37:23.727000",
      "content": "<p>Maybe the key is that effnets were designed to benefit directly from image size, which is not true for all other architectures (? if you find any please post a link to the paper), and in this competition we have to use large image sizes (e.g. &gt;=1024px) to get good performance, maybe it's somehow related to the superiority of effnets in this challenge, by the way, I'm also using a model from effnet family.</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 2079070,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2022-12-28T23:59:02.470000",
      "content": "<p>anyone want to try this:<br>\n<a href=\"https://github.com/BMEII-AI/RadImageNet\" target=\"_blank\">https://github.com/BMEII-AI/RadImageNet</a></p>\n<p>radImageNet pretrained models </p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2077897,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2022-12-27T23:56:04.040000",
      "content": "<p>i read at least 2 papers that claims inception Resnet-v2 is the best. you may want to try that<br>\n(i think it is because of high size variability of size and shape in the abnormality. There can be no \"clear shape\" in some cases if you check BIRADS images from the web)</p>\n<p><img src=\"https://i.ibb.co/FDWp911/Selection-323.png\" alt=\"https://i.ibb.co/FDWp911/Selection-323.png\"></p>\n<p>[1]MULTI-VIEW DEEP EVIDENTIAL FUSION NEURAL NETWORK FOR ASSESSMENT OF SCREENING MAMMOGRAMS<br>\nUnder review as a conference paper at ICLR 202<br>\n<a href=\"https://openreview.net/pdf?id=snjmwYRuqh\" target=\"_blank\">https://openreview.net/pdf?id=snjmwYRuqh</a></p>",
      "votes": 3,
      "replies": [
        {
          "id": 2078380,
          "author_name": "Remek Kinas",
          "author_url": "",
          "post_date": "2022-12-28T09:38:17.447000",
          "content": "<p>I tested both - Inception_Resnet_v2 and Densenet169. So far Effnet is better for me (but I am not happy with my current solution and somebody will be able to train inception and densenet better then me).</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2074647,
      "author_name": "olivepicker",
      "author_url": "",
      "post_date": "2022-12-24T13:14:57.113000",
      "content": "<p>I participated in the same competition and recently had a similar experience. I'm doing research in medical center to classify the conditions of certain cells, find that Transformer model(including ConvNeXT) learns nothing. But SE-ResNeXT, EfficientNet model classifies quite accurately. Interestingly, in Grad-CAM, EfficientNet showed much more detail than SE-ResNeXT. </p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2079566,
      "author_name": "Ali Abdin",
      "author_url": "",
      "post_date": "2022-12-29T12:31:09.013000",
      "content": "<p>Same experience here. Usually in most competitions EfficientNets work best for me. I guess their easily adaptable architecture works well for most simple classification tasks also considering the paper increasing image size benefits accuracy:</p>\n<p><a href=\"https://arxiv.org/pdf/1905.11946.pdf\" target=\"_blank\">https://arxiv.org/pdf/1905.11946.pdf</a> </p>\n<p>What is also great about them is their low parameter size and efficient throughput and latency.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2077974,
      "author_name": "MohamedAmine SAIGHI",
      "author_url": "",
      "post_date": "2022-12-28T02:07:19.130000",
      "content": "<p>Thank you for sharing, very helpful 🙏<br>\nIn the validation set oof my model kept assigning very low probabilities to the two classes.<br>\nMay I ask what I'm doing wrong ? </p>\n<ul>\n<li>Training only on the original dataset.</li>\n<li>Image size 1024*1024</li>\n<li>Batch size 32</li>\n<li>stratifying folds with non-overlapping groups(patient_id)</li>\n<li>Used a BalanceSampler in the dataloader to ensure that every batch has a positive case(s). </li>\n<li>resnet18 with drop_rate = 0.3 and drop_path_rate = 0.2 to prevent overfitting.</li>\n<li>Used BCEWithLogitsLoss (I applied weights also pos_weight)</li>\n<li>fixed lr=3e-4<br>\nI'm working only with kaggle GPU.</li>\n</ul>",
      "votes": 1,
      "replies": [
        {
          "id": 2077977,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-12-28T02:11:23.410000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2075074,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2022-12-25T03:12:29.007000",
      "content": "<p>\"Is this other people's experience? Can they explain it?\"</p>\n<p>you can check the CAM map of Effnet-B2 and another model.<br>\nyou can try other AI explainability lib or paper.</p>\n<p>it is how to conclude without looking at interpretability results.</p>\n<hr>\n<p>in theory, final performance is limited by data. </p>\n<p>when you consider only the model, its performance is limited on how you can prevent overfitting (assuming model capacity &gt; data capacity). </p>\n<p>the classic example is efficient-net. It out performs the initial resnet paper.<br>\nbut in recent resnet paper, they apply modern training techniques (like new augmentation, optimizers, initialisation,  and those found in the efficient net paper), resnet out performs efficient net.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2075321,
          "author_name": "James Howard",
          "author_url": "",
          "post_date": "2022-12-25T08:52:12.170000",
          "content": "<p>Yes, except I am finding SE-ResNeXT struggled to overfit as well as fit appropriately (see the training accuracy and AUC plots on the left)…</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2075070,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2022-12-25T03:01:26.703000",
      "content": "<p>\"In the HuBMAP competition 3 months ago, most people found transformers performed excellently for histology segmentation.\"</p>\n<p>i am now doing experiments on transforms (like HuBMAP) to predict 1/32 scaled heatmap on external data (aux loss) + image label on kaggle. The preliminary results are pretty good. (very fast convergence and good CAM maps/attention maps)</p>\n<p>transformers are very good for this type of problem. the attention captures global interation ( (unlike convolution which is restricted to the receptive field) </p>\n<p>i suspect transformers will be the winner. … more results to come soon in my post later.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2074826,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2022-12-24T16:33:39.070000",
      "content": "<p>The pruned model usually is poor for transfer learning.<br>\nThe pruning is based on Image Net, i.e. the features specialised/biased to Image Net are kept and the rest (important for other domains) are pruned/removed</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2074834,
          "author_name": "James Howard",
          "author_url": "",
          "post_date": "2022-12-24T16:37:56.403000",
          "content": "<p>Whilst I sort of accept that, transfer learning from ImageNet is generally overhyped as it is. Google have shown you lose almost no performance resetting your pretrained model weights to random, if you have as big a domain shift as going into medical imaging:</p>\n<p><a href=\"https://ai.googleblog.com/2019/12/understanding-transfer-learning-for.html\" target=\"_blank\">https://ai.googleblog.com/2019/12/understanding-transfer-learning-for.html</a></p>\n<p>So whilst this might contribute a bit, I don't think it can be the predominant explanation?</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 2075478,
      "author_name": "Leon",
      "author_url": "",
      "post_date": "2022-12-25T13:29:35.090000",
      "content": "<p>My resnext50 model gave me 0.5 LB. It's much faster than effnets.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2075486,
          "author_name": "James Howard",
          "author_url": "",
          "post_date": "2022-12-25T13:46:33.993000",
          "content": "<p>That's very good - is that a single fold or 4/5fold?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2075491,
              "author_name": "Leon",
              "author_url": "",
              "post_date": "2022-12-25T13:55:20.323000",
              "content": "<blockquote>\n  <p>is that a single fold or 4/5fold?</p>\n</blockquote>\n<p>5fold.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2075575,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2022-12-25T15:48:20.360000",
              "content": "<p>single efficientnet-b2 at 2048 for one fold is LB 0.51.<br>\nmaybe the split for this fold is special….</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2078739,
              "author_name": "NitinKshatriya",
              "author_url": "",
              "post_date": "2022-12-28T15:43:59.870000",
              "content": "<p>2048!!! Is this single laterality single view image or have you combined CC and MLO for L and R?</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2075792,
          "author_name": "Ivan Aerlic",
          "author_url": "",
          "post_date": "2022-12-25T22:04:16.787000",
          "content": "<p>Do you think it was just lucky weights, or was the CV good too?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2076590,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2022-12-26T16:11:28.623000",
              "content": "<p>\"Do you think it was just lucky weights,\"</p>\n<p>my suggestion is not to spend time on producing a stable CV-LB. </p>\n<p>I check the CAM of train images. some of them are wrong, i.e.  even the train are wrong so you can forget about the validation. the fact that the train are wrong shows that there isn't enough train data and i would expect LB/CV score to be unstable (which is proven by my observations and others too )</p>\n<p>Rather, try to use external data with mask annotation to build a patch classifier or heat map (coarse segmentation) predictor. Then think of a way to use these results. The prediction will be more stable.<br>\n(e.g. pesudo mask label the kaggle train image, etc …)</p>\n<hr>\n<p>i also check the kaggle pos train images. i feel that some of the mass, calcification are very small. i don't think 1024 is enough. (in fact when i check the literature, people are using image in the near 2400~2600) </p>\n<hr>\n<p>i am expecting score of near 0.65 (using high resolution +  multi-view input + heatmap) in the final kaggle ranking after 2 months</p>",
              "votes": 3,
              "replies": []
            }
          ]
        },
        {
          "id": 2076544,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-12-26T15:36:49.117000",
          "content": "<p>ya thought same.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2077111,
      "author_name": "HAOYI ZHONG",
      "author_url": "",
      "post_date": "2022-12-27T08:40:38.093000",
      "content": "<p>Could you share your LB score for efficientnetb2 and b3?</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2074611": "In the HuBMAP competition 3 months ago, most people found transformers performed excellently for histology segmentation. However, there was an exception; for the slices of lung tissue, nothing could match efficientnets, and I couldn't find a great explanation for this.\n\nLung tissue was more sparse than the other organs, and often involved very fine tissue planes and thin features. I assumed it was that. However, there is a dataset called ImageNet-Sketch, which is like ImageNet for hand drawn images. You'd think if being able to identify fine planes rather than texture was relevant, EfficientNet would do well at this? Well it doesn't - Ross Wightman shows [here](https://github.com/rwightman/pytorch-image-models/blob/main/results/results-sketch.csv) that actually ResNext based models and transformers outperform EfficientNet. This is paralleled by the results of the Bengali challenge, where sketched characters had to be identified. Here, the winners had more success with SE-ResNeXT models rather than EfficientNets.\n\nIn this challenge, where mammograms are made up of spindly ethereal structures in relatively sparse imagines, EfficientNets again are proving unreasonably effective for me. SE-ResNeXT and other novel models just consistently underperform for me, whereas a modest EfficientNet-B2 works fine.\n\nIs this other people's experience? Can they explain it?\n\nLast night I adapted an EfficientNet-B3-pruned model (brown in the image below) from Ross' timm repo and trained it for the challenge with identical hyperparameters (including augmentations and batch size) to my standard EfficientNet-B2 training, and  could not match the standard B2's results (blue line in image below). Bizarre. I also show the performance of a similarly trained seresnext50_32x4d model in green. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3108186%2F27e28e56e34e678505404cbdaeaadfb9%2FScreenshot%202022-12-24%20at%2012.02.21.png?generation=1671883395618303&alt=media)\n\n(Apologies for not logging / plotting the competition metric, pF1, but it's so noisy and I strongly suspect AUC is the validation metric to track).\n\nNote, a STANDARD EfficientNet-B3 model (non pruned) does much better than the pruned model, and similarly to EfficientNet-B2 - there seems to be something subtle here that is of great significance, and the architecture choices almost certainly are going to matter a lot in this competition.\n\nMaybe there is something about my training pipeline where over the years I have subconsciously tweaked it for EfficientNet (LRs, optimizers, schedulers etc.) - but currently I am struggling to find anything that rivals these 3-4 year old models",
    "2074832": "Maybe the key is that effnets were designed to benefit directly from image size, which is not true for all other architectures (? if you find any please post a link to the paper), and in this competition we have to use large image sizes (e.g. >=1024px) to get good performance, maybe it's somehow related to the superiority of effnets in this challenge, by the way, I'm also using a model from effnet family.",
    "2079070": "anyone want to try this:\nhttps://github.com/BMEII-AI/RadImageNet\n\nradImageNet pretrained models ",
    "2077897": "i read at least 2 papers that claims inception Resnet-v2 is the best. you may want to try that\n(i think it is because of high size variability of size and shape in the abnormality. There can be no \"clear shape\" in some cases if you check BIRADS images from the web)\n\n\n![https://i.ibb.co/FDWp911/Selection-323.png](https://i.ibb.co/FDWp911/Selection-323.png)\n\n[1]MULTI-VIEW DEEP EVIDENTIAL FUSION NEURAL NETWORK FOR ASSESSMENT OF SCREENING MAMMOGRAMS\nUnder review as a conference paper at ICLR 202\nhttps://openreview.net/pdf?id=snjmwYRuqh",
    "2074647": "I participated in the same competition and recently had a similar experience. I'm doing research in medical center to classify the conditions of certain cells, find that Transformer model(including ConvNeXT) learns nothing. But SE-ResNeXT, EfficientNet model classifies quite accurately. Interestingly, in Grad-CAM, EfficientNet showed much more detail than SE-ResNeXT. ",
    "2079566": "Same experience here. Usually in most competitions EfficientNets work best for me. I guess their easily adaptable architecture works well for most simple classification tasks also considering the paper increasing image size benefits accuracy:\n \nhttps://arxiv.org/pdf/1905.11946.pdf \n\nWhat is also great about them is their low parameter size and efficient throughput and latency.",
    "2077974": "Thank you for sharing, very helpful 🙏\nIn the validation set oof my model kept assigning very low probabilities to the two classes.\nMay I ask what I'm doing wrong ? \n* Training only on the original dataset.\n* Image size 1024*1024\n* Batch size 32\n* stratifying folds with non-overlapping groups(patient_id)\n* Used a BalanceSampler in the dataloader to ensure that every batch has a positive case(s). \n* resnet18 with drop_rate = 0.3 and drop_path_rate = 0.2 to prevent overfitting.\n* Used BCEWithLogitsLoss (I applied weights also pos_weight)\n* fixed lr=3e-4\nI'm working only with kaggle GPU.",
    "2075074": "\"Is this other people's experience? Can they explain it?\"\n\nyou can check the CAM map of Effnet-B2 and another model.\nyou can try other AI explainability lib or paper.\n\nit is how to conclude without looking at interpretability results.\n\n---\n\nin theory, final performance is limited by data. \n\nwhen you consider only the model, its performance is limited on how you can prevent overfitting (assuming model capacity > data capacity). \n\nthe classic example is efficient-net. It out performs the initial resnet paper.\nbut in recent resnet paper, they apply modern training techniques (like new augmentation, optimizers, initialisation,  and those found in the efficient net paper), resnet out performs efficient net.\n\n \n",
    "2075070": "\"In the HuBMAP competition 3 months ago, most people found transformers performed excellently for histology segmentation.\"\n\ni am now doing experiments on transforms (like HuBMAP) to predict 1/32 scaled heatmap on external data (aux loss) + image label on kaggle. The preliminary results are pretty good. (very fast convergence and good CAM maps/attention maps)\n\ntransformers are very good for this type of problem. the attention captures global interation ( (unlike convolution which is restricted to the receptive field) \n\ni suspect transformers will be the winner. ... more results to come soon in my post later.",
    "2074826": "The pruned model usually is poor for transfer learning.\nThe pruning is based on Image Net, i.e. the features specialised/biased to Image Net are kept and the rest (important for other domains) are pruned/removed",
    "2075478": "My resnext50 model gave me 0.5 LB. It's much faster than effnets.",
    "2077111": "Could you share your LB score for efficientnetb2 and b3?"
  }
}