{
  "id": 391263,
  "title": "[FR Team] 20th place solution",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/391263",
  "author_name": "MPWARE",
  "post_date": "2023-02-28T20:58:03.619000",
  "votes": 14,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hi Kagglers,</p>\n<p>Thanks for this nice competition!</p>\n<p>First, I would like to thank <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> and <a href=\"https://www.kaggle.com/optimo\" target=\"_blank\">@optimo</a> who helped me to make my dream true: Make a French team for one Kaggle competition. My expectation was to finish top #2 like in the FIFA world cup 🙂 but I’m quite happy with the current result. Some insights of our solution:</p>\n<h2>Theo’s Part</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2F858b18feec74c7a59e178265eebcd519%2Ftheo.png?generation=1677617629716416&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>Pretrain for 5 epochs using VinDr data and the BIRADS target</li>\n<li>Finetune for 5 epochs on the competition data + external data<ul>\n<li>bs=8 (6 for v2-s), lr=4e-4 (3e-4 for v2-s), Ranger, Linear Schedule with no warm up</li></ul></li>\n<li>External data varies among models, some only use CBIS. I also trained models with pseudo-labels on VinDr that are used in the final ensemble.</li>\n<li>BCE loss with <strong>no class weight</strong>. Some models use BIRADS as an auxiliary target</li>\n<li>No over/under-sampling</li>\n</ul>\n<p>This is my best submitted ensemble. Models were changed a bit in the final ensemble to maximize CV.</p>\n<h2>Optimo’s Part</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2F845fc6e58654ca5ef2987f5b4a9d79eb%2Foptimo.png?generation=1677617643849717&amp;alt=media\" alt=\"\"></p>\n<p>Likewise my timm backbones were pretrained on VinDr by predicting BIRADS. I used CBIS, Vindr PL as external data (boosted my CV by ~0.02).</p>\n<p>As Theo had better results than mine on ‘standard models’ I tried to provide as much diversity as possible with ideas coming from research papers. I ended up using two different architectures :</p>\n<ul>\n<li>a modified version of GMIC: inspired by this code I implemented a version which allows any timm network as global network and/or local network. I also changed the crop normalization in order to fully use the local brightness. As I was at first only changing the local network I used the public pretrained weights on NYU datasets. Which ended up being forbidden 2 days before the end of the competition. So I trained from scratch a model with 2xeffnet-b0 networks as local and global networks with input size (1472, 960), crop size 384 and 4 extracted patches. CV : <strong>43.69</strong></li>\n<li>a modified version of MVCCL (BiView): inspired by <a href=\"https://www.kaggle.com/heng\" target=\"_blank\">@heng</a> code <a href=\"https://www.google.com/url?q=https://www.kaggle.com/code/hengck23/mvccl-model-for-admani-dataset&amp;sa=D&amp;source=editors&amp;ust=1677621085747449&amp;usg=AOvVaw3ovbCIfqBBwEX0rKY3P-iF\" target=\"_blank\">https://www.kaggle.com/code/hengck23/mvccl-model-for-admani-dataset</a>&nbsp;The main changes where to allow any timm model as main network and use the representation from both main and auxiliary view in the final summation. I used two models in final submission with input size (1536, 768): one b0 with 16 attention heads (<strong>CV: 43.45</strong>) and one b2 with 8 attention heads (<strong>CV 44.99</strong>)</li>\n<li>ensembling those 3 models at laterality level gave <strong>CV 46.94</strong></li>\n</ul>\n<h2>MPWARE’s Part</h2>\n<ul>\n<li>DALI decoder + YoloX ROI followed by crop and resize with aspect ratio = 1 to generate images with height=1024. No windowing.</li>\n<li>External data included: CBIS-DDSM (mass + calcification full images) + PASM</li>\n<li>Training pipeline with limited class oversampling and weighted CrossEntropyLoss on positive labels.</li>\n<li>Augmentations: Random crop, H/V flips, minor RotateShiftResize, Noise/Blur, Random BrightnessContrast and Coarse Dropout</li>\n<li>Backbones: NFNet + NextViT</li>\n<li>No GeM but regular adaptive average pooling.</li>\n<li>Max aggregation for laterality.</li>\n<li>Spent some time to get TensorRT 1.3 working to compile more backbones.</li>\n<li>CV comparable to Optimo’s ensemble at the end (but big CV/LB gap, got LB=0.61)</li>\n</ul>\n<p>What did not work for me:</p>\n<ul>\n<li>Wavelets additional layer</li>\n<li>Age as additional input feature</li>\n<li>Mixup augmentation</li>\n<li>VinDr as external data (no boost compared to CBIS)</li>\n<li>High Resolution with stride=1, it worked at the beginning with EffNet but becomes useless when moving to some different backbones</li>\n<li>Level 2 model based on embeddings.</li>\n</ul>\n<p>Take away: Learn a lot again, great teammates, nice competition, bad metric, hope to make another one with more FR teammates in the coming months.</p>\n<h2>Final submissions &amp; Randomness</h2>\n<p>Our two selected submissions were determined on CV :</p>\n<ul>\n<li><p>A blend of 7 models, with CV 0.555</p>\n<ul>\n<li>Public 0.61, private 0.48</li>\n<li>For this submission we used fullfit models except for MPWARE’s models which were 4 folds. We had a similar submission with a lower public LB that used only fullfit models and that would have ranked #10.</li></ul></li>\n<li><p>A vote of our 3 pipelines with CV 0.547</p>\n<ul>\n<li>Public 0.6, private 0.48</li>\n<li>This vote was updated on the last day, we had a very similar vote with Public 0.61 private 0.52 before that</li></ul></li>\n</ul>\n<p>We got unlucky with submissions selection, our CV 0.5+ blends <strong>all scored 0.48 private&nbsp;or above</strong>&nbsp;(best 0.52, avg 0.495). Unfortunately we chose two subs on the lower end of the gaussian. Even simply selecting our best public LB would’ve put us in the top 8. The metric was too random, next time we’re staying away from F1-score competitions :)</p>",
  "messages": [
    {
      "id": 2163479,
      "postDate": "2023-02-28T20:58:03.620Z",
      "content": "<p>Hi Kagglers,</p>\n<p>Thanks for this nice competition!</p>\n<p>First, I would like to thank <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> and <a href=\"https://www.kaggle.com/optimo\" target=\"_blank\">@optimo</a> who helped me to make my dream true: Make a French team for one Kaggle competition. My expectation was to finish top #2 like in the FIFA world cup 🙂 but I’m quite happy with the current result. Some insights of our solution:</p>\n<h2>Theo’s Part</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2F858b18feec74c7a59e178265eebcd519%2Ftheo.png?generation=1677617629716416&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>Pretrain for 5 epochs using VinDr data and the BIRADS target</li>\n<li>Finetune for 5 epochs on the competition data + external data<ul>\n<li>bs=8 (6 for v2-s), lr=4e-4 (3e-4 for v2-s), Ranger, Linear Schedule with no warm up</li></ul></li>\n<li>External data varies among models, some only use CBIS. I also trained models with pseudo-labels on VinDr that are used in the final ensemble.</li>\n<li>BCE loss with <strong>no class weight</strong>. Some models use BIRADS as an auxiliary target</li>\n<li>No over/under-sampling</li>\n</ul>\n<p>This is my best submitted ensemble. Models were changed a bit in the final ensemble to maximize CV.</p>\n<h2>Optimo’s Part</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2F845fc6e58654ca5ef2987f5b4a9d79eb%2Foptimo.png?generation=1677617643849717&amp;alt=media\" alt=\"\"></p>\n<p>Likewise my timm backbones were pretrained on VinDr by predicting BIRADS. I used CBIS, Vindr PL as external data (boosted my CV by ~0.02).</p>\n<p>As Theo had better results than mine on ‘standard models’ I tried to provide as much diversity as possible with ideas coming from research papers. I ended up using two different architectures :</p>\n<ul>\n<li>a modified version of GMIC: inspired by this code I implemented a version which allows any timm network as global network and/or local network. I also changed the crop normalization in order to fully use the local brightness. As I was at first only changing the local network I used the public pretrained weights on NYU datasets. Which ended up being forbidden 2 days before the end of the competition. So I trained from scratch a model with 2xeffnet-b0 networks as local and global networks with input size (1472, 960), crop size 384 and 4 extracted patches. CV : <strong>43.69</strong></li>\n<li>a modified version of MVCCL (BiView): inspired by <a href=\"https://www.kaggle.com/heng\" target=\"_blank\">@heng</a> code <a href=\"https://www.google.com/url?q=https://www.kaggle.com/code/hengck23/mvccl-model-for-admani-dataset&amp;sa=D&amp;source=editors&amp;ust=1677621085747449&amp;usg=AOvVaw3ovbCIfqBBwEX0rKY3P-iF\" target=\"_blank\">https://www.kaggle.com/code/hengck23/mvccl-model-for-admani-dataset</a>&nbsp;The main changes where to allow any timm model as main network and use the representation from both main and auxiliary view in the final summation. I used two models in final submission with input size (1536, 768): one b0 with 16 attention heads (<strong>CV: 43.45</strong>) and one b2 with 8 attention heads (<strong>CV 44.99</strong>)</li>\n<li>ensembling those 3 models at laterality level gave <strong>CV 46.94</strong></li>\n</ul>\n<h2>MPWARE’s Part</h2>\n<ul>\n<li>DALI decoder + YoloX ROI followed by crop and resize with aspect ratio = 1 to generate images with height=1024. No windowing.</li>\n<li>External data included: CBIS-DDSM (mass + calcification full images) + PASM</li>\n<li>Training pipeline with limited class oversampling and weighted CrossEntropyLoss on positive labels.</li>\n<li>Augmentations: Random crop, H/V flips, minor RotateShiftResize, Noise/Blur, Random BrightnessContrast and Coarse Dropout</li>\n<li>Backbones: NFNet + NextViT</li>\n<li>No GeM but regular adaptive average pooling.</li>\n<li>Max aggregation for laterality.</li>\n<li>Spent some time to get TensorRT 1.3 working to compile more backbones.</li>\n<li>CV comparable to Optimo’s ensemble at the end (but big CV/LB gap, got LB=0.61)</li>\n</ul>\n<p>What did not work for me:</p>\n<ul>\n<li>Wavelets additional layer</li>\n<li>Age as additional input feature</li>\n<li>Mixup augmentation</li>\n<li>VinDr as external data (no boost compared to CBIS)</li>\n<li>High Resolution with stride=1, it worked at the beginning with EffNet but becomes useless when moving to some different backbones</li>\n<li>Level 2 model based on embeddings.</li>\n</ul>\n<p>Take away: Learn a lot again, great teammates, nice competition, bad metric, hope to make another one with more FR teammates in the coming months.</p>\n<h2>Final submissions &amp; Randomness</h2>\n<p>Our two selected submissions were determined on CV :</p>\n<ul>\n<li><p>A blend of 7 models, with CV 0.555</p>\n<ul>\n<li>Public 0.61, private 0.48</li>\n<li>For this submission we used fullfit models except for MPWARE’s models which were 4 folds. We had a similar submission with a lower public LB that used only fullfit models and that would have ranked #10.</li></ul></li>\n<li><p>A vote of our 3 pipelines with CV 0.547</p>\n<ul>\n<li>Public 0.6, private 0.48</li>\n<li>This vote was updated on the last day, we had a very similar vote with Public 0.61 private 0.52 before that</li></ul></li>\n</ul>\n<p>We got unlucky with submissions selection, our CV 0.5+ blends <strong>all scored 0.48 private&nbsp;or above</strong>&nbsp;(best 0.52, avg 0.495). Unfortunately we chose two subs on the lower end of the gaussian. Even simply selecting our best public LB would’ve put us in the top 8. The metric was too random, next time we’re staying away from F1-score competitions :)</p>",
      "rawMarkdown": "Hi Kagglers,\n\nThanks for this nice competition!\n\nFirst, I would like to thank @theoviel and @optimo who helped me to make my dream true: Make a French team for one Kaggle competition. My expectation was to finish top #2 like in the FIFA world cup 🙂 but I’m quite happy with the current result. Some insights of our solution:\n\n## Theo’s Part\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2F858b18feec74c7a59e178265eebcd519%2Ftheo.png?generation=1677617629716416&alt=media)\n\n* Pretrain for 5 epochs using VinDr data and the BIRADS target\n* Finetune for 5 epochs on the competition data + external data\n  * bs=8 (6 for v2-s), lr=4e-4 (3e-4 for v2-s), Ranger, Linear Schedule with no warm up\n* External data varies among models, some only use CBIS. I also trained models with pseudo-labels on VinDr that are used in the final ensemble.\n* BCE loss with **no class weight**. Some models use BIRADS as an auxiliary target\n* No over/under-sampling\n\nThis is my best submitted ensemble. Models were changed a bit in the final ensemble to maximize CV.\n\n## Optimo’s Part\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2F845fc6e58654ca5ef2987f5b4a9d79eb%2Foptimo.png?generation=1677617643849717&alt=media)\n\nLikewise my timm backbones were pretrained on VinDr by predicting BIRADS. I used CBIS, Vindr PL as external data (boosted my CV by ~0.02).\n\nAs Theo had better results than mine on ‘standard models’ I tried to provide as much diversity as possible with ideas coming from research papers. I ended up using two different architectures :\n\n* a modified version of GMIC: inspired by this code I implemented a version which allows any timm network as global network and/or local network. I also changed the crop normalization in order to fully use the local brightness. As I was at first only changing the local network I used the public pretrained weights on NYU datasets. Which ended up being forbidden 2 days before the end of the competition. So I trained from scratch a model with 2xeffnet-b0 networks as local and global networks with input size (1472, 960), crop size 384 and 4 extracted patches. CV : **43.69**\n* a modified version of MVCCL (BiView): inspired by @heng code [https://www.kaggle.com/code/hengck23/mvccl-model-for-admani-dataset](https://www.google.com/url?q=https://www.kaggle.com/code/hengck23/mvccl-model-for-admani-dataset&sa=D&source=editors&ust=1677621085747449&usg=AOvVaw3ovbCIfqBBwEX0rKY3P-iF) The main changes where to allow any timm model as main network and use the representation from both main and auxiliary view in the final summation. I used two models in final submission with input size (1536, 768): one b0 with 16 attention heads (**CV: 43.45**) and one b2 with 8 attention heads (**CV 44.99**)\n* ensembling those 3 models at laterality level gave **CV 46.94**\n\n## MPWARE’s Part\n\n\n* DALI decoder + YoloX ROI followed by crop and resize with aspect ratio = 1 to generate images with height=1024. No windowing.\n* External data included: CBIS-DDSM (mass + calcification full images) + PASM\n* Training pipeline with limited class oversampling and weighted CrossEntropyLoss on positive labels.\n* Augmentations: Random crop, H/V flips, minor RotateShiftResize, Noise/Blur, Random BrightnessContrast and Coarse Dropout\n* Backbones: NFNet + NextViT\n* No GeM but regular adaptive average pooling.\n* Max aggregation for laterality.\n* Spent some time to get TensorRT 1.3 working to compile more backbones.\n* CV comparable to Optimo’s ensemble at the end (but big CV/LB gap, got LB=0.61)\n\nWhat did not work for me:\n\n* Wavelets additional layer\n* Age as additional input feature\n* Mixup augmentation\n* VinDr as external data (no boost compared to CBIS)\n* High Resolution with stride=1, it worked at the beginning with EffNet but becomes useless when moving to some different backbones\n* Level 2 model based on embeddings.\n\nTake away: Learn a lot again, great teammates, nice competition, bad metric, hope to make another one with more FR teammates in the coming months.\n\n## Final submissions & Randomness\n\nOur two selected submissions were determined on CV :\n\n- A blend of 7 models, with CV 0.555\n  - Public 0.61, private 0.48\n  - For this submission we used fullfit models except for MPWARE’s models which were 4 folds. We had a similar submission with a lower public LB that used only fullfit models and that would have ranked #10.\n\n- A vote of our 3 pipelines with CV 0.547\n  - Public 0.6, private 0.48\n  - This vote was updated on the last day, we had a very similar vote with Public 0.61 private 0.52 before that\n\nWe got unlucky with submissions selection, our CV 0.5+ blends **all scored 0.48 private or above** (best 0.52, avg 0.495). Unfortunately we chose two subs on the lower end of the gaussian. Even simply selecting our best public LB would’ve put us in the top 8. The metric was too random, next time we’re staying away from F1-score competitions :)\n ",
      "votes": 14
    },
    {
      "id": 2163497,
      "postDate": "2023-02-28T21:25:46.563Z",
      "content": "<p>Congratulations! Detailed description. I guess a good team is more than the sum of its members. Did you plot any GradCAM's or such for any of your models? It would be interesting to see a GradCAM for microcalcifications, say for some CBIS-DDSM dataset example.</p>",
      "rawMarkdown": "Congratulations! Detailed description. I guess a good team is more than the sum of its members. Did you plot any GradCAM's or such for any of your models? It would be interesting to see a GradCAM for microcalcifications, say for some CBIS-DDSM dataset example.",
      "votes": 3,
      "replies": [
        {
          "id": 2163533,
          "postDate": "2023-02-28T22:30:07.197Z",
          "content": "<p>Thanks! Not tried to plot GradCAMs on my side.</p>",
          "rawMarkdown": "Thanks! Not tried to plot GradCAMs on my side."
        }
      ]
    },
    {
      "id": 2164232,
      "postDate": "2023-03-01T12:32:09.327Z",
      "content": "<p>Béret bas les baguettes !</p>",
      "rawMarkdown": "Béret bas les baguettes !",
      "votes": 1
    },
    {
      "id": 2163989,
      "postDate": "2023-03-01T08:29:02.607Z",
      "content": "<p>Nice solution! BTW could you provide some information about the resolution you used in the final ensemble?</p>",
      "rawMarkdown": "Nice solution! BTW could you provide some information about the resolution you used in the final ensemble?",
      "votes": 2,
      "replies": [
        {
          "id": 2164070,
          "postDate": "2023-03-01T09:28:19.990Z",
          "content": "<p>Every model used its own resolution during final ensembling, starting from the same cropped image every model was doing its own resize/preprocessing.</p>",
          "rawMarkdown": "Every model used its own resolution during final ensembling, starting from the same cropped image every model was doing its own resize/preprocessing.",
          "votes": 3
        }
      ]
    },
    {
      "id": 2251963,
      "postDate": "2023-05-09T19:22:43.347Z",
      "content": "<p>What does \"PASM\" stand for in the sketch and text? Which dataset are you referring to?</p>",
      "rawMarkdown": "What does \"PASM\" stand for in the sketch and text? Which dataset are you referring to?"
    },
    {
      "id": 2163937,
      "postDate": "2023-03-01T07:29:37.990Z",
      "content": "<p>Bravo, et merci pour vos contributions pendant la compétition. Toujours instructif pour un débutant comme moi.</p>",
      "rawMarkdown": "Bravo, et merci pour vos contributions pendant la compétition. Toujours instructif pour un débutant comme moi.",
      "votes": 2,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2163497,
      "author_name": "Antti Isosalo",
      "author_url": "",
      "post_date": "2023-02-28T21:25:46.563000",
      "content": "<p>Congratulations! Detailed description. I guess a good team is more than the sum of its members. Did you plot any GradCAM's or such for any of your models? It would be interesting to see a GradCAM for microcalcifications, say for some CBIS-DDSM dataset example.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2163533,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2023-02-28T22:30:07.197000",
          "content": "<p>Thanks! Not tried to plot GradCAMs on my side.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2164232,
      "author_name": "FabienDaniel",
      "author_url": "",
      "post_date": "2023-03-01T12:32:09.327000",
      "content": "<p>Béret bas les baguettes !</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2163989,
      "author_name": "Chenglu",
      "author_url": "",
      "post_date": "2023-03-01T08:29:02.607000",
      "content": "<p>Nice solution! BTW could you provide some information about the resolution you used in the final ensemble?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2164070,
          "author_name": "Optimo",
          "author_url": "",
          "post_date": "2023-03-01T09:28:19.990000",
          "content": "<p>Every model used its own resolution during final ensembling, starting from the same cropped image every model was doing its own resize/preprocessing.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 2251963,
      "author_name": "Zacharias",
      "author_url": "",
      "post_date": "2023-05-09T19:22:43.347000",
      "content": "<p>What does \"PASM\" stand for in the sketch and text? Which dataset are you referring to?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2163937,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-03-01T07:29:37.990000",
      "content": "<p>Bravo, et merci pour vos contributions pendant la compétition. Toujours instructif pour un débutant comme moi.</p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2163479": "Hi Kagglers,\n\nThanks for this nice competition!\n\nFirst, I would like to thank @theoviel and @optimo who helped me to make my dream true: Make a French team for one Kaggle competition. My expectation was to finish top #2 like in the FIFA world cup 🙂 but I’m quite happy with the current result. Some insights of our solution:\n\n## Theo’s Part\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2F858b18feec74c7a59e178265eebcd519%2Ftheo.png?generation=1677617629716416&alt=media)\n\n* Pretrain for 5 epochs using VinDr data and the BIRADS target\n* Finetune for 5 epochs on the competition data + external data\n  * bs=8 (6 for v2-s), lr=4e-4 (3e-4 for v2-s), Ranger, Linear Schedule with no warm up\n* External data varies among models, some only use CBIS. I also trained models with pseudo-labels on VinDr that are used in the final ensemble.\n* BCE loss with **no class weight**. Some models use BIRADS as an auxiliary target\n* No over/under-sampling\n\nThis is my best submitted ensemble. Models were changed a bit in the final ensemble to maximize CV.\n\n## Optimo’s Part\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2F845fc6e58654ca5ef2987f5b4a9d79eb%2Foptimo.png?generation=1677617643849717&alt=media)\n\nLikewise my timm backbones were pretrained on VinDr by predicting BIRADS. I used CBIS, Vindr PL as external data (boosted my CV by ~0.02).\n\nAs Theo had better results than mine on ‘standard models’ I tried to provide as much diversity as possible with ideas coming from research papers. I ended up using two different architectures :\n\n* a modified version of GMIC: inspired by this code I implemented a version which allows any timm network as global network and/or local network. I also changed the crop normalization in order to fully use the local brightness. As I was at first only changing the local network I used the public pretrained weights on NYU datasets. Which ended up being forbidden 2 days before the end of the competition. So I trained from scratch a model with 2xeffnet-b0 networks as local and global networks with input size (1472, 960), crop size 384 and 4 extracted patches. CV : **43.69**\n* a modified version of MVCCL (BiView): inspired by @heng code [https://www.kaggle.com/code/hengck23/mvccl-model-for-admani-dataset](https://www.google.com/url?q=https://www.kaggle.com/code/hengck23/mvccl-model-for-admani-dataset&sa=D&source=editors&ust=1677621085747449&usg=AOvVaw3ovbCIfqBBwEX0rKY3P-iF) The main changes where to allow any timm model as main network and use the representation from both main and auxiliary view in the final summation. I used two models in final submission with input size (1536, 768): one b0 with 16 attention heads (**CV: 43.45**) and one b2 with 8 attention heads (**CV 44.99**)\n* ensembling those 3 models at laterality level gave **CV 46.94**\n\n## MPWARE’s Part\n\n\n* DALI decoder + YoloX ROI followed by crop and resize with aspect ratio = 1 to generate images with height=1024. No windowing.\n* External data included: CBIS-DDSM (mass + calcification full images) + PASM\n* Training pipeline with limited class oversampling and weighted CrossEntropyLoss on positive labels.\n* Augmentations: Random crop, H/V flips, minor RotateShiftResize, Noise/Blur, Random BrightnessContrast and Coarse Dropout\n* Backbones: NFNet + NextViT\n* No GeM but regular adaptive average pooling.\n* Max aggregation for laterality.\n* Spent some time to get TensorRT 1.3 working to compile more backbones.\n* CV comparable to Optimo’s ensemble at the end (but big CV/LB gap, got LB=0.61)\n\nWhat did not work for me:\n\n* Wavelets additional layer\n* Age as additional input feature\n* Mixup augmentation\n* VinDr as external data (no boost compared to CBIS)\n* High Resolution with stride=1, it worked at the beginning with EffNet but becomes useless when moving to some different backbones\n* Level 2 model based on embeddings.\n\nTake away: Learn a lot again, great teammates, nice competition, bad metric, hope to make another one with more FR teammates in the coming months.\n\n## Final submissions & Randomness\n\nOur two selected submissions were determined on CV :\n\n- A blend of 7 models, with CV 0.555\n  - Public 0.61, private 0.48\n  - For this submission we used fullfit models except for MPWARE’s models which were 4 folds. We had a similar submission with a lower public LB that used only fullfit models and that would have ranked #10.\n\n- A vote of our 3 pipelines with CV 0.547\n  - Public 0.6, private 0.48\n  - This vote was updated on the last day, we had a very similar vote with Public 0.61 private 0.52 before that\n\nWe got unlucky with submissions selection, our CV 0.5+ blends **all scored 0.48 private or above** (best 0.52, avg 0.495). Unfortunately we chose two subs on the lower end of the gaussian. Even simply selecting our best public LB would’ve put us in the top 8. The metric was too random, next time we’re staying away from F1-score competitions :)\n ",
    "2163497": "Congratulations! Detailed description. I guess a good team is more than the sum of its members. Did you plot any GradCAM's or such for any of your models? It would be interesting to see a GradCAM for microcalcifications, say for some CBIS-DDSM dataset example.",
    "2164232": "Béret bas les baguettes !",
    "2163989": "Nice solution! BTW could you provide some information about the resolution you used in the final ensemble?",
    "2251963": "What does \"PASM\" stand for in the sketch and text? Which dataset are you referring to?",
    "2163937": "Bravo, et merci pour vos contributions pendant la compétition. Toujours instructif pour un débutant comme moi."
  }
}