{
  "id": 391286,
  "title": "31st place solution ",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/391286",
  "author_name": "Kirderf",
  "post_date": "2023-02-28T23:40:50.544000",
  "votes": 18,
  "comment_count": 2,
  "views": 0,
  "content": "<p><strong>Summary</strong></p>\n<p>The solutions are based on the Tensorflow framework with Pytorch image preprocessing and XGB GPU/Cuda classifier.</p>\n<p>Best selected submission -  Pr.L. 0.46 Pu.L. 0.59  - 2x4 fold ensemble (100 fold split ~ all data)<br>\nBest private submission with room for improvements - Pr.L 0.48 Pu.L 0.52  - 2x4 fold  ensemble ( 5 fold split) + 2 XGB ensemble. ( CV 0.55 - agg. method median w/ threshold 0.243 )</p>\n<p><strong>Dataset</strong></p>\n<p>Used <a href=\"https://www.kaggle.com/awsaf49\" target=\"_blank\">@awsaf49</a> croped 2048-1024 ds from the original competition dataset aswell as the training code as a base.<br>\n<a href=\"https://www.kaggle.com/datasets/awsaf49/rsna-bcd-roi-2048x1024-png-v2-dataset\" target=\"_blank\">https://www.kaggle.com/datasets/awsaf49/rsna-bcd-roi-2048x1024-png-v2-dataset</a><br>\n<a href=\"https://www.kaggle.com/code/awsaf49/rsna-bcd-efficientnet-tf-tpu-1vm-train\" target=\"_blank\">https://www.kaggle.com/code/awsaf49/rsna-bcd-efficientnet-tf-tpu-1vm-train</a></p>\n<p><strong>Training and Models</strong></p>\n<p>Upsample positive class by x 10<br>\nHeavy augmentation<br>\nSigmoidFocalCrossEntropy</p>\n<p>Model used are B5 and Convnext_base_384_in22ft1k(v1) from tfimm library.<br>\nAdded a classifier head with 32 vs 64 Silu Dense layer.<br>\nAdamW with SWA optimizer and 8 epoch training setup was used.</p>\n<p><strong>Inference</strong>  - T4 x2 in mixed precision.</p>\n<p>Here is where I put majority of the competition time and effort.</p>\n<p>I used the faster inference with NVIDIA Dali for speeding up the image handling.<br>\n<a href=\"https://www.kaggle.com/code/theoviel/rsna-breast-baseline-faster-inference-with-dali\" target=\"_blank\">https://www.kaggle.com/code/theoviel/rsna-breast-baseline-faster-inference-with-dali</a><br>\nMade some changes for handling the croped dataset from training and configurations. As pytorch and TF doesn't run well on same GPU process memory I run this image processing in separate kernel/notebook.</p>\n<p>The above was common for all the solutions but the rest below are where it gets interesting and the score increasing happens.</p>\n<p><strong>Solution 1:</strong></p>\n<p>Used more data for the B5 and ConvNext training and less for validation. It gave a boost vs 5 fold split but restricted the option for the solution no.2. Nevertheless I used it as solution/sub no. 1 if the second more complex solution would fail.<br>\nI increased the dim size for inference vs train, often give a boost in score, and did here as well.</p>\n<p><strong>Solution 2:</strong></p>\n<p>The idea here was to take use of all other information META etc and together with vision model feed features to a XGB model, which is SOTA in handling imbalanced and mix of feature information. We also had information for the negative class that could be used for smoothing the imbalanced situation, like create more classes.<br>\nFor creating extra features and extra classes I got inspirations from 1st place solution in the SIIM-ISIC Melanoma Classification<br>\n<a href=\"https://www.kaggle.com/competitions/siim-isic-melanoma-classification/discussion/175412\" target=\"_blank\">https://www.kaggle.com/competitions/siim-isic-melanoma-classification/discussion/175412</a></p>\n<p>The features for the XGB training and inference where:<br>\n''site_id','laterality','view','age','implant','n_images','image_size' + extracted information from the last dense layer from every model instead of embeddings to minimize the size if the vision information and features to 32 vs 64 features.<br>\nI created a total of 9 classes from the information in the traindata, \"neg\", \"pos\", \"diffneg\", \"rneg\", \"negA\", \"negB\" etc. Data not seen in test data, but it doesn't matter as you only need to train the XGB to classify them instead which help sort the total information within the negative class region of information, leaving it less imbalanced.</p>\n<p>For the traning of the two XGB models per vision model I used a custom version of the Autoxgb framework as base to have a solid standard. I did changed many things in it to fit the problem, updated it with Prune for HPO speed, added comp metric to XGB val at image level and for Optuna metric mimic the final non-image level agg PP for every CV in Optuna trial. Added all XGB parameters to the Optuna HPO and other things like only upsample the cancer class in the 9 classes.</p>\n<p>For Inference I used the best parameters and setup from the HPO and used all 4 fold vision model data merged to a big oof train set,it was better for the XGB to train on all features not only the single fold/model information that it would predict on later. It also speed up the training as all vision folds used the same and single pretrained XGB for prediction.</p>\n<p>So the solution did, end-to-end, in test/inference time for 2 vision model (B5 and Conv base) a 4 fold each:</p>\n<ul>\n<li>Extracted the information from the vision inference of the images to features.</li>\n<li>Created the complete OOF train data with extra features + vision features + 9 classes.</li>\n<li>Trained the 2 XGB in CV 5 fold each.</li>\n<li>Extract the prob. for the cancer class in the 9 classes.</li>\n<li>Run several of optimizations searches for agg image threshold and math method (mean,max and median) together with weighted ensemble threshold for the two XGB.</li>\n<li>Sort the searches for the best setup and use it in the final submission post-processing.</li>\n</ul>\n<p>This solution increased the score from a private score 0.44 with only image inference to 0.48.</p>\n<p>So what could have been improved now looking at the results: </p>\n<ul>\n<li>Run above with higher inference dim size vs train. </li>\n<li>Use more \"different models architectures\"- fold to inference instead of using all the CV folds. <br>\nLooking at the result from other tests that would have increased the score.</li>\n<li>More Feature Engineering.</li>\n</ul>\n<hr>\n<p>That's it! </p>",
  "messages": [
    {
      "id": 2163585,
      "postDate": "2023-02-28T23:40:50.543Z",
      "content": "<p><strong>Summary</strong></p>\n<p>The solutions are based on the Tensorflow framework with Pytorch image preprocessing and XGB GPU/Cuda classifier.</p>\n<p>Best selected submission -  Pr.L. 0.46 Pu.L. 0.59  - 2x4 fold ensemble (100 fold split ~ all data)<br>\nBest private submission with room for improvements - Pr.L 0.48 Pu.L 0.52  - 2x4 fold  ensemble ( 5 fold split) + 2 XGB ensemble. ( CV 0.55 - agg. method median w/ threshold 0.243 )</p>\n<p><strong>Dataset</strong></p>\n<p>Used <a href=\"https://www.kaggle.com/awsaf49\" target=\"_blank\">@awsaf49</a> croped 2048-1024 ds from the original competition dataset aswell as the training code as a base.<br>\n<a href=\"https://www.kaggle.com/datasets/awsaf49/rsna-bcd-roi-2048x1024-png-v2-dataset\" target=\"_blank\">https://www.kaggle.com/datasets/awsaf49/rsna-bcd-roi-2048x1024-png-v2-dataset</a><br>\n<a href=\"https://www.kaggle.com/code/awsaf49/rsna-bcd-efficientnet-tf-tpu-1vm-train\" target=\"_blank\">https://www.kaggle.com/code/awsaf49/rsna-bcd-efficientnet-tf-tpu-1vm-train</a></p>\n<p><strong>Training and Models</strong></p>\n<p>Upsample positive class by x 10<br>\nHeavy augmentation<br>\nSigmoidFocalCrossEntropy</p>\n<p>Model used are B5 and Convnext_base_384_in22ft1k(v1) from tfimm library.<br>\nAdded a classifier head with 32 vs 64 Silu Dense layer.<br>\nAdamW with SWA optimizer and 8 epoch training setup was used.</p>\n<p><strong>Inference</strong>  - T4 x2 in mixed precision.</p>\n<p>Here is where I put majority of the competition time and effort.</p>\n<p>I used the faster inference with NVIDIA Dali for speeding up the image handling.<br>\n<a href=\"https://www.kaggle.com/code/theoviel/rsna-breast-baseline-faster-inference-with-dali\" target=\"_blank\">https://www.kaggle.com/code/theoviel/rsna-breast-baseline-faster-inference-with-dali</a><br>\nMade some changes for handling the croped dataset from training and configurations. As pytorch and TF doesn't run well on same GPU process memory I run this image processing in separate kernel/notebook.</p>\n<p>The above was common for all the solutions but the rest below are where it gets interesting and the score increasing happens.</p>\n<p><strong>Solution 1:</strong></p>\n<p>Used more data for the B5 and ConvNext training and less for validation. It gave a boost vs 5 fold split but restricted the option for the solution no.2. Nevertheless I used it as solution/sub no. 1 if the second more complex solution would fail.<br>\nI increased the dim size for inference vs train, often give a boost in score, and did here as well.</p>\n<p><strong>Solution 2:</strong></p>\n<p>The idea here was to take use of all other information META etc and together with vision model feed features to a XGB model, which is SOTA in handling imbalanced and mix of feature information. We also had information for the negative class that could be used for smoothing the imbalanced situation, like create more classes.<br>\nFor creating extra features and extra classes I got inspirations from 1st place solution in the SIIM-ISIC Melanoma Classification<br>\n<a href=\"https://www.kaggle.com/competitions/siim-isic-melanoma-classification/discussion/175412\" target=\"_blank\">https://www.kaggle.com/competitions/siim-isic-melanoma-classification/discussion/175412</a></p>\n<p>The features for the XGB training and inference where:<br>\n''site_id','laterality','view','age','implant','n_images','image_size' + extracted information from the last dense layer from every model instead of embeddings to minimize the size if the vision information and features to 32 vs 64 features.<br>\nI created a total of 9 classes from the information in the traindata, \"neg\", \"pos\", \"diffneg\", \"rneg\", \"negA\", \"negB\" etc. Data not seen in test data, but it doesn't matter as you only need to train the XGB to classify them instead which help sort the total information within the negative class region of information, leaving it less imbalanced.</p>\n<p>For the traning of the two XGB models per vision model I used a custom version of the Autoxgb framework as base to have a solid standard. I did changed many things in it to fit the problem, updated it with Prune for HPO speed, added comp metric to XGB val at image level and for Optuna metric mimic the final non-image level agg PP for every CV in Optuna trial. Added all XGB parameters to the Optuna HPO and other things like only upsample the cancer class in the 9 classes.</p>\n<p>For Inference I used the best parameters and setup from the HPO and used all 4 fold vision model data merged to a big oof train set,it was better for the XGB to train on all features not only the single fold/model information that it would predict on later. It also speed up the training as all vision folds used the same and single pretrained XGB for prediction.</p>\n<p>So the solution did, end-to-end, in test/inference time for 2 vision model (B5 and Conv base) a 4 fold each:</p>\n<ul>\n<li>Extracted the information from the vision inference of the images to features.</li>\n<li>Created the complete OOF train data with extra features + vision features + 9 classes.</li>\n<li>Trained the 2 XGB in CV 5 fold each.</li>\n<li>Extract the prob. for the cancer class in the 9 classes.</li>\n<li>Run several of optimizations searches for agg image threshold and math method (mean,max and median) together with weighted ensemble threshold for the two XGB.</li>\n<li>Sort the searches for the best setup and use it in the final submission post-processing.</li>\n</ul>\n<p>This solution increased the score from a private score 0.44 with only image inference to 0.48.</p>\n<p>So what could have been improved now looking at the results: </p>\n<ul>\n<li>Run above with higher inference dim size vs train. </li>\n<li>Use more \"different models architectures\"- fold to inference instead of using all the CV folds. <br>\nLooking at the result from other tests that would have increased the score.</li>\n<li>More Feature Engineering.</li>\n</ul>\n<hr>\n<p>That's it! </p>",
      "rawMarkdown": "**Summary**\n\nThe solutions are based on the Tensorflow framework with Pytorch image preprocessing and XGB GPU/Cuda classifier.\n\nBest selected submission -  Pr.L. 0.46 Pu.L. 0.59  - 2x4 fold ensemble (100 fold split ~ all data)\nBest private submission with room for improvements - Pr.L 0.48 Pu.L 0.52  - 2x4 fold  ensemble ( 5 fold split) + 2 XGB ensemble. ( CV 0.55 - agg. method median w/ threshold 0.243 )\n\n**Dataset**\n\nUsed @awsaf49 croped 2048-1024 ds from the original competition dataset aswell as the training code as a base.\nhttps://www.kaggle.com/datasets/awsaf49/rsna-bcd-roi-2048x1024-png-v2-dataset\nhttps://www.kaggle.com/code/awsaf49/rsna-bcd-efficientnet-tf-tpu-1vm-train\n\n**Training and Models**\n\nUpsample positive class by x 10\nHeavy augmentation\nSigmoidFocalCrossEntropy\n\nModel used are B5 and Convnext_base_384_in22ft1k(v1) from tfimm library.\nAdded a classifier head with 32 vs 64 Silu Dense layer.\nAdamW with SWA optimizer and 8 epoch training setup was used.\n\n**Inference**  - T4 x2 in mixed precision.\n\nHere is where I put majority of the competition time and effort.\n\nI used the faster inference with NVIDIA Dali for speeding up the image handling.\nhttps://www.kaggle.com/code/theoviel/rsna-breast-baseline-faster-inference-with-dali\nMade some changes for handling the croped dataset from training and configurations. As pytorch and TF doesn't run well on same GPU process memory I run this image processing in separate kernel/notebook.\n\nThe above was common for all the solutions but the rest below are where it gets interesting and the score increasing happens.\n\n**Solution 1:**\n\nUsed more data for the B5 and ConvNext training and less for validation. It gave a boost vs 5 fold split but restricted the option for the solution no.2. Nevertheless I used it as solution/sub no. 1 if the second more complex solution would fail.\nI increased the dim size for inference vs train, often give a boost in score, and did here as well.\n\n**Solution 2:**\n\nThe idea here was to take use of all other information META etc and together with vision model feed features to a XGB model, which is SOTA in handling imbalanced and mix of feature information. We also had information for the negative class that could be used for smoothing the imbalanced situation, like create more classes.\nFor creating extra features and extra classes I got inspirations from 1st place solution in the SIIM-ISIC Melanoma Classification\nhttps://www.kaggle.com/competitions/siim-isic-melanoma-classification/discussion/175412\n\nThe features for the XGB training and inference where:\n''site_id','laterality','view','age','implant','n_images','image_size' + extracted information from the last dense layer from every model instead of embeddings to minimize the size if the vision information and features to 32 vs 64 features.\nI created a total of 9 classes from the information in the traindata, \"neg\", \"pos\", \"diffneg\", \"rneg\", \"negA\", \"negB\" etc. Data not seen in test data, but it doesn't matter as you only need to train the XGB to classify them instead which help sort the total information within the negative class region of information, leaving it less imbalanced.\n\nFor the traning of the two XGB models per vision model I used a custom version of the Autoxgb framework as base to have a solid standard. I did changed many things in it to fit the problem, updated it with Prune for HPO speed, added comp metric to XGB val at image level and for Optuna metric mimic the final non-image level agg PP for every CV in Optuna trial. Added all XGB parameters to the Optuna HPO and other things like only upsample the cancer class in the 9 classes.\n\nFor Inference I used the best parameters and setup from the HPO and used all 4 fold vision model data merged to a big oof train set,it was better for the XGB to train on all features not only the single fold/model information that it would predict on later. It also speed up the training as all vision folds used the same and single pretrained XGB for prediction.\n\nSo the solution did, end-to-end, in test/inference time for 2 vision model (B5 and Conv base) a 4 fold each:\n\n- Extracted the information from the vision inference of the images to features.\n- Created the complete OOF train data with extra features + vision features + 9 classes.\n- Trained the 2 XGB in CV 5 fold each.\n- Extract the prob. for the cancer class in the 9 classes.\n- Run several of optimizations searches for agg image threshold and math method (mean,max and median) together with weighted ensemble threshold for the two XGB.\n- Sort the searches for the best setup and use it in the final submission post-processing.\n\nThis solution increased the score from a private score 0.44 with only image inference to 0.48.\n\nSo what could have been improved now looking at the results: \n- Run above with higher inference dim size vs train. \n- Use more \"different models architectures\"- fold to inference instead of using all the CV folds. \nLooking at the result from other tests that would have increased the score.\n- More Feature Engineering.\n\n------------------------------------------------------------------------------------------------------------\n\nThat's it! ",
      "votes": 18
    },
    {
      "id": 2164203,
      "postDate": "2023-03-01T12:00:16.953Z",
      "content": "<p>Good work! And seems that also quite a lot of work to accomplish for one person, so very good effort. I liked the systematic approach. 👍</p>",
      "rawMarkdown": "Good work! And seems that also quite a lot of work to accomplish for one person, so very good effort. I liked the systematic approach. 👍",
      "votes": 2,
      "replies": [
        {
          "id": 2164333,
          "postDate": "2023-03-01T13:33:53.433Z",
          "content": "<p>Thanks! Yes it took some time and effort but it's an important area and AI in medical is interesting even though I don't work with it, but maybe some day :)</p>",
          "rawMarkdown": "Thanks! Yes it took some time and effort but it's an important area and AI in medical is interesting even though I don't work with it, but maybe some day :)",
          "votes": 2
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2164203,
      "author_name": "Antti Isosalo",
      "author_url": "",
      "post_date": "2023-03-01T12:00:16.953000",
      "content": "<p>Good work! And seems that also quite a lot of work to accomplish for one person, so very good effort. I liked the systematic approach. 👍</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2164333,
          "author_name": "Kirderf",
          "author_url": "",
          "post_date": "2023-03-01T13:33:53.433000",
          "content": "<p>Thanks! Yes it took some time and effort but it's an important area and AI in medical is interesting even though I don't work with it, but maybe some day :)</p>",
          "votes": 2,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2163585": "**Summary**\n\nThe solutions are based on the Tensorflow framework with Pytorch image preprocessing and XGB GPU/Cuda classifier.\n\nBest selected submission -  Pr.L. 0.46 Pu.L. 0.59  - 2x4 fold ensemble (100 fold split ~ all data)\nBest private submission with room for improvements - Pr.L 0.48 Pu.L 0.52  - 2x4 fold  ensemble ( 5 fold split) + 2 XGB ensemble. ( CV 0.55 - agg. method median w/ threshold 0.243 )\n\n**Dataset**\n\nUsed @awsaf49 croped 2048-1024 ds from the original competition dataset aswell as the training code as a base.\nhttps://www.kaggle.com/datasets/awsaf49/rsna-bcd-roi-2048x1024-png-v2-dataset\nhttps://www.kaggle.com/code/awsaf49/rsna-bcd-efficientnet-tf-tpu-1vm-train\n\n**Training and Models**\n\nUpsample positive class by x 10\nHeavy augmentation\nSigmoidFocalCrossEntropy\n\nModel used are B5 and Convnext_base_384_in22ft1k(v1) from tfimm library.\nAdded a classifier head with 32 vs 64 Silu Dense layer.\nAdamW with SWA optimizer and 8 epoch training setup was used.\n\n**Inference**  - T4 x2 in mixed precision.\n\nHere is where I put majority of the competition time and effort.\n\nI used the faster inference with NVIDIA Dali for speeding up the image handling.\nhttps://www.kaggle.com/code/theoviel/rsna-breast-baseline-faster-inference-with-dali\nMade some changes for handling the croped dataset from training and configurations. As pytorch and TF doesn't run well on same GPU process memory I run this image processing in separate kernel/notebook.\n\nThe above was common for all the solutions but the rest below are where it gets interesting and the score increasing happens.\n\n**Solution 1:**\n\nUsed more data for the B5 and ConvNext training and less for validation. It gave a boost vs 5 fold split but restricted the option for the solution no.2. Nevertheless I used it as solution/sub no. 1 if the second more complex solution would fail.\nI increased the dim size for inference vs train, often give a boost in score, and did here as well.\n\n**Solution 2:**\n\nThe idea here was to take use of all other information META etc and together with vision model feed features to a XGB model, which is SOTA in handling imbalanced and mix of feature information. We also had information for the negative class that could be used for smoothing the imbalanced situation, like create more classes.\nFor creating extra features and extra classes I got inspirations from 1st place solution in the SIIM-ISIC Melanoma Classification\nhttps://www.kaggle.com/competitions/siim-isic-melanoma-classification/discussion/175412\n\nThe features for the XGB training and inference where:\n''site_id','laterality','view','age','implant','n_images','image_size' + extracted information from the last dense layer from every model instead of embeddings to minimize the size if the vision information and features to 32 vs 64 features.\nI created a total of 9 classes from the information in the traindata, \"neg\", \"pos\", \"diffneg\", \"rneg\", \"negA\", \"negB\" etc. Data not seen in test data, but it doesn't matter as you only need to train the XGB to classify them instead which help sort the total information within the negative class region of information, leaving it less imbalanced.\n\nFor the traning of the two XGB models per vision model I used a custom version of the Autoxgb framework as base to have a solid standard. I did changed many things in it to fit the problem, updated it with Prune for HPO speed, added comp metric to XGB val at image level and for Optuna metric mimic the final non-image level agg PP for every CV in Optuna trial. Added all XGB parameters to the Optuna HPO and other things like only upsample the cancer class in the 9 classes.\n\nFor Inference I used the best parameters and setup from the HPO and used all 4 fold vision model data merged to a big oof train set,it was better for the XGB to train on all features not only the single fold/model information that it would predict on later. It also speed up the training as all vision folds used the same and single pretrained XGB for prediction.\n\nSo the solution did, end-to-end, in test/inference time for 2 vision model (B5 and Conv base) a 4 fold each:\n\n- Extracted the information from the vision inference of the images to features.\n- Created the complete OOF train data with extra features + vision features + 9 classes.\n- Trained the 2 XGB in CV 5 fold each.\n- Extract the prob. for the cancer class in the 9 classes.\n- Run several of optimizations searches for agg image threshold and math method (mean,max and median) together with weighted ensemble threshold for the two XGB.\n- Sort the searches for the best setup and use it in the final submission post-processing.\n\nThis solution increased the score from a private score 0.44 with only image inference to 0.48.\n\nSo what could have been improved now looking at the results: \n- Run above with higher inference dim size vs train. \n- Use more \"different models architectures\"- fold to inference instead of using all the CV folds. \nLooking at the result from other tests that would have increased the score.\n- More Feature Engineering.\n\n------------------------------------------------------------------------------------------------------------\n\nThat's it! ",
    "2164203": "Good work! And seems that also quite a lot of work to accomplish for one person, so very good effort. I liked the systematic approach. 👍"
  }
}