{
  "id": 370669,
  "title": "EDA, Density is an important column, plus hillclimbing, transfer learning, and other techniques",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/370669",
  "author_name": "@kaggleqrdl",
  "post_date": "2022-12-05T20:20:52.183000",
  "votes": 9,
  "comment_count": 1,
  "views": 0,
  "content": "<p><strong>EDA</strong></p>\n<p>I explore this in my EDA (feedback appreciated) <a href=\"https://www.kaggle.com/code/kaggleqrdl/non-image-eda?scriptVersionId=113035136\" target=\"_blank\">https://www.kaggle.com/code/kaggleqrdl/non-image-eda?scriptVersionId=113035136</a></p>\n<p><strong>Density</strong></p>\n<p>When controlling for Age, density is a compelling signal.  This is also extensively backed up by the literature. Density C in particular is useful to look at for a positive signal, with Density B and A providing negative signals. Density D is problematic due to low sample space and possibly conflating issues of false negatives.  As per the data dictionary - \"Extremely dense tissue can make diagnosis more difficult.\"</p>\n<p>I was originally concerned that density was not being provided in test, but I believe the reasoning is that a solution is desired which works under less involvement from staff.   Estimating density is possibly tricky.</p>\n<p>I suspect the winning solutions will train on the density column and use that as an additional signal in the submission.</p>\n<p><strong>BIRADS == 0.0 and difficult_negative_case==False</strong><br>\nA curious relationship here.  There are 254 unique patient/breasts in train with BIRADS reported at 0, but didn't turn into a difficult negative case.  They all have cancer.  </p>\n<p>display(origtrain.query(\"BIRADS == 0 and difficult_negative_case == False\").drop_duplicates([\"patient_id\", \"laterality\"]))<br>\ndisplay(origtrain.query(\"BIRADS == 0 and difficult_negative_case == False\").drop_duplicates([\"patient_id\", \"laterality\"]).mean())</p>\n<p><strong>Hill Climbing</strong></p>\n<p>Also worth noting is that a constant theme I see in a lot of winning solutions is the notion of hill climbing.  It can be described by this from giba's great Paw Popularity winning solution) <a href=\"https://www.kaggle.com/competitions/petfinder-pawpularity-score/discussion/301686\" target=\"_blank\">https://www.kaggle.com/competitions/petfinder-pawpularity-score/discussion/301686</a></p>\n<blockquote>\n  <p>Each architecture has differences that can be explored by combining multiple architecture into the same dataset (stacking features side-by-side). The diversity between models can be used that way to boost RMSE. So to select the best models to include in the stacking dataset I used a simple forward feature selection with hill climbing. The algorithm is show below:</p>\n</blockquote>\n<pre><code>Features = [‘tf_efficientnet_l2_ns_475’] \nBest_feat = \nbestRMSE = np.inf\ncurrentRMSE = \n currentRMSE &lt; bestRMSE: \n    bestRMSE = currentRMSE\n    Features.append(best_feat)\n    rmse_scores= [compute_SVR_RMSE(Features + [feat])  feat  all_pretrained_models]\n    currentRMSE = np.(rmse_scores)\n    best_feat = np.argmin(rmse_scores)\n</code></pre>\n<p>The idea is to add as many diverse models as possible - as long as they improve your score.  Admittedly it requires a great deal of compute to train all those models and perform CV folds but that goes with the territory in image comps.  The solution above utilized SVR brilliantly to help deal with the compute problem.  </p>\n<p><strong>CBIS DDSM DataSet and others</strong></p>\n<p><a href=\"https://www.kaggle.com/datasets/awsaf49/cbis-ddsm-breast-cancer-image-dataset\" target=\"_blank\">https://www.kaggle.com/datasets/awsaf49/cbis-ddsm-breast-cancer-image-dataset</a></p>\n<p>Great dataset, note the mass outlines CSV file which can be used to annotate images as to where the salient tumor is.<br>\nMore datasets for this comp listed here - <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369364\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369364</a></p>\n<p><strong>Transfer Learning</strong></p>\n<p>It's worth noting that there is very likely a lot of great transfer learning that can be leveraged for this competition, which can reduce your training overhead.  Google constantly throughout the comp, in particular google scholar and sort / restrict by date.  </p>\n<p><strong>Enriched Dataset</strong></p>\n<blockquote>\n  <p>However, the RSNA team invested a lot of effort in enriching the competition dataset for cancer cases to ensure there are enough of them to use for modeling. From memory, the cancer rate in the general population is closer to 0.1% (@vaillant might have a more accurate number). Since the false negatives weren't enriched there's a reasonable chance that there's literally only one in the entire dataset.</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369262#2056107\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369262#2056107</a></p>\n<p><strong>Ideas for diverse models</strong></p>\n<p>Fine tuning a model such that it can predict BIRADS / difficult negative case, and then seeing if that helps via hill climbing. Also, fine tuning by predicting density. </p>\n<p><strong>Dali as a pipeline for better CPU/GPU optimization</strong></p>\n<p>Currently we have dicomsdl as the fastest way to process images (1.5x faster than dicom), see here - <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/371033#2058156\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/371033#2058156</a></p>\n<p>But another approach might be creating a pipeline which better utilizes the GPU.  Nvidia has dali for this:</p>\n<p><a href=\"https://developer.nvidia.com/blog/rapid-data-pre-processing-with-nvidia-dali/\" target=\"_blank\">https://developer.nvidia.com/blog/rapid-data-pre-processing-with-nvidia-dali/</a></p>\n<p>more about dali from our friends at tds:<br>\n<a href=\"https://towardsdatascience.com/overcoming-data-preprocessing-bottlenecks-with-tensorflow-data-service-nvidia-dali-and-other-d6321917f851\" target=\"_blank\">https://towardsdatascience.com/overcoming-data-preprocessing-bottlenecks-with-tensorflow-data-service-nvidia-dali-and-other-d6321917f851</a></p>\n<blockquote>\n  <p>A CPU bottleneck occurs when the GPU resource is under utilized as a result of one, or more of the CPUs, having reached maximum utilization. In this situation, the GPU will be partially idle while it waits for the CPU to pass in training data. This is an undesired state. Being that the GPU is, typically, the most expensive resource in the system, your goal should always be to maximize its utilization.</p>\n</blockquote>\n<p>notebook here - <br>\n<a href=\"https://www.kaggle.com/code/hirune924/nvidia-dali-the-fastest-data-loading\" target=\"_blank\">https://www.kaggle.com/code/hirune924/nvidia-dali-the-fastest-data-loading</a></p>\n<p><strong>Some breast cancer model links</strong></p>\n<p><a href=\"https://github.com/nyukat/breast_cancer_classifier\" target=\"_blank\">https://github.com/nyukat/breast_cancer_classifier</a><br>\n<a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/372673\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/372673</a></p>",
  "messages": [
    {
      "id": 2056176,
      "postDate": "2022-12-05T20:20:52.183Z",
      "content": "<p><strong>EDA</strong></p>\n<p>I explore this in my EDA (feedback appreciated) <a href=\"https://www.kaggle.com/code/kaggleqrdl/non-image-eda?scriptVersionId=113035136\" target=\"_blank\">https://www.kaggle.com/code/kaggleqrdl/non-image-eda?scriptVersionId=113035136</a></p>\n<p><strong>Density</strong></p>\n<p>When controlling for Age, density is a compelling signal.  This is also extensively backed up by the literature. Density C in particular is useful to look at for a positive signal, with Density B and A providing negative signals. Density D is problematic due to low sample space and possibly conflating issues of false negatives.  As per the data dictionary - \"Extremely dense tissue can make diagnosis more difficult.\"</p>\n<p>I was originally concerned that density was not being provided in test, but I believe the reasoning is that a solution is desired which works under less involvement from staff.   Estimating density is possibly tricky.</p>\n<p>I suspect the winning solutions will train on the density column and use that as an additional signal in the submission.</p>\n<p><strong>BIRADS == 0.0 and difficult_negative_case==False</strong><br>\nA curious relationship here.  There are 254 unique patient/breasts in train with BIRADS reported at 0, but didn't turn into a difficult negative case.  They all have cancer.  </p>\n<p>display(origtrain.query(\"BIRADS == 0 and difficult_negative_case == False\").drop_duplicates([\"patient_id\", \"laterality\"]))<br>\ndisplay(origtrain.query(\"BIRADS == 0 and difficult_negative_case == False\").drop_duplicates([\"patient_id\", \"laterality\"]).mean())</p>\n<p><strong>Hill Climbing</strong></p>\n<p>Also worth noting is that a constant theme I see in a lot of winning solutions is the notion of hill climbing.  It can be described by this from giba's great Paw Popularity winning solution) <a href=\"https://www.kaggle.com/competitions/petfinder-pawpularity-score/discussion/301686\" target=\"_blank\">https://www.kaggle.com/competitions/petfinder-pawpularity-score/discussion/301686</a></p>\n<blockquote>\n  <p>Each architecture has differences that can be explored by combining multiple architecture into the same dataset (stacking features side-by-side). The diversity between models can be used that way to boost RMSE. So to select the best models to include in the stacking dataset I used a simple forward feature selection with hill climbing. The algorithm is show below:</p>\n</blockquote>\n<pre><code>Features = [‘tf_efficientnet_l2_ns_475’] \nBest_feat = \nbestRMSE = np.inf\ncurrentRMSE = \n currentRMSE &lt; bestRMSE: \n    bestRMSE = currentRMSE\n    Features.append(best_feat)\n    rmse_scores= [compute_SVR_RMSE(Features + [feat])  feat  all_pretrained_models]\n    currentRMSE = np.(rmse_scores)\n    best_feat = np.argmin(rmse_scores)\n</code></pre>\n<p>The idea is to add as many diverse models as possible - as long as they improve your score.  Admittedly it requires a great deal of compute to train all those models and perform CV folds but that goes with the territory in image comps.  The solution above utilized SVR brilliantly to help deal with the compute problem.  </p>\n<p><strong>CBIS DDSM DataSet and others</strong></p>\n<p><a href=\"https://www.kaggle.com/datasets/awsaf49/cbis-ddsm-breast-cancer-image-dataset\" target=\"_blank\">https://www.kaggle.com/datasets/awsaf49/cbis-ddsm-breast-cancer-image-dataset</a></p>\n<p>Great dataset, note the mass outlines CSV file which can be used to annotate images as to where the salient tumor is.<br>\nMore datasets for this comp listed here - <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369364\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369364</a></p>\n<p><strong>Transfer Learning</strong></p>\n<p>It's worth noting that there is very likely a lot of great transfer learning that can be leveraged for this competition, which can reduce your training overhead.  Google constantly throughout the comp, in particular google scholar and sort / restrict by date.  </p>\n<p><strong>Enriched Dataset</strong></p>\n<blockquote>\n  <p>However, the RSNA team invested a lot of effort in enriching the competition dataset for cancer cases to ensure there are enough of them to use for modeling. From memory, the cancer rate in the general population is closer to 0.1% (@vaillant might have a more accurate number). Since the false negatives weren't enriched there's a reasonable chance that there's literally only one in the entire dataset.</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369262#2056107\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369262#2056107</a></p>\n<p><strong>Ideas for diverse models</strong></p>\n<p>Fine tuning a model such that it can predict BIRADS / difficult negative case, and then seeing if that helps via hill climbing. Also, fine tuning by predicting density. </p>\n<p><strong>Dali as a pipeline for better CPU/GPU optimization</strong></p>\n<p>Currently we have dicomsdl as the fastest way to process images (1.5x faster than dicom), see here - <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/371033#2058156\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/371033#2058156</a></p>\n<p>But another approach might be creating a pipeline which better utilizes the GPU.  Nvidia has dali for this:</p>\n<p><a href=\"https://developer.nvidia.com/blog/rapid-data-pre-processing-with-nvidia-dali/\" target=\"_blank\">https://developer.nvidia.com/blog/rapid-data-pre-processing-with-nvidia-dali/</a></p>\n<p>more about dali from our friends at tds:<br>\n<a href=\"https://towardsdatascience.com/overcoming-data-preprocessing-bottlenecks-with-tensorflow-data-service-nvidia-dali-and-other-d6321917f851\" target=\"_blank\">https://towardsdatascience.com/overcoming-data-preprocessing-bottlenecks-with-tensorflow-data-service-nvidia-dali-and-other-d6321917f851</a></p>\n<blockquote>\n  <p>A CPU bottleneck occurs when the GPU resource is under utilized as a result of one, or more of the CPUs, having reached maximum utilization. In this situation, the GPU will be partially idle while it waits for the CPU to pass in training data. This is an undesired state. Being that the GPU is, typically, the most expensive resource in the system, your goal should always be to maximize its utilization.</p>\n</blockquote>\n<p>notebook here - <br>\n<a href=\"https://www.kaggle.com/code/hirune924/nvidia-dali-the-fastest-data-loading\" target=\"_blank\">https://www.kaggle.com/code/hirune924/nvidia-dali-the-fastest-data-loading</a></p>\n<p><strong>Some breast cancer model links</strong></p>\n<p><a href=\"https://github.com/nyukat/breast_cancer_classifier\" target=\"_blank\">https://github.com/nyukat/breast_cancer_classifier</a><br>\n<a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/372673\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/372673</a></p>",
      "rawMarkdown": "**EDA**\n\nI explore this in my EDA (feedback appreciated) https://www.kaggle.com/code/kaggleqrdl/non-image-eda?scriptVersionId=113035136\n\n**Density**\n\nWhen controlling for Age, density is a compelling signal.  This is also extensively backed up by the literature. Density C in particular is useful to look at for a positive signal, with Density B and A providing negative signals. Density D is problematic due to low sample space and possibly conflating issues of false negatives.  As per the data dictionary - \"Extremely dense tissue can make diagnosis more difficult.\"\n\nI was originally concerned that density was not being provided in test, but I believe the reasoning is that a solution is desired which works under less involvement from staff.   Estimating density is possibly tricky.\n\nI suspect the winning solutions will train on the density column and use that as an additional signal in the submission.\n\n**BIRADS == 0.0 and difficult_negative_case==False**\nA curious relationship here.  There are 254 unique patient/breasts in train with BIRADS reported at 0, but didn't turn into a difficult negative case.  They all have cancer.  \n\ndisplay(origtrain.query(\"BIRADS == 0 and difficult_negative_case == False\").drop_duplicates([\"patient_id\", \"laterality\"]))\ndisplay(origtrain.query(\"BIRADS == 0 and difficult_negative_case == False\").drop_duplicates([\"patient_id\", \"laterality\"]).mean())\n\n\n\n\n**Hill Climbing**\n\nAlso worth noting is that a constant theme I see in a lot of winning solutions is the notion of hill climbing.  It can be described by this from giba's great Paw Popularity winning solution) https://www.kaggle.com/competitions/petfinder-pawpularity-score/discussion/301686\n\n> Each architecture has differences that can be explored by combining multiple architecture into the same dataset (stacking features side-by-side). The diversity between models can be used that way to boost RMSE. So to select the best models to include in the stacking dataset I used a simple forward feature selection with hill climbing. The algorithm is show below:\n\n```python\nFeatures = [‘tf_efficientnet_l2_ns_475’] # Start using only 1 model\nBest_feat = None\nbestRMSE = np.inf\ncurrentRMSE = 0\nwhile currentRMSE < bestRMSE: # Keep adding models while rmse decreases\n    bestRMSE = currentRMSE\n    Features.append(best_feat)\n    rmse_scores= [compute_SVR_RMSE(Features + [feat]) for feat in all_pretrained_models]\n    currentRMSE = np.min(rmse_scores)\n    best_feat = np.argmin(rmse_scores)\n```\n\nThe idea is to add as many diverse models as possible - as long as they improve your score.  Admittedly it requires a great deal of compute to train all those models and perform CV folds but that goes with the territory in image comps.  The solution above utilized SVR brilliantly to help deal with the compute problem.  \n\n**CBIS DDSM DataSet and others**\n\nhttps://www.kaggle.com/datasets/awsaf49/cbis-ddsm-breast-cancer-image-dataset\n\nGreat dataset, note the mass outlines CSV file which can be used to annotate images as to where the salient tumor is.\nMore datasets for this comp listed here - https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369364\n\n**Transfer Learning**\n\nIt's worth noting that there is very likely a lot of great transfer learning that can be leveraged for this competition, which can reduce your training overhead.  Google constantly throughout the comp, in particular google scholar and sort / restrict by date.  \n\n**Enriched Dataset**\n\n> However, the RSNA team invested a lot of effort in enriching the competition dataset for cancer cases to ensure there are enough of them to use for modeling. From memory, the cancer rate in the general population is closer to 0.1% (@vaillant might have a more accurate number). Since the false negatives weren't enriched there's a reasonable chance that there's literally only one in the entire dataset.\n\nhttps://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369262#2056107\n\n**Ideas for diverse models**\n\nFine tuning a model such that it can predict BIRADS / difficult negative case, and then seeing if that helps via hill climbing. Also, fine tuning by predicting density. \n\n**Dali as a pipeline for better CPU/GPU optimization**\n\nCurrently we have dicomsdl as the fastest way to process images (1.5x faster than dicom), see here - https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/371033#2058156\n\nBut another approach might be creating a pipeline which better utilizes the GPU.  Nvidia has dali for this:\n\nhttps://developer.nvidia.com/blog/rapid-data-pre-processing-with-nvidia-dali/\n\nmore about dali from our friends at tds:\nhttps://towardsdatascience.com/overcoming-data-preprocessing-bottlenecks-with-tensorflow-data-service-nvidia-dali-and-other-d6321917f851\n\n> A CPU bottleneck occurs when the GPU resource is under utilized as a result of one, or more of the CPUs, having reached maximum utilization. In this situation, the GPU will be partially idle while it waits for the CPU to pass in training data. This is an undesired state. Being that the GPU is, typically, the most expensive resource in the system, your goal should always be to maximize its utilization.\n\nnotebook here - \nhttps://www.kaggle.com/code/hirune924/nvidia-dali-the-fastest-data-loading\n\n**Some breast cancer model links**\n\nhttps://github.com/nyukat/breast_cancer_classifier\nhttps://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/372673",
      "votes": 9
    },
    {
      "id": 2126828,
      "postDate": "2023-02-02T13:26:42.933Z",
      "content": "<p>There is a way to get an estimate for the <strong>density</strong> with an acceptaple cost. See <a href=\"https://github.com/nyukat/breast_density_classifier/\" target=\"_blank\">this</a> histogram based method implemented as a baseline.</p>",
      "rawMarkdown": "There is a way to get an estimate for the **density** with an acceptaple cost. See [this](https://github.com/nyukat/breast_density_classifier/) histogram based method implemented as a baseline."
    }
  ],
  "comments": [
    {
      "id": 2126828,
      "author_name": "Antti Isosalo",
      "author_url": "",
      "post_date": "2023-02-02T13:26:42.933000",
      "content": "<p>There is a way to get an estimate for the <strong>density</strong> with an acceptaple cost. See <a href=\"https://github.com/nyukat/breast_density_classifier/\" target=\"_blank\">this</a> histogram based method implemented as a baseline.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2056176": "**EDA**\n\nI explore this in my EDA (feedback appreciated) https://www.kaggle.com/code/kaggleqrdl/non-image-eda?scriptVersionId=113035136\n\n**Density**\n\nWhen controlling for Age, density is a compelling signal.  This is also extensively backed up by the literature. Density C in particular is useful to look at for a positive signal, with Density B and A providing negative signals. Density D is problematic due to low sample space and possibly conflating issues of false negatives.  As per the data dictionary - \"Extremely dense tissue can make diagnosis more difficult.\"\n\nI was originally concerned that density was not being provided in test, but I believe the reasoning is that a solution is desired which works under less involvement from staff.   Estimating density is possibly tricky.\n\nI suspect the winning solutions will train on the density column and use that as an additional signal in the submission.\n\n**BIRADS == 0.0 and difficult_negative_case==False**\nA curious relationship here.  There are 254 unique patient/breasts in train with BIRADS reported at 0, but didn't turn into a difficult negative case.  They all have cancer.  \n\ndisplay(origtrain.query(\"BIRADS == 0 and difficult_negative_case == False\").drop_duplicates([\"patient_id\", \"laterality\"]))\ndisplay(origtrain.query(\"BIRADS == 0 and difficult_negative_case == False\").drop_duplicates([\"patient_id\", \"laterality\"]).mean())\n\n\n\n\n**Hill Climbing**\n\nAlso worth noting is that a constant theme I see in a lot of winning solutions is the notion of hill climbing.  It can be described by this from giba's great Paw Popularity winning solution) https://www.kaggle.com/competitions/petfinder-pawpularity-score/discussion/301686\n\n> Each architecture has differences that can be explored by combining multiple architecture into the same dataset (stacking features side-by-side). The diversity between models can be used that way to boost RMSE. So to select the best models to include in the stacking dataset I used a simple forward feature selection with hill climbing. The algorithm is show below:\n\n```python\nFeatures = [‘tf_efficientnet_l2_ns_475’] # Start using only 1 model\nBest_feat = None\nbestRMSE = np.inf\ncurrentRMSE = 0\nwhile currentRMSE < bestRMSE: # Keep adding models while rmse decreases\n    bestRMSE = currentRMSE\n    Features.append(best_feat)\n    rmse_scores= [compute_SVR_RMSE(Features + [feat]) for feat in all_pretrained_models]\n    currentRMSE = np.min(rmse_scores)\n    best_feat = np.argmin(rmse_scores)\n```\n\nThe idea is to add as many diverse models as possible - as long as they improve your score.  Admittedly it requires a great deal of compute to train all those models and perform CV folds but that goes with the territory in image comps.  The solution above utilized SVR brilliantly to help deal with the compute problem.  \n\n**CBIS DDSM DataSet and others**\n\nhttps://www.kaggle.com/datasets/awsaf49/cbis-ddsm-breast-cancer-image-dataset\n\nGreat dataset, note the mass outlines CSV file which can be used to annotate images as to where the salient tumor is.\nMore datasets for this comp listed here - https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369364\n\n**Transfer Learning**\n\nIt's worth noting that there is very likely a lot of great transfer learning that can be leveraged for this competition, which can reduce your training overhead.  Google constantly throughout the comp, in particular google scholar and sort / restrict by date.  \n\n**Enriched Dataset**\n\n> However, the RSNA team invested a lot of effort in enriching the competition dataset for cancer cases to ensure there are enough of them to use for modeling. From memory, the cancer rate in the general population is closer to 0.1% (@vaillant might have a more accurate number). Since the false negatives weren't enriched there's a reasonable chance that there's literally only one in the entire dataset.\n\nhttps://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369262#2056107\n\n**Ideas for diverse models**\n\nFine tuning a model such that it can predict BIRADS / difficult negative case, and then seeing if that helps via hill climbing. Also, fine tuning by predicting density. \n\n**Dali as a pipeline for better CPU/GPU optimization**\n\nCurrently we have dicomsdl as the fastest way to process images (1.5x faster than dicom), see here - https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/371033#2058156\n\nBut another approach might be creating a pipeline which better utilizes the GPU.  Nvidia has dali for this:\n\nhttps://developer.nvidia.com/blog/rapid-data-pre-processing-with-nvidia-dali/\n\nmore about dali from our friends at tds:\nhttps://towardsdatascience.com/overcoming-data-preprocessing-bottlenecks-with-tensorflow-data-service-nvidia-dali-and-other-d6321917f851\n\n> A CPU bottleneck occurs when the GPU resource is under utilized as a result of one, or more of the CPUs, having reached maximum utilization. In this situation, the GPU will be partially idle while it waits for the CPU to pass in training data. This is an undesired state. Being that the GPU is, typically, the most expensive resource in the system, your goal should always be to maximize its utilization.\n\nnotebook here - \nhttps://www.kaggle.com/code/hirune924/nvidia-dali-the-fastest-data-loading\n\n**Some breast cancer model links**\n\nhttps://github.com/nyukat/breast_cancer_classifier\nhttps://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/372673",
    "2126828": "There is a way to get an estimate for the **density** with an acceptaple cost. See [this](https://github.com/nyukat/breast_density_classifier/) histogram based method implemented as a baseline."
  }
}