{
  "id": 453636,
  "title": "59th Place Solution for the Detect and Classify Traumatic Abdominal Injuries Competition",
  "url": "/competitions/rsna-2023-abdominal-trauma-detection/discussion/453636",
  "author_name": "Jenny Ding",
  "post_date": "2023-11-07T05:20:00.363000",
  "votes": 7,
  "comment_count": 1,
  "views": 0,
  "content": "<p>First and foremost, we extend our gratitude to RSNA for organizing this captivating competition, through which we gained invaluable insights. </p>\n<h1>1. Context section</h1>\n<ul>\n<li>Business context: <a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/overview\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/overview</a></li>\n<li>Data context: <a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/data\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/data</a></li>\n</ul>\n<h1>2. Overview of the Approach</h1>\n<p>In this competition, the motivation is to contribute to the improvement of patient outcomes from traumatic injuries, which account for over 5 million global deaths each year. Blunt-force abdominal trauma, often caused by vehicle accidents, can damage organs and lead to internal bleeding. With traditional methods like physical exams and lab tests often being inconclusive, the emphasis shifts to the need for accurate medical imaging interpretations.</p>\n<p>In this study, our models predict a probability for each of the different possible injury types and degrees using CT scans. While vital for evaluating abdominal trauma, CT scans can be challenging to decipher, especially when injuries are subtle or multiple. Our goal is to harness the power of AI and ML to better interpret CT scans.</p>\n<h1>3.  Details of the submission</h1>\n<h2>3.1 Data Interpretation and Preprocessing</h2>\n<p>The dataset is primarily stored in the .dcm (DICOM) format, which offers a detailed insight into traumatic injuries. The patients underwent CT scans, with each undergoing one to two scans, generally covering from the upper neck down to the region below the anus. However, the specific number of scan images per individual, per scan, remained variable. This variability was due to the unknown intervals between each scan slice and the differing heights of the individuals, introducing a level of complexity in the initial assessment of the data. </p>\n<p>The scans could be classified into two types: complete and incomplete. A complete scan provides a comprehensive view, encompassing all organs within the range from the upper neck to the below the anal region. In contrast, an incomplete scan is a localized examination, focusing on specific organs within this anatomical spectrum. While the training dataset came labeled with indications of completeness, such classification were absent in the test dataset. The training dataset encompassed the CT scan images of 3,147 patients, while the test dataset was composed of approximately 1,300 patents’ scans.</p>\n<p>When these multiple .dcm images from a single scan were stacked together, they coalesced to form a vivid 3D representation of the human body, encapsulating a plethora of details. Such data, especially in the field of medical imaging, is paramount to diagnose and comprehend the intricacies of internal injuries.</p>\n<p>For convenience, we converted the given .dcm (DICOM) files into the more universally recognized .png format, facilitated by a <a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/427427\" target=\"_blank\">method</a> provided by Kaggle platform, streamlining subsequent operations and analyses. </p>\n<h2>3.2 Model Pipeline Design</h2>\n<h3>3.2.1 2D Semantic Segmentation Model</h3>\n<h4>Segmentation with U-Net</h4>\n<p>Our objective with U-Net was twofold: to identify the span of CT images that contained each organ (liver, spleen, kidneys, bowels) and to pinpoint their precise locations within each image. </p>\n<p>We initially employed the Total Segmentator tool (available at: <a href=\"https://github.com/wasserth/TotalSegmentator\" target=\"_blank\">Total Segmentator on GitHub</a>), but faced significant time constraints due to the tool’s processing speed. This constraint necessitated a shift in our approach toward a more efficient solution, leading us to use the U-Net architecture.   </p>\n<p>We utilized the Total Segmentator to generate marked pixel points as input data. These annotated 2D images served as training data for our U-Net model. By harnessing the efficiency of U-Net, we trained a 2D segmenter to efficiently discern the presence and location of the targeted organs across the CT scan images. The refined segmentation process not only indicated which scans contained the organs of interest but also provided their spatial coordinates, thus streamlining the subsequent stages of our analysis. </p>\n<p>The following figures are an illustration of segmentation with U-Net. The black-and-white figure on the left is the original CT scan image, and the right one is the one segmented. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2497266%2F79d5484239de43371fef834bdb1fd493%2Ffigure3.png?generation=1699408525140588&amp;alt=media\" alt=\"U-Net\"></p>\n<h4>Constructing 2.5D Input</h4>\n<p>For this part, our objective was to create a 2.5D input data for the neural network. We implemented a strategy to select 32 evenly spaced images from each scan across the training dataset. Given that the minimum number of slices per scan was 44, this uniform selection process ensured comprehensive coverage of each individual’s scan while maintaining a focus on representation consistency for each organ. Our selection also guaranteed that every organ appeared in at least 4 separate images, thereby capturing the essential anatomical features needed for accurate analysis. </p>\n<p>Then we used EfficientNet architecture. Each image was processed through the network, with the output of the penultimate layer capturing the feature representation extracted by the network after processing. Subsequently, a fully connected layer served as the final decision-making component, distinguishing between health and injury.  </p>\n<p>To synthesize the 32 image-derived insights into a single diagnostic outcome for each scan, we utilized Long Short-Term Memory (LSTM) implemented in TensorFlow. This approach allowed us to analyze the sequence of images capturing the spatial continuity and progression of anatomical structures. </p>\n<p>The following is a prototype of our model.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2497266%2F920be15372ec12783f6e8891e5f34336%2Ffigure2.png?generation=1699408559685906&amp;alt=media\" alt=\"model\"></p>\n<h3>3.2.2 Data Augmentation Techniques</h3>\n<p>To bolster the quality and variety of our dataset, we initially planned to use the <code>albumentations</code> library for data augmentation. However, due to compatibility issues with TPU training, the library could not be directly applied. As a workaround, we drew inspiration from the augmentation techniques available in <code>albumentations</code> and replicated them using TensorFlow. Our customized augmentation pipeline included transformations such as horizontal and vertical flipping, transposition, and various blurring techniques like Gaussian blur, mean blur, and motion blur. We also introduced random Gaussian noise to further augment the dataset.</p>\n<h2>3.3 Loss Function Design</h2>\n<h3>3.3.1 Organ-Specific Loss Calculation</h3>\n<p>For each organ (liver, spleen, kidneys, and bowels), the loss is calculated by taking the slices that correspond to the starting and ending indices of the organ within each scan. For example. If the liver is visualized from slice index 1 to 5, the model will predict the probability of injury for these slices, and the five separate loss values will be computed. The final loss for the liver will be the average of these five loss values. This process is repeated for each organ to obtain individual organ losses. </p>\n<h3>3.3.2 Extravasation Loss Calculation</h3>\n<p>For extravasation, the loss is computed across 32 evenly selected images from each scan. This comprehensive approach ensures that the loss calculation for extravasation is representative of the entire scan.</p>\n<h3>3.3.3 Any Injury Loss Calculation</h3>\n<p>The <code>any_injury</code> loss is calculated automatically on the Kaggle platform based on the previously mentioned loss values. This loss serves as an aggregate indicator of the model’s to detect any form of injury present across the scan.</p>\n<h3>3.3.4 Loss Function Weights</h3>\n<p>Different weights are applied within the loss function: a health prediction is assigned 1 point, and an injury prediction is given 2 points, emphasizing the model’s need to accurately identify injuries. </p>\n<h1>4.  Sources</h1>\n<ul>\n<li>Standardizing Unusual Dicoms: <a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/427217\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/427217</a></li>\n<li>Total Segmentator: <a href=\"https://github.com/wasserth/TotalSegmentator\" target=\"_blank\">https://github.com/wasserth/TotalSegmentator</a></li>\n<li>Data in PNG Format: <a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/427427\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/427427</a></li>\n<li>EfficientNet: Improving Accuracy and Efficiency through AutoML and Model Scaling: <a href=\"https://ai.googleblog.com/2019/05/efficientnet-improving-accuracy-and.html\" target=\"_blank\">https://ai.googleblog.com/2019/05/efficientnet-improving-accuracy-and.html</a></li>\n</ul>",
  "messages": [
    {
      "id": 2515659,
      "postDate": "2023-11-07T05:20:00.363Z",
      "content": "<p>First and foremost, we extend our gratitude to RSNA for organizing this captivating competition, through which we gained invaluable insights. </p>\n<h1>1. Context section</h1>\n<ul>\n<li>Business context: <a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/overview\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/overview</a></li>\n<li>Data context: <a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/data\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/data</a></li>\n</ul>\n<h1>2. Overview of the Approach</h1>\n<p>In this competition, the motivation is to contribute to the improvement of patient outcomes from traumatic injuries, which account for over 5 million global deaths each year. Blunt-force abdominal trauma, often caused by vehicle accidents, can damage organs and lead to internal bleeding. With traditional methods like physical exams and lab tests often being inconclusive, the emphasis shifts to the need for accurate medical imaging interpretations.</p>\n<p>In this study, our models predict a probability for each of the different possible injury types and degrees using CT scans. While vital for evaluating abdominal trauma, CT scans can be challenging to decipher, especially when injuries are subtle or multiple. Our goal is to harness the power of AI and ML to better interpret CT scans.</p>\n<h1>3.  Details of the submission</h1>\n<h2>3.1 Data Interpretation and Preprocessing</h2>\n<p>The dataset is primarily stored in the .dcm (DICOM) format, which offers a detailed insight into traumatic injuries. The patients underwent CT scans, with each undergoing one to two scans, generally covering from the upper neck down to the region below the anus. However, the specific number of scan images per individual, per scan, remained variable. This variability was due to the unknown intervals between each scan slice and the differing heights of the individuals, introducing a level of complexity in the initial assessment of the data. </p>\n<p>The scans could be classified into two types: complete and incomplete. A complete scan provides a comprehensive view, encompassing all organs within the range from the upper neck to the below the anal region. In contrast, an incomplete scan is a localized examination, focusing on specific organs within this anatomical spectrum. While the training dataset came labeled with indications of completeness, such classification were absent in the test dataset. The training dataset encompassed the CT scan images of 3,147 patients, while the test dataset was composed of approximately 1,300 patents’ scans.</p>\n<p>When these multiple .dcm images from a single scan were stacked together, they coalesced to form a vivid 3D representation of the human body, encapsulating a plethora of details. Such data, especially in the field of medical imaging, is paramount to diagnose and comprehend the intricacies of internal injuries.</p>\n<p>For convenience, we converted the given .dcm (DICOM) files into the more universally recognized .png format, facilitated by a <a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/427427\" target=\"_blank\">method</a> provided by Kaggle platform, streamlining subsequent operations and analyses. </p>\n<h2>3.2 Model Pipeline Design</h2>\n<h3>3.2.1 2D Semantic Segmentation Model</h3>\n<h4>Segmentation with U-Net</h4>\n<p>Our objective with U-Net was twofold: to identify the span of CT images that contained each organ (liver, spleen, kidneys, bowels) and to pinpoint their precise locations within each image. </p>\n<p>We initially employed the Total Segmentator tool (available at: <a href=\"https://github.com/wasserth/TotalSegmentator\" target=\"_blank\">Total Segmentator on GitHub</a>), but faced significant time constraints due to the tool’s processing speed. This constraint necessitated a shift in our approach toward a more efficient solution, leading us to use the U-Net architecture.   </p>\n<p>We utilized the Total Segmentator to generate marked pixel points as input data. These annotated 2D images served as training data for our U-Net model. By harnessing the efficiency of U-Net, we trained a 2D segmenter to efficiently discern the presence and location of the targeted organs across the CT scan images. The refined segmentation process not only indicated which scans contained the organs of interest but also provided their spatial coordinates, thus streamlining the subsequent stages of our analysis. </p>\n<p>The following figures are an illustration of segmentation with U-Net. The black-and-white figure on the left is the original CT scan image, and the right one is the one segmented. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2497266%2F79d5484239de43371fef834bdb1fd493%2Ffigure3.png?generation=1699408525140588&amp;alt=media\" alt=\"U-Net\"></p>\n<h4>Constructing 2.5D Input</h4>\n<p>For this part, our objective was to create a 2.5D input data for the neural network. We implemented a strategy to select 32 evenly spaced images from each scan across the training dataset. Given that the minimum number of slices per scan was 44, this uniform selection process ensured comprehensive coverage of each individual’s scan while maintaining a focus on representation consistency for each organ. Our selection also guaranteed that every organ appeared in at least 4 separate images, thereby capturing the essential anatomical features needed for accurate analysis. </p>\n<p>Then we used EfficientNet architecture. Each image was processed through the network, with the output of the penultimate layer capturing the feature representation extracted by the network after processing. Subsequently, a fully connected layer served as the final decision-making component, distinguishing between health and injury.  </p>\n<p>To synthesize the 32 image-derived insights into a single diagnostic outcome for each scan, we utilized Long Short-Term Memory (LSTM) implemented in TensorFlow. This approach allowed us to analyze the sequence of images capturing the spatial continuity and progression of anatomical structures. </p>\n<p>The following is a prototype of our model.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2497266%2F920be15372ec12783f6e8891e5f34336%2Ffigure2.png?generation=1699408559685906&amp;alt=media\" alt=\"model\"></p>\n<h3>3.2.2 Data Augmentation Techniques</h3>\n<p>To bolster the quality and variety of our dataset, we initially planned to use the <code>albumentations</code> library for data augmentation. However, due to compatibility issues with TPU training, the library could not be directly applied. As a workaround, we drew inspiration from the augmentation techniques available in <code>albumentations</code> and replicated them using TensorFlow. Our customized augmentation pipeline included transformations such as horizontal and vertical flipping, transposition, and various blurring techniques like Gaussian blur, mean blur, and motion blur. We also introduced random Gaussian noise to further augment the dataset.</p>\n<h2>3.3 Loss Function Design</h2>\n<h3>3.3.1 Organ-Specific Loss Calculation</h3>\n<p>For each organ (liver, spleen, kidneys, and bowels), the loss is calculated by taking the slices that correspond to the starting and ending indices of the organ within each scan. For example. If the liver is visualized from slice index 1 to 5, the model will predict the probability of injury for these slices, and the five separate loss values will be computed. The final loss for the liver will be the average of these five loss values. This process is repeated for each organ to obtain individual organ losses. </p>\n<h3>3.3.2 Extravasation Loss Calculation</h3>\n<p>For extravasation, the loss is computed across 32 evenly selected images from each scan. This comprehensive approach ensures that the loss calculation for extravasation is representative of the entire scan.</p>\n<h3>3.3.3 Any Injury Loss Calculation</h3>\n<p>The <code>any_injury</code> loss is calculated automatically on the Kaggle platform based on the previously mentioned loss values. This loss serves as an aggregate indicator of the model’s to detect any form of injury present across the scan.</p>\n<h3>3.3.4 Loss Function Weights</h3>\n<p>Different weights are applied within the loss function: a health prediction is assigned 1 point, and an injury prediction is given 2 points, emphasizing the model’s need to accurately identify injuries. </p>\n<h1>4.  Sources</h1>\n<ul>\n<li>Standardizing Unusual Dicoms: <a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/427217\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/427217</a></li>\n<li>Total Segmentator: <a href=\"https://github.com/wasserth/TotalSegmentator\" target=\"_blank\">https://github.com/wasserth/TotalSegmentator</a></li>\n<li>Data in PNG Format: <a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/427427\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/427427</a></li>\n<li>EfficientNet: Improving Accuracy and Efficiency through AutoML and Model Scaling: <a href=\"https://ai.googleblog.com/2019/05/efficientnet-improving-accuracy-and.html\" target=\"_blank\">https://ai.googleblog.com/2019/05/efficientnet-improving-accuracy-and.html</a></li>\n</ul>",
      "rawMarkdown": "First and foremost, we extend our gratitude to RSNA for organizing this captivating competition, through which we gained invaluable insights. \n\n# 1. Context section\n- Business context: [https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/overview](https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/overview)\n- Data context: [https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/data](https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/data)\n\n# 2. Overview of the Approach\nIn this competition, the motivation is to contribute to the improvement of patient outcomes from traumatic injuries, which account for over 5 million global deaths each year. Blunt-force abdominal trauma, often caused by vehicle accidents, can damage organs and lead to internal bleeding. With traditional methods like physical exams and lab tests often being inconclusive, the emphasis shifts to the need for accurate medical imaging interpretations.\n\nIn this study, our models predict a probability for each of the different possible injury types and degrees using CT scans. While vital for evaluating abdominal trauma, CT scans can be challenging to decipher, especially when injuries are subtle or multiple. Our goal is to harness the power of AI and ML to better interpret CT scans.\n\n# 3.  Details of the submission\n## 3.1 Data Interpretation and Preprocessing\nThe dataset is primarily stored in the .dcm (DICOM) format, which offers a detailed insight into traumatic injuries. The patients underwent CT scans, with each undergoing one to two scans, generally covering from the upper neck down to the region below the anus. However, the specific number of scan images per individual, per scan, remained variable. This variability was due to the unknown intervals between each scan slice and the differing heights of the individuals, introducing a level of complexity in the initial assessment of the data. \n\nThe scans could be classified into two types: complete and incomplete. A complete scan provides a comprehensive view, encompassing all organs within the range from the upper neck to the below the anal region. In contrast, an incomplete scan is a localized examination, focusing on specific organs within this anatomical spectrum. While the training dataset came labeled with indications of completeness, such classification were absent in the test dataset. The training dataset encompassed the CT scan images of 3,147 patients, while the test dataset was composed of approximately 1,300 patents’ scans.\n\nWhen these multiple .dcm images from a single scan were stacked together, they coalesced to form a vivid 3D representation of the human body, encapsulating a plethora of details. Such data, especially in the field of medical imaging, is paramount to diagnose and comprehend the intricacies of internal injuries.\n\nFor convenience, we converted the given .dcm (DICOM) files into the more universally recognized .png format, facilitated by a [method](https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/427427) provided by Kaggle platform, streamlining subsequent operations and analyses. \n\n## 3.2 Model Pipeline Design\n\n### 3.2.1 2D Semantic Segmentation Model\n#### Segmentation with U-Net\nOur objective with U-Net was twofold: to identify the span of CT images that contained each organ (liver, spleen, kidneys, bowels) and to pinpoint their precise locations within each image. \n\nWe initially employed the Total Segmentator tool (available at: [Total Segmentator on GitHub](https://github.com/wasserth/TotalSegmentator)), but faced significant time constraints due to the tool’s processing speed. This constraint necessitated a shift in our approach toward a more efficient solution, leading us to use the U-Net architecture.   \n\nWe utilized the Total Segmentator to generate marked pixel points as input data. These annotated 2D images served as training data for our U-Net model. By harnessing the efficiency of U-Net, we trained a 2D segmenter to efficiently discern the presence and location of the targeted organs across the CT scan images. The refined segmentation process not only indicated which scans contained the organs of interest but also provided their spatial coordinates, thus streamlining the subsequent stages of our analysis. \n\nThe following figures are an illustration of segmentation with U-Net. The black-and-white figure on the left is the original CT scan image, and the right one is the one segmented. \n\n![U-Net](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2497266%2F79d5484239de43371fef834bdb1fd493%2Ffigure3.png?generation=1699408525140588&alt=media)\n\n\n#### Constructing 2.5D Input\nFor this part, our objective was to create a 2.5D input data for the neural network. We implemented a strategy to select 32 evenly spaced images from each scan across the training dataset. Given that the minimum number of slices per scan was 44, this uniform selection process ensured comprehensive coverage of each individual’s scan while maintaining a focus on representation consistency for each organ. Our selection also guaranteed that every organ appeared in at least 4 separate images, thereby capturing the essential anatomical features needed for accurate analysis. \n\nThen we used EfficientNet architecture. Each image was processed through the network, with the output of the penultimate layer capturing the feature representation extracted by the network after processing. Subsequently, a fully connected layer served as the final decision-making component, distinguishing between health and injury.  \n\nTo synthesize the 32 image-derived insights into a single diagnostic outcome for each scan, we utilized Long Short-Term Memory (LSTM) implemented in TensorFlow. This approach allowed us to analyze the sequence of images capturing the spatial continuity and progression of anatomical structures. \n\nThe following is a prototype of our model.\n![model](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2497266%2F920be15372ec12783f6e8891e5f34336%2Ffigure2.png?generation=1699408559685906&alt=media)\n\n### 3.2.2 Data Augmentation Techniques\nTo bolster the quality and variety of our dataset, we initially planned to use the `albumentations` library for data augmentation. However, due to compatibility issues with TPU training, the library could not be directly applied. As a workaround, we drew inspiration from the augmentation techniques available in `albumentations` and replicated them using TensorFlow. Our customized augmentation pipeline included transformations such as horizontal and vertical flipping, transposition, and various blurring techniques like Gaussian blur, mean blur, and motion blur. We also introduced random Gaussian noise to further augment the dataset.\n\n## 3.3 Loss Function Design\n### 3.3.1 Organ-Specific Loss Calculation\nFor each organ (liver, spleen, kidneys, and bowels), the loss is calculated by taking the slices that correspond to the starting and ending indices of the organ within each scan. For example. If the liver is visualized from slice index 1 to 5, the model will predict the probability of injury for these slices, and the five separate loss values will be computed. The final loss for the liver will be the average of these five loss values. This process is repeated for each organ to obtain individual organ losses. \n\n### 3.3.2 Extravasation Loss Calculation\nFor extravasation, the loss is computed across 32 evenly selected images from each scan. This comprehensive approach ensures that the loss calculation for extravasation is representative of the entire scan.\n\n### 3.3.3 Any Injury Loss Calculation\nThe `any_injury` loss is calculated automatically on the Kaggle platform based on the previously mentioned loss values. This loss serves as an aggregate indicator of the model’s to detect any form of injury present across the scan.\n\n### 3.3.4 Loss Function Weights\nDifferent weights are applied within the loss function: a health prediction is assigned 1 point, and an injury prediction is given 2 points, emphasizing the model’s need to accurately identify injuries. \n\n# 4.  Sources\n- Standardizing Unusual Dicoms: [https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/427217](https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/427217)\n- Total Segmentator: [https://github.com/wasserth/TotalSegmentator](https://github.com/wasserth/TotalSegmentator)\n- Data in PNG Format: [https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/427427](https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/427427)\n- EfficientNet: Improving Accuracy and Efficiency through AutoML and Model Scaling: [https://ai.googleblog.com/2019/05/efficientnet-improving-accuracy-and.html](https://ai.googleblog.com/2019/05/efficientnet-improving-accuracy-and.html)\n",
      "votes": 7
    },
    {
      "id": 2622566,
      "postDate": "2024-01-27T15:40:39.173Z",
      "content": "<p>can I ask about your accuracy? Thank you!</p>",
      "rawMarkdown": "can I ask about your accuracy? Thank you!"
    }
  ],
  "comments": [
    {
      "id": 2622566,
      "author_name": "Hoài Nhi",
      "author_url": "",
      "post_date": "2024-01-27T15:40:39.173000",
      "content": "<p>can I ask about your accuracy? Thank you!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2515659": "First and foremost, we extend our gratitude to RSNA for organizing this captivating competition, through which we gained invaluable insights. \n\n# 1. Context section\n- Business context: [https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/overview](https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/overview)\n- Data context: [https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/data](https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/data)\n\n# 2. Overview of the Approach\nIn this competition, the motivation is to contribute to the improvement of patient outcomes from traumatic injuries, which account for over 5 million global deaths each year. Blunt-force abdominal trauma, often caused by vehicle accidents, can damage organs and lead to internal bleeding. With traditional methods like physical exams and lab tests often being inconclusive, the emphasis shifts to the need for accurate medical imaging interpretations.\n\nIn this study, our models predict a probability for each of the different possible injury types and degrees using CT scans. While vital for evaluating abdominal trauma, CT scans can be challenging to decipher, especially when injuries are subtle or multiple. Our goal is to harness the power of AI and ML to better interpret CT scans.\n\n# 3.  Details of the submission\n## 3.1 Data Interpretation and Preprocessing\nThe dataset is primarily stored in the .dcm (DICOM) format, which offers a detailed insight into traumatic injuries. The patients underwent CT scans, with each undergoing one to two scans, generally covering from the upper neck down to the region below the anus. However, the specific number of scan images per individual, per scan, remained variable. This variability was due to the unknown intervals between each scan slice and the differing heights of the individuals, introducing a level of complexity in the initial assessment of the data. \n\nThe scans could be classified into two types: complete and incomplete. A complete scan provides a comprehensive view, encompassing all organs within the range from the upper neck to the below the anal region. In contrast, an incomplete scan is a localized examination, focusing on specific organs within this anatomical spectrum. While the training dataset came labeled with indications of completeness, such classification were absent in the test dataset. The training dataset encompassed the CT scan images of 3,147 patients, while the test dataset was composed of approximately 1,300 patents’ scans.\n\nWhen these multiple .dcm images from a single scan were stacked together, they coalesced to form a vivid 3D representation of the human body, encapsulating a plethora of details. Such data, especially in the field of medical imaging, is paramount to diagnose and comprehend the intricacies of internal injuries.\n\nFor convenience, we converted the given .dcm (DICOM) files into the more universally recognized .png format, facilitated by a [method](https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/427427) provided by Kaggle platform, streamlining subsequent operations and analyses. \n\n## 3.2 Model Pipeline Design\n\n### 3.2.1 2D Semantic Segmentation Model\n#### Segmentation with U-Net\nOur objective with U-Net was twofold: to identify the span of CT images that contained each organ (liver, spleen, kidneys, bowels) and to pinpoint their precise locations within each image. \n\nWe initially employed the Total Segmentator tool (available at: [Total Segmentator on GitHub](https://github.com/wasserth/TotalSegmentator)), but faced significant time constraints due to the tool’s processing speed. This constraint necessitated a shift in our approach toward a more efficient solution, leading us to use the U-Net architecture.   \n\nWe utilized the Total Segmentator to generate marked pixel points as input data. These annotated 2D images served as training data for our U-Net model. By harnessing the efficiency of U-Net, we trained a 2D segmenter to efficiently discern the presence and location of the targeted organs across the CT scan images. The refined segmentation process not only indicated which scans contained the organs of interest but also provided their spatial coordinates, thus streamlining the subsequent stages of our analysis. \n\nThe following figures are an illustration of segmentation with U-Net. The black-and-white figure on the left is the original CT scan image, and the right one is the one segmented. \n\n![U-Net](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2497266%2F79d5484239de43371fef834bdb1fd493%2Ffigure3.png?generation=1699408525140588&alt=media)\n\n\n#### Constructing 2.5D Input\nFor this part, our objective was to create a 2.5D input data for the neural network. We implemented a strategy to select 32 evenly spaced images from each scan across the training dataset. Given that the minimum number of slices per scan was 44, this uniform selection process ensured comprehensive coverage of each individual’s scan while maintaining a focus on representation consistency for each organ. Our selection also guaranteed that every organ appeared in at least 4 separate images, thereby capturing the essential anatomical features needed for accurate analysis. \n\nThen we used EfficientNet architecture. Each image was processed through the network, with the output of the penultimate layer capturing the feature representation extracted by the network after processing. Subsequently, a fully connected layer served as the final decision-making component, distinguishing between health and injury.  \n\nTo synthesize the 32 image-derived insights into a single diagnostic outcome for each scan, we utilized Long Short-Term Memory (LSTM) implemented in TensorFlow. This approach allowed us to analyze the sequence of images capturing the spatial continuity and progression of anatomical structures. \n\nThe following is a prototype of our model.\n![model](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2497266%2F920be15372ec12783f6e8891e5f34336%2Ffigure2.png?generation=1699408559685906&alt=media)\n\n### 3.2.2 Data Augmentation Techniques\nTo bolster the quality and variety of our dataset, we initially planned to use the `albumentations` library for data augmentation. However, due to compatibility issues with TPU training, the library could not be directly applied. As a workaround, we drew inspiration from the augmentation techniques available in `albumentations` and replicated them using TensorFlow. Our customized augmentation pipeline included transformations such as horizontal and vertical flipping, transposition, and various blurring techniques like Gaussian blur, mean blur, and motion blur. We also introduced random Gaussian noise to further augment the dataset.\n\n## 3.3 Loss Function Design\n### 3.3.1 Organ-Specific Loss Calculation\nFor each organ (liver, spleen, kidneys, and bowels), the loss is calculated by taking the slices that correspond to the starting and ending indices of the organ within each scan. For example. If the liver is visualized from slice index 1 to 5, the model will predict the probability of injury for these slices, and the five separate loss values will be computed. The final loss for the liver will be the average of these five loss values. This process is repeated for each organ to obtain individual organ losses. \n\n### 3.3.2 Extravasation Loss Calculation\nFor extravasation, the loss is computed across 32 evenly selected images from each scan. This comprehensive approach ensures that the loss calculation for extravasation is representative of the entire scan.\n\n### 3.3.3 Any Injury Loss Calculation\nThe `any_injury` loss is calculated automatically on the Kaggle platform based on the previously mentioned loss values. This loss serves as an aggregate indicator of the model’s to detect any form of injury present across the scan.\n\n### 3.3.4 Loss Function Weights\nDifferent weights are applied within the loss function: a health prediction is assigned 1 point, and an injury prediction is given 2 points, emphasizing the model’s need to accurately identify injuries. \n\n# 4.  Sources\n- Standardizing Unusual Dicoms: [https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/427217](https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/427217)\n- Total Segmentator: [https://github.com/wasserth/TotalSegmentator](https://github.com/wasserth/TotalSegmentator)\n- Data in PNG Format: [https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/427427](https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/427427)\n- EfficientNet: Improving Accuracy and Efficiency through AutoML and Model Scaling: [https://ai.googleblog.com/2019/05/efficientnet-improving-accuracy-and.html](https://ai.googleblog.com/2019/05/efficientnet-improving-accuracy-and.html)\n",
    "2622566": "can I ask about your accuracy? Thank you!"
  }
}