{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":52254,"databundleVersionId":6863140,"sourceType":"competition"},{"sourceId":6983983,"sourceType":"datasetVersion","datasetId":4013814}],"dockerImageVersionId":30587,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Contents\n- [Introduction](#Introduction)\n    - [Prepare the data analysis](#Prepare_the_data_analysis)\n    - [Load packages](#Load_packages)\n    - [Load the data](#Load_the_data)\n- [Data exploration](#Data_exploration)\n    - [Missing data](#Load_the_data)\n    - [Mergeing Lables Tables](#Mergeing_Lables_Tables)\n    - [Explore DICOM data](#Explore_DICOM_data)\n    - [Loading data](#Loading_data)\n    - [DICOM Meat Data](#DICOM_Meat_Data)\n    - [Loading DICOM images](#Loading_DICOM_images)\n- [Conclusions](#Conclusions)\n\n\n","metadata":{}},{"cell_type":"code","source":"!pip install pytz\n!pip install pydicom\n\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from datetime import datetime\nimport pytz\n\n# Get the current time in UTC/GMT\ngmt = pytz.timezone('GMT')\ndt_utc = datetime.now(gmt)\ndt_string_utc = dt_utc.strftime(\"%d/%m/%Y %H:%M:%S\")\n\nprint(f\"Updated {dt_string_utc} (GMT)\")\n\n# Get the current time in GMT+3\ngmt3 = pytz.timezone('Europe/Moscow')  # Adjust the time zone string as needed\ndt_gmt3 = datetime.now(gmt3)\ndt_string_gmt3 = dt_gmt3.strftime(\"%d/%m/%Y %H:%M:%S\")\n\nprint(f\"Updated {dt_string_gmt3} (GMT+3)\")\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# [Introduction](#Introduction)\n\n* RSNA Abdominal Trauma Detection AI Challenge addresses global traumatic injury, a leading cause of death.\n* Focus on blunt force abdominal trauma, often caused by motor vehicle accidents.\n* Goal is to use AI and machine learning to improve the rapid and accurate diagnosis of abdominal injuries.\n* Computed tomography (CT) scans play a crucial role in evaluating patients with suspected abdominal injuries.\n* Challenge involves developing advanced algorithms to assist medical professionals in detecting and grading the severity of injuries to internal abdominal organs (liver, kidneys, spleen, bowel).\n* Emphasis on identifying active internal bleeding for timely interventions.\n* Collaboration with RSNA, American Society of Emergency Radiology (ASER), and Society for Abdominal Radiology (SAR).\n* Potential impact on improving trauma care and patient outcomes in emergency settings worldwide.\n","metadata":{}},{"cell_type":"markdown","source":"# [Prepare the data analysis](Prepare_the_data_analysis)\n\n\n## [Load packages](#Load_packages)\n\n","metadata":{}},{"cell_type":"code","source":"import pandas as pd \nimport numpy as np\nimport matplotlib\nimport matplotlib.pyplot as plt\nfrom tqdm import tqdm_notebook\nfrom matplotlib.patches import Rectangle\nimport seaborn as sns\nimport pydicom as dcm\n%matplotlib inline \nIS_LOCAL = False\nimport os\nif(IS_LOCAL):\n    PATH=\"../input/rsna-2023-abdominal-trauma-detection\"\n    local = True\nelse:\n    PATH=\"../input/\"\n    local = False\n    \nprint(f\"Is it loacal:{local}\")\nprint(os.listdir(PATH))\n\nprint(\"The dataset main files and folders:\")\nprint(os.listdir(\"../input/rsna-2023-abdominal-trauma-detection\"))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"markdown","source":"## [Load Data](load_data)\n\nLet's load the tabular data. There are five files:\n\n* Image lelve lables :\n* train:\n* train seris meta:\n* test series meta :\n* sample submission:","metadata":{}},{"cell_type":"code","source":"PATH = '/kaggle/input/rsna-2023-abdominal-trauma-detection'\n\nimage_level_labels_df = pd.read_csv(PATH+'/image_level_labels.csv')\ntrain_df = pd.read_csv(PATH+'/train.csv')\ntrain_series_meta_df = pd.read_csv(PATH+'/train_series_meta.csv')\ntest_series_meta_df = pd.read_csv(PATH+'/test_series_meta.csv')\nsample_submission_df = pd.read_csv(PATH+'/sample_submission.csv')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(f\"image level lables-  rows: {image_level_labels_df.shape[0]}, columns: {image_level_labels_df.shape[1]}\")\nprint(f\"Train -  rows: {train_df.shape[0]}, columns: {train_df.shape[1]}\")\nprint(f\"train series meta -  rows: {train_series_meta_df.shape[0]}, columns: {train_series_meta_df.shape[1]}\")\n\nprint(f\"Test series meta -  rows: {test_series_meta_df.shape[0]}, columns: {test_series_meta_df.shape[1]}\")\n\nprint(f\"Sample subission -  rows: {sample_submission_df.shape[0]}, columns: {sample_submission_df.shape[1]}\")\n\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Let's explore the two loaded files. We will take out a 5 rows samples from each dataset.\n","metadata":{}},{"cell_type":"code","source":"image_level_labels_df.head(10)\n#image_level_labels_df['injury_name']\n\nprint (f'the length of image_level_labels_df :  {len(image_level_labels_df)}' )\n#patient_id\nprint (f'the number of patients: {len(image_level_labels_df[\"patient_id\"].unique())}' )\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"image_level_labels_df['injury_name'].unique()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\nlen(train_df)\nlen(train_df[\"patient_id\"].unique())\nrepeated_values  = train_df['patient_id'].value_counts()\nrepeated_values = repeated_values[repeated_values > 1]\nprint(repeated_values)\n\n\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_series_meta_df.head(10)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(train_series_meta_df[\"patient_id\"].unique())","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_series_meta_df.head(10)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_submission_df.head(10)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* **Image Level Labels Dataset:** Contains patient_id, series_id, instance_number, and injury_name.\nEach row represents a unique injury for a patient, and patients with multiple injuries are duplicated for each distinct injury_name.\n* **Train Dataset:** Provides the type of injury for each patient's image along with the corresponding patient ID.\nRepresents the presence of the disease or state with binary values.\nAllows for assigning more than one injury_name for each patient ID in a single row, unlike the previous dataset.\n* **Train Series Meta Dataset:** Includes patient_id, series_id, aortic_hu, and incomplete_organ.\n* **Test Series Meta Dataset:** Contains patient_id, series_id, and aortic_hu.\n* **Sample Submission Table:** Serves as a guide for formatting the released results, specifying how the results should be presented.","metadata":{}},{"cell_type":"markdown","source":"folder exploration","metadata":{}},{"cell_type":"code","source":"import os\n\ndef count_folders(path):\n    # Get the list of items in the directory\n    items = os.listdir(path)\n\n    # Filter out only the directories\n    folders = [item for item in items if os.path.isdir(os.path.join(path, item))]\n\n    # Return the count of folders\n    return len(folders)\n\n# Replace 'your_folder_path' with the path to your folder\nfolder_path = '/kaggle/input/rsna-2023-abdominal-trauma-detection/train_images'\nnum_folders = count_folders(folder_path)\n\nprint(f'The number of folders in {folder_path} is: {num_folders}')\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# [Data exploration](#Data_exploration)\nLet's explore the data further.\n\n## [Missing Data](#Missing_Data)\n\nLet's check missing information in the two datasets.\n\n","metadata":{}},{"cell_type":"code","source":"def missing_data(data):\n    total = data.isnull().sum().sort_values(ascending = False)\n    percent = (data.isnull().sum()/data.isnull().count()*100).sort_values(ascending = False)\n    return np.transpose(pd.concat([total, percent], axis=1, keys=['Total', 'Percent']))\n\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"missing_data(image_level_labels_df)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"missing_data(train_df)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"missing_data(train_series_meta_df)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"missing_data(test_series_meta_df)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"missing_data(sample_submission_df)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"So no missing values !!","metadata":{}},{"cell_type":"markdown","source":"**Let's check the class distribution from class detailed info.**","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\nf, ax = plt.subplots(1, 1, figsize=(10, 6))\n\n# Convert 'injury_name' to categorical \nimage_level_labels_df['injury_name'] = pd.Categorical(image_level_labels_df['injury_name'])\n\n# Use the value_counts() to get the order and then create the countplot\nsns.countplot(\n    x='injury_name',\n    data=image_level_labels_df,\n    order=image_level_labels_df['injury_name'].value_counts().index,\n    palette='husl'  # Use a different color palette here\n)\n\n#plotting\nfor p in ax.patches:\n    height = p.get_height()\n    ax.text(p.get_x() + p.get_width() / 2, height + 0.1, height, ha=\"center\")\n\nplt.show()\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def get_feature_distribution(data, feature):\n    # Get the count for each label\n    label_counts = data[feature].value_counts()\n\n    # Get total number of samples\n    total_samples = len(data)\n\n    # Count the number of items in each class\n    print(\"Feature: {}\".format(feature))\n    for i in range(len(label_counts)):\n        label = label_counts.index[i]\n        count = label_counts.values[i]\n        percent = int((count / total_samples) * 10000) / 100\n        print(\"{:<30s}:   {} or {}%\".format(label, count, percent))\n\nget_feature_distribution(image_level_labels_df, 'injury_name')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## [Mergeing Lables Tables](#Mergeing_Lables_Tables)","metadata":{}},{"cell_type":"markdown","source":"**train_class_df , here a new dataframe created which consits of a combnation of image_level_lables_df and train_df**","metadata":{}},{"cell_type":"code","source":"print(f'The train_df patients are:{len(train_df[\"patient_id\"].unique())}')\nprint(f'The train_series_meta_df patients are:{len(train_series_meta_df[\"patient_id\"].unique())}')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"final_train_df = train_series_meta_df.merge(train_df, left_on='patient_id', right_on='patient_id', how='inner')\nfinal_train_df.head(10)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#len(train_class_df)\nlen(final_train_df['patient_id'].unique())\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**If the submission does not contain information about the spleen, the spleen columns were eliminated alongside any injury columns as well.**","metadata":{}},{"cell_type":"markdown","source":"To streamline processing, remove unnecessary columns from the dataset.","metadata":{}},{"cell_type":"code","source":"del final_train_df['spleen_healthy']\ndel final_train_df['spleen_low']\ndel final_train_df['spleen_high']\ndel final_train_df['any_injury']\n\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"final_train_df.head(10)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"here we calculcated the incompleteed scanes \n","metadata":{}},{"cell_type":"code","source":"len(final_train_df.columns)\nfinal_train_df.columns\nlen(final_train_df['patient_id'].unique())\nlen(final_train_df['series_id'].unique())\nlen(final_train_df['incomplete_organ'])\n\ncom = 0\nincom = 0 \n\nfor y in range(len(final_train_df)):\n    if final_train_df['incomplete_organ'][y] == 0:\n        com = com + 1\n    else:\n        incom = incom+1\n        #print('1')\n\nprint (f'The completed images : {com} and  incompleated scans : {incom} --->  total images = {com+incom}')\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We primarily need to exclude those images from this dataset; essentially, we have to remove them.","metadata":{}},{"cell_type":"markdown","source":"## [Explore DICOM data](#Explore_DICOM_data)","metadata":{}},{"cell_type":"code","source":"image_sample_path = os.listdir('/kaggle/input/rsna-2023-abdominal-trauma-detection/train_images')[:5]\nprint(image_sample_path)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's check how many images are in the train and test folders.\n\n1) number of training images\n","metadata":{}},{"cell_type":"code","source":"'''\nimport os\n\ndef count_images_in_folder(folder_path, valid_extensions=('dcm')):\n    image_count = 0\n\n    for root, dirs, files in os.walk(folder_path):\n        for file in files:\n            if file.lower().endswith(valid_extensions):\n                image_count += 1\n\n    return image_count\n\n# Replace 'your_folder_path' with the path to the top-level folder containing subfolders with images\nfolder_path = '/kaggle/input/rsna-2023-abdominal-trauma-detection/train_images'\n\ntotal_images = count_images_in_folder(folder_path)\nprint(f'Total number of images in the folders: {total_images}')\n'''\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"2) number of test ","metadata":{}},{"cell_type":"code","source":"import os\n\ndef count_images_in_folder(folder_path, valid_extensions=('dcm')):\n    image_count = 0\n\n    for root, dirs, files in os.walk(folder_path):\n        for file in files:\n            if file.lower().endswith(valid_extensions):\n                image_count += 1\n\n    return image_count\n\n# Replace 'your_folder_path' with the path to the top-level folder containing subfolders with images\nfolder_path = '/kaggle/input/rsna-2023-abdominal-trauma-detection/test_images'\n\ntotal_images = count_images_in_folder(folder_path)\nprint(f'Total number of images in the folders: {total_images}')\n\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Check duplicates in train dataset\n","metadata":{}},{"cell_type":"code","source":"print(\"Unique patientId in  train_class_df: \", final_train_df['patient_id'].nunique())      \n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**We verified that the quantity of distinct patient IDs and the amount of DICOM images in the training set are equivalent.\n\nLet's examine the entries that are repeated. We wish to ascertain the distribution of these among the classes and the target value.**","metadata":{}},{"cell_type":"markdown","source":"## [DICOM Meta Data](#DICOM_Meta_Data)","metadata":{}},{"cell_type":"code","source":"samplePatientID = list(final_train_df[:3].T.to_dict().values())[0]['patient_id']\ndicom_file_path = os.path.join(\"/kaggle/input/rsna-2023-abdominal-trauma-detection/train_images/10004/21057/1000.dcm\")\ndicom_file_dataset = dcm.read_file(dicom_file_path)\ndicom_file_dataset","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"*  (0018, 5100) Patient Position                    CS: 'HFS'\n* (0028, 0010) Rows                                US: 512\n*  (0028, 0011) Columns                             US: 512\n*  (0028, 0030) Pixel Spacing                       DS: [0.89453125, 0.89453125]","metadata":{}},{"cell_type":"code","source":"final_train_df","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Add a new column 'general_id' by combining 'patient' and 'series'\nfinal_train_df['general_id'] = final_train_df['patient_id'].astype(str) + '_' + final_train_df['series_id'].astype(str)\n# Get the index of the 'series' column\nseries_col_index = final_train_df.columns.get_loc('series_id')\n# Insert the 'general_id' column after the 'series' column\nfinal_train_df.insert(series_col_index + 1, 'general_id', final_train_df.pop('general_id'))\n\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"final_train_df","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## [Loading DICOM images](#Loading_DICOM_images)","metadata":{}},{"cell_type":"code","source":"import os\nimport pydicom\nimport matplotlib.pyplot as plt\nimport numpy as np\n\n# Specify the path to the folder containing DICOM files\ndicom_folder = '/kaggle/input/rsna-2023-abdominal-trauma-detection/train_images/1027/24515'\n\n# Get a list of all DICOM files in the folder\ndicom_files = [os.path.join(dicom_folder, file) for file in os.listdir(dicom_folder) if file.endswith('.dcm')]\n\n# Display details for each sample DICOM image\nnum_samples = 5\nfor i in range(min(num_samples, len(dicom_files))):\n    # Load DICOM image\n    dicom_data = pydicom.dcmread(dicom_files[i])\n\n    # Display image\n    plt.subplot(1, num_samples, i + 1)\n    plt.imshow(dicom_data.pixel_array, cmap='gray')\n    plt.title(f\"Image {i + 1}\")\n\n    # Get image details\n    dimensions = dicom_data.pixel_array.shape\n    pixel_spacing = dicom_data.PixelSpacing if hasattr(dicom_data, 'PixelSpacing') else None\n    slice_thickness = dicom_data.SliceThickness if hasattr(dicom_data, 'SliceThickness') else None\n\n    print(f\"\\nDetails for Image {i + 1}:\")\n    print(f\"Dimensions: {dimensions}\")\n    print(f\"Pixel Spacing: {pixel_spacing}\")\n    print(f\"Slice Thickness: {slice_thickness}\")\n\n# Show the plot\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# [Conclusions](#Conclusions):\n\nAfter exploring the data, both the tabular and DICOM data, we were able to:\n\n* 1) understand the detail of dataset.\n* 2) understand the submitiong system.\n* 3) tracjiogm the duplicaitons and missing data.\n* 5) Meraing tables of data and droping usels information.\n* 6) previong teh dicom images and htier find properties.\n* 7) preparing for how machine elarng dataset will be prepared.\n\n","metadata":{}}]}