{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.11.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":99552,"databundleVersionId":13190393,"sourceType":"competition"}],"dockerImageVersionId":31089,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# RSNA Intracranial Aneurysm Detection\n\n🎯 Objective of the Competition:\n\n- The goal is to develop machine learning models that can automatically detect and localize intracranial (brain) saccular aneurysms across multiple imaging modalities.\n\n\nWhat you must predict ?\nFor each imaging series, output probabilities for the presence of an aneurysm in:\n\n- 13 anatomical locations, and\n- A global label Aneurysm Present (if any aneurysm exists).\n\n---\nHow submissions are scored ?\n\nWeighted multilabel AUC-ROC:\n\n- Each of the 14 targets (13 arteries + Aneurysm Present) is evaluated with AUC-ROC.\n\n- <span style=\"color:red\">Aneurysm Present gets 13× more weight than each artery-specific label.</span>\n\nFinal score = weighted average across all 14 targets.\n\n----\n\nThe big picture\n\n- Model should not only say if an aneurysm is present but also indicate where it is, supporting radiologists in early detection and localization.\n\n----","metadata":{}},{"cell_type":"markdown","source":"colocar conclusions do EDA","metadata":{}},{"cell_type":"markdown","source":"## Libraries","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport os\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nfrom scipy.stats import chi2_contingency\nimport pydicom\nfrom tqdm import tqdm\n","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2025-08-20T17:12:26.665527Z","iopub.execute_input":"2025-08-20T17:12:26.665711Z","iopub.status.idle":"2025-08-20T17:12:31.875321Z","shell.execute_reply.started":"2025-08-20T17:12:26.665691Z","shell.execute_reply":"2025-08-20T17:12:31.874216Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## EDA","metadata":{}},{"cell_type":"markdown","source":"### Preparing the Data","metadata":{}},{"cell_type":"code","source":"# ==========================================\n#             Loading the files\n# ==========================================\ntrain_df = pd.read_csv(\"/kaggle/input/rsna-intracranial-aneurysm-detection/train.csv\")\ntrain_loc = pd.read_csv(\"/kaggle/input/rsna-intracranial-aneurysm-detection/train_localizers.csv\")\ntrain_df.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-20T17:13:13.106896Z","iopub.execute_input":"2025-08-20T17:13:13.10721Z","iopub.status.idle":"2025-08-20T17:13:13.188243Z","shell.execute_reply.started":"2025-08-20T17:13:13.10719Z","shell.execute_reply":"2025-08-20T17:13:13.186586Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"🧠 Extracted DICOM Metadata — Purpose and Rationale\n\nIn this step, we parsed the DICOM headers to extract key metadata that can provide clinical or technical context to each imaging series. These fields are not part of the training labels, but they can offer useful insights for analysis and modeling.\n\n✅ Extracted Fields and Their Purpose\n\n| Field Name                  | Purpose                                                                       |\n| --------------------------- | ----------------------------------------------------------------------------- |\n| `SeriesInstanceUID`         | Unique identifier for the image series (used for joins).                      |\n| `Modality`                  | Indicates the imaging type (e.g., CT, MR, MRA) — helps infer diagnostic flow. |\n| `SeriesDescription`         | Free-text description of the series — may contain protocol hints.             |\n| `ProtocolName`              | Imaging acquisition protocol — useful for understanding scan objectives.      |\n| `BodyPartExamined`          | Specifies the anatomical region imaged (usually brain-related here).          |\n| `Manufacturer`              | Vendor of the imaging equipment — useful for model generalization analysis.   |\n| `ManufacturerModelName`     | Specific scanner model used — may reflect resolution differences.             |\n| `StationName`               | Name of the imaging station — could reflect hospital or scanner unit.         |\n| `SoftwareVersions`          | Imaging software version — potential confounder in image format/quality.      |\n| `SliceThickness`            | Thickness of CT/MRI slices — affects spatial resolution.                      |\n| `PixelSpacing`              | Spacing between pixels in mm — important for voxel calibration.               |\n| `KVP`                       | Tube voltage (for CT scans) — technical parameter for image quality.          |\n| `Exposure`                  | Radiation dose estimate — related to image quality vs. patient safety.        |\n| `ImageType`                 | Indicates image derivation (original, reformatted, etc.).                     |\n| `StudyDate` / `SeriesDate`  | May provide temporal clues (e.g., first vs. follow-up scans).                 |\n| `PatientAge` / `PatientSex` | Patient demographic information — can be predictive of aneurysm risk.         |\n| `LossyImageCompression`     | Indicates if the image was compressed — may reduce quality.                   |\n| `BurnedInAnnotation`        | Indicates if text overlays exist on the image — may affect preprocessing.     |\n\n----\n\nThese metadata fields help us:\n\n- Understand image quality and acquisition protocols\n\n- Identify scanner variability, which can impact model generalization\n\n- Possibly reconstruct patient diagnostic flows (e.g., CT → MR)\n\n- Enrich the dataset with clinical context, useful for a multimodal model","metadata":{}},{"cell_type":"code","source":"# ======================================\n# Joining train metadada with DICOM\n# =======================================\n\n# Root directory for the DICOM series\nroot_dir = \"/kaggle/input/rsna-intracranial-aneurysm-detection/series\"\n\n# Get only the SeriesInstanceUIDs present in train.csv\ntrain_uids = set(train_df[\"SeriesInstanceUID\"].unique())\n\n# Limit processing to series used in training\nseries_dirs = [uid for uid in os.listdir(root_dir) if uid in train_uids]\n\n# Fields to extract from DICOM metadata\nselected_keywords = [\n    \"SeriesInstanceUID\", \"Modality\", \"SeriesDescription\", \"ProtocolName\",\n    \"BodyPartExamined\", \"Manufacturer\", \"ManufacturerModelName\", \"StationName\",\n    \"SoftwareVersions\", \"SliceThickness\", \"PixelSpacing\", \"KVP\", \"Exposure\",\n    \"ImageType\", \"StudyDate\", \"SeriesDate\", \"PatientAge\", \"PatientSex\",\n    \"LossyImageCompression\", \"BurnedInAnnotation\"\n]\n\ndicom_metadata_list = []\n\n# Loop over DICOM series used in training\nfor series_uid in tqdm(series_dirs):\n    series_path = os.path.join(root_dir, series_uid)\n    dicom_files = [f for f in os.listdir(series_path) if f.endswith(\".dcm\")]\n    if not dicom_files:\n        continue\n\n    dicom_path = os.path.join(series_path, dicom_files[0])\n    try:\n        ds = pydicom.dcmread(dicom_path, stop_before_pixels=True)\n        meta = {\"SeriesInstanceUID\": series_uid}\n\n        for keyword in selected_keywords:\n            if keyword == \"SeriesInstanceUID\":\n                continue\n            val = ds.get(keyword, None)\n            meta[keyword] = str(val) if val not in [None, \"\", [], {}] else None\n\n        dicom_metadata_list.append(meta)\n\n    except Exception:\n        continue\n\n# Convert to DataFrame and drop columns with all missing values\ndf_filtered = pd.DataFrame(dicom_metadata_list)\ndf_filtered = df_filtered.dropna(axis=1, how=\"all\")\n\n# Optionally: merge with train_df immediately\ntrain_EDA = train_df.merge(df_filtered, on=\"SeriesInstanceUID\", how=\"left\")\n\n# Preview\ntrain_EDA.head()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-20T17:13:21.588572Z","iopub.execute_input":"2025-08-20T17:13:21.588874Z","iopub.status.idle":"2025-08-20T17:18:47.915875Z","shell.execute_reply.started":"2025-08-20T17:13:21.58885Z","shell.execute_reply":"2025-08-20T17:18:47.914885Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ========================================================\n# Let's create a column to mark the low qiality images\n# ========================================================\n\n# Garantir que os campos estão em formato numérico\ntrain_EDA[\"SliceThickness\"] = pd.to_numeric(train_EDA[\"SliceThickness\"], errors=\"coerce\")\ntrain_EDA[\"Exposure\"] = pd.to_numeric(train_EDA[\"Exposure\"], errors=\"coerce\")\n\n# Critério de baixa qualidade\ntrain_EDA[\"IMG_Low_Quality_CT\"] = (train_EDA[\"SliceThickness\"] > 5) | (train_EDA[\"Exposure\"] < 10)\n\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-20T17:21:06.965688Z","iopub.execute_input":"2025-08-20T17:21:06.967019Z","iopub.status.idle":"2025-08-20T17:21:06.978843Z","shell.execute_reply.started":"2025-08-20T17:21:06.96698Z","shell.execute_reply":"2025-08-20T17:21:06.978083Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_EDA.info()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-20T17:21:10.576033Z","iopub.execute_input":"2025-08-20T17:21:10.57634Z","iopub.status.idle":"2025-08-20T17:21:10.606145Z","shell.execute_reply.started":"2025-08-20T17:21:10.576315Z","shell.execute_reply":"2025-08-20T17:21:10.605254Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Exploring how train metadata are relate to the target label.","metadata":{}},{"cell_type":"code","source":"# --- Compute totals and percentages ---\ntotal_cases = len(train_df)\n\n# Aneurysm prevalence\naneurysm_counts = train_df[\"Aneurysm Present\"].value_counts().sort_index()\naneurysm_pct = (aneurysm_counts / total_cases * 100).round(2)\n\n# Sex distribution by aneurysm\nsex_counts = train_df.groupby([\"PatientSex\", \"Aneurysm Present\"]).size().unstack(fill_value=0)\nsex_pct = (sex_counts.div(sex_counts.sum(axis=1), axis=0) * 100).round(2)\n\n# Age distribution summary\nage_summary = train_df.groupby(\"Aneurysm Present\")[\"PatientAge\"].describe().round(1)\n\n# --- Charts ---\n\n# Aneurysm prevalence\nplt.figure(figsize=(6,4))\nax = sns.countplot(x=\"Aneurysm Present\", data=train_df, palette=\"Set1\")\nplt.title(\"Aneurysm Presence Distribution\")\nplt.xticks([0,1], [\"No Aneurysm\", \"Aneurysm Present\"])\nplt.ylabel(\"Number of Patients\")\nfor p, count, pct in zip(ax.patches, aneurysm_counts, aneurysm_pct):\n    ax.annotate(f'{count}\\n({pct}%)', \n                (p.get_x() + p.get_width() / 2., p.get_height()),\n                ha='center', va='bottom')\nplt.show()\n\n# Sex vs Aneurysm\nplt.figure(figsize=(8,5))\nax = sns.countplot(x=\"PatientSex\", hue=\"Aneurysm Present\", data=train_df, palette=\"Set2\")\nplt.title(\"Sex Distribution by Aneurysm Presence\")\nplt.xlabel(\"Patient Sex\")\nplt.ylabel(\"Number of Patients\")\nfor p in ax.patches:\n    height = p.get_height()\n    ax.annotate(f'{height}', \n                (p.get_x() + p.get_width() / 2., height),\n                ha='center', va='bottom')\nplt.show()\n\n# Age distribution\nplt.figure(figsize=(10,6))\nsns.histplot(data=train_df, x=\"PatientAge\", hue=\"Aneurysm Present\",\n             bins=30, kde=True, palette=\"husl\", element=\"step\", stat=\"count\")\nplt.title(\"Age Distribution by Aneurysm Presence\")\nplt.xlabel(\"Age\")\nplt.ylabel(\"Number of Patients\")\nplt.show()\n\n# --- Textual Summary ---\nprint(\"===== 📊 TEXTUAL SUMMARY =====\\n\")\n\nprint(\"Aneurysm Presence Distribution:\")\nfor label, count, pct in zip(aneurysm_counts.index, aneurysm_counts, aneurysm_pct):\n    label_str = \"Aneurysm Present\" if label == 1 else \"No Aneurysm\"\n    print(f\" - {label_str}: {count} patients ({pct}%)\")\n\nprint(\"\\nSex Distribution by Aneurysm Presence:\")\nprint(sex_counts)\nprint(\"\\nPercentages within each sex:\")\nprint(sex_pct)\n\nprint(\"\\nAge Distribution by Aneurysm Presence:\")\nprint(age_summary)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-20T17:21:23.559544Z","iopub.execute_input":"2025-08-20T17:21:23.559821Z","iopub.status.idle":"2025-08-20T17:21:24.223759Z","shell.execute_reply.started":"2025-08-20T17:21:23.559803Z","shell.execute_reply":"2025-08-20T17:21:24.222902Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"📊 Key Takeaways  \n\n| Observation               | Details                                  |  \n|---------------------------|------------------------------------------|  \n| **Dataset enriched**      | 43% prevalence — higher than real world |  \n| **Sex effect**            | Women more affected (47% vs 34% in men) |  \n| **Age effect**            | Older patients more likely to have aneurysms |  \n\n*Note: Demographic patterns are clinically realistic, but baseline prevalence is artificially inflated.* \n","metadata":{}},{"cell_type":"code","source":"# ============================================== #\n# List of location-specific aneurysm labels      #\n# ============================================== #\n\nlocation_cols = [\n    'Left Infraclinoid Internal Carotid Artery',\n    'Right Infraclinoid Internal Carotid Artery',\n    'Left Supraclinoid Internal Carotid Artery',\n    'Right Supraclinoid Internal Carotid Artery',\n    'Left Middle Cerebral Artery',\n    'Right Middle Cerebral Artery',\n    'Anterior Communicating Artery',\n    'Left Anterior Cerebral Artery',\n    'Right Anterior Cerebral Artery',\n    'Left Posterior Communicating Artery',\n    'Right Posterior Communicating Artery',\n    'Basilar Tip',\n    'Other Posterior Circulation'\n]\n\n# Sum of aneurysms per location grouped by sex\nsex_location_counts = train_df.groupby(\"PatientSex\")[location_cols].sum().T\n\n# Convert counts into prevalence rates (percentage within sex)\nsex_totals = train_df.groupby(\"PatientSex\").size()\nsex_location_pct = sex_location_counts.div(sex_totals, axis=1) * 100\n\n# --- Heatmap of prevalence ---\nplt.figure(figsize=(12,8))\nsns.heatmap(sex_location_pct, annot=True, fmt=\".2f\", cmap=\"coolwarm\")\nplt.title(\"Aneurysm Location Prevalence (%) by Sex\")\nplt.xlabel(\"Sex\")\nplt.ylabel(\"Artery Location\")\nplt.show()\n\n# --- Textual summary ---\nprint(\"===== 📊 TEXTUAL SUMMARY: Aneurysm Location Prevalence by Sex =====\\n\")\nprint(sex_location_pct.round(2))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-20T17:21:53.093486Z","iopub.execute_input":"2025-08-20T17:21:53.093767Z","iopub.status.idle":"2025-08-20T17:21:53.367544Z","shell.execute_reply.started":"2025-08-20T17:21:53.093748Z","shell.execute_reply":"2025-08-20T17:21:53.366877Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"📌 Are those differences are statistically significant or could just be random variation ?","metadata":{}},{"cell_type":"code","source":"\"\"\"\nSTATISTICAL SIGNIFICANCE ANALYSIS: Sex vs. Aneurysm Location\n\nThis cell performs chi-square tests to determine if there are statistically\nsignificant relationships between patient sex and aneurysm location prevalence.\n\nANALYSIS PURPOSE:\n- Tests the null hypothesis that sex and aneurysm location are independent\n- Identifies if certain locations show gender-based predisposition\n- Uses chi-square tests of independence on contingency tables\n\nINTERPRETATION:\n- p-value < 0.05: Statistically significant relationship (reject null hypothesis)\n- p-value >= 0.05: No significant evidence of relationship\n- Chi2 value: Strength of association (higher = stronger relationship)\n\"\"\"\n\nlocation_cols = [\n    'Left Infraclinoid Internal Carotid Artery',\n    'Right Infraclinoid Internal Carotid Artery',\n    'Left Supraclinoid Internal Carotid Artery',\n    'Right Supraclinoid Internal Carotid Artery',\n    'Left Middle Cerebral Artery',\n    'Right Middle Cerebral Artery',\n    'Anterior Communicating Artery',\n    'Left Anterior Cerebral Artery',\n    'Right Anterior Cerebral Artery',\n    'Left Posterior Communicating Artery',\n    'Right Posterior Communicating Artery',\n    'Basilar Tip',\n    'Other Posterior Circulation'\n]\n\nresults = []\n\nfor col in location_cols:\n    # Build contingency table\n    contingency = pd.crosstab(train_df[\"PatientSex\"], train_df[col])\n    \n    # Chi-square test\n    chi2, p, dof, expected = chi2_contingency(contingency)\n    \n    results.append({\n        \"Artery Location\": col,\n        \"Chi2\": chi2,\n        \"p-value\": p,\n        \"Significant (<0.05)\": p < 0.05\n    })\n\n# Convert to DataFrame\nchi_results = pd.DataFrame(results).sort_values(\"p-value\")\n\nprint(\"===== 📊 Chi-Square Test Results by Artery Location =====\\n\")\nprint(chi_results)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-20T17:22:01.13764Z","iopub.execute_input":"2025-08-20T17:22:01.137935Z","iopub.status.idle":"2025-08-20T17:22:01.211271Z","shell.execute_reply.started":"2025-08-20T17:22:01.137915Z","shell.execute_reply":"2025-08-20T17:22:01.210444Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"🧩 Takeaway\n\n**Sex-Specific Risk Distribution**\n- **Female-biased locations**:\n  - Supraclinoid ICA\n  - Middle Cerebral Artery (MCA) \n  - Posterior Communicating Artery (PCom)\n    \n  \n- **Male-biased location**:\n  - Anterior Communicating Artery (ACom) _(slightly)_\n    \n  \n- **Sex-balanced locations** (Posterior circulation):\n  - Basilar Tip\n  - Vertebral arteries\n  - PICA\n  - etc.\n\n### Clinical Correlation\nThis aligns with established clinical literature:\n- **Women** are more prone to ICA and MCA aneurysms (often linked to post-menopausal hormonal changes)\n- **Men** have relatively more ACom aneurysms\n\n<h4 style=\"color:green\">👉 That means including PatientSex and PatientAge in model train isn’t just safe — it’s clinically justified.</h4>","metadata":{}},{"cell_type":"code","source":"\"\"\"\nANALYSIS: Distribution of Imaging Modalities by Aneurysm Label\n\nThis cell analyzes the relationship between imaging modalities (CT, MRI, etc.)\nand the presence of aneurysms in the training dataset.\n\nWHAT IT DOES:\n1. Groups the data by modality and aneurysm label\n2. Calculates counts and percentages for each combination\n3. Creates a bar plot showing the distribution\n4. Annotates bars with both count and percentage values\n\nOUTPUT: Visual comparison of modality usage for positive/negative aneurysm cases\n\"\"\"\n\n# Group by Modality and Aneurysm label, count occurrences\nmodality_label_dist = (\n    train_EDA.groupby([\"Modality_x\", \"Aneurysm Present\"])\n    .size()\n    .reset_index(name=\"count\")\n)\n\n# Calculate percentage within each label group\nmodality_label_dist[\"percent\"] = (\n    modality_label_dist.groupby(\"Aneurysm Present\")[\"count\"]\n    .transform(lambda x: (x / x.sum()) * 100)\n).round(1)\n\n# Plot\nplt.figure(figsize=(10, 6))\nbarplot = sns.barplot(\n    data=modality_label_dist,\n    x=\"Modality_x\",\n    y=\"count\",\n    hue=\"Aneurysm Present\",\n    palette=\"Set2\"\n)\n\n# Add labels with correct count and percentage values\nfor i in range(len(modality_label_dist)):\n    row = modality_label_dist.iloc[i]\n    modality = row[\"Modality_x\"]\n    label = row[\"Aneurysm Present\"]\n    count = row[\"count\"]\n    percent = row[\"percent\"]\n\n    # Find the correct bar\n    for patch in barplot.patches:\n        if patch.get_height() == count and patch.get_x() < barplot.get_xlim()[1]:\n            x = patch.get_x() + patch.get_width() / 2\n            y = patch.get_height()\n            barplot.annotate(\n                f\"{count}\\n({percent:.1f}%)\", \n                (x, y), \n                ha=\"center\", \n                va=\"bottom\", \n                fontsize=9\n            )\n            break\n\nplt.title(\"Distribution of Imaging Modalities by Aneurysm Label\")\nplt.xlabel(\"Imaging Modality\")\nplt.ylabel(\"Number of Exams\")\nplt.xticks(rotation=45)\nplt.legend(title=\"Aneurysm Present\")\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-20T17:22:12.394352Z","iopub.execute_input":"2025-08-20T17:22:12.394637Z","iopub.status.idle":"2025-08-20T17:22:12.602722Z","shell.execute_reply.started":"2025-08-20T17:22:12.394618Z","shell.execute_reply":"2025-08-20T17:22:12.602026Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"💡 Practical Interpretation:\n\n- The predominance of **CTA** and **MRA** reinforces that these modalities are preferred for **initial aneurysm detection**\n  \n- **MRI exams** appear less frequently among positive cases, suggesting a more **complementary role**\n  \n- We can prioritize these modalities when building **initial detection models**  \n  *(e.g., separate models per modality or ensemble approaches)*","metadata":{}},{"cell_type":"markdown","source":"### 🧠 Extracted DICOM Metadata — Purpose and Rationale\n\nIn this step, we parsed the DICOM headers to extract key metadata that can provide clinical or technical context to each imaging series. These fields are not part of the training labels, but they can offer useful insights for analysis and modeling.\n\n✅ Extracted Fields and Their Purpose\n\n| Field Name                  | Purpose                                                                       |\n| --------------------------- | ----------------------------------------------------------------------------- |\n| `SeriesInstanceUID`         | Unique identifier for the image series (used for joins).                      |\n| `Modality`                  | Indicates the imaging type (e.g., CT, MR, MRA) — helps infer diagnostic flow. |\n| `SeriesDescription`         | Free-text description of the series — may contain protocol hints.             |\n| `ProtocolName`              | Imaging acquisition protocol — useful for understanding scan objectives.      |\n| `BodyPartExamined`          | Specifies the anatomical region imaged (usually brain-related here).          |\n| `Manufacturer`              | Vendor of the imaging equipment — useful for model generalization analysis.   |\n| `ManufacturerModelName`     | Specific scanner model used — may reflect resolution differences.             |\n| `StationName`               | Name of the imaging station — could reflect hospital or scanner unit.         |\n| `SoftwareVersions`          | Imaging software version — potential confounder in image format/quality.      |\n| `SliceThickness`            | Thickness of CT/MRI slices — affects spatial resolution.                      |\n| `PixelSpacing`              | Spacing between pixels in mm — important for voxel calibration.               |\n| `KVP`                       | Tube voltage (for CT scans) — technical parameter for image quality.          |\n| `Exposure`                  | Radiation dose estimate — related to image quality vs. patient safety.        |\n| `ImageType`                 | Indicates image derivation (original, reformatted, etc.).                     |\n| `StudyDate` / `SeriesDate`  | May provide temporal clues (e.g., first vs. follow-up scans).                 |\n| `PatientAge` / `PatientSex` | Patient demographic information — can be predictive of aneurysm risk.         |\n| `LossyImageCompression`     | Indicates if the image was compressed — may reduce quality.                   |\n| `BurnedInAnnotation`        | Indicates if text overlays exist on the image — may affect preprocessing.     |\n\n----\n\nThese metadata fields help us:\n\n- Understand image quality and acquisition protocols\n\n- Identify scanner variability, which can impact model generalization\n\n- Possibly reconstruct patient diagnostic flows (e.g., CT → MR)\n\n- Enrich the dataset with clinical context, useful for a multimodal model","metadata":{}},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ============================================\n# EDA: SliceThickness — histogram + boxplot by manufacturer\n# Rule of thumb: thickness > 5 mm → may miss small aneurysms\n# ============================================\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\n\n# -----------------------------\n# 0) Work on a copy; ensure numeric\n# -----------------------------\ndf = train_EDA.copy()\ndf[\"SeriesInstanceUID\"] = df[\"SeriesInstanceUID\"].astype(str)\ndf[\"SliceThickness\"] = pd.to_numeric(df[\"SliceThickness\"], errors=\"coerce\")\n\n# Optional filter: keep only CTA (set to False to use all)\nONLY_CTA = False\nmod_col = \"Modality_x\" if \"Modality_x\" in df.columns else (\"Modality\" if \"Modality\" in df.columns else None)\nif ONLY_CTA and mod_col:\n    df = df[df[mod_col].astype(str).str.upper() == \"CTA\"].copy()\n\n# -----------------------------\n# 1) Create thickness-based quality flag\n# -----------------------------\n# Keep this flag separate from any previous IMG_Low_Quality rule\ndf[\"IMG_Low_Quality_ST\"] = df[\"SliceThickness\"] > 5.0\n\ntotal = df[\"SliceThickness\"].notna().sum()\nn_bad = int(df[\"IMG_Low_Quality_ST\"].sum())\npct_bad = 100.0 * n_bad / max(total, 1)\n\nprint(f\"Total series with SliceThickness available: {total}\")\nprint(f\"Series with SliceThickness > 5 mm: {n_bad} ({pct_bad:.2f}%)\")\n\n# -----------------------------\n# 2) Manufacturer normalization (collapse vendor variants)\n# -----------------------------\ndef normalize_manufacturer(x: str) -> str:\n    if not isinstance(x, str):\n        return \"OTHER\"\n    s = x.strip().upper()\n    if \"SIEMENS\" in s:\n        return \"SIEMENS\"\n    if \"GE\" in s:\n        return \"GE MEDICAL SYSTEMS\"\n    if \"TOSHIBA\" in s:\n        return \"TOSHIBA\"\n    if \"CANON\" in s:\n        return \"CANON\"\n    if \"PHILIPS\" in s:\n        return \"PHILIPS\"\n    return s if s else \"OTHER\"\n\ndf[\"Manufacturer_norm\"] = df[\"Manufacturer\"].apply(normalize_manufacturer)\n\n# -----------------------------\n# 3) Summary table by manufacturer\n# -----------------------------\nsummary = (df.groupby(\"Manufacturer_norm\")\n             .agg(\n                 n=(\"SeriesInstanceUID\",\"nunique\"),\n                 thickness_median=(\"SliceThickness\",\"median\"),\n                 thickness_p75=(\"SliceThickness\", lambda s: np.nanpercentile(s.dropna(), 75) if s.notna().any() else np.nan),\n                 pct_gt_5mm=(\"IMG_Low_Quality_ST\", \"mean\")\n             )\n             .sort_values(\"pct_gt_5mm\", ascending=False))\nsummary[\"pct_gt_5mm\"] = (summary[\"pct_gt_5mm\"] * 100).round(2)\n\nprint(\"\\nSliceThickness summary by Manufacturer (sorted by % > 5 mm):\")\ndisplay(summary)\n\n# -----------------------------\n# 4) Histogram (single axis, matplotlib only)\n# -----------------------------\nplt.figure(figsize=(8,4))\nvals = df[\"SliceThickness\"].dropna().values\nbins = np.linspace(max(0, np.nanmin(vals)), min(10, np.nanmax(vals)), 40)  # clamp to 0–10 mm for readability\nplt.hist(vals, bins=bins)\nplt.axvline(5.0, linestyle=\"--\")  # do not set color per toolbox rules\nplt.title(\"Histogram of SliceThickness (mm)\")\nplt.xlabel(\"SliceThickness (mm)\")\nplt.ylabel(\"Series count\")\nplt.grid(True)\nplt.show()\n\n# -----------------------------\n# 5) Boxplot by manufacturer (top-N vendors with enough data)\n# -----------------------------\nTOP_N = 6\nvc = df[\"Manufacturer_norm\"].value_counts()\ntop_vendors = vc.index[:TOP_N].tolist()\n\ndata = []\nlabels = []\nfor v in top_vendors:\n    arr = df.loc[df[\"Manufacturer_norm\"] == v, \"SliceThickness\"].dropna().values\n    if arr.size >= 5:  # require minimal samples to plot\n        data.append(arr)\n        labels.append(v)\n\nplt.figure(figsize=(10,4))\nplt.boxplot(data, labels=labels, showmeans=True)\nplt.title(\"SliceThickness by Manufacturer (boxplot)\")\nplt.ylabel(\"SliceThickness (mm)\")\nplt.grid(True, axis=\"y\")\nplt.show()\n\n# -----------------------------\n# 6) Attach the new flag back to train_EDA (optional)\n# -----------------------------\ntrain_EDA = train_EDA.merge(\n    df[[\"SeriesInstanceUID\",\"IMG_Low_Quality_ST\",\"Manufacturer_norm\",\"SliceThickness\"]],\n    on=\"SeriesInstanceUID\", how=\"left\"\n)\n\nprint(\"Added columns to train_EDA: IMG_Low_Quality_ST, Manufacturer_norm (and refreshed SliceThickness).\")\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Save with the same name as the DataFrame variable\nout_path = \"/kaggle/working/train_EDA.csv\"  # same name\ntrain_EDA.to_csv(out_path, index=False)\nprint(\"Saved:\", out_path)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Enriquecendo o train.csv com metadados dicom\n\nimport os\nimport pydicom\nimport pandas as pd\nfrom tqdm import tqdm\n\n# Root directory for the DICOM series\nroot_dir = \"/kaggle/input/rsna-intracranial-aneurysm-detection/series\"\n\n# Get only the SeriesInstanceUIDs present in train.csv\ntrain_uids = set(train_df[\"SeriesInstanceUID\"].unique())\n\n# Limit processing to series used in training\nseries_dirs = [uid for uid in os.listdir(root_dir) if uid in train_uids]\n\n# Fields to extract from DICOM metadata\nselected_keywords = [\n    \"SeriesInstanceUID\", \"Modality\", \"SeriesDescription\", \"ProtocolName\",\n    \"BodyPartExamined\", \"Manufacturer\", \"ManufacturerModelName\", \"StationName\",\n    \"SoftwareVersions\", \"SliceThickness\", \"PixelSpacing\", \"KVP\", \"Exposure\",\n    \"ImageType\", \"StudyDate\", \"SeriesDate\", \"PatientAge\", \"PatientSex\",\n    \"LossyImageCompression\", \"BurnedInAnnotation\"\n]\n\ndicom_metadata_list = []\n\n# Loop over DICOM series used in training\nfor series_uid in tqdm(series_dirs):\n    series_path = os.path.join(root_dir, series_uid)\n    dicom_files = [f for f in os.listdir(series_path) if f.endswith(\".dcm\")]\n    if not dicom_files:\n        continue\n\n    dicom_path = os.path.join(series_path, dicom_files[0])\n    try:\n        ds = pydicom.dcmread(dicom_path, stop_before_pixels=True)\n        meta = {\"SeriesInstanceUID\": series_uid}\n\n        for keyword in selected_keywords:\n            if keyword == \"SeriesInstanceUID\":\n                continue\n            val = ds.get(keyword, None)\n            meta[keyword] = str(val) if val not in [None, \"\", [], {}] else None\n\n        dicom_metadata_list.append(meta)\n\n    except Exception:\n        continue\n\n# Convert to DataFrame and drop columns with all missing values\ndf_filtered = pd.DataFrame(dicom_metadata_list)\ndf_filtered = df_filtered.dropna(axis=1, how=\"all\")\n\n# Optionally: merge with train_df immediately\ntrain_EDA = train_df.merge(df_filtered, on=\"SeriesInstanceUID\", how=\"left\")\n\n# Preview\ntrain_EDA.head()\n\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#\n# Preparing a DF to understand the cliical flow(business Process Approach)\n# \n\n# Define the date fields to evaluate\ndate_fields = [\"StudyDate\", \"SeriesDate\"]\n\n# Count non-null values for each date field\nprint(\"🔎 Non-null counts per date field:\")\nprint(train_with_metadata[date_fields].notna().sum())\n\n# Show sample values from each date field\nfor col in date_fields:\n    print(f\"\\n📅 Sample values from {col}:\")\n    print(train_with_metadata[col].dropna().unique()[:10])\n\n# Convert to datetime format for further analysis\nfor col in date_fields:\n    if col in train_with_metadata.columns:\n        train_with_metadata[col + \"_parsed\"] = pd.to_datetime(train_with_metadata[col], errors=\"coerce\", format=\"%Y%m%d\")\n\n# Compare StudyDate and SeriesDate to assess consistency\nif \"StudyDate_parsed\" in train_with_metadata.columns and \"SeriesDate_parsed\" in train_with_metadata.columns:\n    train_with_metadata[\"DateDifferenceDays\"] = (\n        train_with_metadata[\"SeriesDate_parsed\"] - train_with_metadata[\"StudyDate_parsed\"]\n    ).dt.days\n\n    print(\"\\n📊 Distribution of days between StudyDate and SeriesDate:\")\n    print(train_with_metadata[\"DateDifferenceDays\"].value_counts().sort_index())\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Understanding Image Quality from DICOM Metadata\n\n✅ Useful fields for inferring image quality (all available in full_df):\n\n| Field                      | Meaning/Impact |\n|----------------------------|----------------|\n| SliceThickness             | Thinner slices ⇒ higher axial resolution |\n| PixelSpacing               | Smaller values ⇒ better planar resolution |\n| KVP                        | Tube voltage (in CT) ⇒ affects contrast and noise |\n| Exposure                   | Radiation dose ⇒ affects image noise |\n| LossyImageCompression      | If \"01\" ⇒ image was compressed and may have lost quality |\n| BurnedInAnnotation         | If \"YES\" ⇒ may contain overlaid text or labels on the image |\n| ImageType                  | If contains DERIVED or SECONDARY ⇒ may not be the original image |\n\n\n\n------\n\nMRI Image Analysis (Modality = MRA, MRI T1, MRI T2)\n\n🔬 Relevant DICOM Variables:\n\n| Variable                     | Description |\n|------------------------------|-------------|\n| **Modality** + **SeriesDescription**/**ProtocolName** | Defines sequence type (T1, T2, TOF, etc.) |\n| **EchoTime**, **RepetitionTime**, **FlipAngle** (not always available) | Determines tissue contrast characteristics |\n| **PixelSpacing**/**SliceThickness** | Spatial resolution of the image |\n| **ImageType**                | May contain \"M\" (magnitude), \"PHASE\", etc. |\n| **LossyImageCompression**/**BurnedInAnnotation** | Same considerations as CT |\n\n### 📊 Possible Evaluations:\n\n- Sequence type identification (T1/T2/TOF) from description\n- Soft tissue contrast analysis (requires pixel data reading)\n- Verify consistency between sequence type and technical parameters\n\n","metadata":{}},{"cell_type":"code","source":"import seaborn as sns\nimport matplotlib.pyplot as plt\n\n# Agrupa por Modality e Aneurysm label, conta ocorrências\nmodality_label_dist = (\n    train_EDA.groupby([\"Modality_x\", \"Aneurysm Present\"])\n    .size()\n    .reset_index(name=\"count\")\n)\n\n# Calcula o percentual dentro de cada grupo de label\nmodality_label_dist[\"percent\"] = (\n    modality_label_dist.groupby(\"Aneurysm Present\")[\"count\"]\n    .transform(lambda x: (x / x.sum()) * 100)\n).round(1)\n\n# Gráfico\nplt.figure(figsize=(10, 6))\nbarplot = sns.barplot(\n    data=modality_label_dist,\n    x=\"Modality_x\",\n    y=\"count\",\n    hue=\"Aneurysm Present\",\n    palette=\"Set2\"\n)\n\n# Adiciona os rótulos com contagem e percentual corretos\nfor i in range(len(modality_label_dist)):\n    row = modality_label_dist.iloc[i]\n    modality = row[\"Modality_x\"]\n    label = row[\"Aneurysm Present\"]\n    count = row[\"count\"]\n    percent = row[\"percent\"]\n\n    # Encontra a barra correta\n    for patch in barplot.patches:\n        if patch.get_height() == count and patch.get_x() < barplot.get_xlim()[1]:\n            x = patch.get_x() + patch.get_width() / 2\n            y = patch.get_height()\n            barplot.annotate(\n                f\"{count}\\n({percent:.1f}%)\", \n                (x, y), \n                ha=\"center\", \n                va=\"bottom\", \n                fontsize=9\n            )\n            break\n\nplt.title(\"Distribution of Imaging Modalities by Aneurysm Label\")\nplt.xlabel(\"Imaging Modality\")\nplt.ylabel(\"Number of Exams\")\nplt.xticks(rotation=45)\nplt.legend(title=\"Aneurysm Present\")\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-20T12:46:10.224379Z","iopub.execute_input":"2025-08-20T12:46:10.224771Z","iopub.status.idle":"2025-08-20T12:46:10.464808Z","shell.execute_reply.started":"2025-08-20T12:46:10.224746Z","shell.execute_reply":"2025-08-20T12:46:10.464045Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"💡 Practical Interpretation:\n\n- The predominance of **CTA** and **MRA** reinforces that these modalities are preferred for **initial aneurysm detection**\n  \n- **MRI exams** appear less frequently among positive cases, suggesting a more **complementary role**\n  \n- We can prioritize these modalities when building **initial detection models**  \n  *(e.g., separate models per modality or ensemble approaches)*","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n# Lista dos 13 labels arteriais\nartery_labels = [\n    \"Left Infraclinoid Internal Carotid Artery\",\n    \"Right Infraclinoid Internal Carotid Artery\",\n    \"Left Supraclinoid Internal Carotid Artery\",\n    \"Right Supraclinoid Internal Carotid Artery\",\n    \"Left Middle Cerebral Artery\",\n    \"Right Middle Cerebral Artery\",\n    \"Anterior Communicating Artery\",\n    \"Left Anterior Cerebral Artery\",\n    \"Right Anterior Cerebral Artery\",\n    \"Left Posterior Communicating Artery\",\n    \"Right Posterior Communicating Artery\",\n    \"Basilar Tip\",\n    \"Other Posterior Circulation\"\n]\n\n# Transforma para formato longo (1 linha por artéria com aneurisma positivo)\nlong_df = train_EDA.melt(\n    id_vars=[\"SeriesInstanceUID\", \"Modality_x\"],\n    value_vars=artery_labels,\n    var_name=\"Artery\",\n    value_name=\"Aneurysm\"\n)\n\n# Filtra apenas os casos positivos\nlong_df = long_df[long_df[\"Aneurysm\"] == 1]\n\n# Conta quantidade por Artery + Modality\nartery_modality_counts = (\n    long_df.groupby([\"Artery\", \"Modality_x\"])\n    .size()\n    .reset_index(name=\"count\")\n)\n\n# Calcula % por artéria\nartery_modality_counts[\"percent\"] = (\n    artery_modality_counts.groupby(\"Artery\")[\"count\"]\n    .transform(lambda x: 100 * x / x.sum())\n).round(1)\n\n# Gráfico de barras\nplt.figure(figsize=(14, 8))\nbarplot = sns.barplot(\n    data=artery_modality_counts,\n    x=\"Artery\",\n    y=\"count\",\n    hue=\"Modality_x\",\n    palette=\"Set2\"\n)\n\n# Rótulos de % nas barras\nfor container in barplot.containers:\n    for bar in container:\n        height = bar.get_height()\n        if height > 0:\n            x = bar.get_x() + bar.get_width() / 2\n            idx = barplot.patches.index(bar)\n            percent = artery_modality_counts.iloc[idx][\"percent\"]\n            barplot.annotate(f\"{percent:.1f}%\", (x, height), ha=\"center\", va=\"bottom\", fontsize=8)\n\nplt.title(\"Distribution of Imaging Modalities per Aneurysm Artery Label\")\nplt.xlabel(\"Artery Label\")\nplt.ylabel(\"Number of Positive Cases\")\nplt.xticks(rotation=90)\nplt.legend(title=\"Modality\")\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"✅ Confirmation: Modality Trend by Artery\n\n**CTA** (green) dominates in most arteries with positive cases. Key points:\n\n🧠 Main Arteries (highest case counts):\n\n| Artery                          | Dominant Modality | Observations                     |\n|---------------------------------|------------------|----------------------------------|\n| Anterior Communicating Artery   | CTA (64.2%)      | Strong dominance                 |\n| Right Middle Cerebral Artery    | CTA (43%)        | Significant MRA presence (42.6%) |\n| Left Middle Cerebral Artery     | CTA (59.3%)      | Confirms pattern                 |\n| Basilar Tip                     | MRA (55.5%)      | Interesting exception!           |\n\n🧪 Specific Observations:\n\n- **Basilar Tip** and **Right Posterior Communicating Artery** show higher MRA participation\n  - Suggests MRA may be preferred/better for visualizing these regions\n- **MRI T2** and **MRI T1post** appear minimally (<15% in any artery)\n\n📌 Conclusion:\n\n1. Artery distribution follows the general trend seen in previous analysis\n2. **CTA** remains primary diagnostic modality, with **MRA** important for specific arteries\n3. Recommendation options:\n   - Build separate modality-specific models\n   - Test unified model using *Modality* as a feature\n","metadata":{}},{"cell_type":"markdown","source":"### CT Image Analysis (Modality = CTA)\n\n 🔬 Key DICOM Variables:\n\n| Variable                  | Description |\n|---------------------------|-------------|\n| **KVP** (kilovolt peak)   | X-ray beam energy level — affects contrast |\n| **Exposure**              | Related to radiation dose — impacts SNR (signal-to-noise ratio) |\n| **RescaleIntercept** & **RescaleSlope** | Used to convert pixel values to Hounsfield Units (HU) |\n| **PixelSpacing** & **SliceThickness** | Spatial resolution of the image |\n| **LossyImageCompression** | If lossy compression was used, the image may have additional noise |\n\n### 📊 Possible Evaluations:\n\n- HU value distribution (density) in image samples — tissues, vessels, blood, etc.\n- Local SNR/contrast between different image regions (requires loading pixel data)\n- Verify if KVP/Exposure affects visibility/clarity","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport seaborn as sns\nimport matplotlib.pyplot as plt\n\n# Filtrar apenas imagens do tipo CTA\ncta_df = train_EDA[train_EDA[\"Modality_x\"] == \"CTA\"].copy()\n\n# Selecionar os campos de interesse\nquality_fields = [\n    \"SliceThickness\", \"PixelSpacing\", \"KVP\", \"Exposure\",\n    \"LossyImageCompression\", \"BurnedInAnnotation\",\n    \"ImageType\", \"Manufacturer\", \"SoftwareVersions\"\n]\n\n# Converter campos numéricos que ainda estejam como string\ncta_df[\"SliceThickness\"] = pd.to_numeric(cta_df[\"SliceThickness\"], errors=\"coerce\")\ncta_df[\"KVP\"] = pd.to_numeric(cta_df[\"KVP\"], errors=\"coerce\")\ncta_df[\"Exposure\"] = pd.to_numeric(cta_df[\"Exposure\"], errors=\"coerce\")\ncta_df[\"PixelSpacingX\"] = cta_df[\"PixelSpacing\"].apply(lambda x: float(str(x).strip(\"[]\").split(\",\")[0]) if pd.notnull(x) else None)\n\n# Mostrar estatísticas descritivas dos campos quantitativos\nprint(\"📊 Quantitative fields summary:\")\nprint(cta_df[[\"SliceThickness\", \"PixelSpacingX\", \"KVP\", \"Exposure\"]].describe())\n\n# Frequência dos campos categóricos\nfor col in [\"LossyImageCompression\", \"BurnedInAnnotation\", \"ImageType\", \"Manufacturer\"]:\n    print(f\"\\n🔸 Value counts for {col}:\")\n    print(cta_df[col].value_counts(dropna=False))\n\n# Visualizações\nplt.figure(figsize=(14, 4))\nfor i, field in enumerate([\"SliceThickness\", \"PixelSpacingX\", \"KVP\", \"Exposure\"]):\n    plt.subplot(1, 4, i+1)\n    sns.histplot(cta_df[field], kde=True, bins=20)\n    plt.title(field)\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## DICOM Metadata Analysis & Preprocessing Recommendations\n\n### 📏 1. SliceThickness\n**Mean:** 1.23 mm\n**Range:** From 0.47 mm up to 10.2 mm (!)\n\n**Observation:**\n*   Images with a slice thickness **≤ 1 mm** are excellent for detecting small structures like aneurysms.\n*   However, larger thicknesses (**≥ 5 mm**) dilute details and can harm the performance of 3D-based models.\n\n**✅ Recommendation:**\n*   Consider filtering or standardizing exams with thickness outside an ideal range (e.g., > 3 mm).\n*   Normalize the data by resampling everything to a standard 1 mm using `SimpleITK` or `torchio`.\n\n---\n\n### 🖼️ 2. PixelSpacingX (In-plane Axial Resolution)\n**Mean:** 0.47 mm — excellent resolution.\n**Range:** From 0.22 to 0.91 mm — generally quite good.\n\n**✅ Recommendation:**\n*   Resample to an **isotropic resolution** (e.g., 0.5 x 0.5 x 0.5 mm³) for 3D model.\n*   This standardization helps the model generalize better.\n\n---\n\n### ⚡ 3. KVP (X-ray Beam Energy)\n**Mean:** 111.6 kV (typical).\n**Range:** 70–150 kV.\n\n**📌 Importance:**\n*   Higher kV values provide better penetration in patients with higher body mass.\n*   However, differences in kV can change vessel contrast, which may affect AI performance.\n\n**✅ Recommendation:**\n*   This could be useful as an auxiliary feature (metadata) for the model.\n\n---\n\n### ☢️ 4. Exposure (mAs)\n**Mean:** 137.6 mAs, but with a minimum of 0 (!) → likely low-dose images or a tag error.\n*   Wide distribution → models may struggle to learn consistent contrast patterns.\n\n**✅ Recommendation:**\n*   Let's Filter or treat `Exposure = 0` as a missing value.\n*   This variable could be useful in multi-modal models.\n\n---\n\n### 🔧 5. Lossy Image Compression\n*   ~10% of images indicate lossy compression (`'00'`, `'01'`) — this can degrade subtle details.\n\n**✅ Recommendation:**\n*   If possible, avoid using images with lossy compression in your main training/validation sets.\n*   Alternatively, use more aggressive data augmentation on these cases.\n\n---\n\n### 🏷️ 6. BurnedInAnnotation\n*   The majority do **not** have burned-in annotations (`None`), which is great.\n*   ✅ No significant negative impact detected, but remain vigilant for overlaid text.\n\n---\n\n### 🧬 7. Image Types (ImageType)\nConsiderable variety of types:\n*   `'ORIGINAL'`, `'PRIMARY'`, `'AXIAL'` — ideal (raw axial image).\n*   `'DERIVED'`, `'MPR'`, `'SUBTRACTION'`, `'MIP'` — are derived (potentially interpolated or post-processed).\n\n**⚠️ Derived images may lose real anatomical details.**\n\n**✅ Recommendation:**\n*   Prioritize `'ORIGINAL'`, `'PRIMARY'` for main training.\n*   You can test `'DERIVED'` separately or as an experiment.\n\n---\n\n### 🏭 8. Manufacturer\nDiversity of manufacturers: GE, Siemens, Toshiba, Philips, Canon, etc.\n*   📌 This reflects realistic heterogeneity but also differences in image quality and contrast.\n\n**✅ Recommendation:**\n*   Include `Manufacturer` as a metadata feature.\n*   Consider using domain adaptation or `GroupKFold` stratified by manufacturer for better generalization.\n\n---\n\n## ✅ General Conclusion & Recommendations\n\nThe images have globally good technical quality but with significant heterogeneity in:\n*   Slice thickness\n*   Lossy compression\n*   Image types\n*   Equipment\n\n**🛠️ Suggested Best Practices:**\n1.  Normalize voxel spacing (e.g., to 0.5 mm³ or 1 mm³).\n2.  Filter or handle cases with `SliceThickness > 3 mm` or `Exposure = 0`.\n3.  Consider `ImageType`, `Manufacturer`, and `KVP` as auxiliary features in your model.\n```","metadata":{}},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#\n# Count and analyze low-quality CTA images by ImageType\n#\n\nimport pandas as pd\n\n# Filter only CTA modality exams\ncta_df = train_EDA[train_EDA[\"Modality_x\"] == \"CTA\"].copy()\n\n# Convert SliceThickness and Exposure to numeric (in case they are strings)\ncta_df[\"SliceThickness\"] = pd.to_numeric(cta_df[\"SliceThickness\"], errors=\"coerce\")\ncta_df[\"Exposure\"] = pd.to_numeric(cta_df[\"Exposure\"], errors=\"coerce\")\n\n# Define low-quality criteria\nlow_quality_mask = (cta_df[\"SliceThickness\"] > 5) | (cta_df[\"Exposure\"] < 10)\nlow_quality_cta = cta_df[low_quality_mask]\n\n# Count and percentage of low-quality images\ntotal_cta = len(cta_df)\ntotal_low_quality = len(low_quality_cta)\npercent_low_quality = (total_low_quality / total_cta) * 100\n\n# 📊 Summary\nprint(f\"🧪 Total CTA exams: {total_cta}\")\nprint(f\"⚠️ Low-quality CTA exams: {total_low_quality} ({percent_low_quality:.2f}%)\")\n\n# 🔍 View some examples of low-quality images\ndisplay(\n    low_quality_cta[[\"SeriesInstanceUID\", \"SliceThickness\", \"Exposure\", \"ImageType\", \"Aneurysm Present\"]]\n    .sort_values(by=\"SliceThickness\", ascending=False)\n    .head(10)\n)\n\n# 📊 Distribution of low-quality images by ImageType\n# Convert ImageType from string to tuple if needed\nimport ast\nlow_quality_cta[\"ImageType\"] = low_quality_cta[\"ImageType\"].apply(lambda x: tuple(ast.literal_eval(x)) if isinstance(x, str) else x)\n\n# Group by ImageType and count\nimage_type_counts = low_quality_cta.groupby(\"ImageType\").size().sort_values(ascending=False)\nprint(\"\\n📊 Low-quality images by ImageType:\")\nprint(image_type_counts)\n\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## ✅ Current Criteria for Defining Low-Quality Images\n\nAn image is considered low-quality if **any one** of the following two criteria is met:\n\n---\n\n### 🔴 1. SliceThickness > 5 mm\n\n*   **Reason:** Very thick slices reduce spatial resolution and make it difficult to detect small aneurysms.\n*   **Ideal Standard:** For cerebral CT angiography, the ideal slice thickness is **≤ 1 mm**.\n*   **Common Source:** Slices > 5 mm often come from reconstructions that were not optimized for vascular detection.\n\n---\n\n### 🔴 2. Exposure < 10 mAs\n\n*   **Reason:** Exams with very low exposure can have low signal and high noise, impairing contrast and sharpness.\n*   **Common Scenarios:** This often occurs in fragile patients, children, or fast low-dose protocols.\n*   **Note:** Extreme values like `Exposure = 0` may indicate a tag error or an unreliable image.\n\n---\n\n","metadata":{}},{"cell_type":"code","source":"#\n# Let's create a column to mark the low qiality images\n#\n\n# Garantir que os campos estão em formato numérico\ntrain_EDA[\"SliceThickness\"] = pd.to_numeric(train_EDA[\"SliceThickness\"], errors=\"coerce\")\ntrain_EDA[\"Exposure\"] = pd.to_numeric(train_EDA[\"Exposure\"], errors=\"coerce\")\n\n# Critério de baixa qualidade\ntrain_EDA[\"IMG_Low_Quality\"] = (train_EDA[\"SliceThickness\"] > 5) | (train_EDA[\"Exposure\"] < 10)\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Image Quality Analysis by Manufacturer and Protocol","metadata":{}},{"cell_type":"code","source":"# Group by Manufacturer and SeriesDescription to analyze low-quality image distribution\nlow_quality_summary = (\n    train_EDA.groupby([\"Manufacturer\", \"SeriesDescription\", \"IMG_Low_Quality\"])\n    .size()\n    .unstack(fill_value=0)\n    .reset_index()\n)\n\n# Rename columns for clarity\nlow_quality_summary.columns.name = None\nlow_quality_summary = low_quality_summary.rename(columns={False: \"High_Quality\", True: \"Low_Quality\"})\n\n# Add total and percentage columns\nlow_quality_summary[\"Total\"] = low_quality_summary[\"High_Quality\"] + low_quality_summary[\"Low_Quality\"]\nlow_quality_summary[\"Percent_Low_Quality\"] = 100 * low_quality_summary[\"Low_Quality\"] / low_quality_summary[\"Total\"]\n\n# Sort by highest percentage of low-quality images\nlow_quality_summary = low_quality_summary.sort_values(\"Percent_Low_Quality\", ascending=False)\n\n# Display the most problematic combinations\nprint(\"📊 Low-quality image distribution by Manufacturer and SeriesDescription:\")\ndisplay(low_quality_summary.head(20))\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 🧠 Conclusion: Image Quality Analysis by Manufacturer and Protocol\n\nOur analysis reveals that approximately **25% of the CTA images** in the dataset are considered **low quality**, based on the following criteria:\n- **SliceThickness > 5 mm**\n- **Exposure < 10 mAs**\n\nKey findings:\n\n- 💡 **All low-quality images are concentrated in specific protocols and originate almost exclusively from `GE MEDICAL SYSTEMS`.**\n- 📌 Protocols such as `ANGIO`, `VOL ANGIO 0.625`, `AX THIN ARTERIAL`, and `ARTERIAL` account for the majority of the low-quality images, all of which show **100% low-quality rates**.\n- 🧬 Other manufacturers (Siemens, Philips, Toshiba, Canon) contribute virtually no low-quality cases, suggesting consistent imaging protocols and better scanner configurations.\n\n### 🔍 Implications for Model Training:\n- Removing these images would eliminate **a clinically relevant portion** of the dataset.\n- However, using them as-is **without any mitigation** could degrade model performance or introduce bias.\n\n### ✅ Recommendations:\n- Include `IMG_Low_Quality`, `Manufacturer`, and `SeriesDescription` as auxiliary features.\n- Apply **preprocessing normalization** (e.g., resampling to 1 mm, contrast enhancement) to low-quality images.\n- Use **data augmentation** to simulate similar conditions on high-quality images.\n- Consider **curriculum learning** or **fine-tuning stages** where the model is first trained on clean data and later adapted to noisy inputs.\n- Implement **stratified validation** by quality and manufacturer to ensure robust generalization.\n\nThese steps will help the model remain **robust in real-world clinical scenarios**, especially when dealing with heterogeneous scanner outputs and emergency protocols.\n","metadata":{}},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\n### Analysis via Hounsfield Units (HU)\n\n","metadata":{}},{"cell_type":"code","source":"import os\nimport pydicom\nimport numpy as np\nimport matplotlib.pyplot as plt\n\n# Root directory\nroot_dir = \"/kaggle/input/rsna-intracranial-aneurysm-detection/series\"\n\n# Number of samples to load\nN_SAMPLES = 1\n\n# Collect pixel data from N_SAMPLES series\nall_hu = []\n\n# Walk through the DICOM series\nfor patient_id in os.listdir(root_dir)[:N_SAMPLES]:\n    patient_dir = os.path.join(root_dir, patient_id)\n    if not os.path.isdir(patient_dir):\n        continue\n\n    # Load all DICOM slices in the series\n    slices = []\n    for file in os.listdir(patient_dir):\n        path = os.path.join(patient_dir, file)\n        dicom = pydicom.dcmread(path)\n        img = dicom.pixel_array.astype(np.int16)\n\n        # Convert to HU using DICOM metadata\n        intercept = dicom.RescaleIntercept if \"RescaleIntercept\" in dicom else 0\n        slope = dicom.RescaleSlope if \"RescaleSlope\" in dicom else 1\n        img = img * slope + intercept\n\n        slices.append(img)\n\n    if slices:\n        series_hu = np.stack(slices)\n        all_hu.append(series_hu)\n\n# Combine and flatten all HU values\nall_hu = np.concatenate([vol.ravel() for vol in all_hu])\n\n# Plot histogram\nplt.figure(figsize=(12, 6))\nplt.hist(all_hu, bins=500, range=(-1024, 2048), color='navy')\nplt.title(\"Histogram of Hounsfield Units (HU) from Sample Series\")\nplt.xlabel(\"HU Value\")\nplt.ylabel(\"Pixel Count\")\nplt.grid(True)\nplt.show()\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import os\nimport ast\nimport gc\nimport numpy as np\nimport pandas as pd\nimport pydicom\nimport traceback\nfrom collections import defaultdict\nimport matplotlib.pyplot as plt\n\n# -----------------------------\n# Config\n# -----------------------------\nroot_dir = \"/kaggle/input/rsna-intracranial-aneurysm-detection/series\"\n\n# Fixed HU histogram settings (do NOT keep entire volumes in memory)\nHU_MIN, HU_MAX = -1024, 2048\nN_BINS = 512\nBINS = np.linspace(HU_MIN, HU_MAX, N_BINS + 1)\nBIN_CENTERS = (BINS[1:] + BINS[:-1]) / 2\n\n# Optional quick filter: process only CTA in metadata table (recommended)\n# Expecting train_EDA has SeriesInstanceUID and Modality_x (or Modality)\nmodality_col = \"Modality_x\" if \"Modality_x\" in train_EDA.columns else (\"Modality\" if \"Modality\" in train_EDA.columns else None)\nif modality_col:\n    allowed_cta = set(train_EDA.loc[train_EDA[modality_col] == \"CTA\", \"SeriesInstanceUID\"].astype(str))\nelse:\n    allowed_cta = None  # process all if modality column not present\n\n# Ensure numeric types for quality criteria\ntrain_EDA[\"SliceThickness\"] = pd.to_numeric(train_EDA.get(\"SliceThickness\"), errors=\"coerce\")\ntrain_EDA[\"Exposure\"] = pd.to_numeric(train_EDA.get(\"Exposure\"), errors=\"coerce\")\n\n# Create (or reuse) low-quality flag\nif \"IMG_Low_Quality\" not in train_EDA.columns:\n    train_EDA[\"IMG_Low_Quality\"] = (train_EDA[\"SliceThickness\"] > 5) | (train_EDA[\"Exposure\"] < 10)\n\n# Keep only columns we need for merges\nkeep_cols = [\"SeriesInstanceUID\", \"Aneurysm Present\", \"IMG_Low_Quality\", \"Manufacturer\", \"SeriesDescription\", \"ImageType\", \"SliceThickness\", \"Exposure\", \"KVP\"]\nmeta_lookup = train_EDA[keep_cols].drop_duplicates(\"SeriesInstanceUID\").copy()\nmeta_lookup[\"SeriesInstanceUID\"] = meta_lookup[\"SeriesInstanceUID\"].astype(str)\n\n# -----------------------------\n# Helpers\n# -----------------------------\ndef safe_get(ds, key, default=None):\n    \"\"\"Safely fetch a DICOM attribute (tag) with fallback.\"\"\"\n    try:\n        return getattr(ds, key)\n    except Exception:\n        return default\n\ndef parse_imagetype(val):\n    \"\"\"Normalize DICOM ImageType to a tuple for readability.\"\"\"\n    if val is None:\n        return None\n    if isinstance(val, (list, tuple)):\n        return tuple(val)\n    if isinstance(val, str):\n        try:\n            got = ast.literal_eval(val)\n            if isinstance(got, (list, tuple)):\n                return tuple(got)\n        except Exception:\n            pass\n        # Fallback for flattened strings\n        if \"\\\\\" in val:\n            return tuple([x.strip() for x in val.split(\"\\\\\")])\n        if \",\" in val:\n            return tuple([x.strip() for x in val.split(\",\")])\n        return (val,)\n    return (str(val),)\n\ndef iter_series_dirs(root):\n    \"\"\"Yield series directories under root (one-level).\"\"\"\n    for d in os.listdir(root):\n        p = os.path.join(root, d)\n        if os.path.isdir(p):\n            yield d, p\n\ndef stream_histogram_for_series(series_dir):\n    \"\"\"\n    Stream through all DICOM files in a series directory and return:\n      - aggregated HU histogram (counts vector)\n      - lightweight per-series HU stats (mean, std, percentiles) computed online\n    This function NEVER stacks slices -> memory safe.\n    \"\"\"\n    counts = np.zeros(N_BINS, dtype=np.int64)\n    # Running moments for mean/std (Welford)\n    n = 0\n    mean = 0.0\n    M2 = 0.0\n    # Reservoir of samples for percentile estimation (subsample to keep memory low)\n    rng = np.random.default_rng(123)\n    reservoir = []\n    target_reservoir = 200000  # ~200k voxels per series cap\n\n    # Grab minimal DICOM once for meta sanity (optional)\n    first_meta = {}\n\n    for fname in os.listdir(series_dir):\n        if not fname.lower().endswith(\".dcm\"):\n            continue\n        fpath = os.path.join(series_dir, fname)\n        try:\n            ds = pydicom.dcmread(fpath)\n            if not first_meta:\n                first_meta = dict(\n                    Manufacturer=str(safe_get(ds, \"Manufacturer\", \"\")),\n                    SeriesDescription=str(safe_get(ds, \"SeriesDescription\", \"\")),\n                    ImageType=parse_imagetype(safe_get(ds, \"ImageType\", None)),\n                    SliceThickness=pd.to_numeric(safe_get(ds, \"SliceThickness\", np.nan), errors=\"coerce\"),\n                    Exposure=pd.to_numeric(safe_get(ds, \"Exposure\", np.nan), errors=\"coerce\"),\n                    KVP=pd.to_numeric(safe_get(ds, \"KVP\", np.nan), errors=\"coerce\"),\n                )\n\n            arr = ds.pixel_array.astype(np.int16)\n            slope = float(safe_get(ds, \"RescaleSlope\", 1.0))\n            intercept = float(safe_get(ds, \"RescaleIntercept\", 0.0))\n            hu = arr * slope + intercept\n            # Clip HU for stable histograms\n            np.clip(hu, HU_MIN, HU_MAX, out=hu)\n\n            # Update histogram in streaming fashion\n            h, _ = np.histogram(hu, bins=BINS)\n            counts += h\n\n            # Update running mean/std (Welford)\n            flat = hu.ravel()\n            n_new = flat.size\n            if n_new == 0:\n                continue\n            n_old = n\n            n += n_new\n            delta = flat.mean() - mean\n            mean += delta * (n_new / n)\n            M2 += flat.var() * n_new + (delta**2) * (n_old * n_new / n)\n\n            # Subsample for percentiles\n            # Keep at most target_reservoir samples uniformly at random\n            need = max(0, target_reservoir - len(reservoir))\n            if need > 0:\n                take = min(need, flat.size)\n                idx = rng.choice(flat.size, size=take, replace=False)\n                reservoir.append(flat[idx])\n            else:\n                # Occasionally replace a portion to remain uniform-ish\n                if rng.random() < 0.02:\n                    idx = rng.choice(flat.size, size=500, replace=False)\n                    replace_idx = rng.choice(len(reservoir), size=500, replace=False)\n                    for i, r_i in enumerate(replace_idx):\n                        reservoir[r_i] = flat[idx[i]]\n\n            # Free slice ASAP\n            del ds, arr, hu, flat, h\n        except Exception:\n            # Keep going even if one file is corrupt\n            traceback.print_exc()\n            continue\n\n    if n == 0:\n        return None, None, None\n\n    # Finalize stats\n    var = M2 / max(1, (n - 1))\n    std = np.sqrt(max(var, 0.0))\n    if reservoir:\n        reservoir = np.concatenate(reservoir)\n        p10, p50, p90 = np.percentile(reservoir, [10, 50, 90])\n    else:\n        p10 = p50 = p90 = np.nan\n\n    stats = dict(hu_mean=float(mean), hu_std=float(std), hu_p10=float(p10), hu_p50=float(p50), hu_p90=float(p90))\n    return counts, stats, first_meta\n\n# -----------------------------\n# Streaming over ALL series\n# -----------------------------\n# Global aggregators (histograms)\nhist_by_label = {0: np.zeros(N_BINS, dtype=np.int64), 1: np.zeros(N_BINS, dtype=np.int64)}\nhist_by_quality = {False: np.zeros(N_BINS, dtype=np.int64), True: np.zeros(N_BINS, dtype=np.int64)}\nhist_by_label_quality = defaultdict(lambda: np.zeros(N_BINS, dtype=np.int64))\n\n# Per-series summary for later analysis\nseries_summaries = []\n\nn_series = 0\nfor sid, sdir in iter_series_dirs(root_dir):\n    # Optional: skip non-CTA if we know modality\n    if allowed_cta is not None and sid not in allowed_cta:\n        continue\n\n    counts, stats, first_meta = stream_histogram_for_series(sdir)\n    if counts is None:\n        continue\n\n    # Merge with train_EDA for label and quality\n    row = meta_lookup.loc[meta_lookup[\"SeriesInstanceUID\"] == sid]\n    if row.empty:\n        # If not in the table, still keep stats with unknown label/quality\n        label = np.nan\n        lowq = np.nan\n        manufacturer = first_meta.get(\"Manufacturer\", \"\")\n        series_desc = first_meta.get(\"SeriesDescription\", \"\")\n        img_type = first_meta.get(\"ImageType\", None)\n        slice_thk = first_meta.get(\"SliceThickness\", np.nan)\n        exposure = first_meta.get(\"Exposure\", np.nan)\n        kvp = first_meta.get(\"KVP\", np.nan)\n    else:\n        r = row.iloc[0]\n        label = r.get(\"Aneurysm Present\", np.nan)\n        lowq = bool(r.get(\"IMG_Low_Quality\", False))\n        manufacturer = r.get(\"Manufacturer\", first_meta.get(\"Manufacturer\", \"\"))\n        series_desc = r.get(\"SeriesDescription\", first_meta.get(\"SeriesDescription\", \"\"))\n        img_type = r.get(\"ImageType\", first_meta.get(\"ImageType\", None))\n        slice_thk = r.get(\"SliceThickness\", first_meta.get(\"SliceThickness\", np.nan))\n        exposure = r.get(\"Exposure\", first_meta.get(\"Exposure\", np.nan))\n        kvp = r.get(\"KVP\", first_meta.get(\"KVP\", np.nan))\n\n    # Update global histograms guardedly\n    if label in (0, 1):\n        hist_by_label[int(label)] += counts\n        if lowq in (True, False):\n            hist_by_label_quality[(int(label), bool(lowq))] += counts\n    if lowq in (True, False):\n        hist_by_quality[bool(lowq)] += counts\n\n    # Save per-series summary row\n    series_summaries.append({\n        \"SeriesInstanceUID\": sid,\n        \"Aneurysm Present\": label,\n        \"IMG_Low_Quality\": lowq,\n        \"Manufacturer\": manufacturer,\n        \"SeriesDescription\": series_desc,\n        \"ImageType\": img_type,\n        \"SliceThickness\": slice_thk,\n        \"Exposure\": exposure,\n        \"KVP\": kvp,\n        **stats\n    })\n\n    n_series += 1\n    if n_series % 50 == 0:\n        print(f\"Processed {n_series} series...\")\n        gc.collect()\n\nprint(f\"✅ Done. Total processed series: {n_series}\")\n\nseries_df = pd.DataFrame(series_summaries)\n\n# -----------------------------\n# Plots (aggregated, still light on memory)\n# -----------------------------\n# 1) By label\nplt.figure(figsize=(12,5))\nplt.plot(BIN_CENTERS, hist_by_label[0], label=\"Label 0 (No Aneurysm)\")\nplt.plot(BIN_CENTERS, hist_by_label[1], label=\"Label 1 (Aneurysm)\")\nplt.title(\"HU Histogram (streaming, ALL series) by Label\")\nplt.xlabel(\"HU\")\nplt.ylabel(\"Voxel Count\")\nplt.legend(); plt.grid(True); plt.show()\n\n# 2) By quality flag\nplt.figure(figsize=(12,5))\nplt.plot(BIN_CENTERS, hist_by_quality[False], label=\"High Quality (IMG_Low_Quality=False)\")\nplt.plot(BIN_CENTERS, hist_by_quality[True],  label=\"Low Quality (IMG_Low_Quality=True)\")\nplt.title(\"HU Histogram (streaming, ALL series) by Quality\")\nplt.xlabel(\"HU\")\nplt.ylabel(\"Voxel Count\")\nplt.legend(); plt.grid(True); plt.show()\n\n# 3) By (label × quality)\nfor k, h in hist_by_label_quality.items():\n    lbl, q = k\n    plt.plot(BIN_CENTERS, h, label=f\"Label={lbl}, LowQ={q}\")\nplt.title(\"HU Histogram (streaming) by Label × Quality\")\nplt.xlabel(\"HU\"); plt.ylabel(\"Voxel Count\")\nplt.legend(); plt.grid(True); plt.show()\n\n# -----------------------------\n# Useful tables for the notebook\n# -----------------------------\nprint(\"Per-series HU summary (head):\")\ndisplay(series_df.head())\n\nprint(\"Quality rate by manufacturer (from processed series):\")\ntbl = (series_df.groupby([\"Manufacturer\",\"IMG_Low_Quality\"])\n              .size().unstack(fill_value=0))\nif True in tbl.columns and False in tbl.columns:\n    tbl[\"Percent_Low_Quality\"] = 100 * tbl[True] / (tbl[True] + tbl[False])\ndisplay(tbl.sort_values(by=tbl.columns[-1], ascending=False))\n\nprint(\"Correlation between HU stats and quality (sanity check):\")\ndisplay(series_df[[\"hu_mean\",\"hu_std\",\"hu_p10\",\"hu_p50\",\"hu_p90\",\"IMG_Low_Quality\"]]\n        .groupby(\"IMG_Low_Quality\").describe().T)\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}