{"metadata":{"kernelspec":{"name":"ir","display_name":"R","language":"R"},"language_info":{"name":"R","codemirror_mode":"r","pygments_lexer":"r","mimetype":"text/x-r-source","file_extension":".r","version":"4.4.0"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Competition Overview\n\nThe knee is the most commonly injured and imaged joint in the body, however, \nthe ways in which radiologists interpret MRI imaging scans differ.  The \ngoal of the competition is to develop ML models that detect clinically\nimportant knee abnormalities, which could provide useful decision support\ntools to radiologists in practice.\n\n## Competition Training Data & Prior Inventory Findings\n\nThe competition provides the following training data that can be used to create\na model that predicts anatomical abnormalities from imaging series alone:\n\n|File or Folder|Description|Notes|\n|--------------|-----------|-----|\n|`train.csv`|Contains one row per study with the following labels: `StudyInstanceUID`(Study Unique ID), `Report`, `ACL`, `MCL`, `Medial Meniscus`, `Lateral Meniscus`, `Medial OA`, `Lateral_OA`, `PF OA`,`Effusion`,`Synovitis`, `Baker's`, `Contusion`, `Fracture`|All are binary except `StudyInstanceUID`, `PatientSex`, and `Report`. `Report` is the clinicians' report and includes reports in multiple languages. `PatientSex` is described in the docs, but missing|\n|`train_series.csv` | Contains one row per training series (each study containing several series) with the following labels: `StudyInstanceUID` (Study Unique ID), `SeriesInstanceUID` (Series Unique ID), `Fluid_Sensitive`, `Fat_Suppression`, `Anotomical Plane`| `Fluid Sensitive` and `Fat Suppression` are binary and indicate different MRI weightings. Values for `Anatomical Plan` are 'Sagittal', 'Coronal', or 'Axial'|\n|`train_series/`|Folder containing MRI series associated with the training studies, with files in dcm (DICOM) format). Files are organized as  `train_series/<StudyInstanceUID>/<SeriesInstanceUID>/<SOPInstanceUID>.dcm`| Series may have 20–45 slices (median 30), with a long tail out to a few hundred|\n\nOur previous inventory of the training data data in the `RSNA Data Inventory Notebook` \nevaluated the integrity of the provided data, with the following findings:\n- 58 records 'gold'-labeled for anatomical abnormalities  of 4407 in train.csv; \n  all record have `Report`s `StudyInstanceUID`s \n- `StudyInstanceUID`s and `SeriesInstanceUIDs` are linked with file folders\n  as expected\n- The shape of data within each csv is as expected, and captured in the table above,\n  with the noted missingness of the PatientSex\n\n## EDA Objectives:\nThe immediate next step is to label the remain studies based on the information \nin the reports, which the preliminary EDA can support by:\n- Helping understand the frequency and correlations of different types of knee injuries\n  in the labeled set, as well as how diagnoses relate to types/number of imaging studies\n- Evaluating the data in `Report` column - quantitative features and differences between\n  labeled and unlabeled reports\n- Preliminary evaluation of the .dcm files (primarily number, orientations, and weightings);\n  follow-up deep-dive on the files themselvs when I understand the format better","metadata":{}},{"cell_type":"code","source":"install.packages('cld3', lib = '/usr/local/lib/R/site-library')\ninstall.packages('oro.dicom', lib = '/usr/local/lib/R/site-library')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-21T19:24:06.493514Z","iopub.execute_input":"2026-08-21T19:24:06.495197Z","iopub.status.idle":"2026-08-21T19:25:14.613696Z","shell.execute_reply":"2026-08-21T19:25:14.610911Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"library(dplyr)\nlibrary(tidyr)\nlibrary(purrr)\nlibrary(ggplot2)\nlibrary(forcats)\nlibrary(ggcorrplot)\nlibrary(cld3)\nlibrary(stringr)\nlibrary(gt)\nlibrary(oro.dicom)\nlibrary(IRdisplay)\n\noptions(verbose = FALSE)\noptions(repr.plot.width = 12, repr.plot.height = 8)\n\nDATA_LOC <- '/kaggle/input/competitions/rsna-knee-abnormality-detection'\nTRAIN_FILES_LOC <- paste(DATA_LOC, 'train_series', sep = \"/\")\n\ntrain_df <- read.csv(paste(DATA_LOC, 'train.csv', sep = \"/\"))\ncolnames(train_df) <- c(\n    'StudyInstanceUID', 'Report', 'ACL', 'MCL', 'Medial_Meniscus', \n    'Lateral_Meniscus', 'Medial_OA', 'Lateral_OA', 'PF_OA', 'Effusion', 'Synovitis', \n    'Baker', 'Contusion', 'Fracture'\n)\ncat(\"Read in 'train.csv':\", nrow(train_df), \"rows by\", ncol(train_df), \"columns. \\n\")\n\ntrain_series_df <- read.csv(paste(DATA_LOC, 'train_series.csv', sep = \"/\"))\ncolnames(train_series_df) <- c(\n    'StudyInstanceUID', 'SeriesInstanceUID', 'Fluid_Sensitive', 'Fat_Suppression',\n    'Anatomical_Plane'\n)\ncat(\n    \"Read in 'train_series.csv':\", nrow(train_series_df), \"rows by\", \n    ncol(train_series_df), \"columns. \\n\"\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-21T19:25:14.617489Z","iopub.execute_input":"2026-08-21T19:25:14.619762Z","iopub.status.idle":"2026-08-21T19:25:14.968903Z","shell.execute_reply":"2026-08-21T19:25:14.966853Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"## General Data Tidying\n\nfindings_df <- train_df |>\n    select(!Report) |>\n    pivot_longer(!StudyInstanceUID, names_to = \"Finding\", values_to = \"Evaluation\")\n\ngold_labeled_df <- findings_df |>\n    filter(!is.na(Evaluation))\n\nreports_df <- train_df |>\n    select(c(StudyInstanceUID, Report)) |>\n    mutate(\n        Gold_Label = case_when(\n            StudyInstanceUID %in% gold_labeled_df$StudyInstanceUID ~ TRUE,\n            !(StudyInstanceUID %in% gold_labeled_df$StudyInstanceUID) ~ FALSE\n        ),\n        Language = detect_language(Report),\n        NumChars = str_length(Report)\n    )","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-21T19:25:14.972206Z","iopub.execute_input":"2026-08-21T19:25:14.974008Z","iopub.status.idle":"2026-08-21T19:25:17.855161Z","shell.execute_reply":"2026-08-21T19:25:17.85273Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Knee Injuries in the 'Gold Labeled' Set\n- How many abnormalities are present for each patient?\n- What are the most common abnormalities?\n- Are any abnormalities correlated?","metadata":{}},{"cell_type":"code","source":"pt_findings_df <- gold_labeled_df |> \n    summarize(\n        num = as.integer((sum(Evaluation))),\n        .by = StudyInstanceUID\n    )\n\npt_findings_ss <- pt_findings_df |>\n    summarize(\n        avg = mean(num),\n        median = median(num),\n        Q1 = quantile(num)[2],\n        Q4 = quantile(num)[4]\n    )\n\npt_findings_df |> ggplot(aes(x = num)) +\n    geom_histogram(binwidth = 1, fill = \"blue\", col = \"white\") +\n    coord_cartesian(xlim = c(0,10), ylim = c(0,15)) +\n    labs(\n        title = \"Number of Knee Findings Per Patient in Gold-Labeled Data\",\n        x = 'Findings per Patient', \n        y = 'Frequency'\n    ) +\n    scale_x_continuous(\n        breaks = seq(0, 10, by = 2),\n        labels = seq(0, 10, by = 2),\n        minor_breaks = NULL\n    ) +\n    geom_vline(xintercept = pt_findings_ss$median, linetype = 2, color = \"gray\") +\n    geom_vline(xintercept = pt_findings_ss$avg, linetype = 2, color = \"gray\") +\n    geom_vline(xintercept = pt_findings_ss$Q1, linetype = 2, color = \"gray\") +\n    geom_vline(xintercept = pt_findings_ss$Q4, linetype = 2, color = \"gray\") +\n    annotate(\n        \"text\", x = 2, y = 15, \n        label = paste(\"Q1:\", pt_findings_ss$Q1, sep = \" \"), \n        hjust = 1, fontface = 2\n    ) +\n    annotate(\n        \"text\", x = 3.9, y = 15, \n        label = paste(\"Median:\", pt_findings_ss$median, sep = \" \"), \n        hjust = 1, fontface = 2\n    ) +\n    annotate(\n        \"text\", x = 4.25, y = 15, \n        label = paste(\"Mean:\", round(pt_findings_ss$avg, digits = 2), sep = \" \"), \n        hjust = 0, fontface = 2\n    ) +\n    annotate(\n        \"text\", x = 5.9, y = 15, \n        label = paste(\"Q4:\", pt_findings_ss$Q4, sep = \" \"), \n        hjust = -1, fontface = 2\n    ) +\n    theme_minimal() +\n    theme(plot.title = element_text(hjust = 0.5, face = \"bold\", size = 15)) +\n    theme(panel.grid.major.x = element_blank()) +\n    theme(axis.title = element_text(face = \"bold\", size = 12))\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-21T19:25:17.859964Z","iopub.execute_input":"2026-08-21T19:25:17.862746Z","iopub.status.idle":"2026-08-21T19:25:18.273053Z","shell.execute_reply":"2026-08-21T19:25:18.27064Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Conlusions: Findings per Patient\n- For labeled cases, most patients have multiple knee findings (median of 4, mean of 4.14)\n- Slight right skew as expected","metadata":{}},{"cell_type":"code","source":"gold_labeled_df |>\n    summarize(\n        num = sum(Evaluation),\n        .by = Finding\n    ) |>\n    arrange(desc(num)) |>\n    mutate(\n      pct = num / sum(num) * 100,\n      Finding = fct_inorder(Finding)\n    ) |>\n    ggplot(aes(x = Finding, y = num)) +\n        geom_col(fill = \"blue\") +\n        coord_cartesian(ylim = c(0,40), clip = \"off\") +\n        labs(\n            title = \"Knee-Finding Frequency in Gold-Labeled Data\",\n            x = 'Finding Type', \n            y = 'Frequency, n (%)'\n        ) +\n        scale_y_continuous(\n            breaks = seq(0, 40, by = 10),\n            labels = seq(0, 40, by = 10),\n            minor_breaks = NULL\n        ) +\n        geom_text(\n            aes(\n                label = paste(num, \" (\", round(pct, digits = 1), \")\", sep = \"\"), \n                vjust = -1, \n                fontface = 2\n            )\n        ) +\n        theme_minimal() + \n        theme(plot.title = element_text(hjust = 0.5, face = \"bold\", size = 15)) +\n        theme(axis.title = element_text(face = \"bold\", size = 12)) +\n        theme(axis.text.x = element_text(angle = 45, vjust = 1, hjust=1)) +\n        theme(panel.grid.major.x = element_blank()) +\n        theme(plot.margin = margin(t = 20, r = 20, b = 20, l = 20, unit = \"pt\"))\n    ","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-21T19:25:18.276507Z","iopub.execute_input":"2026-08-21T19:25:18.278241Z","iopub.status.idle":"2026-08-21T19:25:18.626729Z","shell.execute_reply":"2026-08-21T19:25:18.624395Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Conlusions: Knee Finding Frequency\n- Finding prevalence in the labeled data ranges from 14.6% to 3.8%\n- Effusion is the most common finding, present at ~3-4 times the frequency of the least (MCL)","metadata":{}},{"cell_type":"code","source":"gold_labeled_df |>\n    pivot_wider(\n        id_cols = StudyInstanceUID, \n        names_from = Finding, \n        values_from = Evaluation\n    ) |> \n    select(!StudyInstanceUID) |>\n    cor() |>\n    ggcorrplot(\n        title = \"Pairwise Correlation Between Finding Types in Gold-Labeled Data\",\n        legend.title = \"Phi\",\n        hc.order = TRUE, \n        lab = TRUE\n    ) +\n    theme(plot.title = element_text(hjust = 0.5, vjust = 1, face = \"bold\", size = 15)) +\n    theme(plot.margin = margin(t = 20, r = 20, b = 20, l = 20, unit = \"pt\"))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-21T19:25:18.629963Z","iopub.execute_input":"2026-08-21T19:25:18.631772Z","iopub.status.idle":"2026-08-21T19:25:19.046045Z","shell.execute_reply":"2026-08-21T19:25:19.043362Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Conlusions: Knee Finding Relationships\n- Simple pairwise correlation with heirarchical clustering consistent with relationships between findings\n- One cluster appears to be something like *traumatic* injury (MCL, Fracture, ACL, Contusion)\n- One cluster appears to be more degenerative [OA (medial, PF, lateral), Effusion, Synovitis, and Baker's Cyst]","metadata":{}},{"cell_type":"markdown","source":"## Report Characteristics for Labeled and Unlabeled Records\n- Only 58 records are fully labeled; additional labeling will have to be derived from the information in reports\n- Are there any obvious differences in size or language distribution that must be accounted for?","metadata":{}},{"cell_type":"code","source":"reports_df |>\n    ggplot(aes(y = NumChars, fill = Gold_Label)) + \n        coord_cartesian(ylim = c(0,5000), clip = \"off\") +\n        labs(\n            title = \"Radiology Report Length (Labeled vs Unlabeled)\",\n            x = 'Labeling Status', \n            y = 'Character Count, n',\n            fill = \"Labeling Status\"\n        ) +\n        scale_fill_manual(\n            values = c(\"grey\", \"blue\"),\n            labels = c(\"Unlabeled\", \"Labeled\")\n        ) +\n        scale_y_continuous(\n            breaks = seq(0, 5000, by = 1000),\n            labels = seq(0, 5000, by = 1000),\n            minor_breaks = NULL\n        ) +\n        geom_boxplot() +\n        theme_minimal() +\n        theme(plot.title = element_text(hjust = 0.5, face = \"bold\", size = 15)) +\n        theme(axis.title = element_text(face = \"bold\", size = 12)) +\n        theme(axis.text.x = element_blank()) +\n        theme(panel.grid.major.x = element_blank()) +\n        theme(panel.grid.minor.x = element_blank()) +\n        theme(plot.margin = margin(t = 20, r = 20, b = 20, l = 20, unit = \"pt\"))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-21T19:25:19.050827Z","iopub.execute_input":"2026-08-21T19:25:19.052885Z","iopub.status.idle":"2026-08-21T19:25:19.496125Z","shell.execute_reply":"2026-08-21T19:25:19.493791Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"labeled_langs_df <- reports_df |>\n    filter(Gold_Label & !is.na(Language)) |>\n    summarize(\n      LblLangCount = n(), \n      .by = Language\n    ) \n\nall_langs_df <- reports_df |>\n    filter(!is.na(Language)) |>\n    summarize(\n      LangCount = n(), \n      .by = Language\n    ) |>\n    left_join(labeled_langs_df, by = \"Language\") |>\n    mutate(\n        LblLangCount = case_when(\n            is.na(LblLangCount) ~ as.integer(0), \n            !is.na(LblLangCount) ~ LblLangCount\n        ),\n        LangPct = (LangCount / sum(LangCount)) * 100,\n        LblLangPct = (LblLangCount / sum(LblLangCount)) * 100\n    ) |>\n    arrange(desc(LangCount)) |>\n    mutate(Language = fct_inorder(Language)) \nall_langs_df |> \n    select(!c(LangCount, LblLangCount)) |>\n    pivot_longer(cols = !Language, names_to = \"LangParam\", values_to = \"Value\") |>\n    ggplot(aes(x = Language, y = Value, fill = LangParam)) +\n        geom_col(position = \"dodge\") +\n        scale_fill_manual(\n            values = c(\"grey\", \"blue\"),\n            labels = c(\"All (Labeled + Unlabeled)\", \"Labeled\")\n        ) +\n        labs(\n            title = \"Report Language: Percentage by Labeling Status\",\n            x = 'Language', \n            y = 'Reports, %',\n            fill = \"Labeling Status\"\n        ) +\n        theme_minimal() +\n        theme(plot.title = element_text(hjust = 0.5, face = \"bold\", size = 15)) +\n        theme(axis.title = element_text(face = \"bold\", size = 12)) +\n        theme(panel.grid.minor.y = element_blank()) +\n        theme(panel.grid.minor.x = element_blank()) +\n        theme(panel.grid.major.x = element_blank()) +\n        theme(plot.margin = margin(t = 20, r = 20, b = 20, l = 20, unit = \"pt\"))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-21T19:25:19.499522Z","iopub.execute_input":"2026-08-21T19:25:19.501367Z","iopub.status.idle":"2026-08-21T19:25:19.913979Z","shell.execute_reply":"2026-08-21T19:25:19.911633Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Conclusions: Report Characteristics\n- No obvious/strong differences between labeled and ulabeled reports\n- Most reports are in English or Spanish - a smattering of languages are not present in the labeled data","metadata":{}},{"cell_type":"markdown","source":"## Imaging Series for Each Record\n- How many series per case?\n- What types of series are available for each case?\n- How are weightings used?\n- What information is in the raw DICOM files?","metadata":{}},{"cell_type":"code","source":"imaging_series_tbl <- train_series_df |>\n    select(!c(StudyInstanceUID, SeriesInstanceUID)) |>\n    summarize(\n        plane_count = n(),\n        fluid_sens_count = sum(Fluid_Sensitive),\n        fluid_sens_pct = (fluid_sens_count / plane_count) * 100,\n        fat_supp_count = sum(Fat_Suppression),\n        fat_supp_pct = (fat_supp_count / plane_count) * 100,\n        .by = Anatomical_Plane\n    ) |>\n    mutate(plane_pct = round(plane_count / sum(plane_count) * 100), digits = 2) |>\n    select(Anatomical_Plane, plane_count, plane_pct, fluid_sens_count, fluid_sens_pct, fat_supp_count, fat_supp_pct) |>\n    gt() |>\n    cols_label(\n        Anatomical_Plane = \"Anatomical Plane\",\n        plane_count      = \"Series, n\",\n        plane_pct        = \"Series, %\",\n        fluid_sens_count = \"Fluid Sensitive, n\",\n        fluid_sens_pct   = \"Fluid Sensitive, %\",\n        fat_supp_count   = \"Fat Suppression, n\",\n        fat_supp_pct     = \"Fat Suppression, %\"\n    ) |>\n    fmt_number(columns = ends_with(\"pct\"), decimals = 2) |>\n    grand_summary_rows(\n        columns = ends_with(\"count\"),\n        fns = list(Total ~ sum(.)),\n        fmt = ~ fmt_number(., decimals = 0)\n    ) |>\n    grand_summary_rows(\n        columns = plane_pct,\n        fns = list(Total ~ sum(plane_count) / sum(plane_count) * 100),\n        fmt = ~ fmt_number(., decimals = 2)\n    ) |>\n    grand_summary_rows(\n        columns = fluid_sens_pct,\n        fns = list(Total ~ sum(fluid_sens_count) / sum(plane_count) * 100),\n        fmt = ~ fmt_number(., decimals = 2)\n    ) |>\n    grand_summary_rows(\n        columns = fat_supp_pct,\n        fns = list(Total ~ sum(fat_supp_count) / sum(plane_count) * 100),\n        fmt = ~ fmt_number(., decimals = 2)\n    ) |>\n    cols_align(align = \"center\", columns = !Anatomical_Plane) |>\n    tab_options(column_labels.font.weight = \"bold\") |>\n    tab_style(\n        style = cell_text(weight = \"bold\"),\n        locations = cells_grand_summary()\n    ) |>\n    as_raw_html()\n\ndisplay_html(imaging_series_tbl)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-21T19:25:19.917723Z","iopub.execute_input":"2026-08-21T19:25:19.919671Z","iopub.status.idle":"2026-08-21T19:25:20.34889Z","shell.execute_reply":"2026-08-21T19:25:20.346736Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Conclusions: Imaging Series Breakdown\n- Sagittal series are the most common, axial the least\n- Weightings for Fat Suppression and Fluid Sensitive are dramatically higher in Axial series\n- Notable that in the training data there are **no** instances of fluid sensitive without fat suppression weightings (or vice versa)","metadata":{}},{"cell_type":"code","source":"series_details_df <- train_series_df |>\n    select(!c(Fluid_Sensitive, Fat_Suppression)) |>\n    mutate(\n      APValue = 1\n    ) |>\n    pivot_wider(names_from = Anatomical_Plane, values_from = APValue) |>\n    mutate(\n        across(\n            c(Sagittal, Coronal, Axial), \n            ~ case_when(\n                is.na(.x) ~ 0L,\n                TRUE ~ as.integer(.x)\n            )\n          )\n    ) |>\n    summarize(\n        Sagittal_Series = sum(Sagittal),\n        Axial_Series = sum(Axial),\n        Coronal_Series = sum(Coronal),\n        Total_Series = Sagittal_Series + Axial_Series + Coronal_Series,\n        .by = StudyInstanceUID\n    ) \n\nseries_details_df |>\n    pivot_longer(\n        cols = c(Sagittal_Series, Axial_Series, Coronal_Series, Total_Series),\n        names_to = \"Series_Param\",\n        values_to = \"Series_Count\"\n    )|>\n    mutate(\n       Series_Param = factor(\n          Series_Param, \n          labels = c(\"Total\", \"Sagittal\", \"Coronal\", \"Axial\"),\n          levels = c(\"Total_Series\", \"Sagittal_Series\", \"Coronal_Series\", \"Axial_Series\")\n        ) \n    ) |>\n    ggplot(aes(x = Series_Param, y = Series_Count)) +\n        geom_count(color = \"blue\") +\n        labs(\n            title = \"MRI Series per Study: Overall and by Series Type\",\n            y = \"Series per Study, n\",\n            x = \"Series Type\"\n        )+\n        scale_size_area(\n            name = \"Series Count\",\n            max_size = 15,\n            breaks = scales::breaks_pretty(n = 4)\n        ) +\n        coord_cartesian(ylim = c(0, 14), clip = 'off') +\n        scale_y_continuous(\n            breaks = seq(0, 14, by = 2),\n            labels = seq(0, 14, by = 2),\n            minor_breaks = NULL\n        ) +\n        theme_minimal() +\n        theme(plot.title = element_text(hjust = 0.5, face = \"bold\", size = 15)) +\n        theme(axis.title = element_text(face = \"bold\", size = 12)) +\n        theme(panel.grid.minor.y = element_blank()) +\n        theme(panel.grid.minor.x = element_blank()) +\n        theme(panel.grid.major.x = element_blank()) +\n        theme(plot.margin = margin(t = 20, r = 20, b = 20, l = 20, unit = \"pt\"))\n  \n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-21T19:25:20.352089Z","iopub.execute_input":"2026-08-21T19:25:20.353905Z","iopub.status.idle":"2026-08-21T19:25:20.999102Z","shell.execute_reply":"2026-08-21T19:25:20.996729Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Conclusions: Number of Series per Patient\n- All cases appear to be associated with multiple series in the training data; at least one from each plane\n- Notable that the *norm* for Sagittal and Coronal planes is multiple series per patient; even for the Axial plane it is common\n- Should look into this more deeply - how do series of the same plane differ (eg, various positions, +/- weights?) ","metadata":{}},{"cell_type":"markdown","source":"## Evaluation of Sample DICOM Data\n- The DICOM format combines imaging data with header information containing a lot of additional information\n- What's in the header for the data we have?\n- Let's also try to take a peek at a few images!","metadata":{}},{"cell_type":"code","source":"grab_a_slice <- function(study, series, base = TRAIN_FILES_LOC){\n    loc <- file.path(base, study, series)\n    dcm_file <- sample(list.files(loc, pattern = \"\\\\.dcm$\"), 1)\n    dcm_data <- readDICOMFile(file.path(loc, dcm_file))\n    return (\n        list(\n            file_name = dcm_file,\n            hdr = dcm_data$hdr,\n            img = dcm_data$img\n        )\n    )\n}\n\n\nset.seed(42)\ndicom_sample_df <- train_series_df |>\n    group_by(Anatomical_Plane) |>\n    slice_sample(n = 5) |>\n    ungroup()|>\n    mutate(\n        DICOMData = map2(\n            StudyInstanceUID, \n            SeriesInstanceUID, \n            grab_a_slice\n        )\n    ) \n\ncat(\"Grabbed\", nrow(dicom_sample_df), \"slices from\", n_distinct(dicom_sample_df$StudyInstanceUID), \"studies:\\n\",\n    \"Coronal:\", nrow(dicom_sample_df[dicom_sample_df$Anatomical_Plane == \"Coronal\",]), \"slices\\n\",\n    \"Axial:\", nrow(dicom_sample_df[dicom_sample_df$Anatomical_Plane == \"Axial\",]), \"slices\\n\",\n    \"Sagittal:\", nrow(dicom_sample_df[dicom_sample_df$Anatomical_Plane == \"Sagittal\",]), \"slices\\n\"\n)\n    ","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-21T19:25:21.003293Z","iopub.execute_input":"2026-08-21T19:25:21.005143Z","iopub.status.idle":"2026-08-21T19:25:22.848662Z","shell.execute_reply":"2026-08-21T19:25:22.845587Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"smpl_hdrs_df <- dicom_sample_df |>\n  unnest_wider(DICOMData) |>\n  select(c(StudyInstanceUID, SeriesInstanceUID, file_name, hdr)) |>\n  unnest_wider(hdr) |>\n  unnest_longer(c(group, element, name, code, length, value, sequence))\n\noptions(repr.plot.width = 15, repr.plot.height = 20)\nsmpl_hdrs_df |>\n   summarize(\n       n_files = n_distinct(file_name), \n       .by = c(group, element, name)\n   ) |>\n   mutate(label = fct_reorder(paste0(\"(\", group, \",\", element, \") \", name), n_files)) |>\n   ggplot(aes(x = n_files, y = label, fill = group)) +\n        geom_col() + \n        labs(\n            title = \"DICOM Header Variables in Sample Set\",\n            x = \"Number of Files Present\",\n            y = \"Variable, (Group No., Element No.) Name\"\n        ) +\n        scale_fill_brewer(palette = \"Blues\", direction = -1, name = \"Group Number\")+\n        theme_minimal() +\n        theme(plot.title = element_text(hjust = 0.5,face = \"bold\", size = 20)) +\n        theme(axis.title = element_text(face = \"bold\", size = 15)) +\n        theme(panel.grid.minor.x = element_blank()) +\n        theme(panel.grid.major.y = element_blank()) +\n        theme(plot.margin = margin(t = 20, r = 20, b = 20, l = 20, unit = \"pt\"))\n   ","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-21T19:25:22.852787Z","iopub.execute_input":"2026-08-21T19:25:22.854683Z","iopub.status.idle":"2026-08-21T19:25:23.574152Z","shell.execute_reply":"2026-08-21T19:25:23.57156Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Conclusions: DICOM Header Data\n- There is a lot here potentially - in the sample, 34 elements from 5 group numbers were available in all 15 sampled DICOMs\n- Worth looking into further whether any of these might be useful; some are certainly redundant; but we can see the missing `PatientSex` might   be findable here, in addition to the report!","metadata":{}},{"cell_type":"code","source":"img_samples_df <- dicom_sample_df |>\n    unnest_wider(DICOMData)|>\n    select(c(StudyInstanceUID, SeriesInstanceUID, Anatomical_Plane, file_name, img))\n\noptions(repr.plot.width = 15, repr.plot.height = 9)\npar(mfrow = c(3, 5))\n\nfor (i in seq_len(nrow(img_samples_df))) {\n    m <- img_samples_df$img[[i]]\n    image(\n        t(m[nrow(m):1, ]),\n        col = gray.colors(256),\n        axes = FALSE,\n        asp = 1,\n        main = img_samples_df$Anatomical_Plane[i]\n    )\n}","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-21T19:39:45.668637Z","iopub.execute_input":"2026-08-21T19:39:45.670499Z","iopub.status.idle":"2026-08-21T19:39:55.868942Z","shell.execute_reply":"2026-08-21T19:39:55.865749Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Conclusions: Actual Image Samples\n- Looks to be a fair amount of variability in the images themselves, not suprisingly\n- Look deeper into this (including what info can be gotten from the header information) for normalization","metadata":{}}]}