{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.12.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[],"dockerImageVersionId":28755,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<p style='font: 600 36px \"Lato\", \"Open Sans\", \"Helvetica Neue\", Helvetica, Arial, sans-serif'>\n    <img align=\"right\" width=\"300\" height=\"150\" style=\"border-radius: 12px;\" src=\"https://www.kaggle.com/competitions/154281/images/header\">\n    RSNA Knee Abnormality Detection [EDA]\n    <p style='font: 500 22px \"Lato\", \"Open Sans\", \"Helvetica Neue\", Helvetica, Arial, sans-serif'>\n        Create a model that can detect knee abnormalities based on multimodal imaging data\n    </p>\n</p>\n\n<p style='font: 300 18px/22px \"Lato\", \"Open Sans\", \"Helvetica Neue\", Helvetica, Arial, sans-serif'>\n    In this notebook, we'll try to understand what data we have to work with. We'll examine the data provided in the <code style=\"background-color: #f4f4f4; color: #d63384; padding: 2px 6px; border-radius: 4px; font-family: monospace; font-size: 15px;\">train.csv</code> and <code style=\"background-color: #f4f4f4; color: #d63384; padding: 2px 6px; border-radius: 4px; font-family: monospace; font-size: 15px;\">train_series.csv</code> files, and we'll also analyze the <code style=\"background-color: #f4f4f4; color: #d63384; padding: 2px 6px; border-radius: 4px; font-family: monospace; font-size: 15px;\">metadata</code> content of the <code style=\"background-color: #f4f4f4; color: #d63384; padding: 2px 6px; border-radius: 4px; font-family: monospace; font-size: 15px;\">DICOM</code> files\n</p>\n        \n","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nfrom kaggle_secrets import UserSecretsClient\nimport os\nimport re\nimport glob\nfrom transformers import pipeline\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nfrom wordcloud import WordCloud\nimport pydicom\nimport warnings\nwarnings.filterwarnings(\"ignore\", category=DeprecationWarning)\npd.set_option('display.max_colwidth', 300)","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2026-09-04T23:50:28.113589Z","iopub.execute_input":"2026-09-04T23:50:28.113879Z","iopub.status.idle":"2026-09-04T23:50:58.164247Z","shell.execute_reply.started":"2026-09-04T23:50:28.113853Z","shell.execute_reply":"2026-09-04T23:50:58.163134Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"MAIN_PATH = '/kaggle/input/competitions/rsna-knee-abnormality-detection/'\nTRAIN_PATH = '/kaggle/input/competitions/rsna-knee-abnormality-detection/train_series/'","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-04T23:50:58.165961Z","iopub.execute_input":"2026-09-04T23:50:58.166652Z","iopub.status.idle":"2026-09-04T23:50:58.171373Z","shell.execute_reply.started":"2026-09-04T23:50:58.166621Z","shell.execute_reply":"2026-09-04T23:50:58.170426Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"user_secrets = UserSecretsClient()\nhf_token = user_secrets.get_secret(\"HF_TOKEN\")\nos.environ[\"HF_TOKEN\"] = hf_token","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-04T23:50:58.172634Z","iopub.execute_input":"2026-09-04T23:50:58.173041Z","iopub.status.idle":"2026-09-04T23:50:58.273596Z","shell.execute_reply.started":"2026-09-04T23:50:58.173014Z","shell.execute_reply":"2026-09-04T23:50:58.272694Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Analyzing the train.csv file","metadata":{}},{"cell_type":"code","source":"train_df = pd.read_csv(MAIN_PATH + 'train.csv')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-04T23:50:58.275545Z","iopub.execute_input":"2026-09-04T23:50:58.275859Z","iopub.status.idle":"2026-09-04T23:50:58.451728Z","shell.execute_reply.started":"2026-09-04T23:50:58.275834Z","shell.execute_reply":"2026-09-04T23:50:58.450394Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_df.info()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-04T23:50:58.452941Z","iopub.execute_input":"2026-09-04T23:50:58.453319Z","iopub.status.idle":"2026-09-04T23:50:58.491362Z","shell.execute_reply.started":"2026-09-04T23:50:58.453289Z","shell.execute_reply":"2026-09-04T23:50:58.49001Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Labels","metadata":{}},{"cell_type":"markdown","source":"Labels are available for only 58 studies out of 4,500! Nevertheless, let's take a look at the class distribution","metadata":{}},{"cell_type":"code","source":"feature_means = train_df.dropna()[train_df.columns[3:]].mean().sort_values(ascending=False)\n\nplt.figure(figsize=(8, 4))\nsns.barplot(x=feature_means.values, y=feature_means.index)\nplt.axvline(0.5, color=\"red\", linestyle=\"--\", alpha=0.5, label=\"50% series\")\n\nplt.title(\"What proportion of series have each feature?\")\nplt.xlabel(\"rate\")\nplt.ylabel(\"feature\")\nplt.xlim(0, 1)\nplt.legend();","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-04T23:50:58.492452Z","iopub.execute_input":"2026-09-04T23:50:58.493224Z","iopub.status.idle":"2026-09-04T23:50:58.84775Z","shell.execute_reply.started":"2026-09-04T23:50:58.493183Z","shell.execute_reply":"2026-09-04T23:50:58.846852Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Reports statistic","metadata":{}},{"cell_type":"markdown","source":"Each study has a description (report). Based on these descriptions, the dataset will need to be annotated. An additional complication is that the descriptions may be in different languages.\n\nLet's check how many descriptions are repeated and how extensive they are.","metadata":{}},{"cell_type":"code","source":"tst_uniq = train_df[['StudyInstanceUID', 'Report']].groupby('Report').size().reset_index(name='count')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-04T23:50:58.849322Z","iopub.execute_input":"2026-09-04T23:50:58.849663Z","iopub.status.idle":"2026-09-04T23:50:58.885796Z","shell.execute_reply.started":"2026-09-04T23:50:58.849609Z","shell.execute_reply":"2026-09-04T23:50:58.88484Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"tst_uniq[tst_uniq['count']>1].sort_values(by='count', ascending=False).head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-04T23:50:58.88698Z","iopub.execute_input":"2026-09-04T23:50:58.887635Z","iopub.status.idle":"2026-09-04T23:50:58.915506Z","shell.execute_reply.started":"2026-09-04T23:50:58.887599Z","shell.execute_reply":"2026-09-04T23:50:58.914417Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"There are not many repetitions, and they mainly concern studies of healthy organs","metadata":{}},{"cell_type":"code","source":"# reports length statistic\ntrain_df['Report'].str.len().describe()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-04T23:50:58.916823Z","iopub.execute_input":"2026-09-04T23:50:58.917467Z","iopub.status.idle":"2026-09-04T23:50:58.931899Z","shell.execute_reply.started":"2026-09-04T23:50:58.917439Z","shell.execute_reply":"2026-09-04T23:50:58.93084Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"On average, a description contains about 1000 characters","metadata":{}},{"cell_type":"markdown","source":"## Report languages","metadata":{}},{"cell_type":"markdown","source":"Let's determine what languages we will have to deal with","metadata":{}},{"cell_type":"code","source":"import fasttext\nfrom huggingface_hub import hf_hub_download\n\nmodel_path = hf_hub_download(repo_id=\"facebook/fasttext-language-identification\", filename=\"model.bin\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-04T23:50:58.934903Z","iopub.execute_input":"2026-09-04T23:50:58.935326Z","iopub.status.idle":"2026-09-04T23:51:07.58412Z","shell.execute_reply.started":"2026-09-04T23:50:58.935298Z","shell.execute_reply":"2026-09-04T23:51:07.58324Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"model = fasttext.load_model(model_path)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-04T23:51:07.585195Z","iopub.execute_input":"2026-09-04T23:51:07.58544Z","iopub.status.idle":"2026-09-04T23:51:08.67399Z","shell.execute_reply.started":"2026-09-04T23:51:07.585419Z","shell.execute_reply":"2026-09-04T23:51:08.672831Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"texts = train_df['Report'].str.replace('\\n', ' ', regex=False).replace('\\r', '', regex=False).to_list()\nres = model.predict(texts)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-04T23:51:08.676657Z","iopub.execute_input":"2026-09-04T23:51:08.677104Z","iopub.status.idle":"2026-09-04T23:51:13.235894Z","shell.execute_reply.started":"2026-09-04T23:51:08.677043Z","shell.execute_reply":"2026-09-04T23:51:13.234786Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"languages = [(re.findall(r\"__label__(.*?)_\", tuple_lang_prob[0][0])[0], \n              tuple_lang_prob[1][0],\n              re.findall(r\"__label__(.*?)$\", tuple_lang_prob[0][0])[0], \n             ) \n             for tuple_lang_prob in zip(res[0], res[1])\n            ]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-04T23:51:13.237825Z","iopub.execute_input":"2026-09-04T23:51:13.238276Z","iopub.status.idle":"2026-09-04T23:51:13.258021Z","shell.execute_reply.started":"2026-09-04T23:51:13.238248Z","shell.execute_reply":"2026-09-04T23:51:13.256955Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"lang_mapping = {\n    'bos': 'Bosnian',  'bul': 'Bulgarian',   'deu': 'German',\n    'ell': 'Greek',    'eng': 'English',     'fra': 'French',\n    'glg': 'Galician', 'hrv': 'Croatian',    'kor': 'Korean',\n    'nld': 'Dutch',    'por': 'Portuguese',  'spa': 'Spanish',\n    'tur': 'Turkish',  'yue': 'Cantonese'\n}","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-04T23:51:13.259339Z","iopub.execute_input":"2026-09-04T23:51:13.259789Z","iopub.status.idle":"2026-09-04T23:51:13.281969Z","shell.execute_reply.started":"2026-09-04T23:51:13.259761Z","shell.execute_reply":"2026-09-04T23:51:13.280365Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"wordcloud = WordCloud(\n    width=800, \n    height=400, \n    background_color=\"white\"\n).generate(' '.join([lang_mapping[id[0]] for id in languages]))\n\nplt.imshow(wordcloud, interpolation=\"bilinear\")\nplt.axis(\"off\");","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-04T23:51:13.28348Z","iopub.execute_input":"2026-09-04T23:51:13.283752Z","iopub.status.idle":"2026-09-04T23:51:13.592695Z","shell.execute_reply.started":"2026-09-04T23:51:13.28373Z","shell.execute_reply":"2026-09-04T23:51:13.591699Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Anything from Cantonese and minor is not. Below, we'll consider situations where the model is less than 50% confident in its answer. All of these rare languages ​​will be included in this sample","metadata":{}},{"cell_type":"code","source":"# if someone perceives it better in percentages )\nlang_df = pd.DataFrame({\n    'StudyInstanceUID': train_df['StudyInstanceUID'],\n    'report': texts, \n    'language': [res[2] for res in languages], \n    'prob': [res[1] for res in languages],\n})\n\n(lang_df['language'].value_counts(normalize=True) * 100).round(1).astype(str) + '%'","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-04T23:51:13.593728Z","iopub.execute_input":"2026-09-04T23:51:13.594028Z","iopub.status.idle":"2026-09-04T23:51:13.610013Z","shell.execute_reply.started":"2026-09-04T23:51:13.594001Z","shell.execute_reply":"2026-09-04T23:51:13.609153Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# here the model doubts the result\nlang_doubt_df = lang_df[lang_df['prob']<0.5]\nlang_doubt_df['language'].value_counts()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-04T23:51:13.611158Z","iopub.execute_input":"2026-09-04T23:51:13.611539Z","iopub.status.idle":"2026-09-04T23:51:13.637059Z","shell.execute_reply.started":"2026-09-04T23:51:13.611489Z","shell.execute_reply":"2026-09-04T23:51:13.635863Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"lang_doubt_df[lang_doubt_df['language']=='kor_Hang']","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-04T23:51:13.638416Z","iopub.execute_input":"2026-09-04T23:51:13.638874Z","iopub.status.idle":"2026-09-04T23:51:13.665355Z","shell.execute_reply.started":"2026-09-04T23:51:13.638831Z","shell.execute_reply":"2026-09-04T23:51:13.664226Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"For example, this \"Korean\" turned out to be English. You can check the others. It seems the reasons for this uncertainty lie in the short phrases and abundance of specialized terminology\n\nIn any case, we can confirm that we are dealing with a dozen European languages","metadata":{}},{"cell_type":"markdown","source":"We'll leave English and Spanish as is. We'll change Cantonese, Korean, and Croatian to English, and Portuguese and Galician to Spanish. We'll only change all of the above for the sample below 50%, and we'll change Cantonese entirely, as one row with a 57% probability was missing from the sample, but it's also incorrect.","metadata":{}},{"cell_type":"code","source":"lang_df.loc[lang_df['language'] == 'yue_Hant', 'language'] = 'eng_Latn'\nlang_df.loc[(lang_df['prob'] < 0.5) & (lang_df['language'].isin(['kor_Hang', 'hrv_Latn'])), 'language'] = 'eng_Latn'\nlang_df.loc[(lang_df['prob'] < 0.5) & (lang_df['language'].isin(['por_Latn', 'glg_Latn'])), 'language'] = 'spa_Latn'","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-04T23:51:13.666617Z","iopub.execute_input":"2026-09-04T23:51:13.666966Z","iopub.status.idle":"2026-09-04T23:51:13.695507Z","shell.execute_reply.started":"2026-09-04T23:51:13.666931Z","shell.execute_reply":"2026-09-04T23:51:13.694349Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# let's save it for the future\nlang_df.drop('prob', axis=1).to_parquet('language.parquet', index=False)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-04T23:51:13.696701Z","iopub.execute_input":"2026-09-04T23:51:13.696948Z","iopub.status.idle":"2026-09-04T23:51:13.914171Z","shell.execute_reply.started":"2026-09-04T23:51:13.696926Z","shell.execute_reply":"2026-09-04T23:51:13.913029Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Analyzing the train_series.csv file","metadata":{}},{"cell_type":"markdown","source":"The train_series table contains three additional attributes for each series: 'Fluid_Sensitive', 'Fat_Suppression', and 'Anatomical_Plane'. The last attribute is simple—it's the imaging projection, which can be one of three options: `Coronal`, `Sagittal`, or `Axial`. The other two are helpful in image preprocessing, as they indicate that fluid and fat are being displayed in a specific way.\n\nLet's determine the basic minimum:\n1. How many series are there per study?\n2. Are all projections available for each study?\n3. What combinations of the binary attributes 'Fluid_Sensitive' and 'Fat_Suppression' are available for each projection?","metadata":{}},{"cell_type":"code","source":"train_series_df = pd.read_csv(MAIN_PATH + 'train_series.csv')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-04T23:58:27.730295Z","iopub.execute_input":"2026-09-04T23:58:27.730658Z","iopub.status.idle":"2026-09-04T23:58:27.82725Z","shell.execute_reply.started":"2026-09-04T23:58:27.73063Z","shell.execute_reply":"2026-09-04T23:58:27.826124Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# there arn't any nulls\ntrain_series_df.info()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-04T23:58:27.829295Z","iopub.execute_input":"2026-09-04T23:58:27.829683Z","iopub.status.idle":"2026-09-04T23:58:27.847265Z","shell.execute_reply.started":"2026-09-04T23:58:27.829656Z","shell.execute_reply":"2026-09-04T23:58:27.845886Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"First, let's determine what options for the number of series are available for all studies","metadata":{}},{"cell_type":"code","source":"# number of Series per Study\nax = train_series_df[['StudyInstanceUID', 'SeriesInstanceUID']].groupby('StudyInstanceUID').count()['SeriesInstanceUID'].value_counts().plot.bar()\n\nfor container in ax.containers:\n    ax.bar_label(container, padding=3)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-04T23:58:27.848662Z","iopub.execute_input":"2026-09-04T23:58:27.849156Z","iopub.status.idle":"2026-09-04T23:58:28.145802Z","shell.execute_reply.started":"2026-09-04T23:58:27.849116Z","shell.execute_reply":"2026-09-04T23:58:28.144807Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"More than half of the studies contained five series. It's also worth noting that no study contained fewer than three series, suggesting that each study contained images in all three projections. However, this is only an assumption that needs to be verified","metadata":{}},{"cell_type":"code","source":"train_series_df[['StudyInstanceUID', 'Anatomical_Plane']].groupby('StudyInstanceUID').nunique()['Anatomical_Plane'].value_counts()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-04T23:58:28.148044Z","iopub.execute_input":"2026-09-04T23:58:28.148405Z","iopub.status.idle":"2026-09-04T23:58:28.169811Z","shell.execute_reply.started":"2026-09-04T23:58:28.148366Z","shell.execute_reply":"2026-09-04T23:58:28.168944Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"We've confirmed that for all 4,407 studies, there's at least one series of each of the three projections. At least one, but it's clear there could be more. Let's explore the options!","metadata":{}},{"cell_type":"code","source":"# how many variants of one projection can there be in a series?\nnumb_of_planes = train_series_df[['StudyInstanceUID', 'Anatomical_Plane']].value_counts().reset_index(name='counts')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-04T23:58:28.171067Z","iopub.execute_input":"2026-09-04T23:58:28.171496Z","iopub.status.idle":"2026-09-04T23:58:28.194846Z","shell.execute_reply.started":"2026-09-04T23:58:28.171456Z","shell.execute_reply":"2026-09-04T23:58:28.193604Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plt.figure(figsize=(10, 6))\nsns.countplot(data=numb_of_planes, x='Anatomical_Plane', hue='counts', palette=\"muted\");","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-04T23:58:28.195901Z","iopub.execute_input":"2026-09-04T23:58:28.19624Z","iopub.status.idle":"2026-09-04T23:58:28.519238Z","shell.execute_reply.started":"2026-09-04T23:58:28.196212Z","shell.execute_reply":"2026-09-04T23:58:28.518031Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"In the diagram, we see that the maximum number of series in a single projection in a study can be up to seven! It's unclear why... They were probably just there ) We also see that there's usually one axial projection, and two coronal and sagittal projections.\n\nWe also have two binary attributes: `Fluid_Sensitive` and `Fat_Suppression`, which could be the reason for multiple series in a single projection. Since these are two binary attributes, they can explain the maximum of four series in a single projection. But there can be up to seven! Well, at least we'll try to understand something.\n\nFirst, let's look at the possible combinations of `Fluid_Sensitive` and `Fat_Suppression`","metadata":{}},{"cell_type":"code","source":"# number of fluid/fat signals (both signals occur only simultaneously)\n(train_series_df[['Fluid_Sensitive', 'Fat_Suppression']].value_counts(normalize=True) * 100).round(1).astype(str) + '%'","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-04T23:58:28.520589Z","iopub.execute_input":"2026-09-04T23:58:28.520935Z","iopub.status.idle":"2026-09-04T23:58:28.535389Z","shell.execute_reply.started":"2026-09-04T23:58:28.520897Z","shell.execute_reply":"2026-09-04T23:58:28.534186Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"And here's a surprise! There are only two combinations, not four as we initially assumed. Either both parameters are True, or both are False. The fat suppression and fluid sensitivity series account for just over half.\n\nLet's look at this distribution by projection:","metadata":{}},{"cell_type":"code","source":"((train_series_df[['Anatomical_Plane', 'Fluid_Sensitive', 'Fat_Suppression']].value_counts() / \n train_series_df[['Anatomical_Plane']].value_counts()) * 100).round(1).astype(str) + '%'","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-04T23:58:28.536877Z","iopub.execute_input":"2026-09-04T23:58:28.537336Z","iopub.status.idle":"2026-09-04T23:58:28.574988Z","shell.execute_reply.started":"2026-09-04T23:58:28.5373Z","shell.execute_reply":"2026-09-04T23:58:28.573904Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Here we see that the distribution is roughly 50/50 for the coronal and sagittal projections, while for the axial projection, the images are predominantly (80%) fat-suppressed and fluid-emphasized.\n\nNext, we'll calculate how many combinations of additional attributes are available for each study","metadata":{}},{"cell_type":"code","source":"numb_of_planes_and_ext = train_series_df[['StudyInstanceUID', 'Anatomical_Plane', 'Fluid_Sensitive']].groupby(['StudyInstanceUID', 'Anatomical_Plane', 'Fluid_Sensitive']).size().reset_index(name='counts')\nnumb_of_planes_and_ext.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-04T23:58:28.576402Z","iopub.execute_input":"2026-09-04T23:58:28.576762Z","iopub.status.idle":"2026-09-04T23:58:28.605636Z","shell.execute_reply.started":"2026-09-04T23:58:28.576727Z","shell.execute_reply":"2026-09-04T23:58:28.604741Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Let's look at the series that have all the same additional features, i.e. the same projection and the same combination of fat suppression and fluid sensitivity","metadata":{}},{"cell_type":"code","source":"print('the number of repetitions of series with the same parameters in studies:')\nprint(f' - doubled             {len(numb_of_planes_and_ext[numb_of_planes_and_ext['counts']==2]):>4}')\nprint(f' - tripled             {len(numb_of_planes_and_ext[numb_of_planes_and_ext['counts']==3]):>4}')\nprint(f' - for and more times  {len(numb_of_planes_and_ext[numb_of_planes_and_ext['counts']>3]):>4}')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-04T23:58:28.608744Z","iopub.execute_input":"2026-09-04T23:58:28.609152Z","iopub.status.idle":"2026-09-04T23:58:28.618849Z","shell.execute_reply.started":"2026-09-04T23:58:28.609117Z","shell.execute_reply":"2026-09-04T23:58:28.617868Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Well, quite a few... Let's look at an example. Let's select studies with more than three repetitions and examine images from one of them","metadata":{}},{"cell_type":"code","source":"numb_of_planes_and_ext[numb_of_planes_and_ext['counts']>3]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-04T23:58:28.620036Z","iopub.execute_input":"2026-09-04T23:58:28.620343Z","iopub.status.idle":"2026-09-04T23:58:28.641635Z","shell.execute_reply.started":"2026-09-04T23:58:28.620322Z","shell.execute_reply":"2026-09-04T23:58:28.640643Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Let's take as an example the first study in the list with index 4966, where there are five images in axial projection without fat suppression and fluid accentuation","metadata":{}},{"cell_type":"code","source":"for_print_df = train_series_df[train_series_df['StudyInstanceUID'] == numb_of_planes_and_ext['StudyInstanceUID'][4966]]\nfor_print_df","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-04T23:58:28.642853Z","iopub.execute_input":"2026-09-04T23:58:28.643504Z","iopub.status.idle":"2026-09-04T23:58:28.670706Z","shell.execute_reply.started":"2026-09-04T23:58:28.643475Z","shell.execute_reply":"2026-09-04T23:58:28.669634Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"We see that the study contains six axial images, five of which are without fat suppression, and one with suppression. Let's look at the images.","metadata":{}},{"cell_type":"code","source":"def feat_by_series_uid(df, series_uid, prefix=TRAIN_PATH):\n    \"\"\"\n    Returns a dict of features for a series\n    df - dataframe based on train(test)_series.csv\n    series_uid - unique series identifier\n    prefix - part of the path to files\n    \"\"\"\n    \n    feat_dict =  df[df['SeriesInstanceUID']==series_uid].to_dict(orient='records')[0]\n    feat_dict['path'] = TRAIN_PATH + '/'.join([feat_dict['StudyInstanceUID'], feat_dict['SeriesInstanceUID']]) + '/*.dcm'\n    return feat_dict\n\ndef print_series_images(data, clmns_numb = 5, figsize=(16,20)):\n    \"\"\"\n    Prints images of a series in the order they were scanned\n    data - dict from feat_by_series_uid()\n    clmns_numb - number of columns\n    figsize - tuple for plt.subplots\n    \"\"\"\n    series_path = glob.glob(data['path'])\n    fig, axs = plt.subplots(int(np.ceil(len(series_path) / clmns_numb)), \n                            clmns_numb,\n                            figsize=figsize,\n                            constrained_layout=True,\n                            subplot_kw={'xticks':[], 'yticks':[]})\n    \n    for i, path in enumerate(series_path):\n        \n        ds = pydicom.dcmread(path)\n        img = ds.pixel_array\n        idx = ds['InstanceNumber'].value - 1 # for sort by number in series\n        axs[idx // clmns_numb, idx % clmns_numb].imshow(img, cmap='bone')\n\n    # let's get some dicom metadata\n    manufacturer = getattr(ds, 'Manufacturer', \"Unknown\")\n    magn_field_strng = getattr(ds, 'MagneticFieldStrength', \"Unknown\")\n    tr = getattr(ds, 'RepetitionTime', \"Unknown\")\n    te = getattr(ds, 'EchoTime', \"Unknown\")\n    fig.suptitle('{0}\\n{1}, Fluid_Sensitive - {2}, Fat_Suppression - {3}\\nManufacturer: {4}, Field strength - {5}\\nRepetition Time - {6}, Echo Time {7}'.format(\n        data['SeriesInstanceUID'],   # 0\n        data['Anatomical_Plane'],    # 1\n        data['Fluid_Sensitive'],     # 2\n        data['Fat_Suppression'],     # 3\n        manufacturer,                # 4\n        magn_field_strng,            # 5\n        tr,                          # 6\n        te,                          # 7\n    )\n                )","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-04T23:58:28.671956Z","iopub.execute_input":"2026-09-04T23:58:28.672422Z","iopub.status.idle":"2026-09-04T23:58:28.694468Z","shell.execute_reply.started":"2026-09-04T23:58:28.67238Z","shell.execute_reply":"2026-09-04T23:58:28.693357Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The first series in the axial projection without fat suppression and fluid sensitivity (index 5678)","metadata":{}},{"cell_type":"code","source":"print_series_images(feat_by_series_uid(train_series_df, for_print_df['SeriesInstanceUID'][5678]))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-04T23:58:28.69579Z","iopub.execute_input":"2026-09-04T23:58:28.696501Z","iopub.status.idle":"2026-09-04T23:58:35.738431Z","shell.execute_reply.started":"2026-09-04T23:58:28.696453Z","shell.execute_reply":"2026-09-04T23:58:35.736953Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The next series uses the same parameters. Note that, despite the similar parameters, the color gamut is significantly different. It's obvious that these images have different weighting. Unfortunately, it wasn't provided in the train_series file, so we'll have to decipher the dicom file metadata.","metadata":{}},{"cell_type":"code","source":"print_series_images(feat_by_series_uid(train_series_df, for_print_df['SeriesInstanceUID'][5679]))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-04T23:58:35.740169Z","iopub.execute_input":"2026-09-04T23:58:35.740711Z","iopub.status.idle":"2026-09-04T23:58:42.116665Z","shell.execute_reply.started":"2026-09-04T23:58:35.740656Z","shell.execute_reply":"2026-09-04T23:58:42.115457Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Well, here's another series of this study, with the same parameters. This one is more similar to the first series, but there are still some differences","metadata":{}},{"cell_type":"code","source":"print_series_images(feat_by_series_uid(train_series_df, for_print_df['SeriesInstanceUID'][5680]))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-04T23:58:42.117952Z","iopub.execute_input":"2026-09-04T23:58:42.118406Z","iopub.status.idle":"2026-09-04T23:58:48.49289Z","shell.execute_reply.started":"2026-09-04T23:58:42.118369Z","shell.execute_reply":"2026-09-04T23:58:48.491647Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Let's take another look at the same projection, but with fat suppression and fluid emphasis. It's clearly different from all three previous ones...","metadata":{}},{"cell_type":"code","source":"print_series_images(feat_by_series_uid(train_series_df, for_print_df['SeriesInstanceUID'][5683]))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-04T23:58:48.494208Z","iopub.execute_input":"2026-09-04T23:58:48.494562Z","iopub.status.idle":"2026-09-04T23:58:55.456262Z","shell.execute_reply.started":"2026-09-04T23:58:48.494535Z","shell.execute_reply":"2026-09-04T23:58:55.454774Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# DICOM metadata","metadata":{}},{"cell_type":"markdown","source":"Next, we'll move on to examining a broader set of metadata that we'll extract from DICOM files. First, let's determine what this metadata set is and whether it's the same for all images","metadata":{}},{"cell_type":"code","source":"import polars as pl\nimport random\nfrom tqdm.notebook import tqdm\nfrom tqdm.contrib.concurrent import thread_map\nfrom tqdm.asyncio import tqdm as async_tqdm","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-04T23:58:55.457822Z","iopub.execute_input":"2026-09-04T23:58:55.458725Z","iopub.status.idle":"2026-09-04T23:58:56.294377Z","shell.execute_reply.started":"2026-09-04T23:58:55.458654Z","shell.execute_reply":"2026-09-04T23:58:56.29331Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"%%time\nfile_list = glob.glob(TRAIN_PATH + '/**/*.dcm', recursive=True)\nprint(f'Total train images - {len(file_list):,d}')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-04T23:58:56.2958Z","iopub.execute_input":"2026-09-04T23:58:56.296198Z","iopub.status.idle":"2026-09-05T00:10:25.846514Z","shell.execute_reply.started":"2026-09-04T23:58:56.29616Z","shell.execute_reply":"2026-09-05T00:10:25.845209Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def attr_list(dcm_file):\n\n    dcm_ds = pydicom.dcmread(dcm_file, stop_before_pixels=True)\n    df = pl.DataFrame({'file':dcm_file, 'feat_list':[dcm_ds.dir()], 'hash_fl':hash(tuple(dcm_ds.dir()))})\n    meta_data.vstack(df, in_place=True)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-05T00:10:25.848006Z","iopub.execute_input":"2026-09-05T00:10:25.848411Z","iopub.status.idle":"2026-09-05T00:10:25.855025Z","shell.execute_reply.started":"2026-09-05T00:10:25.84836Z","shell.execute_reply":"2026-09-05T00:10:25.853897Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"rate = 0.05 # if you want to take only sample \n\nmeta_data = pl.DataFrame(schema={\n    'file':pl.String, \n    'feat_list':pl.List(pl.String),\n    'hash_fl':pl.Int64\n})\n\nif rate < 1:\n    k = int(len(file_list) * rate)\n    file_sample = random.sample(file_list, k)\nelse:\n    file_sample = file_list\n    \n_ = thread_map(attr_list, [dcm_file for dcm_file in file_sample],\n               tqdm_class=async_tqdm, max_workers=32, total=len(file_sample), ncols=100)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-05T00:10:25.856284Z","iopub.execute_input":"2026-09-05T00:10:25.856551Z","iopub.status.idle":"2026-09-05T00:13:19.611458Z","shell.execute_reply.started":"2026-09-05T00:10:25.856529Z","shell.execute_reply":"2026-09-05T00:13:19.610199Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# let's get some stat of features sets\nfeat_sets_stat = (\n    meta_data.group_by('feat_list', 'hash_fl').agg(pl.col('hash_fl').len().alias('count'))\n    .sort('count', descending=True)\n    .with_row_index(name=\"index\")\n    .with_columns(\n        pl.col('feat_list').list.len().alias('feat_len'), \n        (pl.col(\"count\").cum_sum() / pl.col(\"count\").sum()).alias(\"cum_percentage\")\n    )\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-05T00:13:19.613232Z","iopub.execute_input":"2026-09-05T00:13:19.613564Z","iopub.status.idle":"2026-09-05T00:13:20.007518Z","shell.execute_reply.started":"2026-09-05T00:13:19.613529Z","shell.execute_reply":"2026-09-05T00:13:20.00625Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"feat_sets_stat","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-05T00:13:20.008725Z","iopub.execute_input":"2026-09-05T00:13:20.009255Z","iopub.status.idle":"2026-09-05T00:13:20.027343Z","shell.execute_reply.started":"2026-09-05T00:13:20.009225Z","shell.execute_reply":"2026-09-05T00:13:20.026376Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(f'Number of metadata list variants - {feat_sets_stat.height}')\nprint(f'Number of list length options - {feat_sets_stat['feat_len'].unique().len()}')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-05T00:13:20.028481Z","iopub.execute_input":"2026-09-05T00:13:20.028805Z","iopub.status.idle":"2026-09-05T00:13:20.054733Z","shell.execute_reply.started":"2026-09-05T00:13:20.02878Z","shell.execute_reply":"2026-09-05T00:13:20.053676Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"To speed things up, I took a random sample of 5% and got 83 metadata set variants. Someone didn't pay much attention to consistency... If you don't mind the time, you can set `rate=1`, and in 45 minutes you'll get all 86 variants. Not much more )","metadata":{}},{"cell_type":"code","source":"feat_sets_stat['feat_len'].describe()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-05T00:13:20.056305Z","iopub.execute_input":"2026-09-05T00:13:20.056751Z","iopub.status.idle":"2026-09-05T00:13:20.082171Z","shell.execute_reply.started":"2026-09-05T00:13:20.056723Z","shell.execute_reply":"2026-09-05T00:13:20.081181Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The sets of props also vary greatly in length: a minimum of 30, a maximum of 74. Therefore, these sets will intersect at a maximum of 30 elements. Let's check...","metadata":{}},{"cell_type":"code","source":"def feat_set_intersection(df):\n    common_feat = list(set.intersection(*map(set, df)))\n    \n    print(f\"The number of common features for all {df.shape[0]} lists is: {len(common_feat)}\")\n    print(\"List of common features:\", common_feat)\n    \nfeat_set_intersection(feat_sets_stat['feat_list'])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-05T00:13:20.08343Z","iopub.execute_input":"2026-09-05T00:13:20.083782Z","iopub.status.idle":"2026-09-05T00:13:20.093114Z","shell.execute_reply.started":"2026-09-05T00:13:20.083745Z","shell.execute_reply":"2026-09-05T00:13:20.092122Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Only these 28 features are guaranteed to be present for all series.\n\nPerhaps it's also worth looking at the quantitative ratio of all the sets. For this, it's best to create a set for all files, but let's trust the statistics of a random selection.","metadata":{}},{"cell_type":"code","source":"feat_sets_stat.select(['index', 'cum_percentage']).plot.line(x='index', y='cum_percentage')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-05T00:13:20.094574Z","iopub.execute_input":"2026-09-05T00:13:20.09503Z","iopub.status.idle":"2026-09-05T00:13:22.361948Z","shell.execute_reply.started":"2026-09-05T00:13:20.095005Z","shell.execute_reply":"2026-09-05T00:13:22.360755Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The first thirty sets represent 95% of the total sample size. Let's check if they intersect in more elements","metadata":{}},{"cell_type":"code","source":"feat_set_intersection(feat_sets_stat['feat_list'][:30])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-05T00:13:22.36677Z","iopub.execute_input":"2026-09-05T00:13:22.367261Z","iopub.status.idle":"2026-09-05T00:13:22.374188Z","shell.execute_reply.started":"2026-09-05T00:13:22.367229Z","shell.execute_reply":"2026-09-05T00:13:22.373026Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"No, the same 28 features. Let's try to group them by their usefulness for further pre-processing","metadata":{}},{"cell_type":"markdown","source":"1. **Geometry and resampling (may be required)**\n- PixelSpacing: Pixel size. Needed for resampling along the X and Y axes.\n- SliceThickness: Slice thickness. Needed for resampling along the Z axis when assembling the 3D volume.\n- ImagePositionPatient: Coordinates of the upper left corner of the image. Allows for proper slice ordering.\n- ImageOrientationPatient: Coordinate direction vectors. Needed for projection verification, but we've already defined one.\n- SliceLocation: Relative slice position. Used as an alternative or supplement to slice sorting.\n2. **Brightness and contrast normalization (not needed, as these nuances will be taken into account when reading using pixel_array)**\n- WindowCenter and WindowWidth: Center and width of the visualization window\n- BitsAllocated: Number of bits allocated per pixel\n- BitsStored: Number of bits actually used\n- HighBit: The most significant bit, which determines the sign or significant part\n- PixelRepresentation: Determines whether the data is signed (1) or unsigned (0)\n3. **Data format and dimensions (may be needed)**\n- Rows and Columns: Height and width of the matrix in pixels (for resizing)\n- SamplesPerPixel: Number of color channels (not needed, the dataset only has one channel)\n- PhotometricInterpretation: Color model. Specifies how to interpret the channels (the train data only shows MONOCHROME2, but MONOCHROME1 should also be considered in case it appears in the test data).\n4. **Data Filtering and Grouping**\n- Modality: Study Type (train only shows MR)\n- ImageType: Image Type. Allows you to filter out secondary (DERIVED) or service images, leaving only original (ORIGINAL) images.\n- SeriesInstanceUID: Series ID (not needed, the structure is organized by series)\n- InstanceNumber: Image number in the series. I used this above to sort images during visualization (ImagePositionPatient is said to be more reliable).\n5. **Less useful preprocessing attributes:** PatientID, StudyInstanceUID, SOPInstanceUID, SOPClassUID, FrameOfReferenceUID, Manufacturer, ManufacturerModelName, AcquisitionNumber, and SeriesNumber. Manufacturer and ManufacturerModelName can be used for specific equipment, as we have images from different MRI machines. However, this isn't for preprocessing, but for troubleshooting.","metadata":{}},{"cell_type":"markdown","source":"But I'd like to see something else... We've already noted the different weightings above, which can be worked with using `Repetition Time`, `Echo Time`, and `ScanningSequence`. Let's see how fully they are represented in the dataset.","metadata":{}},{"cell_type":"code","source":"def repr_feats(feat_list: list):\n    \n    feat_extr = [pl.col('feat_list').list.contains(name) for name in feat_list]\n    intercept_slope_sets = feat_sets_stat.filter(pl.all_horizontal(feat_extr))\n\n    print(f'Number of sets with this feature - {intercept_slope_sets.shape[0]}')\n    print(f'Share of the total number of images in the sample - {(intercept_slope_sets['count'].sum() / feat_sets_stat['count'].sum()):.1%}')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-05T00:13:22.375663Z","iopub.execute_input":"2026-09-05T00:13:22.37605Z","iopub.status.idle":"2026-09-05T00:13:22.399693Z","shell.execute_reply.started":"2026-09-05T00:13:22.376012Z","shell.execute_reply":"2026-09-05T00:13:22.398508Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"repr_feats(['RepetitionTime', 'EchoTime'])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-05T00:13:22.401261Z","iopub.execute_input":"2026-09-05T00:13:22.401813Z","iopub.status.idle":"2026-09-05T00:13:22.433383Z","shell.execute_reply.started":"2026-09-05T00:13:22.401776Z","shell.execute_reply":"2026-09-05T00:13:22.432223Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"repr_feats(['ScanningSequence'])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-05T00:13:22.434672Z","iopub.execute_input":"2026-09-05T00:13:22.435334Z","iopub.status.idle":"2026-09-05T00:13:22.44169Z","shell.execute_reply.started":"2026-09-05T00:13:22.43528Z","shell.execute_reply":"2026-09-05T00:13:22.440664Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Very good! These features are available for almost all series. Let's also grab the `SeriesDescription`; it often contains a lot of useful information, but it needs to be parsed","metadata":{}},{"cell_type":"code","source":"repr_feats(['SeriesDescription'])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-05T00:13:22.442953Z","iopub.execute_input":"2026-09-05T00:13:22.443417Z","iopub.status.idle":"2026-09-05T00:13:22.462291Z","shell.execute_reply.started":"2026-09-05T00:13:22.443364Z","shell.execute_reply":"2026-09-05T00:13:22.461169Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"We're trying to collect metadata. We could try doing it on a sample, but I decided to be savvy and did it on all 820,000 images! )","metadata":{}},{"cell_type":"code","source":"# let's stop at this list, we have to stop somewhere )\nfeat_list = [\n    'Manufacturer',\n    'ManufacturerModelName',\n    'PixelSpacing', \n    'SliceThickness', \n    'ImagePositionPatient', \n    'Rows',\n    'Columns',\n    'PhotometricInterpretation',\n    'InstanceNumber',\n    # not all fields below are available for all images\n    'RepetitionTime',\n    'EchoTime',\n    'ScanningSequence',\n    'SeriesDescription',\n]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-05T00:13:22.46363Z","iopub.execute_input":"2026-09-05T00:13:22.464025Z","iopub.status.idle":"2026-09-05T00:13:22.477904Z","shell.execute_reply.started":"2026-09-05T00:13:22.463984Z","shell.execute_reply":"2026-09-05T00:13:22.476882Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# take some of the data from train_series\next_data = (train_series_df\n            [['SeriesInstanceUID', 'Fluid_Sensitive', 'Fat_Suppression', 'Anatomical_Plane']]\n            .set_index('SeriesInstanceUID').to_dict(orient='index')\n           )","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-05T00:13:22.479164Z","iopub.execute_input":"2026-09-05T00:13:22.479645Z","iopub.status.idle":"2026-09-05T00:13:22.555105Z","shell.execute_reply.started":"2026-09-05T00:13:22.479606Z","shell.execute_reply":"2026-09-05T00:13:22.553767Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# about an hour...\ndef extract_metadata(file_path, features):\n    \n    global series_uid, ext_feat\n    \n    ds = pydicom.dcmread(file_path, stop_before_pixels=True)\n        \n    file_metadata  = {\n        'path': file_path,\n    }\n\n    cur_series_uid = file_path.split('/')[-2]\n    \n    if series_uid == cur_series_uid:\n        file_metadata |= ext_feat\n    else:\n        series_uid = cur_series_uid\n        ext_feat = ext_data[cur_series_uid]\n        \n        file_metadata |= ext_feat\n    \n    for feat in features:\n        file_metadata[feat]  =  getattr(ds, feat, None)\n    \n    metadata_lst.append(file_metadata)    \n\n\nmetadata_lst = []\nseries_uid = ''\next_feat = {}\n\n_ = thread_map(\n    lambda path: extract_metadata(path, feat_list), [dcm_file for dcm_file in file_list],\n    tqdm_class = async_tqdm, \n    max_workers = 32,\n    total = len(file_list),\n    ncols = 100\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-05T00:13:22.556733Z","iopub.execute_input":"2026-09-05T00:13:22.557131Z","iopub.status.idle":"2026-09-05T00:56:35.037861Z","shell.execute_reply.started":"2026-09-05T00:13:22.557095Z","shell.execute_reply":"2026-09-05T00:56:35.036693Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"metadata_df = pd.DataFrame(metadata_lst)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-05T00:56:35.092747Z","iopub.execute_input":"2026-09-05T00:56:35.093717Z","iopub.status.idle":"2026-09-05T00:56:41.538113Z","shell.execute_reply.started":"2026-09-05T00:56:35.093675Z","shell.execute_reply":"2026-09-05T00:56:41.53702Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"You can see the distribution of categorical features","metadata":{}},{"cell_type":"code","source":"\ndef feat_values_rate(name: str, prc_sign=True):\n    if prc_sign:\n        return (metadata_df[name].value_counts(normalize=True) * 100).round(1).astype(str) + '%'\n    else:\n        return (metadata_df[name].value_counts(normalize=True) * 100).round(3)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-05T00:56:41.539865Z","iopub.execute_input":"2026-09-05T00:56:41.540897Z","iopub.status.idle":"2026-09-05T00:56:41.547545Z","shell.execute_reply.started":"2026-09-05T00:56:41.540861Z","shell.execute_reply":"2026-09-05T00:56:41.546451Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"feat_values_rate('Manufacturer')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-05T00:56:41.548856Z","iopub.execute_input":"2026-09-05T00:56:41.549631Z","iopub.status.idle":"2026-09-05T00:56:41.77515Z","shell.execute_reply.started":"2026-09-05T00:56:41.549602Z","shell.execute_reply":"2026-09-05T00:56:41.77413Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Siemens, Philips, GE - undisputed leaders","metadata":{}},{"cell_type":"code","source":"feat_values_rate('SeriesDescription')[:10]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-05T00:56:41.776423Z","iopub.execute_input":"2026-09-05T00:56:41.777352Z","iopub.status.idle":"2026-09-05T00:56:41.992158Z","shell.execute_reply.started":"2026-09-05T00:56:41.777309Z","shell.execute_reply":"2026-09-05T00:56:41.991253Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"13 percent is a mockery! And there's no data for another 5 percent. Otherwise, you can check it out occasionally if you have any doubts about the weightening, etc. Parsing this unbridled creativity is best done only as a last resort )","metadata":{}},{"cell_type":"code","source":"feat_values_rate('ScanningSequence')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-05T00:56:41.993678Z","iopub.execute_input":"2026-09-05T00:56:41.994054Z","iopub.status.idle":"2026-09-05T00:56:42.175273Z","shell.execute_reply.started":"2026-09-05T00:56:41.994028Z","shell.execute_reply":"2026-09-05T00:56:42.174096Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"3.4% is \"Unknown\", that's about 28k images","metadata":{}},{"cell_type":"markdown","source":"Since I plan to use the `Anatomical_Plane`, `Fat_Suppression`, `ScanningSequence`, `RepetitionTime`, and `EchoTime` features during preprocessing to distribute images across subnets, let's take a closer look at their contents","metadata":{}},{"cell_type":"code","source":"metadata_df.loc[:, 'ScanningSequence'] = metadata_df['ScanningSequence'].astype(str)\nres = (metadata_df\n       .groupby(['Anatomical_Plane', 'Fat_Suppression', 'ScanningSequence'])\n       .agg({'RepetitionTime': ['min', 'max'], 'EchoTime': ['min', 'max', 'size']})\n      )\nres.style.format(\n    {\n        ('RepetitionTime', 'min'): '{:,.2f}'.format,\n        ('RepetitionTime', 'max'): '{:,.2f}'.format,\n        ('EchoTime', 'min'): '{:,.2f}'.format,\n        ('EchoTime', 'max'): '{:,.2f}'.format,\n        ('EchoTime', 'size'): '{:,d}'.format,\n    }\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-05T00:56:42.176866Z","iopub.execute_input":"2026-09-05T00:56:42.177831Z","iopub.status.idle":"2026-09-05T00:56:42.6879Z","shell.execute_reply.started":"2026-09-05T00:56:42.1778Z","shell.execute_reply":"2026-09-05T00:56:42.686832Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Well, we'll analyze this in more detail during preprocessing, but for now, let's note the missing data for approximately 28,000 images. We'll also consider how this can be filled by parsing SeriesDescription.","metadata":{}},{"cell_type":"code","source":"metadata_df[metadata_df['ScanningSequence']=='None'][['Fat_Suppression', 'SeriesDescription']].value_counts()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-05T00:56:42.689427Z","iopub.execute_input":"2026-09-05T00:56:42.689845Z","iopub.status.idle":"2026-09-05T00:56:42.86766Z","shell.execute_reply.started":"2026-09-05T00:56:42.689818Z","shell.execute_reply":"2026-09-05T00:56:42.866733Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"That's not bad! At least the missing `ScanningSequence`, `RepetitionTime`, and `EchoTime` don't overlap with the missing `SeriesDescription`","metadata":{}},{"cell_type":"code","source":"# we need to convert the MultiValue type to a list, before saving\nmv_cols = ['PixelSpacing', 'ImagePositionPatient']\n\nfor col in mv_cols:\n    metadata_df[col] = metadata_df[col].apply(\n        lambda x: list(x) if x is not None and hasattr(x, '__iter__') else x\n    )\n\n# let's save it for the future\nmetadata_df.to_parquet('metadata.parquet', index=False)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-05T00:56:42.869598Z","iopub.execute_input":"2026-09-05T00:56:42.869985Z","iopub.status.idle":"2026-09-05T00:56:49.942309Z","shell.execute_reply.started":"2026-09-05T00:56:42.869958Z","shell.execute_reply":"2026-09-05T00:56:49.941384Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"We've gotten a bit deep, time to move on to pre-processing. Stay tuned!","metadata":{}}]}