{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Intro\n<font size=3.5>\n\nThis notebook contains an exploratory analysis of the dataset from the RSNA Breast Cancer Detection challenge. The organizers created a dataset of mammographies toghether with a table that contains information related to the patients, images and the cancer (where it's present). The notebook goes more or less from general to particular. It starts with simple statistics that is counting different informations given in the csv, and then goes in more detail analysis especially related to the images. \n    \nIt does not explain the code or how the plots are generated. I believe code should be self-explanatory :D , but if you have questions feel free to ask. I'll try to explain as best I can.\n    \nThis is how the notebook is structured. Of course it starts with the library imports :))\n</font>\n\n\n* [Libraries import](#import)\n* [General overview](#overview)\n* [In more depth](#detailed)\n    * [Types of views](#views)\n    * [Positive cases](#positive)\n    * [Implants](#implants)\n* [ToDo List](#todo)\n* [Sources](#sources)\n    \n \n","metadata":{"_kg_hide-input":false}},{"cell_type":"code","source":"! pip install -qU \"python-gdcm\" pydicom pylibjpeg \"opencv-python-headless\"","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2023-01-22T17:37:23.691864Z","iopub.execute_input":"2023-01-22T17:37:23.692324Z","iopub.status.idle":"2023-01-22T17:37:35.397411Z","shell.execute_reply.started":"2023-01-22T17:37:23.692289Z","shell.execute_reply":"2023-01-22T17:37:35.39598Z"},"_kg_hide-input":true,"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Libraries import\n<a id=\"import\"></a>","metadata":{}},{"cell_type":"code","source":"import pathlib as pt\nimport random\n\nimport matplotlib.pyplot as plt\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport seaborn as sns\nimport pydicom\n\n\nfrom tqdm import tqdm\n\nsns.set_style('darkgrid')","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-01-22T17:37:45.456282Z","iopub.execute_input":"2023-01-22T17:37:45.456738Z","iopub.status.idle":"2023-01-22T17:37:45.464357Z","shell.execute_reply.started":"2023-01-22T17:37:45.456697Z","shell.execute_reply":"2023-01-22T17:37:45.46329Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Simple overview\n<a id='overview'></a>\n\n\n<font size=3.5>\nIn this part I'll be looking at each piece of information we are provided with by the organizers and see what can I extract that would be helpful to me in order to better understand the data and to better design an AI algorithm later on. Usually in this first part I just count things which in some cases it might be difficult than you might think, especially when part of one row is duplicated. For example <strong> site_id </strong> and <strong> patient_id </strong> have the same value for multiple rows because there are more than one image associated with a patient. <strong> One must be careful not to count more than once the same thing. </strong>\n</font>","metadata":{}},{"cell_type":"code","source":"data_dir = pt.Path(\"/kaggle/input/rsna-breast-cancer-detection/\")\ntrain_csv = data_dir/\"train.csv\"\ntrain_df = pd.read_csv(train_csv, dtype=str)\ntrain_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-01-22T16:41:33.982657Z","iopub.execute_input":"2023-01-22T16:41:33.983102Z","iopub.status.idle":"2023-01-22T16:41:34.122379Z","shell.execute_reply.started":"2023-01-22T16:41:33.983068Z","shell.execute_reply":"2023-01-22T16:41:34.121083Z"},"jupyter":{"source_hidden":true},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"images_path = data_dir/\"train_images\"\ntrain_df['img_path'] = \"\"\nfor idx, row in tqdm(train_df.iterrows(), total=len(train_df), unit='row'):\n    train_df.iloc[idx]['img_path'] = str(images_path/row[\"patient_id\"]/f'{row[\"image_id\"]}.dcm')","metadata":{"execution":{"iopub.status.busy":"2023-01-22T16:41:34.125281Z","iopub.execute_input":"2023-01-22T16:41:34.125798Z","iopub.status.idle":"2023-01-22T16:41:43.268875Z","shell.execute_reply.started":"2023-01-22T16:41:34.125748Z","shell.execute_reply":"2023-01-22T16:41:43.267793Z"},"jupyter":{"source_hidden":true},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-01-22T16:41:43.270493Z","iopub.execute_input":"2023-01-22T16:41:43.271742Z","iopub.status.idle":"2023-01-22T16:41:43.293655Z","shell.execute_reply.started":"2023-01-22T16:41:43.271691Z","shell.execute_reply":"2023-01-22T16:41:43.292313Z"},"jupyter":{"source_hidden":true},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.info()","metadata":{"execution":{"iopub.status.busy":"2023-01-22T16:41:43.296114Z","iopub.execute_input":"2023-01-22T16:41:43.296616Z","iopub.status.idle":"2023-01-22T16:41:43.347942Z","shell.execute_reply.started":"2023-01-22T16:41:43.296568Z","shell.execute_reply":"2023-01-22T16:41:43.347049Z"},"_kg_hide-input":true,"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font size=3.5>\n    Columns with missing values: <strong> age, BIRADS, density </strong>\n</font>","metadata":{}},{"cell_type":"code","source":"nr_patients = len(train_df['patient_id'].unique())\nprint(f\"# unique patients {nr_patients}\")\n\nnr_images = len(train_df['image_id'].unique())\nprint(f\"# unique images {nr_images}\")\n\ntrain_df['site_id'].value_counts(dropna=False)","metadata":{"execution":{"iopub.status.busy":"2023-01-22T16:41:43.349424Z","iopub.execute_input":"2023-01-22T16:41:43.34978Z","iopub.status.idle":"2023-01-22T16:41:43.375814Z","shell.execute_reply.started":"2023-01-22T16:41:43.349749Z","shell.execute_reply":"2023-01-22T16:41:43.374645Z"},"jupyter":{"source_hidden":true},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"_df = train_df[['site_id', 'patient_id']].drop_duplicates().groupby('site_id').count().reset_index()\nfig, ax = plt.subplots(1, 2, figsize=(16,5))\nsns.barplot(ax=ax[0], data=_df, x='site_id', y='patient_id', palette=['silver', 'goldenrod'])\nax[0].bar_label(ax[0].containers[0])\nax[0].set_xlabel(\"Site ID\")\nax[0].set_ylabel(\"Nr patients\")\nax[0].set_title(\"Nr unique patients in each site\")\n\n_df = train_df[['site_id', 'image_id']].drop_duplicates().groupby('site_id').count().reset_index()\nsns.barplot(ax=ax[1], data=_df, x='site_id', y='image_id', palette=['silver', 'goldenrod'])\nax[1].bar_label(ax[1].containers[0])\nax[1].set_xlabel(\"Site ID\")\nax[1].set_ylabel(\"Nr images\")\nax[1].set_title(\"Nr unique images in each site\")","metadata":{"execution":{"iopub.status.busy":"2023-01-22T16:41:43.377664Z","iopub.execute_input":"2023-01-22T16:41:43.378007Z","iopub.status.idle":"2023-01-22T16:41:43.825135Z","shell.execute_reply.started":"2023-01-22T16:41:43.377976Z","shell.execute_reply":"2023-01-22T16:41:43.823976Z"},"jupyter":{"source_hidden":true},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"target_ratio = {\"patient_id\":[], \"image_id\": []}\n\n_df = train_df[['site_id', 'patient_id', 'cancer']].drop_duplicates().groupby('cancer').count().reset_index()\ntarget_ratio['patient_id'] = _df['patient_id'].to_list()\n\nfig, ax = plt.subplots(1, 2, figsize=(16,5))\nsns.barplot(ax=ax[0], data=_df, x='cancer', y='patient_id', palette=['cornflowerblue', 'darkorange'])\nax[0].bar_label(ax[0].containers[0])\nax[0].set_xlabel(\"Cancer: Negative/Positive\")\nax[0].set_ylabel(\"Nr patients\")\nax[0].set_title(\"Nr patients by label\")\n\n_df = train_df[['site_id', 'patient_id', 'image_id', 'cancer']].drop_duplicates().groupby('cancer').count().reset_index()\ntarget_ratio['image_id'] = _df['image_id'].to_list()\n\n# display(_df)\nsns.barplot(ax=ax[1], data=_df, x='cancer', y='image_id', palette=['steelblue', 'goldenrod'])\nax[1].bar_label(ax[1].containers[0])\nax[1].set_xlabel(\"Cancer: Negative/Positive\")\nax[1].set_ylabel(\"Nr images\")\nax[1].set_title(\"Nr images by label\")","metadata":{"execution":{"iopub.status.busy":"2023-01-22T16:41:43.827021Z","iopub.execute_input":"2023-01-22T16:41:43.827356Z","iopub.status.idle":"2023-01-22T16:41:44.266962Z","shell.execute_reply.started":"2023-01-22T16:41:43.827327Z","shell.execute_reply":"2023-01-22T16:41:44.266151Z"},"jupyter":{"source_hidden":true},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pat_neg2pos = target_ratio['patient_id'][0] // target_ratio['patient_id'][1]\nimg_neg2pos = target_ratio['image_id'][0] // target_ratio['image_id'][1]\nprint(f\"Patient wise: neg:pos = {pat_neg2pos}:1\")\nprint(f\"Image wise: neg:pos = {img_neg2pos}:1\")","metadata":{"execution":{"iopub.status.busy":"2023-01-22T16:41:44.27042Z","iopub.execute_input":"2023-01-22T16:41:44.270977Z","iopub.status.idle":"2023-01-22T16:41:44.276464Z","shell.execute_reply.started":"2023-01-22T16:41:44.270942Z","shell.execute_reply":"2023-01-22T16:41:44.275695Z"},"jupyter":{"source_hidden":true},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"_df = train_df[['site_id', 'patient_id', 'implant']].drop_duplicates().groupby('implant').count().reset_index()\n# display(_df)\nfig, ax = plt.subplots(2, 2, figsize=(20,12))\n\n\nsns.barplot(ax=ax[0][0], data=_df, x='implant', y='patient_id', palette=['royalblue', 'coral'])\nax[0][0].bar_label(ax[0][0].containers[0])\nax[0][0].set_xlabel(\"Implant\")\nax[0][0].set_ylabel(\"Nr patients\")\nax[0][0].set_title(\"Nr patients with or without an implant\")\n\n\n_df = train_df[['site_id', 'patient_id', 'biopsy']].drop_duplicates().groupby('biopsy').count().reset_index()\n# display(_df)\nsns.barplot(ax=ax[0][1], data=_df, x='biopsy', y='patient_id', palette=['seagreen', 'maroon'])\nax[0][1].bar_label(ax[0][1].containers[0])\nax[0][1].set_xlabel(\"Biopsy\")\nax[0][1].set_ylabel(\"Nr patients\")\nax[0][1].set_title(\"Nr patients with or without a biopsy\")\n\n\n\n_df = train_df[['site_id', 'patient_id', 'density']].drop_duplicates().dropna().groupby('density').count().reset_index()\n# display(_df)\nsns.barplot(ax=ax[1][0], data=_df, x='density', y='patient_id', color='cornflowerblue')\nax[1][0].bar_label(ax[1][0].containers[0])\nax[1][0].set_xlabel(\"Density\")\nax[1][0].set_ylabel(\"Nr patients\")\nax[1][0].set_title(\"Density frequency\")\n\n\n_df = train_df[['site_id', 'patient_id', 'BIRADS']].drop_duplicates().dropna().groupby('BIRADS').count().reset_index()\n# display(_df)\nsns.barplot(ax=ax[1][1], data=_df, x='BIRADS', y='patient_id', color='cornflowerblue')\nax[1][1].bar_label(ax[1][1].containers[0])\nax[1][1].set_xlabel(\"BIRADS\")\nax[1][1].set_ylabel(\"Nr patients\")\nax[1][1].set_title(\"BIRADS frequency\")","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-01-22T17:38:26.496272Z","iopub.execute_input":"2023-01-22T17:38:26.49674Z","iopub.status.idle":"2023-01-22T17:38:27.325367Z","shell.execute_reply.started":"2023-01-22T17:38:26.496699Z","shell.execute_reply":"2023-01-22T17:38:27.324193Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df['machine_id'] = train_df['machine_id'].astype(int)\n_df = train_df[['image_id', 'machine_id']].drop_duplicates().groupby('machine_id').count().sort_values(by='machine_id').reset_index()\n# display(_df)\nfig, ax = plt.subplots(1, 1, figsize=(15,5))\nsns.barplot(ax=ax, data=_df, x='machine_id', y='image_id', color='cornflowerblue')\nax.bar_label(ax.containers[0])\nax.set_xlabel(\"Machine ID\")\nax.set_ylabel(\"Nr images\")\nax.set_title(\"Nr images taken with each machine\")","metadata":{"execution":{"iopub.status.busy":"2023-01-22T16:41:45.872647Z","iopub.execute_input":"2023-01-22T16:41:45.873011Z","iopub.status.idle":"2023-01-22T16:41:46.262911Z","shell.execute_reply.started":"2023-01-22T16:41:45.872958Z","shell.execute_reply":"2023-01-22T16:41:46.26174Z"},"jupyter":{"source_hidden":true},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"_df = train_df['patient_id'].value_counts().reset_index()\n# display(_df)\nfig, ax = plt.subplots(1, 1, figsize=(10,5))\nsns.countplot(data=_df, x='patient_id', ax=ax, palette='Blues_r')\nfor bars in ax.containers:\n    ax.bar_label(bars)\nax.set_xlabel(\"Nr images/patient\")\nax.set_ylabel(\"Nr patients\")\nax.set_title(\"Nr images for each patient\")","metadata":{"execution":{"iopub.status.busy":"2023-01-22T16:41:46.264345Z","iopub.execute_input":"2023-01-22T16:41:46.264699Z","iopub.status.idle":"2023-01-22T16:41:46.643433Z","shell.execute_reply.started":"2023-01-22T16:41:46.264666Z","shell.execute_reply":"2023-01-22T16:41:46.642534Z"},"jupyter":{"source_hidden":true},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"_df = train_df[['site_id', 'patient_id', 'age']].drop_duplicates().dropna().copy()\n_df['age'] = _df['age'].astype(int)\nfig, ax = plt.subplots(1, 1, figsize=(10, 5))\nsns.histplot(_df, x='age', bins=60, ax=ax)\nax.set_title(\"Patient age distribution\")","metadata":{"execution":{"iopub.status.busy":"2023-01-22T16:41:46.644725Z","iopub.execute_input":"2023-01-22T16:41:46.645233Z","iopub.status.idle":"2023-01-22T16:41:47.093132Z","shell.execute_reply.started":"2023-01-22T16:41:46.6452Z","shell.execute_reply":"2023-01-22T16:41:47.092164Z"},"jupyter":{"source_hidden":true},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"_df = train_df[['site_id', 'patient_id', 'age', 'cancer']].drop_duplicates().dropna().copy()\n_df = _df.loc[_df['cancer'] == '1', :].copy()\n_df['age'] = _df['age'].astype(int)\nfig, ax = plt.subplots(1, 1, figsize=(10, 5))\nsns.histplot(_df, x='age', bins=60, ax=ax)\nax.set_title(\"Cancer patient age distribution\")","metadata":{"execution":{"iopub.status.busy":"2023-01-22T16:41:47.097051Z","iopub.execute_input":"2023-01-22T16:41:47.09744Z","iopub.status.idle":"2023-01-22T16:41:47.532521Z","shell.execute_reply.started":"2023-01-22T16:41:47.097404Z","shell.execute_reply":"2023-01-22T16:41:47.531348Z"},"jupyter":{"source_hidden":true},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Insights\n\n<font size=3.5>\nThese are the insight that I gained and what might mean for future development.\n\n \n* heavily imbalanced dataset => to apply techniques that can reduce/alleviate this imbalance (ex. under/over sampling, data augmentation, weighted sampling, loss weights, loss function specialized that overcome this issue etc)\n* implant, biopsy, density and BIRADS are provided only for train => useful only for debugging and gaining an a clinical understanding of the dataset\n* multiple scanners used => pixel intensities might be very different between scanners => to be carefull at normalization\n* multiple images per patient; most common 4 images/patient => how to use 4+ images during training, and during testing\n* patients with cancer tent to be older, 50+\n</font>","metadata":{}},{"cell_type":"markdown","source":"# In detail view\n<a id='detailed'></a>","metadata":{}},{"cell_type":"markdown","source":"## Types of views\n<a id='views'></a>\n","metadata":{}},{"cell_type":"code","source":"train_df['view'].value_counts(dropna=False)\n","metadata":{"execution":{"iopub.status.busy":"2023-01-22T16:41:47.534352Z","iopub.execute_input":"2023-01-22T16:41:47.534726Z","iopub.status.idle":"2023-01-22T16:41:47.546084Z","shell.execute_reply.started":"2023-01-22T16:41:47.534692Z","shell.execute_reply":"2023-01-22T16:41:47.544949Z"},"_kg_hide-input":true,"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font size=\"3.5\">As we can see from the column 'view' there are 2 predominant types of views present in this dataset: **cranio-caudal (CC)** and **mediolateral-oblique (MLO)**. These are the standard views that are taken in a routine screening. They are both images taken from above (cranio-caudal = \"head to tail\"), but the latter is at an angle and includes some pectoral (chest) tissue. Another thing to remember about these 2 views is that they are not orthogonal which means if you want to see the same tumor in both images one must think a little bit and have some breast anatomy knowledge.</font>\n\n<div style=\"width:100%;text-align: center;\"> <img align=middle src=\"https://www.imaginis.com/files/media/transfer/img/breasthealth/mammography_imaging/1.jpg\" alt=\"views\" style=\"height:300px;margin-top:3rem;\"> </div>\n\n\n<font size=\"3.5\">\nWhen in doubt radiologist can use additional views, such as: <strong> mediolateral (ML), lateromedial (LM), lateromedial oblique (LMO) and axillary tail (AT)</strong>. \n\nBasically, the difference between all these views is the direction from which the scan is taken (top to bottom, left to right, etc), which tissue from the breast is in focus and sometimes tissue from the chest or axillary (armpit) is visible in the scan.\n\nWhat this means when training a network? Usually there are 4 images per patient, 2 views for each breast, which can translate to a multiple input network. However, a patient might have one or two extra views which can contain useful information for detection.\n\nLet's see how they look!\n</font>","metadata":{}},{"cell_type":"code","source":"def convert_to_hu(img, interp, slope):\n    # Convert raw pixels to Hounsfield Units\n    # It's usually done for CT images, \n    # but for completeness it's also used here\n    res_img = img.copy()\n    res_img = res_img * slope + interp\n    return res_img\n\ndef display_img(img, ax=None, figsize=(7, 5)):\n    if ax is None:\n        fig, ax = plt.subplots(1, 1, figsize=figsize)\n    ax.imshow(img, cmap='gray')\n    ax.axis('off')","metadata":{"execution":{"iopub.status.busy":"2023-01-22T16:41:47.547474Z","iopub.execute_input":"2023-01-22T16:41:47.548518Z","iopub.status.idle":"2023-01-22T16:41:47.558053Z","shell.execute_reply.started":"2023-01-22T16:41:47.548483Z","shell.execute_reply":"2023-01-22T16:41:47.556953Z"},"jupyter":{"source_hidden":true},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def display_grid(sample_df, title=None, nrows=2, ncols=5, figsize=(20, 10)):\n    fig, ax = plt.subplots(nrows, ncols, figsize=figsize)\n    i,j = 0, 0\n    for i in range(nrows):\n        for j in range(ncols):\n            row = sample_df.iloc[i*ncols+j]\n            p = row['img_path']\n            dcm_img = pydicom.dcmread(p)\n        #     print(im)\n            img = dcm_img.pixel_array\n            # (0028, 0004) Photometric Interpretation\n            if dcm_img['PhotometricInterpretation'].value == 'MONOCHROME1': # intensities values are inverted\n                img = np.invert(img)\n            display_img(img, ax=ax[i][j])\n            label = 'POS' if row['cancer'] == '1' else 'NEG'\n            ax[i][j].set_title(f\"{row['laterality']}_{row['view']}: {label}\")\n    if title is not None:\n        fig.suptitle(title, size=16, y=1.)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-01-22T17:39:07.766424Z","iopub.execute_input":"2023-01-22T17:39:07.766838Z","iopub.status.idle":"2023-01-22T17:39:07.777259Z","shell.execute_reply.started":"2023-01-22T17:39:07.766806Z","shell.execute_reply":"2023-01-22T17:39:07.776076Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"seed = 0\ncommon_views = ['MLO', 'CC']\n\nsample_df = train_df.loc[train_df['view'].isin(common_views), :].sample(n=10, random_state=seed).copy().reset_index()\ndisplay_grid(sample_df, title=\"Common views\")","metadata":{"execution":{"iopub.status.busy":"2023-01-22T16:41:47.571524Z","iopub.execute_input":"2023-01-22T16:41:47.572185Z","iopub.status.idle":"2023-01-22T16:42:05.459852Z","shell.execute_reply.started":"2023-01-22T16:41:47.57214Z","shell.execute_reply":"2023-01-22T16:42:05.45869Z"},"jupyter":{"source_hidden":true},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"other_views = ['AT', 'ML', 'LM', 'LMO']\nother_df = train_df.loc[train_df['view'].isin(other_views), :].copy().reset_index()\nsample_df2 = other_df.sample(n=10, random_state=seed).copy()\ndisplay_grid(sample_df2, title=\"Uncommon views\")","metadata":{"execution":{"iopub.status.busy":"2023-01-22T16:42:05.461221Z","iopub.execute_input":"2023-01-22T16:42:05.461585Z","iopub.status.idle":"2023-01-22T16:42:22.2691Z","shell.execute_reply.started":"2023-01-22T16:42:05.46153Z","shell.execute_reply":"2023-01-22T16:42:22.267811Z"},"jupyter":{"source_hidden":true},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font size=\"3.5\">We can also look at views in terms of which breast they are used for: left or right. The right breast seems to have more type of views than the left one.</font>","metadata":{}},{"cell_type":"code","source":"print(\"Laterality, nr of images\")\ndisplay(train_df['laterality'].value_counts(dropna=False))","metadata":{"execution":{"iopub.status.busy":"2023-01-22T17:08:20.557789Z","iopub.execute_input":"2023-01-22T17:08:20.558912Z","iopub.status.idle":"2023-01-22T17:08:20.571604Z","shell.execute_reply.started":"2023-01-22T17:08:20.55887Z","shell.execute_reply":"2023-01-22T17:08:20.570253Z"},"_kg_hide-input":true,"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df[['site_id', 'patient_id', 'laterality', 'view']].drop_duplicates().groupby(['laterality', 'view']).count()","metadata":{"execution":{"iopub.status.busy":"2023-01-22T17:08:28.153172Z","iopub.execute_input":"2023-01-22T17:08:28.153629Z","iopub.status.idle":"2023-01-22T17:08:28.209719Z","shell.execute_reply.started":"2023-01-22T17:08:28.15359Z","shell.execute_reply":"2023-01-22T17:08:28.20854Z"},"_kg_hide-input":true,"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Positive cases\n<a id='positive'></a>\n<font size=\"3.5\">\nLet's now look at positive cases and what other information are we given in the csv about these cases:\n\n\n* **implant** - Whether or not the patient had breast implants. Site 1 only provides breast implant information at the patient level, not at the breast level.\n* **density** - A rating for how dense the breast tissue is, with A being the least dense and D being the most dense. Extremely dense tissue can make diagnosis more difficult. Only provided for train.\n* **biopsy** - Whether or not a follow-up biopsy was performed on the breast. Only provided for train.\n* **invasive** - If the breast is positive for cancer, whether or not the cancer proved to be invasive. Only provided for train.\n* **BIRADS** - 0 if the breast required follow-up, 1 if the breast was rated as negative for cancer, and 2 if the breast was rated as normal. Only provided for train.\n* **difficult_negative_case** - True if the case was unusually difficult. Only provided for train.\n\n\nExcept for the implant which should be visible in the image (and might affect the detection of tumors) the other 5 values are \"extracted\" by a human from the scans and from other sources of information: patient's history, follow-up investigation, radiologyst's years of training and experience etc. One might think of a scenario where these values could serve as additional inputs (features) to a machine learning model, but they are provided only for training. They could be useful, however for understanding the data, in particular difficult cases that might pose a problem to a model.\n</font>","metadata":{}},{"cell_type":"code","source":"sample_df = train_df.loc[train_df['cancer'] == '1', :].sample(n=10, random_state=seed).copy().reset_index()\ndisplay_grid(sample_df, title=\"Cancer positive\")","metadata":{"execution":{"iopub.status.busy":"2023-01-22T16:42:22.350984Z","iopub.execute_input":"2023-01-22T16:42:22.351314Z","iopub.status.idle":"2023-01-22T16:42:42.488446Z","shell.execute_reply.started":"2023-01-22T16:42:22.351284Z","shell.execute_reply":"2023-01-22T16:42:42.48736Z"},"jupyter":{"source_hidden":true},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_df = train_df.loc[train_df['cancer'] == '0', :].sample(n=10, random_state=seed).copy().reset_index()\ndisplay_grid(sample_df, title=\"Cancer negative\")","metadata":{"execution":{"iopub.status.busy":"2023-01-22T16:42:42.490205Z","iopub.execute_input":"2023-01-22T16:42:42.490918Z","iopub.status.idle":"2023-01-22T16:43:00.833121Z","shell.execute_reply.started":"2023-01-22T16:42:42.49087Z","shell.execute_reply":"2023-01-22T16:43:00.832224Z"},"_kg_hide-input":true,"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font size=\"3.5\">As we can see it is very difficult to distingush with the naked eye positive cases from negative ones. But maybe, some views are more informative than others for a particular breast or patient. Let's look at all the views  for several patients.</font> ","metadata":{}},{"cell_type":"code","source":"pos_patients = train_df.loc[train_df['cancer'] == '1', 'patient_id'].unique().tolist()\nrandom.seed(seed)\nrandom.shuffle(pos_patients)\n\nex_patients = pos_patients[:3]\nprint(\"Example patients:\", ex_patients)\n\nsample_df = train_df.loc[train_df['patient_id'].isin(ex_patients), :].reset_index()\n# sample_df.groupby(['patient_id', 'laterality']).count()","metadata":{"execution":{"iopub.status.busy":"2023-01-22T17:07:42.867931Z","iopub.execute_input":"2023-01-22T17:07:42.868337Z","iopub.status.idle":"2023-01-22T17:07:42.888722Z","shell.execute_reply.started":"2023-01-22T17:07:42.868303Z","shell.execute_reply":"2023-01-22T17:07:42.887634Z"},"_kg_hide-input":true,"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for i, pat in enumerate(ex_patients):\n    _df = sample_df.loc[sample_df['patient_id'] == pat, :].reset_index()\n#     display(_df)\n    ncols = len(_df)\n    width, height = 4*ncols, 1.5*ncols\n    fig, ax = plt.subplots(1, ncols, figsize=(width, height))\n    fig.tight_layout()\n    fig.subplots_adjust(top=0.97)\n    fig.suptitle(f\"Patient id = {pat}\", size=16)\n    for j, row in _df.iterrows():\n        p = row['img_path']\n        dcm_img = pydicom.dcmread(p)\n    #     print(im)\n        img = dcm_img.pixel_array\n        # (0028, 0004) Photometric Interpretation\n        if dcm_img['PhotometricInterpretation'].value == 'MONOCHROME1': # intensities values are inverted\n            img = np.invert(img)\n        display_img(img, ax=ax[j])\n        label = 'POS' if row['cancer'] == '1' else 'NEG'\n        ax[j].set_title(f\"{row['laterality']} - {row['view']}\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-01-22T16:43:00.855925Z","iopub.execute_input":"2023-01-22T16:43:00.856292Z","iopub.status.idle":"2023-01-22T16:43:34.057334Z","shell.execute_reply.started":"2023-01-22T16:43:00.856258Z","shell.execute_reply":"2023-01-22T16:43:34.056182Z"},"jupyter":{"source_hidden":true},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font size=\"3.5\">Still, it's difficult to asses which breast contains a malignat tumor (\"bad\" tumor), that is positive for cancer. There are also benign tumors (\"good\" tumors) which a potential AI might pick up and confuse it with a malignant tumor. The challenge does not provide us with localication of the tumors or which type there are or with an experienced radiologyst :) These type of knowledge would be useful while debugging the AI model or while setting the limitations of the model. But let's leave that for now and continue our exploration of the dataset.</font>","metadata":{}},{"cell_type":"markdown","source":"# Implants\n<a id='implants'></a>\n\n<font size=\"3.5\">\nAnother information we are provided with in the csv is whether a patient has breast implants or not. Below we can see how they look  like. Since the purpose of this competition is to detect breast cancer I would like to know how many patients diagnosed with cancer have an implant, since in my opinion the implant might hinder the detection of tumor for an AI model. As would for a radiologist, I think.\n</font>","metadata":{}},{"cell_type":"code","source":"_df = train_df[['site_id', 'patient_id', 'cancer', 'implant']].drop_duplicates().groupby(['cancer', 'implant']).count().reset_index()\n# display(_df)\nfig, ax = plt.subplots(1, 1, figsize=(10,5))\nsns.barplot(ax=ax, data=_df, x='cancer', y='patient_id', hue='implant', ci=None)\nfor bars in ax.containers:\n    ax.bar_label(bars)\nax.set_xlabel(\"Cancer\")\nax.set_ylabel(\"Nr patients\")\nax.set_title(\"Nr patients having an implant by label\")","metadata":{"execution":{"iopub.status.busy":"2023-01-22T16:43:34.05896Z","iopub.execute_input":"2023-01-22T16:43:34.059383Z","iopub.status.idle":"2023-01-22T16:43:34.371057Z","shell.execute_reply.started":"2023-01-22T16:43:34.05934Z","shell.execute_reply":"2023-01-22T16:43:34.369856Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font size='3.5'>\nAs we can see from the graph above, most patients with implant do not have cancer, which is a good thing for our future algorithm. Let's look at these 3 patients.\n</font>","metadata":{}},{"cell_type":"code","source":"implant_df = train_df.loc[train_df['implant'] == '1', :].copy().reset_index()\npos_patients = implant_df.loc[implant_df['cancer'] == '1', 'patient_id'].unique()\nprint(\"Patients positive for cancer and with implants\", pos_patients)","metadata":{"execution":{"iopub.status.busy":"2023-01-22T17:07:14.823302Z","iopub.execute_input":"2023-01-22T17:07:14.823804Z","iopub.status.idle":"2023-01-22T17:07:14.841153Z","shell.execute_reply.started":"2023-01-22T17:07:14.823738Z","shell.execute_reply":"2023-01-22T17:07:14.839652Z"},"_kg_hide-input":true,"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_df = train_df.loc[train_df['patient_id'].isin(pos_patients), :]\n\nfor i, pat in enumerate(pos_patients):\n    _df = sample_df.loc[sample_df['patient_id'] == pat, :].reset_index()\n#     display(_df)\n    ncols = len(_df)\n    width, height = 3*ncols, ncols\n    fig, ax = plt.subplots(1, ncols, figsize=(width, height))\n#     fig.tight_layout(rect=[0, 0.03, 1, 0.97])\n#     fig.subplots_adjust(top=0.97)\n    fig.suptitle(f\"Patient id = {pat}\", y=0.85, size=16)\n    for j, row in _df.iterrows():\n        p = row['img_path']\n        dcm_img = pydicom.dcmread(p)\n    #     print(im)\n        img = dcm_img.pixel_array\n        # (0028, 0004) Photometric Interpretation\n        if dcm_img['PhotometricInterpretation'].value == 'MONOCHROME1': # intensities values are inverted\n            img = np.invert(img)\n        display_img(img, ax=ax[j])\n        label = 'POS' if row['cancer'] == '1' else 'NEG'\n        ax[j].set_title(f\"{row['laterality']} - {row['view']}\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-01-22T16:45:00.842074Z","iopub.execute_input":"2023-01-22T16:45:00.842494Z","iopub.status.idle":"2023-01-22T16:45:30.25543Z","shell.execute_reply.started":"2023-01-22T16:45:00.842461Z","shell.execute_reply":"2023-01-22T16:45:30.254609Z"},"jupyter":{"source_hidden":true},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font size=3.5>\nThe first thing I notice is that these patients have more than 4 views, and that there are duplicated types of views per breast. For example, the first patient (id = 12725) has 2 MLO scans and 3 CC scans for the left breast. Only in 2 of them I can clearly see the implant. This is similar for the remaining 2 patients. I am not sure at this point if images with a visible implant should be removed when training an AI. My guess is that radiologysts use those images as well, so why shouldn't an AI see them as well?\n</font>","metadata":{}},{"cell_type":"markdown","source":"# TODO List\n<a id='todo'></a>\n* <font size=3.5> pixel intensities and resolution per machine </font>","metadata":{}},{"cell_type":"markdown","source":"# Sources\n<a id='sources'></a>\n\n* [https://www.imaginis.com/mammography/how-mammography-is-performed-imaging-and-positioning-2?r](https://www.imaginis.com/mammography/how-mammography-is-performed-imaging-and-positioning-2?r)\n* [https://radiopaedia.org/articles/mammography-views](https://radiopaedia.org/articles/mammography-views)","metadata":{}},{"cell_type":"code","source":"","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"execution_count":null,"outputs":[]}]}