{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Table of contents\n\n* [Loading packages](#packages)\n* [Exploring the meta data](#eda)\n    * [Missing values](#missing_vals)\n    * [Understanding patients](#patients)\n    * [Image features](#image_features)\n    * [Inspecting target features](#target_features)\n    * [Machine Ids](#machine_ids)\n* [Exploring the images](#images)\n    * [Inspecting dicom files](#dicom)\n    * [Pixelarray distributions](#raw_values)","metadata":{}},{"cell_type":"code","source":"!pip install -U pylibjpeg pylibjpeg-openjpeg pylibjpeg-libjpeg pydicom python-gdcm","metadata":{"execution":{"iopub.status.busy":"2022-12-14T12:58:08.823703Z","iopub.execute_input":"2022-12-14T12:58:08.824185Z","iopub.status.idle":"2022-12-14T12:58:25.960853Z","shell.execute_reply.started":"2022-12-14T12:58:08.824087Z","shell.execute_reply":"2022-12-14T12:58:25.959625Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Loading packages <a class=\"anchor\" id=\"packages\"></a>","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\n\nimport pydicom\nfrom os import listdir\n\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nsns.set_style('darkgrid')\nsns.set_color_codes('bright')\n\nfrom scipy.stats import mode\n\n\nimport warnings\nwarnings.filterwarnings(\"ignore\", category=DeprecationWarning)\nwarnings.filterwarnings(\"ignore\", category=UserWarning)\nwarnings.filterwarnings(\"ignore\", category=FutureWarning)","metadata":{"execution":{"iopub.status.busy":"2022-12-14T12:58:25.963588Z","iopub.execute_input":"2022-12-14T12:58:25.96395Z","iopub.status.idle":"2022-12-14T12:58:26.755162Z","shell.execute_reply.started":"2022-12-14T12:58:25.963916Z","shell.execute_reply":"2022-12-14T12:58:26.754023Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Exploring at the meta data <a class=\"anchor\" id=\"eda\"></a>","metadata":{}},{"cell_type":"code","source":"train = pd.read_csv('/kaggle/input/rsna-breast-cancer-detection/train.csv')\ntrain.head()","metadata":{"execution":{"iopub.status.busy":"2022-12-14T12:58:26.756916Z","iopub.execute_input":"2022-12-14T12:58:26.757263Z","iopub.status.idle":"2022-12-14T12:58:26.900776Z","shell.execute_reply.started":"2022-12-14T12:58:26.757231Z","shell.execute_reply":"2022-12-14T12:58:26.899677Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.shape","metadata":{"execution":{"iopub.status.busy":"2022-12-14T12:58:26.903847Z","iopub.execute_input":"2022-12-14T12:58:26.904685Z","iopub.status.idle":"2022-12-14T12:58:26.912206Z","shell.execute_reply.started":"2022-12-14T12:58:26.904637Z","shell.execute_reply":"2022-12-14T12:58:26.910988Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Ok, 54706 rows and 14 columns. Do we have missing values?","metadata":{}},{"cell_type":"code","source":"test = pd.read_csv('/kaggle/input/rsna-breast-cancer-detection/test.csv')\ntest.head()","metadata":{"execution":{"iopub.status.busy":"2022-12-14T12:58:26.913776Z","iopub.execute_input":"2022-12-14T12:58:26.914486Z","iopub.status.idle":"2022-12-14T12:58:26.939521Z","shell.execute_reply.started":"2022-12-14T12:58:26.914437Z","shell.execute_reply":"2022-12-14T12:58:26.938243Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test.shape","metadata":{"execution":{"iopub.status.busy":"2022-12-14T12:58:26.940832Z","iopub.execute_input":"2022-12-14T12:58:26.941693Z","iopub.status.idle":"2022-12-14T12:58:26.948224Z","shell.execute_reply.started":"2022-12-14T12:58:26.941658Z","shell.execute_reply":"2022-12-14T12:58:26.946897Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Aha! There are a lot of columns in train that are not present in test. Some of them are related to the target feature **cancer** and it's obvious that they are not given (biopsy, invasive, BIRADS and difficult_negative_case). But what makes me wonder is that the density feature is also not given in test... ","metadata":{}},{"cell_type":"markdown","source":"## Missing values <a class=\"anchor\" id=\"missing_vals\"></a>","metadata":{}},{"cell_type":"code","source":"train.info()","metadata":{"execution":{"iopub.status.busy":"2022-12-14T12:58:26.94991Z","iopub.execute_input":"2022-12-14T12:58:26.950822Z","iopub.status.idle":"2022-12-14T12:58:26.986568Z","shell.execute_reply.started":"2022-12-14T12:58:26.950778Z","shell.execute_reply":"2022-12-14T12:58:26.985285Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Ok, we do have missing values for the age, BIRADS and density. ","metadata":{}},{"cell_type":"markdown","source":"## Understanding patients <a class=\"anchor\" id=\"patients\"></a>","metadata":{}},{"cell_type":"markdown","source":"How many patients do we have?","metadata":{}},{"cell_type":"code","source":"train.patient_id.nunique()","metadata":{"execution":{"iopub.status.busy":"2022-12-14T12:58:26.988491Z","iopub.execute_input":"2022-12-14T12:58:26.988901Z","iopub.status.idle":"2022-12-14T12:58:26.997792Z","shell.execute_reply.started":"2022-12-14T12:58:26.988867Z","shell.execute_reply":"2022-12-14T12:58:26.996913Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"How old are the patients?","metadata":{}},{"cell_type":"code","source":"train.groupby('patient_id').age.nunique().unique()","metadata":{"execution":{"iopub.status.busy":"2022-12-14T12:58:26.999239Z","iopub.execute_input":"2022-12-14T12:58:27.000401Z","iopub.status.idle":"2022-12-14T12:58:27.022817Z","shell.execute_reply.started":"2022-12-14T12:58:27.00035Z","shell.execute_reply":"2022-12-14T12:58:27.021622Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Ok, that's good! We only have one age value per patient or none at all. ","metadata":{}},{"cell_type":"code","source":"ages = train[train.age.isnull() == False].groupby(\n    'patient_id').age.apply(lambda l: np.unique(l)[0])\nplt.figure(figsize=(20,8))\n\nsns.histplot(ages, color='orange', bins=60)\nplt.title('Age distribution of patients');","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-12-14T12:58:27.026757Z","iopub.execute_input":"2022-12-14T12:58:27.027154Z","iopub.status.idle":"2022-12-14T12:58:27.94862Z","shell.execute_reply.started":"2022-12-14T12:58:27.027117Z","shell.execute_reply":"2022-12-14T12:58:27.946999Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Insights\n\n* Most of the patients are older than 40 years. \n* It seems that we have two peaks around the age of 50 and close to 70. \n* There is a drop of patients counts after the age of 70. ","metadata":{}},{"cell_type":"markdown","source":"How many patients do have cancer and how many images were taken per patient?","metadata":{}},{"cell_type":"code","source":"def has_cancer(l):\n    if len(l) == 1:\n        if l[0] == 0:\n            return False\n        elif l[0] == 1:\n            return True\n        else:\n            raise Exception\n    elif len(l) == 2:\n        return True\n    else:\n        raise Exception\n\npatient_cancer_map = train.groupby('patient_id').cancer.unique().apply(lambda l: has_cancer(l))","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-12-14T12:58:27.950499Z","iopub.execute_input":"2022-12-14T12:58:27.951005Z","iopub.status.idle":"2022-12-14T12:58:28.607935Z","shell.execute_reply.started":"2022-12-14T12:58:27.950957Z","shell.execute_reply":"2022-12-14T12:58:28.606784Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(1,2,figsize=(20,5))\n\nsns.countplot(patient_cancer_map, palette='Paired', ax=ax[0]);\nax[0].set_title('Number of patients with cancer');\n\nsns.countplot(train.groupby('patient_id').size(), ax=ax[1])\nax[1].set_title('Number of images per patient')\nax[1].set_xlabel('Number of images')\nax[1].set_ylabel('Counts of patients');","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-12-14T12:58:28.609623Z","iopub.execute_input":"2022-12-14T12:58:28.610661Z","iopub.status.idle":"2022-12-14T12:58:29.061127Z","shell.execute_reply.started":"2022-12-14T12:58:28.61062Z","shell.execute_reply":"2022-12-14T12:58:29.059788Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Insights\n\n* It's an imbalanced classification problem! \n* In most cases we have given 4 images per patient but there are a few cases with more than 10 images as well! Why?","metadata":{}},{"cell_type":"markdown","source":"## Image features <a class=\"anchor\" id=\"image_features\"></a>","metadata":{}},{"cell_type":"code","source":"fig, ax = plt.subplots(2,2,figsize=(20,10))\nsns.countplot(train.laterality, ax=ax[0,0], palette='Greens_r')\nsns.countplot(train.view, ax=ax[0,1], palette='Reds_r')\nsns.countplot(train.implant, ax=ax[1,0], palette='Blues_r')\nsns.countplot(train.density, ax=ax[1,1], palette='Purples_r');","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-12-14T12:58:29.062357Z","iopub.execute_input":"2022-12-14T12:58:29.062718Z","iopub.status.idle":"2022-12-14T12:58:29.780713Z","shell.execute_reply.started":"2022-12-14T12:58:29.062688Z","shell.execute_reply":"2022-12-14T12:58:29.779515Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Insights\n\n* The laterality is quite balanced. \n* In the data description we can find that there are usually two views per breast. The most common views are CC and MLO and given these 2 views for the right and left breast we end up with 4 images that most of the patients show.  \n* Only a very few images show implants. \n* Most of the images show medium dense images of category B and C. Nonetheless there are also cases that are very dense (D) or less dense (A). Given the information that it could be more difficult to identify cancer in dense tissues this could be an interesting feature when thinking about validation strategies. ","metadata":{}},{"cell_type":"markdown","source":"## Inspecting target features <a class=\"anchor\" id=\"target_features\"></a>\n\nWe have already seen that the number of patients with cancer is quiet low compared to the patients without. Let's see how it looks like on the image-level.","metadata":{}},{"cell_type":"code","source":"biopsy_counts = train.groupby('cancer').biopsy.value_counts().unstack().fillna(0) \nbiopsy_perc = biopsy_counts.transpose() / biopsy_counts.sum(axis=1)","metadata":{"execution":{"iopub.status.busy":"2022-12-14T12:58:29.782524Z","iopub.execute_input":"2022-12-14T12:58:29.783239Z","iopub.status.idle":"2022-12-14T12:58:29.797905Z","shell.execute_reply.started":"2022-12-14T12:58:29.783192Z","shell.execute_reply":"2022-12-14T12:58:29.796686Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(1,2,figsize=(20,5))\nsns.countplot(train.cancer, palette='Reds', ax=ax[0])\nax[0].set_title('Number of images displaying cancer');\nsns.heatmap(biopsy_perc.transpose(), ax=ax[1], annot=True, cmap='Oranges');","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-12-14T12:58:29.799549Z","iopub.execute_input":"2022-12-14T12:58:29.800041Z","iopub.status.idle":"2022-12-14T12:58:30.224057Z","shell.execute_reply.started":"2022-12-14T12:58:29.799991Z","shell.execute_reply":"2022-12-14T12:58:30.222835Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Insights\n\n* The number of images displaying cancer is very low. It should be even of lower percentage than the number of patients with cancer. \n* Looking at the biopsy feature we can say that all patients with cancer had a biopsy. But only around 3 % of images without cancer had resulted in a follow-up biopsy. Maybe we should better have a look at this feature on the patient-level. ","metadata":{}},{"cell_type":"code","source":"train.cancer.value_counts()/train.shape[0]","metadata":{"execution":{"iopub.status.busy":"2022-12-14T12:58:30.225682Z","iopub.execute_input":"2022-12-14T12:58:30.226039Z","iopub.status.idle":"2022-12-14T12:58:30.236499Z","shell.execute_reply.started":"2022-12-14T12:58:30.226007Z","shell.execute_reply":"2022-12-14T12:58:30.235362Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"patient_cancer_map.value_counts() / patient_cancer_map.shape[0]","metadata":{"execution":{"iopub.status.busy":"2022-12-14T12:58:30.238198Z","iopub.execute_input":"2022-12-14T12:58:30.238692Z","iopub.status.idle":"2022-12-14T12:58:30.252743Z","shell.execute_reply.started":"2022-12-14T12:58:30.238645Z","shell.execute_reply":"2022-12-14T12:58:30.251612Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As expected, the percentage of patients with cancer is higher than the percentage of images displaying cancer. How many images show invasive cancer that has spread beyond the layer of tissue in which it developed?","metadata":{}},{"cell_type":"code","source":"fig, ax = plt.subplots(2,2,figsize=(20,10))\nsns.countplot(train[train.cancer==1].invasive, ax=ax[0,0], palette='Reds')\nsns.countplot(train.BIRADS, order=[0., 1., 2.], ax=ax[0,1], palette='Blues')\nax[0,0].set_title('Images showing invasive cancer');\n\nsns.countplot(train[train.cancer==0].BIRADS, order=[0., 1., 2.], ax=ax[1,0], palette='Purples')\nsns.countplot(train[train.cancer==1].BIRADS, order=[0., 1., 2.], ax=ax[1,1], palette='Purples')","metadata":{"execution":{"iopub.status.busy":"2022-12-14T12:58:30.25406Z","iopub.execute_input":"2022-12-14T12:58:30.254582Z","iopub.status.idle":"2022-12-14T12:58:30.918208Z","shell.execute_reply.started":"2022-12-14T12:58:30.254544Z","shell.execute_reply":"2022-12-14T12:58:30.916939Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train[train.cancer==1].invasive.value_counts() / train[train.cancer==1].shape[0]","metadata":{"execution":{"iopub.status.busy":"2022-12-14T12:58:30.92015Z","iopub.execute_input":"2022-12-14T12:58:30.920561Z","iopub.status.idle":"2022-12-14T12:58:30.936113Z","shell.execute_reply.started":"2022-12-14T12:58:30.920524Z","shell.execute_reply":"2022-12-14T12:58:30.93503Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Insights\n\n* Roughly 70 % of all images displaying cancer are showing invasive cancer thus the one that spreads into other tissues as well. :(\n* For BIRADS we need to remember that its value is\n    * 0 if the breast required follow-up\n    * 1 if the breast was rated as negative for cancer \n    * 2 if the breast was rated as normal\n* Given that information, we can see that most images were rated as negative which suites to the fact that most images are not positive for cancer. But nonetheless there is still a high number of images that lead into a follow-up. I'm a bit confused... what's the difference between 1 and 2? From the ordering I would suspect that 2 is better than 1. \n* Going one level deeper and showing the BIRADS counts for 'no cancer'-images and 'cancer'-images, we can at least say that all images showing cancer needed a follow up (what we expected! ;))","metadata":{}},{"cell_type":"markdown","source":"## Machine Ids <a class=\"anchor\" id=\"machine_ids\"></a>","metadata":{}},{"cell_type":"markdown","source":"How many different imaging devices were used?","metadata":{}},{"cell_type":"code","source":"train.machine_id.nunique()","metadata":{"execution":{"iopub.status.busy":"2022-12-14T12:58:30.93737Z","iopub.execute_input":"2022-12-14T12:58:30.937859Z","iopub.status.idle":"2022-12-14T12:58:30.945526Z","shell.execute_reply.started":"2022-12-14T12:58:30.937823Z","shell.execute_reply":"2022-12-14T12:58:30.944209Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(1,1,figsize=(20,5))\nsns.countplot(train.machine_id, palette='tab10', ax=ax)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-12-14T12:58:30.947418Z","iopub.execute_input":"2022-12-14T12:58:30.9478Z","iopub.status.idle":"2022-12-14T12:58:31.376117Z","shell.execute_reply.started":"2022-12-14T12:58:30.947767Z","shell.execute_reply":"2022-12-14T12:58:31.374642Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Insights\n\n* 10 different machines but most of the images were take with device of id 49, 21, 29 and 48. \n* Personally I'm not sure whether this feature has some importance for us. To be honest, I think that the patient_id, the age and the density will be much more important than the machine. ","metadata":{}},{"cell_type":"markdown","source":"# Exploring the images <a class=\"anchor\" id=\"images\"></a>\n\nBrowsing through the train and test folder of the images, we can see that the data is given as dicom files. Long time ago I wrote a tutorial notebook on dicom files while I was learning more about it myself. Please take a look at it if dicom is unknown to you. I hope that it will help to get started with the data structure. ;) \n\nhttps://www.kaggle.com/code/allunia/pulmonary-dicom-preprocessing\n","metadata":{}},{"cell_type":"code","source":"train_path = '/kaggle/input/rsna-breast-cancer-detection/train_images/'\ntest_path = '/kaggle/input/rsna-breast-cancer-detection/test_images/'","metadata":{"execution":{"iopub.status.busy":"2022-12-14T12:58:31.377613Z","iopub.execute_input":"2022-12-14T12:58:31.37802Z","iopub.status.idle":"2022-12-14T12:58:31.383738Z","shell.execute_reply.started":"2022-12-14T12:58:31.377981Z","shell.execute_reply":"2022-12-14T12:58:31.382492Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"First of all, we need a method to get all scans of one patient:","metadata":{}},{"cell_type":"code","source":"def load_scans(path, patient_id):\n    dcm_path = path + str(patient_id)\n    slices = [pydicom.dcmread(dcm_path + '/' + file, force=True) for file in listdir(dcm_path)]\n    return slices","metadata":{"execution":{"iopub.status.busy":"2022-12-14T12:58:31.386488Z","iopub.execute_input":"2022-12-14T12:58:31.387477Z","iopub.status.idle":"2022-12-14T12:58:31.394313Z","shell.execute_reply.started":"2022-12-14T12:58:31.387407Z","shell.execute_reply":"2022-12-14T12:58:31.393094Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Pixelarray distributions <a class=\"anchor\" id=\"raw values\"></a>","metadata":{}},{"cell_type":"markdown","source":"* Usually we need to transform the raw pixelarray distributions to Hounsfield Units. The raw values depend on the measurement settings like acquisition parameters and tube voltage of the scanner. \n* By normalizing to values of water and air (water has HU 0 and air -1000) the images of different measurements are becoming comparable. \n* We also need to understand how background values were treated! In previous competitions it was often the case that background values were set to values smaller than -1000, but we need to check whether this was done in this competition as well!","metadata":{}},{"cell_type":"markdown","source":"Let's load an example:","metadata":{}},{"cell_type":"code","source":"patient_ids = train.patient_id.unique()","metadata":{"execution":{"iopub.status.busy":"2022-12-14T12:58:31.395992Z","iopub.execute_input":"2022-12-14T12:58:31.396609Z","iopub.status.idle":"2022-12-14T12:58:31.406551Z","shell.execute_reply.started":"2022-12-14T12:58:31.396571Z","shell.execute_reply":"2022-12-14T12:58:31.405583Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"scans = load_scans(train_path, patient_ids[0])","metadata":{"execution":{"iopub.status.busy":"2022-12-14T12:58:31.40808Z","iopub.execute_input":"2022-12-14T12:58:31.408933Z","iopub.status.idle":"2022-12-14T12:58:31.651899Z","shell.execute_reply.started":"2022-12-14T12:58:31.408896Z","shell.execute_reply":"2022-12-14T12:58:31.650886Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"And let's plot the raw pixelarray distribution for this example and display the related image:","metadata":{}},{"cell_type":"code","source":"fig, ax = plt.subplots(1,2,figsize=(20,5))\nsns.histplot(scans[0].pixel_array.flatten(), ax=ax[0])\nax[1].imshow(scans[0].pixel_array)","metadata":{"execution":{"iopub.status.busy":"2022-12-14T12:58:31.653045Z","iopub.execute_input":"2022-12-14T12:58:31.653838Z","iopub.status.idle":"2022-12-14T12:58:55.137351Z","shell.execute_reply.started":"2022-12-14T12:58:31.653801Z","shell.execute_reply":"2022-12-14T12:58:55.136141Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Insights\n\n* Uhh! That's interesting! The background seems to be above 3000 this time! So it's competely different to previous competitions.\n* As the raw values depend on the machine, it could be worth it to explore background values dependent on the machine id. ","metadata":{}},{"cell_type":"markdown","source":"Let's get an impression by extracting the mode per scan for a few patients (let's say 50) for each machine id.","metadata":{}},{"cell_type":"code","source":"machine_ids = train.machine_id.unique()\nmachine_ids","metadata":{"execution":{"iopub.status.busy":"2022-12-14T12:58:55.139201Z","iopub.execute_input":"2022-12-14T12:58:55.139679Z","iopub.status.idle":"2022-12-14T12:58:55.148865Z","shell.execute_reply.started":"2022-12-14T12:58:55.139634Z","shell.execute_reply":"2022-12-14T12:58:55.147767Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"scans[0].Columns","metadata":{"execution":{"iopub.status.busy":"2022-12-14T12:58:55.154339Z","iopub.execute_input":"2022-12-14T12:58:55.154766Z","iopub.status.idle":"2022-12-14T12:58:55.17357Z","shell.execute_reply.started":"2022-12-14T12:58:55.154718Z","shell.execute_reply":"2022-12-14T12:58:55.172476Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"background_values = []\nrows = []\ncolumns = []\nfor mid in train.machine_id.unique():\n    mid_background_values = []\n    mid_rows = []\n    mid_columns = []\n    print(f'machine id {mid} in progress')\n    for n in range(50):\n        try:\n            scans = load_scans(train_path, train[train.machine_id==mid].patient_id.unique()[n])\n            mid_background_values.append(mode(scans[0].pixel_array.flatten())[0][0])\n            mid_rows.append(scans[0].Rows)\n            mid_columns.append(scans[0].Columns)\n        except IndexError:\n            break\n    background_values.append(mid_background_values)\n    rows.append(mid_rows)\n    columns.append(mid_columns)","metadata":{"execution":{"iopub.status.busy":"2022-12-14T13:10:09.034112Z","iopub.execute_input":"2022-12-14T13:10:09.035185Z","iopub.status.idle":"2022-12-14T13:19:51.953671Z","shell.execute_reply.started":"2022-12-14T13:10:09.035143Z","shell.execute_reply":"2022-12-14T13:19:51.951869Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for n in range(len(background_values)):\n    print(machine_ids[n], np.median(background_values[n]), np.std(background_values[n]))","metadata":{"execution":{"iopub.status.busy":"2022-12-14T13:19:51.956732Z","iopub.execute_input":"2022-12-14T13:19:51.957185Z","iopub.status.idle":"2022-12-14T13:19:51.967408Z","shell.execute_reply.started":"2022-12-14T13:19:51.957136Z","shell.execute_reply":"2022-12-14T13:19:51.96586Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Insights\n\n* It depends on the machine id, how background values are treated. In most of the cases they seem to be 0. But for id 29 it's above 3000 and for id 210 it seems to be 1017. \n* Id 197 looks interesting. Currently I don't know, what it means. Let's take a look at the values","metadata":{}},{"cell_type":"code","source":"background_values[-1]","metadata":{"execution":{"iopub.status.busy":"2022-12-14T13:19:51.969294Z","iopub.execute_input":"2022-12-14T13:19:51.969763Z","iopub.status.idle":"2022-12-14T13:19:51.979807Z","shell.execute_reply.started":"2022-12-14T13:19:51.969728Z","shell.execute_reply":"2022-12-14T13:19:51.97874Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Lets visualize the scans taken from each machine","metadata":{}},{"cell_type":"code","source":"values = []\nfor mid in train.machine_id.unique():\n    scans = load_scans(train_path, train[train.machine_id==mid].patient_id.unique()[0])\n    \n#     for n in range(1):\n    values.append(scans[0].pixel_array)\n    \n# ax[0].imshow(values[0])\nprint(len(values))\n# ax[1].imshow(values[1])\n# plt.imshow(values[0])\nfig = plt.figure(figsize=(10, 7))\nax1=fig.add_subplot(2,2,1)\nax1.title.set_text(train.machine_id.unique()[0])\nplt.imshow(values[0])\n\nax2=fig.add_subplot(2,2,2)\nax2.title.set_text(train.machine_id.unique()[1])\nplt.imshow(values[1])\n\nax3=fig.add_subplot(2,2,3)\nax3.title.set_text(train.machine_id.unique()[2])\nplt.imshow(values[2])\n\nax4=fig.add_subplot(2,2,4)\nax4.title.set_text(train.machine_id.unique()[3])\nplt.imshow(values[3])","metadata":{"execution":{"iopub.status.busy":"2022-12-14T13:23:00.503972Z","iopub.execute_input":"2022-12-14T13:23:00.504407Z","iopub.status.idle":"2022-12-14T13:23:11.071736Z","shell.execute_reply.started":"2022-12-14T13:23:00.504369Z","shell.execute_reply":"2022-12-14T13:23:11.07045Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = plt.figure(figsize=(10, 7))\nax1=fig.add_subplot(2,2,1)\nax1.title.set_text(train.machine_id.unique()[4])\nplt.imshow(values[4])\n\nax2=fig.add_subplot(2,2,2)\nax2.title.set_text(train.machine_id.unique()[5])\nplt.imshow(values[5])\n\nax3=fig.add_subplot(2,2,3)\nax3.title.set_text(train.machine_id.unique()[6])\nplt.imshow(values[6])\n\nax4=fig.add_subplot(2,2,4)\nax4.title.set_text(train.machine_id.unique()[7])\nplt.imshow(values[7])","metadata":{"execution":{"iopub.status.busy":"2022-12-14T13:23:11.073961Z","iopub.execute_input":"2022-12-14T13:23:11.074337Z","iopub.status.idle":"2022-12-14T13:23:15.453103Z","shell.execute_reply.started":"2022-12-14T13:23:11.074305Z","shell.execute_reply":"2022-12-14T13:23:15.451871Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = plt.figure(figsize=(10, 7))\nax1=fig.add_subplot(2,2,1)\nax1.title.set_text(train.machine_id.unique()[8])\nplt.imshow(values[8])\n\nax2=fig.add_subplot(2,2,2)\nax2.title.set_text(train.machine_id.unique()[9])\n### Insights\n\n* It depends on the machine id, how background values are treated. In most of the cases they seem to be 0. But for id 29 it's above 3000 and for id 210 it seems to be 1017. \n* Id 197 looks interesting. Currently I don't know, what it means. Let's take a look at the valuesplt.imshow(values[9])","metadata":{"execution":{"iopub.status.busy":"2022-12-14T13:23:15.454902Z","iopub.execute_input":"2022-12-14T13:23:15.455281Z","iopub.status.idle":"2022-12-14T13:23:16.699057Z","shell.execute_reply.started":"2022-12-14T13:23:15.455247Z","shell.execute_reply":"2022-12-14T13:23:16.697889Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Insights\n\n* From viewing the images we can infer that for machine ID 210 and 29, we have a whitish background but have black background for the rest of the images","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Ah, ok, only 4 patients and it could be that the background values for patient 2 and 4 are not the same as the mode.\n\n### Conclusion\n\nWe need to be careful during preprocessing as background values were set differently depending on the machine id.\n","metadata":{}}]}