{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Table of contents - WIP (being adapted from https://www.kaggle.com/code/allunia/rsna-breast-cancer-eda)\n\n* [Loading packages and presets](#packages_presets)\n* [Exploring the data](#eda)\n    * [Checking missing values per column](#missing_vals)\n    * [Investigating hospital and scanner provenance](#hospital_scanner)\n    * [Understanding patients](#patients)\n    * [Image features](#image_features)\n    * [Inspecting target features](#target_features)","metadata":{}},{"cell_type":"markdown","source":"# Loading packages and definition of presets<a class=\"anchor\" id=\"packages_presets\"></a>","metadata":{}},{"cell_type":"code","source":"# Import Packages\nimport numpy as np\nimport pandas as pd\nimport seaborn as sns\nimport matplotlib.pyplot as plt\nimport warnings\n\nfrom sklearn.model_selection import StratifiedGroupKFold\n\n# Presets\nsns.set_style('darkgrid')\nsns.set_color_codes('bright')\n\nwarnings.filterwarnings(\"ignore\", category=DeprecationWarning)\nwarnings.filterwarnings(\"ignore\", category=UserWarning)\nwarnings.filterwarnings(\"ignore\", category=FutureWarning)","metadata":{"execution":{"iopub.status.busy":"2022-11-30T21:37:18.694701Z","iopub.execute_input":"2022-11-30T21:37:18.695419Z","iopub.status.idle":"2022-11-30T21:37:20.224971Z","shell.execute_reply.started":"2022-11-30T21:37:18.695296Z","shell.execute_reply":"2022-11-30T21:37:20.223769Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Exploring at the data <a class=\"anchor\" id=\"eda\"></a>","metadata":{}},{"cell_type":"code","source":"train = pd.read_csv('/kaggle/input/rsna-breast-cancer-detection/train.csv')\ntrain.head()","metadata":{"execution":{"iopub.status.busy":"2022-11-30T21:37:20.22706Z","iopub.execute_input":"2022-11-30T21:37:20.227657Z","iopub.status.idle":"2022-11-30T21:37:20.399557Z","shell.execute_reply.started":"2022-11-30T21:37:20.227624Z","shell.execute_reply":"2022-11-30T21:37:20.398426Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(f'Training dataset is comprised of {train.shape[0]} rows and {train.shape[1]} columns')","metadata":{"execution":{"iopub.status.busy":"2022-11-30T21:37:20.401564Z","iopub.execute_input":"2022-11-30T21:37:20.401983Z","iopub.status.idle":"2022-11-30T21:37:20.409665Z","shell.execute_reply.started":"2022-11-30T21:37:20.401948Z","shell.execute_reply":"2022-11-30T21:37:20.407931Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Let's get that info for the test","metadata":{}},{"cell_type":"code","source":"test = pd.read_csv('/kaggle/input/rsna-breast-cancer-detection/test.csv')\ntest.head()","metadata":{"execution":{"iopub.status.busy":"2022-11-30T21:37:20.413557Z","iopub.execute_input":"2022-11-30T21:37:20.414042Z","iopub.status.idle":"2022-11-30T21:37:20.440347Z","shell.execute_reply.started":"2022-11-30T21:37:20.413984Z","shell.execute_reply":"2022-11-30T21:37:20.438937Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(f'Test dataset is comprised of {test.shape[0]} rows and {test.shape[1]} columns')","metadata":{"execution":{"iopub.status.busy":"2022-11-30T21:37:20.442147Z","iopub.execute_input":"2022-11-30T21:37:20.442648Z","iopub.status.idle":"2022-11-30T21:37:20.45132Z","shell.execute_reply.started":"2022-11-30T21:37:20.442601Z","shell.execute_reply":"2022-11-30T21:37:20.449988Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There are 5 columns that do not exist on the test set. Let's see which are those","metadata":{}},{"cell_type":"code","source":"list(set(train.columns.to_list()).difference(test.columns.to_list()))","metadata":{"execution":{"iopub.status.busy":"2022-11-30T21:37:20.45313Z","iopub.execute_input":"2022-11-30T21:37:20.453851Z","iopub.status.idle":"2022-11-30T21:37:20.464903Z","shell.execute_reply.started":"2022-11-30T21:37:20.453796Z","shell.execute_reply":"2022-11-30T21:37:20.463566Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The **cancer** is the target feature and biopsy, invasive, BIRADS  and difficult_negative_case are related to the **cancer** feature. As for the density, it known that breast density affect lesion detection, with denser breasts being harder to detect cancer using mammography. This may be a reason for the density feature to not be present on the test. The final may need to in a first step to estimate the breast density and then proceed to predict the **cancer** feature according to the breast density.\n\nAdditional information on:\n- BIRADS may be found [here](https://radiologyassistant.nl/breast/bi-rads/bi-rads-for-mammography-and-ultrasound-2013)\n- Breast Density is know to affect cancer detection as \"cancer and dense breast tissue both appear white on a mammogram\" [more here](https://www.mayoclinic.org/tests-procedures/mammogram/in-depth/dense-breast-tissue/art-20123968) - ***TLDR - Lesions may be harder to detect in denser breasts!!!!***\n![Alt text](https://www.mayoclinic.org/-/media/kcms/gbs/patient-consumer/images/2013/08/26/11/07/an01137_im03415_ans7_dense_breaststhu_jpg.jpg \"a title\")\n- The column view is use to indicate the acquisition plane: mediolateral oblique (MLO) view or cranial caudal (CC). Example below shows different acquisition planes where MLO and CC are the most common acquisition planes\n![Alt text](https://radiologykey.com/wp-content/uploads/2016/03/B9780323073226500228_u23-002a-9780323073226.jpg \"a title\")\n","metadata":{}},{"cell_type":"markdown","source":"## Checking missing values per column <a class=\"anchor\" id=\"missing_vals\"></a>","metadata":{}},{"cell_type":"code","source":"train.isna().sum()","metadata":{"execution":{"iopub.status.busy":"2022-11-30T21:37:20.466319Z","iopub.execute_input":"2022-11-30T21:37:20.467451Z","iopub.status.idle":"2022-11-30T21:37:20.488241Z","shell.execute_reply.started":"2022-11-30T21:37:20.467417Z","shell.execute_reply":"2022-11-30T21:37:20.487013Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(f\"We have missing values for the age (no. {train['age'].isna().sum()}), BIRADS (no. {train['BIRADS'].isna().sum()}), and density (no. {train['density'].isna().sum()})\")","metadata":{"execution":{"iopub.status.busy":"2022-11-30T21:37:20.490083Z","iopub.execute_input":"2022-11-30T21:37:20.491055Z","iopub.status.idle":"2022-11-30T21:37:20.506396Z","shell.execute_reply.started":"2022-11-30T21:37:20.491012Z","shell.execute_reply":"2022-11-30T21:37:20.504966Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Investigating hospital and scanner provenance <a class=\"anchor\" id=\"hospital_scanner\"></a>","metadata":{}},{"cell_type":"markdown","source":"A source of generalizability issues is the acquisition of medical images using different scanners and acquisition settings. Eventhough acquisition protocols may be somewhat standardized there are plenty of adjustments and choices that can be made by mammography scanner vendors or even the end-user. Thus we are interested in checking the number of different hospitals and scanners contributing for this dataset.","metadata":{}},{"cell_type":"code","source":"print(f\"We have {train.site_id.nunique()} different hospitals contributing to the dataset\")","metadata":{"execution":{"iopub.status.busy":"2022-11-30T21:37:20.508011Z","iopub.execute_input":"2022-11-30T21:37:20.508393Z","iopub.status.idle":"2022-11-30T21:37:20.519068Z","shell.execute_reply.started":"2022-11-30T21:37:20.508361Z","shell.execute_reply":"2022-11-30T21:37:20.517843Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(f\"and we have images from {train.machine_id.nunique()} different scanners included in the dataset\")","metadata":{"execution":{"iopub.status.busy":"2022-11-30T21:37:20.524097Z","iopub.execute_input":"2022-11-30T21:37:20.524488Z","iopub.status.idle":"2022-11-30T21:37:20.532843Z","shell.execute_reply.started":"2022-11-30T21:37:20.524453Z","shell.execute_reply":"2022-11-30T21:37:20.53194Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(f\"Of these, we have {train[train.site_id==1].machine_id.nunique()} scanners in hospital 1 and {train[train.site_id==2].machine_id.nunique()} scanners in hospital 2.\")","metadata":{"execution":{"iopub.status.busy":"2022-11-30T21:37:20.534089Z","iopub.execute_input":"2022-11-30T21:37:20.534925Z","iopub.status.idle":"2022-11-30T21:37:20.557468Z","shell.execute_reply.started":"2022-11-30T21:37:20.534891Z","shell.execute_reply":"2022-11-30T21:37:20.556148Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Understanding patients <a class=\"anchor\" id=\"patients\"></a>","metadata":{}},{"cell_type":"markdown","source":"How many patients do we have?","metadata":{}},{"cell_type":"code","source":"train.patient_id.nunique()","metadata":{"execution":{"iopub.status.busy":"2022-11-30T21:37:20.55889Z","iopub.execute_input":"2022-11-30T21:37:20.55922Z","iopub.status.idle":"2022-11-30T21:37:20.577335Z","shell.execute_reply.started":"2022-11-30T21:37:20.559181Z","shell.execute_reply":"2022-11-30T21:37:20.575093Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"How many patients are from hospitals (site_id) 1 and how many from hospital 2?","metadata":{}},{"cell_type":"code","source":"plt.bar([1,2],[train[train.site_id == 1].patient_id.nunique(), train[train.site_id == 2].patient_id.nunique()])\nplt.ylabel(\"No Patients\")\n#plt.yticks(values * value_increment, ['%d' % val for val in values])\nplt.xticks([1,2])\nplt.title('Number of patients per hospital')\nplt.show()\n\nprint(f\"We have {train[train.site_id == 1].patient_id.nunique()} from hospital 1 and {train[train.site_id == 2].patient_id.nunique()} from hospital 2\")","metadata":{"execution":{"iopub.status.busy":"2022-11-30T21:37:20.579951Z","iopub.execute_input":"2022-11-30T21:37:20.580931Z","iopub.status.idle":"2022-11-30T21:37:20.816277Z","shell.execute_reply.started":"2022-11-30T21:37:20.580867Z","shell.execute_reply":"2022-11-30T21:37:20.815139Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"How old are the patients?","metadata":{}},{"cell_type":"code","source":"ages = train[train.age.isnull() == False].groupby(\n    'patient_id').age.apply(lambda l: np.unique(l)[0])\nplt.figure(figsize=(20,8))\n\nsns.histplot(ages, color='orange', bins=60)\nplt.title('Age distribution of patients');","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-11-30T21:37:20.817715Z","iopub.execute_input":"2022-11-30T21:37:20.818551Z","iopub.status.idle":"2022-11-30T21:37:21.751461Z","shell.execute_reply.started":"2022-11-30T21:37:20.818502Z","shell.execute_reply":"2022-11-30T21:37:21.750097Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Are there age distribution differences between the two hospitals","metadata":{}},{"cell_type":"code","source":"ages_site_1 = train[train.site_id==1][train[train.site_id==1].age.isnull() == False].groupby(\n    'patient_id').age.apply(lambda l: np.unique(l)[0])\nages_site_2 = train[train.site_id==2][train[train.site_id==2].age.isnull() == False].groupby(\n    'patient_id').age.apply(lambda l: np.unique(l)[0])\n#plt.figure(figsize=(20,8))\nfig, ax = plt.subplots(figsize=(20,8))\nsns.histplot(ages_site_1, bins=60, ax=ax)\nsns.histplot(ages_site_2, color='orange', bins=60, ax=ax)\n\nplt.legend(loc='upper left', labels=['Hospital 1', 'Hospital 2'], fontsize=20)\nplt.title('Age distribution of patients per Hospital', fontsize=20);","metadata":{"execution":{"iopub.status.busy":"2022-11-30T21:37:21.752914Z","iopub.execute_input":"2022-11-30T21:37:21.754124Z","iopub.status.idle":"2022-11-30T21:37:22.891503Z","shell.execute_reply.started":"2022-11-30T21:37:21.754059Z","shell.execute_reply":"2022-11-30T21:37:22.890286Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can see that there is a slight difference in age distributions with hospital 1 have distribution being a little more skewed towards younger patients. Furthermore, in Hospital 2 there seem to be some missing values in the agg distribution. Could these be related to the missing values?","metadata":{}},{"cell_type":"markdown","source":"Let's see what is the youger and older patients for each hospital","metadata":{}},{"cell_type":"code","source":"print(f\"Age minimum for hospital 1: {train[train.site_id==1][train[train.site_id==1].age.isnull() == False].age.min()} \\nAge maximum for hospital 1: {train[train.site_id==1][train[train.site_id==1].age.isnull() == False].age.max()} \\nAge minimum for hospital 2: {train[train.site_id==2][train[train.site_id==2].age.isnull() == False].age.min()} \\nAge maximum for hospital 2: {train[train.site_id==2][train[train.site_id==2].age.isnull() == False].age.max()}\")","metadata":{"execution":{"iopub.status.busy":"2022-11-30T21:37:22.893243Z","iopub.execute_input":"2022-11-30T21:37:22.894615Z","iopub.status.idle":"2022-11-30T21:37:22.94604Z","shell.execute_reply.started":"2022-11-30T21:37:22.894535Z","shell.execute_reply":"2022-11-30T21:37:22.945166Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"How many patients do have cancer and how many images were taken per patient?","metadata":{}},{"cell_type":"code","source":"def has_cancer(l):\n    if len(l) == 1:\n        if l[0] == 0:\n            return False\n        elif l[0] == 1:\n            return True\n        else:\n            raise Exception\n    elif len(l) == 2:\n        return True\n    else:\n        raise Exception\n\npatient_cancer_map = train.groupby('patient_id').cancer.unique().apply(lambda l: has_cancer(l))","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-11-30T21:37:22.947083Z","iopub.execute_input":"2022-11-30T21:37:22.947897Z","iopub.status.idle":"2022-11-30T21:37:23.553101Z","shell.execute_reply.started":"2022-11-30T21:37:22.94786Z","shell.execute_reply":"2022-11-30T21:37:23.552019Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(2,1,figsize=(20,20))\nsns.set(font_scale=3)\nsns.countplot(patient_cancer_map, palette='Paired', ax=ax[0]);\nax[0].set_title('Number of patients with cancer', fontsize=20);\n\nsns.countplot(train.groupby('patient_id').size(), ax=ax[1])\nax[1].set_title('Number of images per patient', fontsize=20)\nax[1].set_xlabel('Number of images', fontsize=20)\nax[1].set_ylabel('Counts of patients', fontsize=20);","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-11-30T21:37:23.554282Z","iopub.execute_input":"2022-11-30T21:37:23.554613Z","iopub.status.idle":"2022-11-30T21:37:24.097242Z","shell.execute_reply.started":"2022-11-30T21:37:23.554567Z","shell.execute_reply":"2022-11-30T21:37:24.095363Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Comments\n- There classification problem is imbalanced\n- The majority of patients have 4 images (likely 2 (MLO view + CC view) for the left breast + 2 (MLO view + CC view) for the right breast). There are several cases with more than 4 images per patient. This may be related to additional views necessary to ensure a more confident radiological diagnosis (radiological diagnosis is different than the final diagnosis as it may need to be confirmed by additional radiological exams after a certain period of time - less suspicious cases- or biopsy in more suspicious cases). This will be confirmed next by looking at the image features like the **view**.","metadata":{}},{"cell_type":"markdown","source":"## Image features <a class=\"anchor\" id=\"image_features\"></a>","metadata":{}},{"cell_type":"markdown","source":"In terms of image features we have:\n- laterality - reflecting if image is from the left (L) or (R) breast\n- view - reflecting the image acquisition plane CC and MLO are the standard views. Other views present: ML - Mediolateral; LM - Lateromedial; AT - Mediolateral Oblique for Axillary Tail; LMO - Lateromedial Oblique; others (shown in first image of notebook).\n- density - discrete scale ranging from A to D. \n    - **A: Almost entirely fatty indicates that the breasts are almost entirely composed of fat.** About 1 in 10 women has this result; \n    - **B: Scattered areas of fibroglandular density indicates there are some scattered areas of density, but the majority of the breast tissue is nondense.** About 4 in 10 women have this result; \n    - **C: Heterogeneously dense indicates that there are some areas of nondense tissue, but that the majority of the breast tissue is dense.** About 4 in 10 women have this result; \n    - **D: Extremely dense indicates that nearly all of the breast tissue is dense.** About 1 in 10 women has this result.\n- implant - whether the patient has breast implants or not [additional information for breast mammography with implants](https://radiologyassistant.nl/breast/breast-prosthesis/breast-prosthesis-imaging#mammography)","metadata":{}},{"cell_type":"code","source":"fig, ax = plt.subplots(2,2,figsize=(20,20))\nsns.set(font_scale=2.5)\nsns.countplot(train.laterality, ax=ax[0,0], palette='Greens_r')\nsns.countplot(train.view, ax=ax[0,1], palette='Reds_r')\nsns.countplot(train.implant, ax=ax[1,0], palette='Blues_r')\nsns.countplot(train.density, ax=ax[1,1], palette='Purples_r');","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-11-30T21:37:24.099085Z","iopub.execute_input":"2022-11-30T21:37:24.099505Z","iopub.status.idle":"2022-11-30T21:37:24.981023Z","shell.execute_reply.started":"2022-11-30T21:37:24.099465Z","shell.execute_reply":"2022-11-30T21:37:24.980022Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Comments\n\n- Dataset ballanced in terms of laterality. \n- CC and MLO are the most common views with some cases of ML, LM, AT, LMO. \n- Small number of images obtained from patients with breast implants. \n- Majority of patients have breast densities in the categories of B and C (corresponding to cases of breast with scattered areas of fibroglandular density indicates there are some scattered areas of density, but the majority of the breast tissue is nondense, and breast predominantly dense with some areas of nondense tissue)\n\nAs intended before this may be of value and an estimation of the breast density may be important for the model. ","metadata":{}},{"cell_type":"markdown","source":"## Inspecting target features <a class=\"anchor\" id=\"target_features\"></a>\n\nWe have already seen that the number of patients with cancer is quiet low compared to the patients without. Let's see how it looks like on the image-level.","metadata":{}},{"cell_type":"code","source":"biopsy_counts = train.groupby('cancer').biopsy.value_counts().unstack().fillna(0) \nbiopsy_perc = biopsy_counts.transpose() / biopsy_counts.sum(axis=1)","metadata":{"execution":{"iopub.status.busy":"2022-11-30T21:37:24.982308Z","iopub.execute_input":"2022-11-30T21:37:24.983245Z","iopub.status.idle":"2022-11-30T21:37:24.998141Z","shell.execute_reply.started":"2022-11-30T21:37:24.983207Z","shell.execute_reply":"2022-11-30T21:37:24.996716Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(1,2,figsize=(20,5))\nsns.countplot(train.cancer, palette='Reds', ax=ax[0])\nax[0].set_title('Number of images displaying cancer');\nsns.heatmap(biopsy_perc.transpose(), ax=ax[1], annot=True, cmap='Oranges');","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-11-30T21:37:25.000321Z","iopub.execute_input":"2022-11-30T21:37:25.000718Z","iopub.status.idle":"2022-11-30T21:37:25.447153Z","shell.execute_reply.started":"2022-11-30T21:37:25.000683Z","shell.execute_reply":"2022-11-30T21:37:25.445806Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Insights\n\n* The number of images displaying cancer is very low. It should be even of lower percentage than the number of patients with cancer. \n* Looking at the biopsy feature we can say that all patients with cancer had a biopsy. But only around 3 % of images without cancer had resulted in a follow-up biopsy. Maybe we should better have a look at this feature on the patient-level. ","metadata":{}},{"cell_type":"code","source":"train.cancer.value_counts()/train.shape[0]","metadata":{"execution":{"iopub.status.busy":"2022-11-30T21:37:25.448862Z","iopub.execute_input":"2022-11-30T21:37:25.449223Z","iopub.status.idle":"2022-11-30T21:37:25.460776Z","shell.execute_reply.started":"2022-11-30T21:37:25.449192Z","shell.execute_reply":"2022-11-30T21:37:25.459487Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"patient_cancer_map.value_counts() / patient_cancer_map.shape[0]","metadata":{"execution":{"iopub.status.busy":"2022-11-30T21:37:25.462834Z","iopub.execute_input":"2022-11-30T21:37:25.463712Z","iopub.status.idle":"2022-11-30T21:37:25.474689Z","shell.execute_reply.started":"2022-11-30T21:37:25.463663Z","shell.execute_reply":"2022-11-30T21:37:25.473697Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As expected, the percentage of patients with cancer is higher than the percentage of images displaying cancer. ","metadata":{}},{"cell_type":"markdown","source":"## Create k-folds to train and validate models\n\n#### To correctly assess the model performance and given that patients have multiple images we need to create the folds at the patient-level and ideally the fold partitioning should be stratified","metadata":{}},{"cell_type":"markdown","source":"To achieve this we can use the StratifiedGroupKFold() class and create a new column in the train dataframe to indicate to which fold it belongs","metadata":{}},{"cell_type":"code","source":"n_folds = 5\nshuffle = True\nrandom_state_var = 1\ntrain['fold']= -1\nsgkf_cv = StratifiedGroupKFold(n_splits=n_folds, shuffle=shuffle, random_state=random_state_var)\n\nfor fold, (train_indx, valid_indx) in enumerate(sgkf_cv.split(X=train[\"image_id\"], y=train[\"cancer\"], groups=train[\"patient_id\"])):\n    train.loc[valid_indx, \"fold\"] = fold\n    \nassert train.groupby(['fold', 'cancer']).size().sum() == train.shape[0]","metadata":{"execution":{"iopub.status.busy":"2022-11-30T21:37:25.47597Z","iopub.execute_input":"2022-11-30T21:37:25.476568Z","iopub.status.idle":"2022-11-30T21:37:30.539627Z","shell.execute_reply.started":"2022-11-30T21:37:25.476535Z","shell.execute_reply":"2022-11-30T21:37:30.53836Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"display(train.fold.unique())","metadata":{"execution":{"iopub.status.busy":"2022-11-30T21:37:30.54112Z","iopub.execute_input":"2022-11-30T21:37:30.541487Z","iopub.status.idle":"2022-11-30T21:37:30.549407Z","shell.execute_reply.started":"2022-11-30T21:37:30.541448Z","shell.execute_reply":"2022-11-30T21:37:30.548079Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Get DICOM Images Properties\n","metadata":{}},{"cell_type":"code","source":"import pydicom\nimport os\n\ndcm_voxel_spacings_x = []\ndcm_voxel_spacings_y = []\ndcm_image_sizes_x = []\ndcm_image_sizes_y = []\n\npath_to_dcms = \"/kaggle/input/rsna-breast-cancer-detection/train_images\"\n\nfor indx in range(len(train.image_id)):\n    dcm_path = os.path.join(path_to_dcms, str(train.loc[indx, \"patient_id\"]), f\"{train.loc[indx, 'image_id']}.dcm\")\n    dcm = pydicom.dcmread(dcm_path)\n    print(dcm.Rows)\n    print(dcm.Columns)\n    print(dcm)\n    break","metadata":{"execution":{"iopub.status.busy":"2022-11-30T22:24:02.95507Z","iopub.execute_input":"2022-11-30T22:24:02.955773Z","iopub.status.idle":"2022-11-30T22:24:02.973201Z","shell.execute_reply.started":"2022-11-30T22:24:02.955734Z","shell.execute_reply":"2022-11-30T22:24:02.972229Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### It seems that we don't have voxel/pixel spacing information\n\nThis may change a bit the transformations that we can and cannot apply to the image (in specific cases one doesn't want to distord the image) ","metadata":{}},{"cell_type":"markdown","source":"# Stuff below is WIP!!!!! ","metadata":{}},{"cell_type":"code","source":"# WIP","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import torch\nfrom torch import nn\nfrom torch.utils.data import Dataset, DataLoader\nfrom torchvision import transforms, models","metadata":{"execution":{"iopub.status.busy":"2022-11-30T22:02:19.692293Z","iopub.execute_input":"2022-11-30T22:02:19.692982Z","iopub.status.idle":"2022-11-30T22:02:19.699057Z","shell.execute_reply.started":"2022-11-30T22:02:19.692943Z","shell.execute_reply":"2022-11-30T22:02:19.697728Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import torch\nfrom torch import optim\nfrom torch.utils.data import DataLoader, SequentialSampler\n\nimport monai\nfrom monai.optimizers.lr_scheduler import WarmupCosineSchedule\nfrom monai.utils import set_determinism\nfrom monai.data import DataLoader, Dataset\nfrom monai.data.utils import no_collation\nfrom monai.transforms import (\n    Compose,\n    EnsureChannelFirstd,\n    EnsureTyped,\n    LoadImaged,\n    Orientationd,\n    SaveImaged,\n    Spacingd,\n)\n\nprocess_transforms = Compose(\n        [\n            LoadImaged(\n                keys=[\"image\"],\n                meta_key_postfix=\"meta_dict\",\n                reader=\"itkreader\",\n                affine_lps_to_ras=False,\n            ),\n            EnsureChannelFirstd(keys=[\"image\"]),\n            EnsureTyped(keys=[\"image\"], dtype=torch.float16),\n            #Orientationd(keys=[\"image\"], axcodes=\"RAS\"),\n            Spacingd(keys=[\"image\"], pixdim=args.spacing, padding_mode=\"border\"),\n            SpacialPad()\n        ]\n    )\n\nclass CancerDataset(Dataset):\n    def __init__(self, df, transform=None):\n        super(CancerDataset, self).__init__()\n        self.df = df.copy()\n        self.transform = transform\n        self.path_to_dcms = \"/kaggle/input/rsna-breast-cancer-detection/train_images\"\n        \n    def __getitem__(self, idx):\n        dcm_path = os.path.join(self.path_to_dcms, str(self.df.loc[idx, \"patient_id\"]), f\"{self.df.loc[idx, 'image_id']}.dcm\")\n        dcm = dicom.dcmread(dcm_path)\n        dcm = dcm.pixel_array.astype(np.float32)\n        dcm = cv2.resize(dcm, (224,224))\n        if self.transform:\n            dcm = self.transform(dcm)\n        label = self.df.loc[idx, \"cancer\"]\n        label = torch.tensor(label, dtype=torch.int32)\n        return dcm, label\n    \n    def __len__(self):\n        return len(self.df)\n\ndef set_seed(seed):\n    # use monai's function to set the seed.\n    # since the function will also change the deterministic settings, which are unnecessary here,\n    # we need to modify the values back.\n    set_determinism(seed=seed)\n    torch.backends.cudnn.deterministic = False\n    torch.backends.cudnn.benchmark = True\n\ndef get_train_dataloader(train_dataset, cfg):\n\n    train_dataloader = DataLoader(\n        train_dataset,\n        sampler=None,\n        shuffle=True,\n        batch_size=cfg.batch_size,\n        num_workers=cfg.num_workers,\n        pin_memory=False,\n        collate_fn=None,\n        drop_last=cfg.drop_last,\n    )\n    print(f\"train: dataset {len(train_dataset)}, dataloader {len(train_dataloader)}\")\n    return train_dataloader","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for fold in range(n_folds):\n    print(f'------------fold no---------{fold}----------------------')\n    train_dataset_image_ids = ","metadata":{},"execution_count":null,"outputs":[]}]}