{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# RSNA Screening Mammography Breast Cancer Detection\n\n*Author: Eda AYDIN*","metadata":{}},{"cell_type":"markdown","source":"# A. Business Understanding\n\nAccording to the WHO, breast cancer is the most commonly occurring cancer worldwide. In 2020 alone, there were 2.3 million new breast cancer diagnoses and 685,000 deaths. Yet breast cancer mortality in high-income countries has dropped by 40% since the 1980s when health authorities implemented regular mammography screening in age groups considered at risk. Early detection and treatment are critical to reducing cancer fatalities, and your machine learning skills could help streamline the process radiologists use to evaluate screening mammograms.\n\nCurrently, early detection of breast cancer requires the expertise of highly-trained human observers, making screening mammography programs expensive to conduct. A looming shortage of radiologists in several countries will likely worsen this problem. Mammography screening also leads to a high incidence of false positive results. This can result in unnecessary anxiety, inconvenient follow-up care, extra imaging tests, and sometimes a need for tissue sampling (often a needle biopsy).\n\nThe competition host, the Radiological Society of North America (RSNA) is a non-profit organization that represents 31 radiologic subspecialties from 145 countries around the world. RSNA promotes excellence in patient care and health care delivery through education, research, and technological innovation.\n\nYour efforts in this competition could help extend the benefits of early detection to a broader population. Greater access could further reduce breast cancer mortality worldwide.\n\n\nThe goal of this competition is to indentify breast cancer. You'll train your model with screening mammograms obtained from regular screening.\n\nYour work improving the automation of detection in screening mammography may enable radiologists to be more accurate and efficient, improving the quality and safety of patient care. It could also help reduce costs and unnecessary medical procedures.\n\n*Resources from [website](https://www.kaggle.com/competitions/rsna-breast-cancer-detection)*","metadata":{}},{"cell_type":"markdown","source":"# B. Data Understanding\n\n**[train/test]_images / [patient_id] / [image_id].dcm** : The mammograms, in dicom format. You can expect roughly 8,000 patients in the hidden test. There are usually but not always 4 images per patient. Note that many of the image use the jpeg 2000 format which you may need special libraries to load.\n\n**sample_submission.csv** : A valid sample submission. Only the first few rows are available for downlaod.\n\n**[train/test].csv** : Metadata for each patient and image. Only the first few rows of the test set are available for download.\n\n* `site_id`: ID code for the source hospital.\n* `patient_id`: ID code for the patient.\n* `image_id`: ID code for the image.\n* `laterality`: Whether the image is of the left or bright breast.\n* `view`: The orientation of the image. The default for a screening exam is to capture two views per breast.\n    * `MLO` : Mediolateral oblique which is a type of mammogram view that shows the breast from the side and slightly from above.\n    * `CC`: Cranio-caudal which is a type of mammogram view that shows the breast from the top to bottom direction\n    * `AT`: Axillary tail which is an extension of the breast tissue that extends into the armpit.\n    * `LM`: Lateral Medial which is a type of mammogram view that shows the breast from the side and slightly from below.\n    * `ML`: Medial Lateral which is a type of mammogram view that shows the breast from the side and slightly from above.\n    * `LMO`: Lateromedial oblique which is a type of mammogram view that shows the breast from the side and slightly from below, opposite to the MLO view.\n* `age`: The patient's age in years.\n* `implant`: Whether or not the patient had breast implants. Site 1 only provides breast implant information at the patient level, not at the breast level.\n* `density`: A rating for how dense the breast tissue is, with A being the least dense and D being the most dense. Extremely dense tissue can make diagnosis more difficult. Only provided for train.\n    * `A`: Almost all fatty tissue, which makes it easy to detect abnormalities. About 10% of women have this density.\n    * `B`: Scattered areas of fibroglangular density. About 40% of women have this density.\n    * `C`: Heterogeneously dense breast tissue, which can make it more difficult to detect abnormalities. About 40% of women have this density.\n    * `D`: Extremely dense breast tissue, which can make it very difficult to detect abnormalities. About 10% of women have this density.\n* `machine_id`: An ID code for the imaging device.\n* `cancer`: Whether or not the breast was positive for malignant cancer. The target value. Only provided for train.\n* `biopsy`: Whether or not a follow-up biopsy was performed on the breast. Only provided for train.\n* `invasive`: If the breast is positive for cancer, whether or not the cancer proved to be invasive. Only provided for train.\n* `BIRADS`:\n    * 0 if the breast required follow-up,\n    * 1 if the breast was rated as negative for cancer, and\n    * 2 if the breast was rated as normal. Only provided for train.\n* `prediction_id`: The ID for the matching submission row. Multiple images will share the same prediction ID. Test only.\n* `difficult_negative_case`: True if the case was unusually difficult. Only provided for train.\n","metadata":{}},{"cell_type":"markdown","source":"### Import Libraries","metadata":{}},{"cell_type":"code","source":"!cp /kaggle/input/gdcm-conda-install/gdcm.tar .\n!tar -xvzf gdcm.tar\n!conda install --offline ./gdcm/gdcm-2.8.9-py37h71b2a6d_0.tar.bz2\n!rm -rf ./gdcm.tar","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:35:29.908442Z","iopub.execute_input":"2023-02-27T11:35:29.908947Z","iopub.status.idle":"2023-02-27T11:35:40.246813Z","shell.execute_reply.started":"2023-02-27T11:35:29.908898Z","shell.execute_reply":"2023-02-27T11:35:40.245383Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"try:\n    import pylibjpeg\nexcept:\n    !pip install /kaggle/input/rsna-2022-whl/{pydicom-2.3.0-py3-none-any.whl,pylibjpeg-1.4.0-py3-none-any.whl,python_gdcm-3.0.15-cp37-cp37m-manylinux_2_17_x86_64.manylinux2014_x86_64.whl}","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:36:04.626919Z","iopub.execute_input":"2023-02-27T11:36:04.628454Z","iopub.status.idle":"2023-02-27T11:36:40.125517Z","shell.execute_reply.started":"2023-02-27T11:36:04.628388Z","shell.execute_reply":"2023-02-27T11:36:40.123997Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#! pip install -U pylibjpeg pylibjpeg-openjpeg pylibjpeg-libjpeg pydicom python-gdcm\n#! pip install --upgrade pydicom","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:29:52.0723Z","iopub.execute_input":"2023-02-27T11:29:52.07275Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport pandas_profiling\nimport warnings\nimport os\n\n# visualization\nimport matplotlib.pyplot as plt\nimport matplotlib.ticker as ticker\nimport seaborn as sns\n%matplotlib inline\n\n# install gdcm to open the DICOM file\n# !pip install -qU python-gdcm pydicom pylibjpeg\n\nimport pydicom\nimport pylibjpeg\n\nimport glob\nimport cv2\n\nfrom path import Path\nfrom tqdm import tqdm,trange\nimport pydicom as dicom\n\nfrom sklearn.model_selection import train_test_split\n\nimport tensorflow as tf\nfrom tensorflow import keras\nfrom tensorflow.keras import layers\nfrom tensorflow.keras.preprocessing.image import ImageDataGenerator\nfrom keras_preprocessing.image.dataframe_iterator import DataFrameIterator\n\n\nDEVICE='GPU'\nwarnings.simplefilter(action=\"ignore\")\nos.environ[\"TF_CPP_MIN_LOG_LEVEL\"] = \"3\"","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:37:06.666045Z","iopub.execute_input":"2023-02-27T11:37:06.666506Z","iopub.status.idle":"2023-02-27T11:37:16.257207Z","shell.execute_reply.started":"2023-02-27T11:37:06.666466Z","shell.execute_reply":"2023-02-27T11:37:16.256328Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"url = \"/kaggle/input/rsna-breast-cancer-detection/\"\n\ntrain_dir = url+\"train_images/\"\ntrain_df = pd.read_csv(url+\"train.csv\")\ntrain_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:37:26.571867Z","iopub.execute_input":"2023-02-27T11:37:26.572838Z","iopub.status.idle":"2023-02-27T11:37:26.725057Z","shell.execute_reply.started":"2023-02-27T11:37:26.572801Z","shell.execute_reply":"2023-02-27T11:37:26.723704Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_dir = url+\"test_images/\"\ntest_df = pd.read_csv(url+\"test.csv\")\ntest_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:37:28.968686Z","iopub.execute_input":"2023-02-27T11:37:28.969652Z","iopub.status.idle":"2023-02-27T11:37:28.988377Z","shell.execute_reply.started":"2023-02-27T11:37:28.969612Z","shell.execute_reply":"2023-02-27T11:37:28.987391Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Shape\n\ntrain_df.shape","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:37:40.756245Z","iopub.execute_input":"2023-02-27T11:37:40.756655Z","iopub.status.idle":"2023-02-27T11:37:40.763685Z","shell.execute_reply.started":"2023-02-27T11:37:40.756623Z","shell.execute_reply":"2023-02-27T11:37:40.762479Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Types\ntrain_df.dtypes","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:37:41.929745Z","iopub.execute_input":"2023-02-27T11:37:41.930669Z","iopub.status.idle":"2023-02-27T11:37:41.939225Z","shell.execute_reply.started":"2023-02-27T11:37:41.930625Z","shell.execute_reply":"2023-02-27T11:37:41.937982Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# NA\ntrain_df.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:37:42.891502Z","iopub.execute_input":"2023-02-27T11:37:42.891911Z","iopub.status.idle":"2023-02-27T11:37:42.911029Z","shell.execute_reply.started":"2023-02-27T11:37:42.891879Z","shell.execute_reply":"2023-02-27T11:37:42.910111Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df[\"difficult_negative_case\"] = train_df[\"difficult_negative_case\"].astype(int)\ntrain_df.quantile([0, 0.05, 0.50, 0.95, 0.99, 1]).T","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:37:43.86092Z","iopub.execute_input":"2023-02-27T11:37:43.861848Z","iopub.status.idle":"2023-02-27T11:37:43.907078Z","shell.execute_reply.started":"2023-02-27T11:37:43.861806Z","shell.execute_reply":"2023-02-27T11:37:43.906258Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def grab_col_names(dataframe, categorical_threshold=10, cardinal_threshold=20):\n    \"\"\"\n    It gives the names of categorical, numerical and categorical but cardinal,nominal variables in the data set.\n    Note: Categorical variables but numerical variables are also included in categorical variables.\n\n    Parameters\n    ----------\n    dataframe : dataframe\n        The dataframe from which variables names are to be retrieved.\n    categorical_threshold : int, optional\n        class threshold for numeric but categorical variables\n    cardinal_threshold : int, optional\n        Class threshold for categorical but cardinal variables\n\n    Returns\n    -------\n        categorical_cols : list\n            Categorical variable list\n        numerical_cols : list\n            Numerical variable list\n        cardinal_cols : list\n            Categorical looking cardinal variable list\n\n    Examples\n    -------\n        import seaborn as sns\n        df = sns.load_titanic_dataset(\"iris\")\n        print(grab_col_names(df))\n\n    Notes\n    -------\n        categorical_cols + numerical_cols + cardinal_cols = total number of variables.\n        nominal_cols is inside categorical_cols\n        The sum of the 3 returned lists equals the total number of variables: categorical_cols + cardinal_cols = number of variables\n\n    \"\"\"\n\n    categorical_cols = [col for col in dataframe.columns if dataframe[col].dtypes == \"O\"]\n    nominal_cols = [col for col in dataframe.columns if\n                    dataframe[col].nunique() < categorical_threshold and dataframe[col].dtypes != \"O\"]\n    cardinal_cols = [col for col in dataframe.columns if\n                     dataframe[col].nunique() > cardinal_threshold and dataframe[col].dtypes == \"O\"]\n    categorical_cols = categorical_cols + nominal_cols\n    categorical_cols = [col for col in categorical_cols if col not in cardinal_cols]\n\n    # numerical_cols\n    numerical_cols = [col for col in dataframe.columns if dataframe[col].dtypes != \"O\"]\n    numerical_cols = [col for col in numerical_cols if col not in categorical_cols]\n\n    print(f\"Observations: {dataframe.shape[0]}\")\n    print(f\"Variables: {dataframe.shape[1]}\")\n    print(f'categorical_cols: {len(categorical_cols)}')\n    print(f'numerical_cols: {len(numerical_cols)}')\n    print(f'cardinal_cols: {len(cardinal_cols)}')\n    print(f'nominal_cols: {len(nominal_cols)}')\n    return categorical_cols, numerical_cols, cardinal_cols, nominal_cols","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:37:44.772471Z","iopub.execute_input":"2023-02-27T11:37:44.773227Z","iopub.status.idle":"2023-02-27T11:37:44.783796Z","shell.execute_reply.started":"2023-02-27T11:37:44.773183Z","shell.execute_reply":"2023-02-27T11:37:44.782667Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Değişken türlerin ayrıştırılması\ncategorical_cols, numerical_cols, cardinal_cols, nominal_cols = grab_col_names(train_df, categorical_threshold=5, cardinal_threshold=20)","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:37:45.613063Z","iopub.execute_input":"2023-02-27T11:37:45.6135Z","iopub.status.idle":"2023-02-27T11:37:45.659676Z","shell.execute_reply.started":"2023-02-27T11:37:45.613468Z","shell.execute_reply":"2023-02-27T11:37:45.658774Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Categorical column names: {}\".format(categorical_cols))\nprint(\"\\n\")\nprint(\"Numerical column names: {}\".format(numerical_cols))\nprint(\"\\n\")\nprint(\"Cardinal column names: {}\".format(cardinal_cols))\nprint(\"\\n\")\nprint(\"Nominal column names: {}\".format(nominal_cols))","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:37:46.466555Z","iopub.execute_input":"2023-02-27T11:37:46.467319Z","iopub.status.idle":"2023-02-27T11:37:46.475192Z","shell.execute_reply.started":"2023-02-27T11:37:46.467267Z","shell.execute_reply":"2023-02-27T11:37:46.473579Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\"\"\"\n\nDescription:\n-----------\n\nAlgorithm print out comprises missing ratios and unique values of each column i a given dataframe\n\n\nR&D:\n---\n\nAdd '#_infinity_' column to the dataframe\n\n\"\"\"\n\ndef MissingUniqueStatistics(df):\n\n  import io\n  import pandas as pd\n  import psutil, os, gc, time\n  import seaborn as sns\n  from IPython.display import display, HTML\n  # pd.set_option('display.max_colwidth', -1)\n  from io import BytesIO\n  import base64\n\n  print(\"MissingUniqueStatistics process has began:\\n\")\n  proc = psutil.Process(os.getpid())\n  gc.collect()\n  mem_0 = proc.memory_info().rss\n  start_time = time.time()\n\n  variable_name_list = []\n  total_entry_list = []\n  data_type_list = []\n  unique_values_list = []\n  number_of_unique_values_list = []\n  missing_value_number_list = []\n  missing_value_ratio_list = []\n  mean_list=[]\n  std_list=[]\n  min_list=[]\n  Q1_list=[]\n  Q2_list=[]\n  Q3_list=[]\n  max_list=[]\n\n  df_statistics = df.describe().copy()\n\n  for col in df.columns:\n\n    variable_name_list.append(col)\n    total_entry_list.append(df.loc[:,col].shape[0])\n    data_type_list.append(df.loc[:,col].dtype)\n    unique_values_list.append(list(df.loc[:,col].unique()))\n    number_of_unique_values_list.append(len(list(df.loc[:,col].unique())))\n    missing_value_number_list.append(df.loc[:,col].isna().sum())\n    missing_value_ratio_list.append(round((df.loc[:,col].isna().sum()/df.loc[:,col].shape[0]),4))\n\n    try:\n      mean_list.append(df_statistics.loc[:,col][1])\n      std_list.append(df_statistics.loc[:,col][2])\n      min_list.append(df_statistics.loc[:,col][3])\n      Q1_list.append(df_statistics.loc[:,col][4])\n      Q2_list.append(df_statistics.loc[:,col][5])\n      Q3_list.append(df_statistics.loc[:,col][6])\n      max_list.append(df_statistics.loc[:,col][7])\n    except:\n      mean_list.append('NaN')\n      std_list.append('NaN')\n      min_list.append('NaN')\n      Q1_list.append('NaN')\n      Q2_list.append('NaN')\n      Q3_list.append('NaN')\n      max_list.append('NaN')\n\n  data_info_df = pd.DataFrame({'Variable': variable_name_list,\n                               '#_Total_Entry':total_entry_list,\n                               '#_Missing_Value': missing_value_number_list,\n                               '%_Missing_Value':missing_value_ratio_list,\n                               'Data_Type': data_type_list,\n                               'Unique_Values': unique_values_list,\n                               '#_Unique_Values':number_of_unique_values_list,\n                               'Mean':mean_list,\n                               'STD':std_list,\n                               'Min':min_list,\n                               'Q1':Q1_list,\n                               'Q2':Q2_list,\n                               'Q3':Q3_list,\n                               'Max':max_list\n                               })\n\n  data_info_df = data_info_df.set_index(\"Variable\", inplace=False)\n\n\n  print('MissingUniqueStatistics process has been completed!')\n  print(\"--- in %s minutes ---\" % ((time.time() - start_time)/60))\n\n  return data_info_df.sort_values(by='%_Missing_Value', ascending=False)#, HTML(df.to_html(escape=False, formatters=dict(col=mapping)))","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:37:47.430335Z","iopub.execute_input":"2023-02-27T11:37:47.431117Z","iopub.status.idle":"2023-02-27T11:37:47.455054Z","shell.execute_reply.started":"2023-02-27T11:37:47.431074Z","shell.execute_reply":"2023-02-27T11:37:47.453899Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_info = MissingUniqueStatistics(train_df)\ndata_info","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:37:48.255911Z","iopub.execute_input":"2023-02-27T11:37:48.256375Z","iopub.status.idle":"2023-02-27T11:37:48.643255Z","shell.execute_reply.started":"2023-02-27T11:37:48.256333Z","shell.execute_reply":"2023-02-27T11:37:48.642113Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Veri setindeki sütunlar\nall_columns = train_df.columns.tolist()\n\n# Her sütunun özelliklerini kaydedeceğimiz bir sözlük\ncolumn_types = {}\n\nfor column in all_columns:\n    if column in nominal_cols:\n        column_types[column] = \"nominal\"\n    elif column in categorical_cols: \n        column_types[column] = \"categorical\"\n    elif column in numerical_cols:\n        column_types[column] = \"numeric\"\n    elif column in cardinal_cols:\n        column_types[column] = \"cardinal\"\n    else:\n        column_types[column] = \"unknown\"\n\ncolumn_types    ","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:37:49.59662Z","iopub.execute_input":"2023-02-27T11:37:49.59706Z","iopub.status.idle":"2023-02-27T11:37:49.608067Z","shell.execute_reply.started":"2023-02-27T11:37:49.597015Z","shell.execute_reply":"2023-02-27T11:37:49.606798Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Her sütunun özelliklerini data_info içerisine kaydet\ndf_dict = pd.DataFrame.from_dict(column_types, orient=\"index\",columns=[\"data_type\"] )","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:37:50.413727Z","iopub.execute_input":"2023-02-27T11:37:50.414468Z","iopub.status.idle":"2023-02-27T11:37:50.419851Z","shell.execute_reply.started":"2023-02-27T11:37:50.414429Z","shell.execute_reply":"2023-02-27T11:37:50.418964Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_info = pd.merge(data_info, df_dict, how=\"left\", left_on =\"Variable\", right_index = True)","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:37:51.81466Z","iopub.execute_input":"2023-02-27T11:37:51.81512Z","iopub.status.idle":"2023-02-27T11:37:51.828691Z","shell.execute_reply.started":"2023-02-27T11:37:51.815082Z","shell.execute_reply":"2023-02-27T11:37:51.827691Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_info","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:37:52.829617Z","iopub.execute_input":"2023-02-27T11:37:52.830774Z","iopub.status.idle":"2023-02-27T11:37:52.861103Z","shell.execute_reply.started":"2023-02-27T11:37:52.830733Z","shell.execute_reply":"2023-02-27T11:37:52.859653Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# C. Data Analysis ","metadata":{}},{"cell_type":"code","source":"%matplotlib inline\n# Histogram of the target categories\ndef histogram(df,feature):\n    #df = input(\"Enter a DataFrame name: \")\n    #col = input(\"Enter a target column name: \")\n    #df=eval(df)\n    ncount = len(df)\n    ax = sns.countplot(x = feature, data=df ,palette=\"hls\")\n    sns.set(font_scale=1)\n    ax.set_xlabel('Target Segments')\n    plt.xticks()\n    ax.set_ylabel('Number of Observations')\n    fig = plt.gcf()\n    fig.set_size_inches(12,10)\n    # Make twin axis\n    ax2=ax.twinx()\n    # Switch so count axis is on right, frequency on left\n    ax2.yaxis.tick_left()\n    ax.yaxis.tick_right()\n    # Also switch the labels over\n    ax.yaxis.set_label_position('right')\n    ax2.yaxis.set_label_position('left')\n    ax2.set_ylabel('Frequency [%]')\n    for p in ax.patches:\n        x=p.get_bbox().get_points()[:,0]\n        y=p.get_bbox().get_points()[1,1]\n        ax.annotate('{:.2f}%'.format(100.*y/ncount), (x.mean(), y),\n                ha='center', va='bottom') # set the alignment of the text\n    # Use a LinearLocator to ensure the correct number of ticks\n    ax.yaxis.set_major_locator(ticker.LinearLocator(11))\n    # Fix the frequency range to 0-100\n    ax2.set_ylim(0,100)\n    ax.set_ylim(0,ncount)\n    # And use a MultipleLocator to ensure a tick spacing of 10\n    ax2.yaxis.set_major_locator(ticker.MultipleLocator(10))\n    # Need to turn the grid on ax2 off, otherwise the gridlines end up on top of the bars\n    ax2.grid(None)\n    plt.title('Histogram of Binary Target Categories', fontsize=20, y=1.08)\n    plt.show()\n    #plt.savefig('col.png')\n    del ncount, x, y","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:37:54.657362Z","iopub.execute_input":"2023-02-27T11:37:54.657759Z","iopub.status.idle":"2023-02-27T11:37:54.676663Z","shell.execute_reply.started":"2023-02-27T11:37:54.657728Z","shell.execute_reply":"2023-02-27T11:37:54.675684Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"target = \"cancer\"\n\nhistogram(train_df, target)","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:37:55.505216Z","iopub.execute_input":"2023-02-27T11:37:55.505665Z","iopub.status.idle":"2023-02-27T11:37:55.89449Z","shell.execute_reply.started":"2023-02-27T11:37:55.505627Z","shell.execute_reply":"2023-02-27T11:37:55.893209Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"When we look at the histogram graph, we understand that we have imbalanced dataset. We will do oversampling to solve this problem.","metadata":{}},{"cell_type":"markdown","source":"# C.1 Preparatory Data Analysis","metadata":{}},{"cell_type":"code","source":"# Outlier Handling\n\ndef outlier_thresholds(dataframe, col_name, q1=0.25, q3=0.75):\n    quartile1 = dataframe[col_name].quantile(q1)\n    quartile3 = dataframe[col_name].quantile(q3)\n    interquartile_range = quartile3 - quartile1\n    up_limit = quartile3 + (1.5 * interquartile_range)\n    low_limit = quartile1 - (1.5 * interquartile_range)\n    return low_limit, up_limit","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:37:58.280787Z","iopub.execute_input":"2023-02-27T11:37:58.281198Z","iopub.status.idle":"2023-02-27T11:37:58.287992Z","shell.execute_reply.started":"2023-02-27T11:37:58.281166Z","shell.execute_reply":"2023-02-27T11:37:58.286882Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def check_outlier(dataframe, col_name):\n    low_limit, up_limit = outlier_thresholds(dataframe, col_name)\n    if dataframe[(dataframe[col_name] > up_limit) | (dataframe[col_name] < low_limit)].any(axis=None):\n        return True\n    else:\n        return False","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:38:00.183658Z","iopub.execute_input":"2023-02-27T11:38:00.184079Z","iopub.status.idle":"2023-02-27T11:38:00.190862Z","shell.execute_reply.started":"2023-02-27T11:38:00.184044Z","shell.execute_reply":"2023-02-27T11:38:00.189601Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for col in numerical_cols:\n    print(col, check_outlier(train_df, col))","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:38:01.178123Z","iopub.execute_input":"2023-02-27T11:38:01.178577Z","iopub.status.idle":"2023-02-27T11:38:01.209408Z","shell.execute_reply.started":"2023-02-27T11:38:01.17854Z","shell.execute_reply":"2023-02-27T11:38:01.20822Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Accessing Outliers\n\ndef grab_outliers(dataframe, col_name, index=False):\n    low, up = outlier_thresholds(dataframe, col_name)\n    if dataframe[((dataframe[col_name] < low) | (dataframe[col_name] > up))].shape[0] > 10:\n        print(dataframe[((dataframe[col_name] < low) | (dataframe[col_name] > up))].head())\n    else:\n        print(dataframe[((dataframe[col_name] < low) | (dataframe[col_name] > up))])\n\n    if index:\n        outlier_index = dataframe[((dataframe[col_name] < low) | (dataframe[col_name] > up))].index\n        return outlier_index","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:38:01.976422Z","iopub.execute_input":"2023-02-27T11:38:01.976886Z","iopub.status.idle":"2023-02-27T11:38:01.985719Z","shell.execute_reply.started":"2023-02-27T11:38:01.976842Z","shell.execute_reply":"2023-02-27T11:38:01.984411Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for col in numerical_cols:\n  grab_outliers(train_df,col,True)","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:38:02.721016Z","iopub.execute_input":"2023-02-27T11:38:02.722062Z","iopub.status.idle":"2023-02-27T11:38:02.767073Z","shell.execute_reply.started":"2023-02-27T11:38:02.722023Z","shell.execute_reply":"2023-02-27T11:38:02.765943Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Missing Data Handling\n\ndef missing_values_table(dataframe, na_name=False):\n\n    # Column Names with Missing Values\n    na_columns = [col for col in dataframe.columns if dataframe[col].isnull().sum() > 0]\n\n    # Number of Missing Values of One Column\n    number_of_missing_values = dataframe[na_columns].isnull().sum().sort_values(ascending=False)\n\n    # Percentage Distribution of Missing Data\n    percentage_ratio = (dataframe[na_columns].isnull().sum() / dataframe.shape[0] * 100).sort_values(ascending=False)\n\n    # Dataframe with Missing Data\n    missing_df = pd.concat([number_of_missing_values, np.round(percentage_ratio, 2)], axis=1, keys=['number_of_missing_values', 'percentage_ratio'])\n\n    print(missing_df, end=\"\\n\")\n\n    if na_name:\n        return na_columns","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:38:03.777242Z","iopub.execute_input":"2023-02-27T11:38:03.777695Z","iopub.status.idle":"2023-02-27T11:38:03.786653Z","shell.execute_reply.started":"2023-02-27T11:38:03.777659Z","shell.execute_reply":"2023-02-27T11:38:03.784955Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"missing_values_table(train_df)","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:38:04.601563Z","iopub.execute_input":"2023-02-27T11:38:04.60197Z","iopub.status.idle":"2023-02-27T11:38:04.631718Z","shell.execute_reply.started":"2023-02-27T11:38:04.601938Z","shell.execute_reply":"2023-02-27T11:38:04.630453Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Feature Engineering\n\ndef one_hot_encoder(dataframe, categorical_cols, drop_first=False):\n    \"\"\"\n    Apply One Hot Encoding to all specified categorical columns.\n\n    Args:\n        dataframe (dataframe): The dataframe from which variables names are to be retrieved.\n        categorical_col (string): The numerical column names are to be retrieved.\n        drop_first (bool, optional): Remove the first column after one hot encoding process to prevent overfitting. Defaults to False.\n\n    Returns:\n        dataframe: Return the new dataframe\n    \"\"\"\n    dataframe = pd.get_dummies(dataframe, columns=categorical_cols, drop_first=drop_first)\n    return dataframe","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:38:05.295897Z","iopub.execute_input":"2023-02-27T11:38:05.297026Z","iopub.status.idle":"2023-02-27T11:38:05.303324Z","shell.execute_reply.started":"2023-02-27T11:38:05.296984Z","shell.execute_reply":"2023-02-27T11:38:05.301847Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"one_hot_encoding_cols = [col for col in categorical_cols if 10 >= train_df[col].nunique() > 2]","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:38:06.128091Z","iopub.execute_input":"2023-02-27T11:38:06.12854Z","iopub.status.idle":"2023-02-27T11:38:06.149289Z","shell.execute_reply.started":"2023-02-27T11:38:06.128502Z","shell.execute_reply":"2023-02-27T11:38:06.147738Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"one_hot_encoding_cols","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:38:07.066863Z","iopub.execute_input":"2023-02-27T11:38:07.06746Z","iopub.status.idle":"2023-02-27T11:38:07.076286Z","shell.execute_reply.started":"2023-02-27T11:38:07.067415Z","shell.execute_reply":"2023-02-27T11:38:07.074794Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# C.2 Exploratory Data Analysis","metadata":{}},{"cell_type":"code","source":"# Analysis of Categorical Columns\n\ndef cat_summary(dataframe, col_name, plot=False,savefig=False):\n    \"\"\"\n    It gives summary of categorical columns with a plot.\n\n    Args:\n        dataframe (dataframe): The dataframe from which variables names are to be retrieved.\n        col_name (string): The column names from which features names are to be retrieved\n        plot (bool, optional): Plot the figure of the specified column. Defaults to False.\n        savefig(bool, optional): Save the figure of the specific column to the folder. Defaults to False\n    \"\"\"\n    print(pd.DataFrame({col_name: dataframe[col_name].value_counts(),\n                        \"Ratio\": 100 * dataframe[col_name].value_counts() / len(dataframe)}))\n    print(\"########################################## \\n\")\n\n    if plot:\n        ax = sns.countplot(x=dataframe[col_name], data=dataframe,\n                           order = dataframe[col_name].value_counts().index,\n                          hue = target)\n\n        ncount = len(dataframe)\n        sns.set(font_scale = 1)\n\n        for p in ax.patches:\n            x = p.get_bbox().get_points()[:, 0]\n            y = p.get_bbox().get_points()[1, 1]\n            ax.annotate('{:.2f}%'.format(100.*y/ncount), (x.mean(), y),ha='center', va='bottom')  # set the alignment of the text\n\n        # Use a LinearLocator to ensure the correct number of ticks\n        ax.yaxis.set_major_locator(ticker.LinearLocator(11))\n\n        plt.xticks(rotation=45)\n        plt.title(\"{} Count Graph.png\".format(col_name.capitalize()))\n        if savefig:\n            plt.savefig(\"{} Count Graph.png\".format(col_name.capitalize()))\n        plt.show(block=True)","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:38:13.348852Z","iopub.execute_input":"2023-02-27T11:38:13.34936Z","iopub.status.idle":"2023-02-27T11:38:13.36118Z","shell.execute_reply.started":"2023-02-27T11:38:13.349323Z","shell.execute_reply":"2023-02-27T11:38:13.359847Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Kategorik değişkenlerin incelenmesi\n\nfor col in categorical_cols:\n    cat_summary(train_df, col, plot=True, savefig=False)","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:38:13.363112Z","iopub.execute_input":"2023-02-27T11:38:13.36343Z","iopub.status.idle":"2023-02-27T11:38:15.806414Z","shell.execute_reply.started":"2023-02-27T11:38:13.363402Z","shell.execute_reply":"2023-02-27T11:38:15.805381Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Hedef Değişkenin Kategorik Değişkenlerle Analizi\n\ndef target_summary_with_categorical_data(dataframe, target, categorical_col):\n    \"\"\"\n    It gives the summary of specified categorical column name according to target column.\n\n    Args:\n        dataframe (dataframe): The dataframe from which variables names are to be retrieved.\n        target (string): The target column name are to be retrieved.\n        categorical_col (string): The categorical column names are to be retrieved.\n    \"\"\"\n    print(pd.DataFrame({\"TARGET_MEAN\": dataframe.groupby(categorical_col)[target].mean()}), end=\"\\n\\n\\n\")\n","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:38:15.808043Z","iopub.execute_input":"2023-02-27T11:38:15.808758Z","iopub.status.idle":"2023-02-27T11:38:15.815823Z","shell.execute_reply.started":"2023-02-27T11:38:15.808715Z","shell.execute_reply":"2023-02-27T11:38:15.814869Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for col in categorical_cols:\n    target_summary_with_categorical_data(dataframe=train_df, target = target, categorical_col=col)","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:38:15.819548Z","iopub.execute_input":"2023-02-27T11:38:15.820048Z","iopub.status.idle":"2023-02-27T11:38:15.879947Z","shell.execute_reply.started":"2023-02-27T11:38:15.819991Z","shell.execute_reply":"2023-02-27T11:38:15.878689Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Analysis of Result of Categorical Variables**\n\nThe given dataset has several categorical columns, and we can analyze each of them to gain insights into the data.\n\n* **Laterality**: This column shows the side of the breast that was imaged in the mammogram, either right or left. The ratio is almost equal, with a slight difference of 0.3% more images for the right side.\n\n* **View**: This column indicates the type of image, which includes MLO, CC, AT, LM, ML, and LMO. MLO and CC views are the most common types, with 51% and 49% respectively. The remaining views make up a very small percentage of the dataset.\n\n* **Density**: This column shows the density of the breast tissue, which is categorized as A, B, C, or D. The density level B is the most frequent, followed by C, A, and D, respectively.\n\n* **Site_id**: This column indicates the facility where the mammogram was taken, and there are only two facilities in the dataset. The majority of the mammograms were taken at site_id 1, with a ratio of 53.9%.\n\n* **Biopsy**: This column indicates whether a biopsy was performed on the patient. The vast majority of the mammograms did not have a biopsy performed, with a ratio of 94.6%.\n\n* **Invasive**: This column indicates whether the biopsy performed was invasive or not. Only 1.5% of the biopsies were invasive, with the majority being non-invasive.\n\n* **Implant**: This column indicates whether the patient had a breast implant. The majority of the patients did not have an implant, with a ratio of 97.3%.\n\n* **Difficult_negative_case**: This column indicates whether the mammogram was considered difficult to read, and if it was a negative case. The majority of the mammograms were not considered difficult negative cases, with a ratio of 85.9%.","metadata":{}},{"cell_type":"code","source":"# Analysis of Numerical Variables\n\ntrain_df[numerical_cols].describe().T","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:38:15.881947Z","iopub.execute_input":"2023-02-27T11:38:15.882425Z","iopub.status.idle":"2023-02-27T11:38:15.922019Z","shell.execute_reply.started":"2023-02-27T11:38:15.882379Z","shell.execute_reply":"2023-02-27T11:38:15.921012Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def num_summary(dataframe, col_name, plot=False, savefig=False):\n    \"\"\"\n    It gives the summary of numerical columns with a plot.\n\n    Args:\n        dataframe (dataframe): The dataframe from which variables names are to be retrieved.\n        col_name (string): The column names from which features names are to be retrieved\n        plot (bool, optional): Plot the figure of the specified column. Defaults to False.\n        savefig(bool, optional): Save the figure of the specific column to the folder. Defaults to False\n\n    \"\"\"\n    quantiles = [0.05, 0.10, 0.20, 0.30, 0.40, 0.50, 0.60, 0.70, 0.80, 0.90, 0.95, 0.99]\n    print(dataframe[col_name].describe(quantiles).T)\n\n    if plot:\n        dataframe[col_name].hist()\n        plt.xlabel(col_name)\n        plt.title(\"{} Histogram Graph.png\".format(col_name.capitalize()))\n        if savefig:\n            plt.savefig(\"{} Histogram Graph.png\".format(col_name.capitalize()))\n        plt.show(block=True)","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:38:15.923136Z","iopub.execute_input":"2023-02-27T11:38:15.923437Z","iopub.status.idle":"2023-02-27T11:38:15.931334Z","shell.execute_reply.started":"2023-02-27T11:38:15.923409Z","shell.execute_reply":"2023-02-27T11:38:15.93031Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for col in numerical_cols:\n    num_summary(train_df, col, True)","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:38:15.93241Z","iopub.execute_input":"2023-02-27T11:38:15.932754Z","iopub.status.idle":"2023-02-27T11:38:17.023721Z","shell.execute_reply.started":"2023-02-27T11:38:15.932725Z","shell.execute_reply":"2023-02-27T11:38:17.022677Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Analysis of Target Variable with Numerical Variables\n\ndef target_summary_with_numerical_data(dataframe, target, numerical_col):\n    \"\"\"\n    It gives the summary of specified numerical column name according to target column.\n\n    Args:\n        dataframe (dataframe): The dataframe from which variables names are to be retrieved.\n        target (string): The target column name are to be retrieved.\n        numerical_col (string): The numerical column names are to be retrieved.\n    \"\"\"\n    print(dataframe.groupby(target).agg({numerical_col: \"mean\"}), end=\"\\n\\n\\n\")","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:38:17.02521Z","iopub.execute_input":"2023-02-27T11:38:17.025758Z","iopub.status.idle":"2023-02-27T11:38:17.030804Z","shell.execute_reply.started":"2023-02-27T11:38:17.025726Z","shell.execute_reply":"2023-02-27T11:38:17.029795Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for col in numerical_cols:\n    target_summary_with_numerical_data(train_df, target ,col)","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:38:17.035563Z","iopub.execute_input":"2023-02-27T11:38:17.035901Z","iopub.status.idle":"2023-02-27T11:38:17.061944Z","shell.execute_reply.started":"2023-02-27T11:38:17.035872Z","shell.execute_reply":"2023-02-27T11:38:17.060955Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Analysis of Numerical Columns**\n\nThe dataset contains four numerical columns: patient_id, image_id, age, and machine_id.\n\n* **age**: The average age of patients with cancer (1) is higher than patients without cancer (0), with cancer patients having an average age of 63.678756 compared to 58.432808 for non-cancer patients.\n    * The '**age**' column has a mean value of 58.54 and a standard deviation of 10.05. The minimum and maximum values are 26 and 89, respectively. The majority of values fall within the 10th and 90th percentile range of 45 and 71, respectively.","metadata":{}},{"cell_type":"code","source":"# Multivariate Plots (Scatter Plot / Histogram)\n\n# scatter plot\nplt.figure(figsize=(8,6))\nsns.scatterplot(data=train_df, x='age', y='density', hue='cancer')\nplt.title('Breast Cancer Detection - Age vs Density', fontsize=14)\nplt.xlabel('Age', fontsize=12)\nplt.ylabel('Density', fontsize=12)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:38:17.063572Z","iopub.execute_input":"2023-02-27T11:38:17.063898Z","iopub.status.idle":"2023-02-27T11:38:18.376511Z","shell.execute_reply.started":"2023-02-27T11:38:17.063868Z","shell.execute_reply":"2023-02-27T11:38:18.375355Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# scatter plot\nplt.figure(figsize=(8,6))\nsns.scatterplot(data=train_df, x='age', y='BIRADS', hue='cancer')\nplt.title('Breast Cancer Detection - Age vs BIRADS', fontsize=14)\nplt.xlabel('Age', fontsize=12)\nplt.ylabel('BIRADS', fontsize=12)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:38:18.377685Z","iopub.execute_input":"2023-02-27T11:38:18.378075Z","iopub.status.idle":"2023-02-27T11:38:19.458311Z","shell.execute_reply.started":"2023-02-27T11:38:18.378039Z","shell.execute_reply":"2023-02-27T11:38:19.457193Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Resource: https://www.kaggle.com/code/shujaat91/tensorflow-basic-eda-and-training-vit?scriptVersionId=115678766&cellId=7\nfig, ax1 = plt.subplots()\n\ncolor = 'tab:blue'\nax1.set_xlabel('age')\nax1.set_ylabel('cancer = 0 count', color= color)\nax1.hist(train_df.loc[train_df['cancer'] == 0, 'age'].dropna(), bins = 30, alpha = 0.5, label = '0', color = color)\nax1.tick_params(axis='y', labelcolor=color)\n\nax2 = ax1.twinx()\n\ncolor = 'tab:green'\nax2.set_ylabel('cancer = 1 count', color=color) \nax2.hist(train_df.loc[train_df['cancer'] == 1, 'age'].dropna(), bins=30, alpha=0.8, label='1', color=color)\nax2.tick_params(axis='y', labelcolor=color)\n\nfig.tight_layout()  \nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:38:19.459572Z","iopub.execute_input":"2023-02-27T11:38:19.459876Z","iopub.status.idle":"2023-02-27T11:38:20.065106Z","shell.execute_reply.started":"2023-02-27T11:38:19.459849Z","shell.execute_reply":"2023-02-27T11:38:20.063986Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax1 = plt.subplots()\n\ncolor = 'tab:blue'\nax1.set_xlabel('density')\nax1.set_ylabel('cancer = 0 count', color= color)\nax1.hist(train_df.loc[train_df['cancer'] == 0, 'density'].dropna(), bins = 30, alpha = 0.5, label = '0', color = color)\nax1.tick_params(axis='y', labelcolor=color)\n\nax2 = ax1.twinx()\n\ncolor = 'tab:green'\nax2.set_ylabel('cancer = 1 count', color=color) \nax2.hist(train_df.loc[train_df['cancer'] == 1, 'density'].dropna(), bins=30, alpha=0.8, label='1', color=color)\nax2.tick_params(axis='y', labelcolor=color)\n\nfig.tight_layout()  \nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:38:20.066881Z","iopub.execute_input":"2023-02-27T11:38:20.067596Z","iopub.status.idle":"2023-02-27T11:38:20.580213Z","shell.execute_reply.started":"2023-02-27T11:38:20.067562Z","shell.execute_reply":"2023-02-27T11:38:20.579196Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Calculate the correlation matrix\ncorr = train_df[['cancer', 'age', 'view', 'laterality', 'biopsy', 'BIRADS', 'implant', 'invasive', 'density', 'difficult_negative_case']].corr()\n\n# Heatmap of the correlation matrix\nsns.heatmap(corr)","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:38:20.582102Z","iopub.execute_input":"2023-02-27T11:38:20.583243Z","iopub.status.idle":"2023-02-27T11:38:20.980315Z","shell.execute_reply.started":"2023-02-27T11:38:20.583197Z","shell.execute_reply":"2023-02-27T11:38:20.979275Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# C.3 Confirmatory Data Analysis\n\n","metadata":{}},{"cell_type":"code","source":"# Chi-Square test for Nominal Features\n\nnominal_cols","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:38:20.981909Z","iopub.execute_input":"2023-02-27T11:38:20.982504Z","iopub.status.idle":"2023-02-27T11:38:20.989949Z","shell.execute_reply.started":"2023-02-27T11:38:20.982457Z","shell.execute_reply":"2023-02-27T11:38:20.989199Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def chi2_by_hand(df, col1, col2):\n    #---create the contingency table---\n    df_cont = pd.crosstab(index = df[col1], columns = df[col2])\n    display(df_cont)\n    #---calculate degree of freedom---\n    degree_f = (df_cont.shape[0]-1) * (df_cont.shape[1]-1)\n    #---sum up the totals for row and columns---\n    df_cont.loc[:,'Total']= df_cont.sum(axis=1)\n    df_cont.loc['Total']= df_cont.sum()\n\n    #---create the expected value dataframe---\n    df_exp = df_cont.copy()\n    df_exp.iloc[:,:] = np.multiply.outer(\n        df_cont.sum(1).values,df_cont.sum().values) / df_cont.sum().sum()\n\n    # calculate chi-square values\n    df_chi2 = ((df_cont - df_exp)**2) / df_exp\n    df_chi2.loc[:,'Total']= df_chi2.sum(axis=1)\n    df_chi2.loc['Total']= df_chi2.sum()\n\n    #---get chi-square score---\n    chi_square_score = df_chi2.iloc[:-1,:-1].sum().sum()\n\n    #---calculate the p-value---\n    from scipy import stats\n    p = stats.distributions.chi2.sf(chi_square_score, degree_f)\n\n    return chi_square_score, degree_f, p","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:38:20.991228Z","iopub.execute_input":"2023-02-27T11:38:20.991568Z","iopub.status.idle":"2023-02-27T11:38:21.001718Z","shell.execute_reply.started":"2023-02-27T11:38:20.991528Z","shell.execute_reply":"2023-02-27T11:38:21.000801Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for col in nominal_cols:\n    chi_score, degree_f, p = chi2_by_hand(train_df,col,target)\n    print(f'Column: {col}, Chi2_score: {chi_score}, Degrees of freedom: {degree_f}, p-value: {p}')\n","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:38:21.003445Z","iopub.execute_input":"2023-02-27T11:38:21.004409Z","iopub.status.idle":"2023-02-27T11:38:21.209475Z","shell.execute_reply.started":"2023-02-27T11:38:21.004292Z","shell.execute_reply":"2023-02-27T11:38:21.208424Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The chi-square test was applied to the nominal features of the dataset to assess the statistical significance of their association with the target variable. The results showed that all nominal features, except for \"site_id\" and \"implant\", were strongly associated with the target variable \"cancer\", \"biopsy\", \"invasive\", \"BIRADS\", and \"difficult_negative_case\". The \"BIRADS\" feature had a p-value of 0.0, indicating a highly significant association with the target variable. The \"site_id\" and \"implant\" features also had a statistically significant association, but to a lesser extent compared to the other features. These results suggest that the nominal features can be useful in predicting the target variable and should be included in the model building process.","metadata":{}},{"cell_type":"code","source":"# ANOVA test for Numerical Features\n\nimport statsmodels.api as sm\nfrom statsmodels.formula.api import ols\n\nfor col in numerical_cols:\n    model = ols(target + '~' + col, data = train_df).fit() #Oridnary least square method\n    result_anova = sm.stats.anova_lm(model) # ANOVA Test\n    print(result_anova)\n    print(\"\\n\")","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:38:21.210957Z","iopub.execute_input":"2023-02-27T11:38:21.211321Z","iopub.status.idle":"2023-02-27T11:38:21.92227Z","shell.execute_reply.started":"2023-02-27T11:38:21.211289Z","shell.execute_reply":"2023-02-27T11:38:21.921052Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The ANOVA test was conducted to investigate the statistical significance of the difference between the groups within each numerical feature. The null hypothesis states that there is no significant difference between the groups, while the alternative hypothesis suggests that there is a difference. For the 'patient_id' and 'image_id' columns, the p-values were greater than the significance level of 0.05, indicating that there was no significant difference between the groups. On the other hand, for the 'age' and 'machine_id' columns, the p-values were less than 0.05, indicating a statistically significant difference between the groups. Therefore, we can reject the null hypothesis for 'age' and 'machine_id' and conclude that there is a significant difference between the groups.","metadata":{}},{"cell_type":"markdown","source":"# C.4 Data Analysis of DICOM images","metadata":{}},{"cell_type":"code","source":"train_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:38:21.923778Z","iopub.execute_input":"2023-02-27T11:38:21.925002Z","iopub.status.idle":"2023-02-27T11:38:21.955411Z","shell.execute_reply.started":"2023-02-27T11:38:21.924957Z","shell.execute_reply":"2023-02-27T11:38:21.954208Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import os\nimport pandas as pd\n\ntrain_dir = '../input/rsna-breast-cancer-detection/train_images/'\n\n# Create a new DataFrame with patient_id, image_id, cancer, and image_path columns\nimage_df = train_df[['patient_id', 'image_id', 'cancer']].copy()\nimage_df['long_image_path'] = train_dir + image_df['patient_id'].astype(str) + '/' + image_df['image_id'].astype(str) + '.dcm'\nimage_df[\"short_image_path\"] = image_df['patient_id'].astype(str) + '/' + image_df['image_id'].astype(str) + '.dcm'\n\n# Check the resulting DataFrame\nimage_df.head()\n","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:38:21.957023Z","iopub.execute_input":"2023-02-27T11:38:21.958119Z","iopub.status.idle":"2023-02-27T11:38:22.209923Z","shell.execute_reply.started":"2023-02-27T11:38:21.958071Z","shell.execute_reply":"2023-02-27T11:38:22.209168Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"image_df.loc[image_df[\"cancer\"] == 0, \"cancer\"] = \"negative\"\nimage_df.loc[image_df[\"cancer\"] == 1, \"cancer\"] = \"positive\"","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:38:22.211338Z","iopub.execute_input":"2023-02-27T11:38:22.212471Z","iopub.status.idle":"2023-02-27T11:38:22.229011Z","shell.execute_reply.started":"2023-02-27T11:38:22.212423Z","shell.execute_reply":"2023-02-27T11:38:22.227851Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"training_df = image_df.head(1000)\ntraining_df","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:38:22.230445Z","iopub.execute_input":"2023-02-27T11:38:22.231331Z","iopub.status.idle":"2023-02-27T11:38:22.246654Z","shell.execute_reply.started":"2023-02-27T11:38:22.231296Z","shell.execute_reply":"2023-02-27T11:38:22.245402Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import random\nimport matplotlib.pyplot as plt\nimport pydicom\nimport numpy as np\n\n# Set seed for randomization\nrandom.seed(1337)\n\n# Select 5 random patients with cancer = 0\nrandom_patients_0 = image_df.loc[image_df['cancer'] == \"negative\", 'patient_id'].unique()\nrandom_patients_0 = random.sample(list(random_patients_0), 5)\n\n# Select 5 random patients with cancer = 1\nrandom_patients_1 = image_df.loc[image_df['cancer'] == \"positive\", 'patient_id'].unique()\nrandom_patients_1 = random.sample(list(random_patients_1), 5)\n\n# Create figure and axes\nfig, axs = plt.subplots(nrows=2, ncols=5, figsize=(15, 6))\naxs = axs.ravel()\n\narthopod_types = {0: 'Not Malignant', 1: 'Malignant Cancer'}\n\n# Iterate over random patients and show random images for cancer = negative\nfor i, patient_id in enumerate(random_patients_0):\n    # Get all images for the patient with cancer = negative\n    patient_images = image_df.loc[(image_df['patient_id'] == patient_id) & (image_df['cancer'] == \"negative\")]\n    # Select a random image\n    random_image = patient_images.sample()\n    # Load the DICOM file\n    dcm_file = random_image['long_image_path'].values[0]\n    ds = pydicom.dcmread(dcm_file)\n    # Convert pixel data to numpy array\n    img = ds.pixel_array.astype(float)\n    # Rescale pixel values to 0-1 range\n    img /= np.max(img)\n    # Show the image\n    # cmap='viridis'\n    axs[i].imshow(img, cmap=\"viridis\")\n    axs[i].axis('off')\n    axs[i].set_title(f'Cancer: {random_image[\"cancer\"].values[0]}')\n\n\n# Iterate over random patients and show random images for cancer = positive\nfor i, patient_id in enumerate(random_patients_1):\n    # Get all images for the patient with cancer = positive\n    patient_images = image_df.loc[(image_df['patient_id'] == patient_id) & (image_df['cancer'] == \"positive\")]\n    # Select a random image\n    random_image = patient_images.sample()\n    # Load the DICOM file\n    dcm_file = random_image['long_image_path'].values[0]\n    ds = pydicom.dcmread(dcm_file)\n    # Convert pixel data to numpy array\n    img = ds.pixel_array.astype(float)\n    # Rescale pixel values to 0-1 range\n    img /= np.max(img)\n    # Show the image\n    axs[i+5].imshow(img, cmap=\"viridis\")\n    axs[i+5].axis('off')\n    axs[i+5].set_title(f'Cancer: {random_image[\"cancer\"].values[0]}')\n\n# Add supertitle for the rows\nplt.figtext(0.5,0.99, \"Random DICOM Breast Images for Patients with Not Malignant\", ha=\"center\", va=\"top\", fontsize=14, fontweight='bold')\nplt.figtext(0.5,0.52, \"Random DICOM Breast Images for Patients with Malignant Cancer\", ha=\"center\", va=\"top\", fontsize=14, fontweight='bold')\n\n# Adjust spacing between subplots\nplt.subplots_adjust(wspace=0.1, hspace=0.5)\n\n# Show the plot\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:38:22.247936Z","iopub.execute_input":"2023-02-27T11:38:22.248292Z","iopub.status.idle":"2023-02-27T11:38:37.170462Z","shell.execute_reply.started":"2023-02-27T11:38:22.24826Z","shell.execute_reply":"2023-02-27T11:38:37.169422Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# D. Data Augmentation","metadata":{}},{"cell_type":"code","source":"# https://www.kaggle.com/code/shujaat91/tensorflow-basic-eda-and-training-vit?scriptVersionId=115678766&cellId=23\n\n# Creating a custom generator class extends DataFrameIterator\nclass DCMDataFrameIterator(DataFrameIterator):\n    \n    # Constructor for the class that initializes\n    def __init__(self, *arg, **kwargs):\n        self.white_list_formats = ('dcm')\n        super(DCMDataFrameIterator, self).__init__(*arg, **kwargs)\n        self.dataframe = kwargs['dataframe']\n        self.x = self.dataframe[kwargs['x_col']]\n        self.y = self.dataframe[kwargs['y_col']]\n        self.color_mode = kwargs['color_mode']\n        self.target_size = kwargs['target_size']\n        self.directory = kwargs['directory']\n    \n    # Method returns a batch of transformed samples\n    def _get_batches_of_transformed_samples(self, indices_array):\n        batch_x = np.array([self.read_dcm_as_array(self.directory + dcm_path, self.target_size, color_mode=self.color_mode) for dcm_path in self.x.iloc[indices_array]])\n        batch_y = np.array([np.array(0) if i == \"negative\" else np.array(1) for i in self.y.iloc[indices_array]])\n        batch_y = batch_y.reshape((-1, 1))\n        \n        # Apply random transformations to the batch if an image data generator is provided\n        if self.image_data_generator is not None:\n            for i, (x, y) in enumerate(zip(batch_x, batch_y)):\n                transform_params = self.image_data_generator.get_random_transform(x.shape)\n                batch_x[i] = self.image_data_generator.apply_transform(x, transform_params)\n                \n        return batch_x, batch_y\n    \n    # Method that reads DICOM images as numpy array\n    @staticmethod\n    def read_dcm_as_array(dcm_path, target_size=(256, 256), color_mode='rgb'):\n        image_array = pydicom.dcmread(dcm_path).pixel_array\n        image_array = cv2.resize(image_array, target_size, interpolation=cv2.INTER_NEAREST)\n        image_array = np.expand_dims(image_array, -1)\n        if color_mode == 'rgb':\n            image_array = cv2.cvtColor(image_array, cv2.COLOR_GRAY2RGB)\n        return image_array\n    \n\n# Setting up augmentation parameters for the train, validation and test data generators\ntrain_augmentation_parameters = {\n    'rescale': 1.0/255.0,\n    'rotation_range': 10,\n    'zoom_range': 0.2,\n    'horizontal_flip': True,\n    'fill_mode': 'nearest',\n    'brightness_range': [0.8, 1.2],\n    'validation_split': 0.2\n}\n\nvalid_augmentation_parameters = {\n    'rescale': 1.0/255.0,\n    'validation_split': 0.2\n}\n\ntest_augmentation_parameters = {\n    'rescale': 1.0/255.0\n}\n\n# Setting up constants for the train, validation and test data generators\nBATCH_SIZE = 32\nCLASS_MODE = 'binary'\nCOLOR_MODE = 'grayscale'\nTARGET_SIZE = (48, 48)\nEPOCHS = 10\nSEED = 1337\n\ntrain_consts = {\n    'seed': SEED,\n    'batch_size': BATCH_SIZE,\n    'class_mode': CLASS_MODE,\n    'color_mode': COLOR_MODE,\n    'target_size': TARGET_SIZE,  \n    'subset': 'training'\n}\n\nvalid_consts = {\n    'seed': SEED,\n    'batch_size': BATCH_SIZE,\n    'class_mode': CLASS_MODE,\n    'color_mode': COLOR_MODE,\n    'target_size': TARGET_SIZE, \n    'subset': 'validation'\n}\n\ntest_consts = {\n    'batch_size': 1,\n    'class_mode': CLASS_MODE,\n    'color_mode': COLOR_MODE,\n    'target_size': TARGET_SIZE,\n    'shuffle': False\n}\n\n# Initializing an ImageDataGenerator object for augmenting the training and validation data\ntrain_augmenter = ImageDataGenerator(**train_augmentation_parameters)\nvalid_augmenter = ImageDataGenerator(**valid_augmentation_parameters)\n\n\n# Creating a custom iterator for the training data using DCMDataFrameIterator\ntrain_generator = DCMDataFrameIterator(dataframe=training_df,\n                             x_col='short_image_path',\n                             y_col='cancer',\n                             directory=train_dir,\n                             image_data_generator=train_augmenter,\n                             **train_consts)\n\nvalid_generator = DCMDataFrameIterator(dataframe=training_df,\n                             x_col='short_image_path',\n                             y_col='cancer',\n                             directory=train_dir,\n                             image_data_generator=valid_augmenter,\n                             **valid_consts)\n","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:38:37.175843Z","iopub.execute_input":"2023-02-27T11:38:37.176228Z","iopub.status.idle":"2023-02-27T11:38:38.954729Z","shell.execute_reply.started":"2023-02-27T11:38:37.176196Z","shell.execute_reply":"2023-02-27T11:38:38.953378Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Set the figure size and spacing between subplots\nfig = plt.figure(figsize=(10, 4))\nfig.subplots_adjust(wspace=0.4, hspace=0.6)\n\n# Loop through the first row (train data)\nfor i in range(5):\n    # Get the batch of images and labels\n    batch_x, batch_y = next(train_generator)\n    # Plot the images in a subplot\n    ax = fig.add_subplot(2, 5, i+1)\n    ax.imshow(batch_x[0], cmap='gray')\n    ax.axis('off')\n    # Set the title of the subplot to the label\n    ax.set_title(f\"Cancer: {str(batch_y[0][0])}\")\n    \n# Loop through the second row (validation data)\nfor i in range(5):\n    # Get the batch of images and labels\n    batch_x, batch_y = next(valid_generator)\n    # Plot the images in a subplot\n    ax = fig.add_subplot(2, 5, i+6)\n    ax.imshow(batch_x[0], cmap='gray')\n    ax.axis('off')\n    # Set the title of the subplot to the label\n    ax.set_title(f\"Cancer: {str(batch_y[0][0])}\")\n\n    \n# Add supertitle for the rows\nplt.figtext(0.5,0.99, \"Genereated Training Images\", ha=\"center\", va=\"top\", fontsize=14, fontweight='bold')\nplt.figtext(0.5,0.52, \"Generated Validation Images\", ha=\"center\", va=\"top\", fontsize=14, fontweight='bold')\n\n\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:38:38.956727Z","iopub.execute_input":"2023-02-27T11:38:38.957553Z","iopub.status.idle":"2023-02-27T11:42:45.4302Z","shell.execute_reply.started":"2023-02-27T11:38:38.957502Z","shell.execute_reply":"2023-02-27T11:42:45.429341Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# E. Modeling","metadata":{}},{"cell_type":"code","source":"# Parameters for model\n\n# The number of training examples in one forward/backward pass.\n# Larger batch size require more memory but can result in faster training times and more stale gradients.\nbatch_size = 256\n\n# The sşze (in pixels) to which the input images will be resized.\n# This can affect the performance of the model, as larger images can provide more detail but require more computational resources.\nimage_size = 224\n\n# The size (in pixels) of the patches that will be extracted from the input images.\n#  These patches are then processed by the transformer layers in the model.\npatch_size = 16\n\n# The dimensionality of the embeddings that will be learned by the model.\n# These embeddings represent each path of the input image and are used as input to the transformer layers.\nprojection_dim = 64\n\n# The number of units in each layer of the transformer.\n# The first value in the list corresponds to the number of units in the first hidden layer, \n# the second value correspond to the number of units in the second layer, and so on.\ntransformer_units = [projection_dim * 2, projection_dim,]\n\n# The number of transformers layer in the model\n# More layers can capture more complex relationships in the data but require more computation.\ntransformer_layers = 8\n\n# The number of units in each layer of the MLP (multi-layer perceptron) head of the model.\n# This head takes the output of the transformer layers and processes it to produce the final classification output.\nmlp_head_units = [2048, 1024]","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:42:45.431415Z","iopub.execute_input":"2023-02-27T11:42:45.432002Z","iopub.status.idle":"2023-02-27T11:42:45.438857Z","shell.execute_reply.started":"2023-02-27T11:42:45.431966Z","shell.execute_reply":"2023-02-27T11:42:45.437706Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\"\"\"\nA functions called \"MLP\" is defined which takes in three arguments:\n- x : input tensor to the MLP layers\n- hidden_units: a list of integers representing the number of units in each hidden layer\n- dropout_rate : a float representing the dropout rate to be applied after each hidden layer\n\"\"\"\ndef MLPBlock(x, hidden_units, dropout_rate):\n    \"\"\"\n    A for loop is defined to iterate over each element (integer) in the \"hidden units\" list.\n    Inside the for loop:\n    - a fully connected Dense layer is created with the specified number of units and gelu activation function\n    - a dropout layer is applied to the output of the Dense layer with specified dropout rate\n    \"\"\"\n    for units in hidden_units:\n        x = layers.Dense(units, activation=tf.nn.gelu)(x)\n        x = layers.Dropout(dropout_rate)(x)\n    return x","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:42:45.440517Z","iopub.execute_input":"2023-02-27T11:42:45.440901Z","iopub.status.idle":"2023-02-27T11:42:45.458168Z","shell.execute_reply.started":"2023-02-27T11:42:45.440869Z","shell.execute_reply":"2023-02-27T11:42:45.457361Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Define a custom layer for patch extraction from input image\nclass Patches(layers.Layer):\n    def __init__(self, patch_size):\n        super(Patches, self).__init__()\n        self.patch_size = patch_size\n\n    # The call function extracts patches from the input images and reshapes them into a 2D tensor\n    def call(self, images):\n        batch_size = tf.shape(images)[0]\n        patches = tf.image.extract_patches(\n            images=images,\n            sizes=[1, self.patch_size, self.patch_size, 1],\n            strides=[1, self.patch_size, self.patch_size, 1],\n            rates=[1, 1, 1, 1],\n            padding=\"VALID\",\n        )\n        patch_dims = patches.shape[-1]\n        patches = tf.reshape(patches, [batch_size, -1, patch_dims])\n        return patches\n        ","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:42:45.45951Z","iopub.execute_input":"2023-02-27T11:42:45.460648Z","iopub.status.idle":"2023-02-27T11:42:45.476246Z","shell.execute_reply.started":"2023-02-27T11:42:45.46061Z","shell.execute_reply":"2023-02-27T11:42:45.475052Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Define a custome layer for encoding patches with projection and position embedding\nclass PatchEncoder(layers.Layer):\n    \n    def __init__(self, num_patches, projection_dim):\n        super(PatchEncoder, self).__init__()\n        self.num_patches = num_patches\n        self.projection = layers.Dense(units=projection_dim)\n        self.position_embedding = layers.Embedding(\n            input_dim=num_patches, output_dim=projection_dim\n        )\n    \n    # The call function encodes each patch with projection and position embedding\n    def call(self, patch):\n        positions = tf.range(start=0, limit=self.num_patches, delta=1)\n        encoded = self.projection(patch) + self.position_embedding(positions)\n        return encoded","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:42:45.477666Z","iopub.execute_input":"2023-02-27T11:42:45.478017Z","iopub.status.idle":"2023-02-27T11:42:45.488112Z","shell.execute_reply.started":"2023-02-27T11:42:45.477986Z","shell.execute_reply":"2023-02-27T11:42:45.487015Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Define a function called 'VITModel' that takes input shape and number of classes as inputs\ndef VITModel(input_shape=(image_size, image_size, 3), classes=3):\n    \n    # Create input layer with input_shape\n    inputs = layers.Input(shape=input_shape)\n\n    # Extract patches from the input using 'Patches' function and 'patch_size'\n    patches = Patches(patch_size)(inputs)\n\n    # Calculate the number of patches\n    num_patches = (input_shape[0] // patch_size) ** 2\n\n    # Encode the patches using 'PatchEncoder' function and 'projection_dim'\n    encoded_patches = PatchEncoder(num_patches, projection_dim)(patches)\n\n    # Implement 'transformer_layers' number of transformer layers\n    for _ in range(transformer_layers):\n        # Normalize the encoded patches\n        x1 = layers.LayerNormalization(epsilon=1e-6)(encoded_patches)\n        # Apply multi-head attention to the normalized patches\n        attention_output = layers.MultiHeadAttention(\n            num_heads=4, key_dim=projection_dim, dropout=0.1\n        )(x1, x1)\n        # Add the attention output to the encoded patches\n        x2 = layers.Add()([attention_output, encoded_patches])\n        # Normalize the output of the previous step\n        x3 = layers.LayerNormalization(epsilon=1e-6)(x2)\n        # Apply a multi-layer perceptron (MLP) to the normalized output\n        x3 = MLPBlock(x3, hidden_units=transformer_units, dropout_rate=0.1)\n        # Add the output of the MLP to the output of the previous step\n        encoded_patches = layers.Add()([x3, x2])\n\n    # Normalize the output of the final transformer layer\n    representation = layers.LayerNormalization(epsilon=1e-6)(encoded_patches)\n    # Flatten the representation\n    representation = layers.Flatten()(representation)\n    # Apply dropout to the representation\n    representation = layers.Dropout(0.5)(representation)\n    # Apply an MLP to the representation to obtain the features\n    features = MLPBlock(representation, hidden_units=mlp_head_units, dropout_rate=0.5)\n    # Apply a sigmoid activation function to the output of the MLP to obtain the logits\n    logits = layers.Dense(classes, activation=\"sigmoid\")(features)\n    # Create a Keras model with inputs and logits as outputs\n    model = keras.models.Model(inputs=inputs, outputs=logits)\n    \n    # Return the model\n    return model\n","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:42:45.489546Z","iopub.execute_input":"2023-02-27T11:42:45.48987Z","iopub.status.idle":"2023-02-27T11:42:45.505825Z","shell.execute_reply.started":"2023-02-27T11:42:45.489839Z","shell.execute_reply":"2023-02-27T11:42:45.504737Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model = VITModel((48,48,1),1)","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:42:45.507171Z","iopub.execute_input":"2023-02-27T11:42:45.508033Z","iopub.status.idle":"2023-02-27T11:42:46.893524Z","shell.execute_reply.started":"2023-02-27T11:42:45.507997Z","shell.execute_reply":"2023-02-27T11:42:46.89236Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model.summary()","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:42:46.895026Z","iopub.execute_input":"2023-02-27T11:42:46.895359Z","iopub.status.idle":"2023-02-27T11:42:47.097991Z","shell.execute_reply.started":"2023-02-27T11:42:46.89533Z","shell.execute_reply":"2023-02-27T11:42:47.09474Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"EPOCHS = 5\nmodel.compile(\"adam\", \"binary_crossentropy\",[\"accuracy\"])\nearly_stop_callback = tf.keras.callbacks.EarlyStopping(patience=2,\n                                                      restore_best_weights=True)\ncallbacks = [early_stop_callback]\nmodel.fit(train_generator,\n         epochs=EPOCHS,\n         validation_data = valid_generator,\n         callbacks = callbacks)","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:42:47.100366Z","iopub.execute_input":"2023-02-27T11:42:47.100708Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"AUTOTUNE = tf.data.experimental.AUTOTUNE\nfrom joblib import Parallel, delayed\ntest_images = glob.glob(\"/kaggle/input/rsna-breast-cancer-detection/test_images/*/*.dcm\")\n\nimage_dir = '/kaggle/tmp/test_images'\nos.makedirs(image_dir, exist_ok=True)\n\nimage_size = 48\ndcm_dir  = '/kaggle/input/rsna-breast-cancer-detection/test_images'\n\ndef dicom_to_png(dcm_file, image_size=image_size, image_dir=''):\n    patient_id = dcm_file.split('/')[-2]\n    image_id   = dcm_file.split('/')[-1][:-4]\n\n    dicom = pydicom.dcmread(dcm_file)\n    img = dicom.pixel_array\n    img = (img - img.min()) / (img.max() - img.min())\n    if dicom.PhotometricInterpretation == 'MONOCHROME1':\n        img = 1 - img\n\n    img = cv2.resize(img, (image_size, image_size), interpolation=cv2.INTER_LINEAR)\n    img = (img * 255).astype(np.uint8)\n    cv2.imwrite(image_dir +'/'+ f'{patient_id}_{image_id}.png', img)\n\n\ndcm_file = dcm_dir + '/' + test_df.patient_id.astype(str) + '/'  + test_df.image_id.astype(str) + '.dcm'\nParallel(n_jobs=2)(\n        delayed(dicom_to_png)(f, image_size=image_size, image_dir=image_dir)\n        for f in tqdm(dcm_file))\n\ndef load_image(image_path):\n    img = tf.io.read_file(image_path)\n    img = tf.image.decode_jpeg(img, channels = 1)\n    img = tf.image.resize(img, [image_size, image_size])\n    img = tf.cast(img, dtype = tf.float32)\n    img = img/255.0\n    return img\n\ntest_image_paths = []\ntest_dir = os.listdir('/kaggle/tmp/test_images')\nfor i in range(len(test_dir)):\n    img_path = '/kaggle/tmp/test_images' + '/' + test_dir[i]\n    test_image_paths.append(img_path)\ntest_img_path_ds = tf.data.Dataset.from_tensor_slices(test_image_paths)\ntest_ds = test_img_path_ds.map(load_image, num_parallel_calls = AUTOTUNE)\ntest_ds = test_ds.batch(BATCH_SIZE)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"preds = model.predict(test_ds)\npreds","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df2 = pd.read_csv(\"/kaggle/input/rsna-breast-cancer-detection/test.csv\")\ntest_df2['cancer'] = 0\n\nTHRESHOLD = 0.02\n\npreds = (preds > THRESHOLD).astype(int)\ntest_df2[\"cancer\"] = preds","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# F. Submission","metadata":{}},{"cell_type":"code","source":"test_df2['prediction_id'] = test_df2['patient_id'].astype(str) + \"_\" + test_df2['laterality']\n\nsub = test_df2[['prediction_id', 'cancer']].groupby(\"prediction_id\").mean().reset_index()\n\nsub.to_csv('/kaggle/working/submission.csv', index=False)\n\nsub.head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]}]}