{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.7.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":39272,"databundleVersionId":4629629,"sourceType":"competition"},{"sourceId":104036025,"sourceType":"kernelVersion"}],"dockerImageVersionId":30381,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Introduction\n\n**In this notebook, I would like to describe a baseline for medical image classification using PyTorch.**\nI am a physician and I am studying AI and data science. I am working on public health, such as health screening. Mammography plays a great role in public health and preventive medicine to detect breast cancer at an early stage.\nThis time we try machine learning to detect breast cancer using PyTorch.","metadata":{"papermill":{"duration":0.019148,"end_time":"2023-02-13T09:00:58.341378","exception":false,"start_time":"2023-02-13T09:00:58.32223","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"# Install and Import Necessary Libraries\n\nFirst, we install and import necessary libraries. Generally, medical image data is stored as the DICOM format. Reading medical images requires a special step. **In case of submission, the internet must be turned off and the special package must be downloaded and uploaded in advance.**","metadata":{"papermill":{"duration":0.013078,"end_time":"2023-02-13T09:00:58.368059","exception":false,"start_time":"2023-02-13T09:00:58.354981","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Install pydicom + pylibjpeg components\n!pip install -U pydicom pylibjpeg pylibjpeg-libjpeg pylibjpeg-openjpeg\n\n# Try installing an older python-gdcm ONLY if a wheel exists.\n# The --only-binary=:all: flag prevents source builds (which cause freezing).\n!pip install \"python-gdcm==3.0.22\" --only-binary=:all: || echo \"No wheel for python-gdcm; skipping (safe to continue).\"","metadata":{"papermill":{"duration":22.534607,"end_time":"2023-02-13T09:01:20.916251","exception":false,"start_time":"2023-02-13T09:00:58.381644","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T08:44:41.167306Z","iopub.execute_input":"2025-11-28T08:44:41.167582Z","iopub.status.idle":"2025-11-28T08:45:02.576691Z","shell.execute_reply.started":"2025-11-28T08:44:41.16752Z","shell.execute_reply":"2025-11-28T08:45:02.575726Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"In case of submission, we need, for example, to run the cell below after accessing [rsna-2022-whl](https://www.kaggle.com/code/vslaykovsky/rsna-2022-whl).","metadata":{}},{"cell_type":"code","source":"try:\n    import pylibjpeg\nexcept:\n    !pip install /kaggle/input/rsna-2022-whl/{pydicom-2.3.0-py3-none-any.whl,pylibjpeg-1.4.0-py3-none-any.whl,python_gdcm-3.0.15-cp37-cp37m-manylinux_2_17_x86_64.manylinux2014_x86_64.whl}","metadata":{"papermill":{"duration":0.030413,"end_time":"2023-02-13T09:01:20.962274","exception":false,"start_time":"2023-02-13T09:01:20.931861","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T08:45:02.57934Z","iopub.execute_input":"2025-11-28T08:45:02.580112Z","iopub.status.idle":"2025-11-28T08:45:02.590242Z","shell.execute_reply.started":"2025-11-28T08:45:02.580071Z","shell.execute_reply":"2025-11-28T08:45:02.58954Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# basic libraries\nimport numpy as np\nimport pandas as pd\nfrom matplotlib import pyplot as plt\nimport seaborn as sns\nimport os\nimport cv2\nfrom sklearn.metrics import confusion_matrix\nimport random\nfrom PIL import Image\n\n# specific for medical image data\nimport pydicom\npydicom.__version__\n\n# PyTorch libraries\nimport torch\nimport torch.nn as nn\nimport torch.nn.functional as F\nimport torch.optim as optim\nimport torchvision.transforms as transforms\nimport torchvision.models as models\nfrom torch.utils.data import Dataset","metadata":{"papermill":{"duration":3.033682,"end_time":"2023-02-13T09:01:24.010944","exception":false,"start_time":"2023-02-13T09:01:20.977262","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T08:45:02.591164Z","iopub.execute_input":"2025-11-28T08:45:02.59145Z","iopub.status.idle":"2025-11-28T08:45:05.071084Z","shell.execute_reply.started":"2025-11-28T08:45:02.591417Z","shell.execute_reply":"2025-11-28T08:45:05.070385Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Data Access and Analysis\n\nNext, we get the paths to access the image data and csv data. The csv data include various information about patients, such as biopsy, malignant cancer, and invasive cancer.","metadata":{"papermill":{"duration":0.014453,"end_time":"2023-02-13T09:01:24.041715","exception":false,"start_time":"2023-02-13T09:01:24.027262","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# the path to the image data\nRSNA_2022_path = '/kaggle/input/rsna-breast-cancer-detection/train_images'","metadata":{"papermill":{"duration":0.023374,"end_time":"2023-02-13T09:01:24.079866","exception":false,"start_time":"2023-02-13T09:01:24.056492","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T08:45:05.072149Z","iopub.execute_input":"2025-11-28T08:45:05.07263Z","iopub.status.idle":"2025-11-28T08:45:05.076603Z","shell.execute_reply.started":"2025-11-28T08:45:05.072603Z","shell.execute_reply":"2025-11-28T08:45:05.075718Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Read the csv data.\ndf_train = pd.read_csv('/kaggle/input/rsna-breast-cancer-detection/train.csv')\ndf_train.head()","metadata":{"papermill":{"duration":0.140708,"end_time":"2023-02-13T09:01:24.235195","exception":false,"start_time":"2023-02-13T09:01:24.094487","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T08:45:05.077998Z","iopub.execute_input":"2025-11-28T08:45:05.078915Z","iopub.status.idle":"2025-11-28T08:45:05.230763Z","shell.execute_reply.started":"2025-11-28T08:45:05.078878Z","shell.execute_reply":"2025-11-28T08:45:05.229825Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Attention\n\nGenerally, mammography is conducted with 2 views (from 2 angles), because **the position of breast cancer is various, although the upper outer quadrant of the breast is the most common site of breast cancer occurrence**. But the shape of the breast is almost the same regardless of different views. Moreover, the left and right breast generally have the same view. Therefore, we ignore the laterality and view for the machine learning purpose. The patient of the test data has no implant, so patients having implants can be excluded from the train data set. However, this patient has no risk of false positive for cancer because of the implant. Although it is unknown whether the other patients in the hidden test data set have an implant, generally few patients have implants. Thus, implant appears to have little influence on this machine learning.","metadata":{}},{"cell_type":"code","source":"# the number of total patients\nlen(df_train)","metadata":{"papermill":{"duration":0.024468,"end_time":"2023-02-13T09:01:24.275418","exception":false,"start_time":"2023-02-13T09:01:24.25095","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T08:45:05.231844Z","iopub.execute_input":"2025-11-28T08:45:05.232118Z","iopub.status.idle":"2025-11-28T08:45:05.238305Z","shell.execute_reply.started":"2025-11-28T08:45:05.232092Z","shell.execute_reply":"2025-11-28T08:45:05.237346Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# the number of patients having implant\nlen(df_train[df_train['implant'] == 1])","metadata":{"papermill":{"duration":0.025413,"end_time":"2023-02-13T09:01:24.406329","exception":false,"start_time":"2023-02-13T09:01:24.380916","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T08:45:05.241944Z","iopub.execute_input":"2025-11-28T08:45:05.242486Z","iopub.status.idle":"2025-11-28T08:45:05.254389Z","shell.execute_reply.started":"2025-11-28T08:45:05.24246Z","shell.execute_reply":"2025-11-28T08:45:05.253704Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# the number of patients without malignant cancer\nlen(df_train[df_train['cancer'] == 0])","metadata":{"papermill":{"duration":0.033745,"end_time":"2023-02-13T09:01:24.32428","exception":false,"start_time":"2023-02-13T09:01:24.290535","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T08:45:05.255442Z","iopub.execute_input":"2025-11-28T08:45:05.255687Z","iopub.status.idle":"2025-11-28T08:45:05.264878Z","shell.execute_reply.started":"2025-11-28T08:45:05.255659Z","shell.execute_reply":"2025-11-28T08:45:05.264137Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# the number of patient who took biopsy\nlen(df_train[df_train['biopsy'] == 1])","metadata":{"papermill":{"duration":0.026251,"end_time":"2023-02-13T09:01:24.365749","exception":false,"start_time":"2023-02-13T09:01:24.339498","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T08:45:05.26572Z","iopub.execute_input":"2025-11-28T08:45:05.266006Z","iopub.status.idle":"2025-11-28T08:45:05.273151Z","shell.execute_reply.started":"2025-11-28T08:45:05.265984Z","shell.execute_reply":"2025-11-28T08:45:05.27244Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# the number of patients having malignant cancer\nlen(df_train[df_train['cancer'] == 1])","metadata":{"papermill":{"duration":0.025413,"end_time":"2023-02-13T09:01:24.406329","exception":false,"start_time":"2023-02-13T09:01:24.380916","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T08:45:05.274103Z","iopub.execute_input":"2025-11-28T08:45:05.274366Z","iopub.status.idle":"2025-11-28T08:45:05.283217Z","shell.execute_reply.started":"2025-11-28T08:45:05.274345Z","shell.execute_reply":"2025-11-28T08:45:05.282401Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# the number of patients whose malignant cancer is invasive\nlen(df_train[df_train['invasive'] == 1])","metadata":{"papermill":{"duration":0.025425,"end_time":"2023-02-13T09:01:24.44691","exception":false,"start_time":"2023-02-13T09:01:24.421485","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T08:45:05.284373Z","iopub.execute_input":"2025-11-28T08:45:05.285023Z","iopub.status.idle":"2025-11-28T08:45:05.293087Z","shell.execute_reply.started":"2025-11-28T08:45:05.284998Z","shell.execute_reply":"2025-11-28T08:45:05.292203Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# the same as above\nlen(df_train[(df_train['cancer'] == 1) & (df_train['invasive'] == 1)])","metadata":{"papermill":{"duration":0.026607,"end_time":"2023-02-13T09:01:24.488965","exception":false,"start_time":"2023-02-13T09:01:24.462358","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T08:45:05.293959Z","iopub.execute_input":"2025-11-28T08:45:05.294242Z","iopub.status.idle":"2025-11-28T08:45:05.303957Z","shell.execute_reply.started":"2025-11-28T08:45:05.294219Z","shell.execute_reply":"2025-11-28T08:45:05.303031Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Most of the cases are normal or not-malignant cancer. Thus, physicians sometimes overlook cancer.\ndata = pd.DataFrame(np.concatenate([['Total'] * len(df_train) , ['Maglignant Cancer'] *  len(df_train[df_train['cancer'] == 1]), ['Invasive Cancer'] *  len(df_train[(df_train['cancer'] == 1) & (df_train['invasive'] == 1)])]), columns = [\"class\"])\n\nsns.countplot(x = 'class', data = data)","metadata":{"papermill":{"duration":0.252709,"end_time":"2023-02-13T09:01:24.757149","exception":false,"start_time":"2023-02-13T09:01:24.50444","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T08:45:05.30495Z","iopub.execute_input":"2025-11-28T08:45:05.30521Z","iopub.status.idle":"2025-11-28T08:45:05.540862Z","shell.execute_reply.started":"2025-11-28T08:45:05.305166Z","shell.execute_reply":"2025-11-28T08:45:05.539898Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Around 3000 patient took biopsy and malignant cancer was found from some of them.\ndata = pd.DataFrame(np.concatenate([['Biopsy'] * len(df_train[df_train['biopsy'] == 1]) , ['Malignant Cancer'] *  len(df_train[df_train['cancer'] == 1]), ['Invasive Cancer'] *  len(df_train[(df_train['cancer'] == 1) & (df_train['invasive'] == 1)])]), columns = [\"class\"])\n\nsns.countplot(x = 'class', data = data)","metadata":{"papermill":{"duration":0.196823,"end_time":"2023-02-13T09:01:24.970072","exception":false,"start_time":"2023-02-13T09:01:24.773249","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T08:45:05.541869Z","iopub.execute_input":"2025-11-28T08:45:05.542142Z","iopub.status.idle":"2025-11-28T08:45:05.712116Z","shell.execute_reply.started":"2025-11-28T08:45:05.542119Z","shell.execute_reply":"2025-11-28T08:45:05.711236Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# the number of not-malignant cancer cases from biopsy\nlen(df_train[(df_train['biopsy'] == 1) & (df_train['cancer'] == 0)])","metadata":{"papermill":{"duration":0.027406,"end_time":"2023-02-13T09:01:25.013858","exception":false,"start_time":"2023-02-13T09:01:24.986452","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T08:45:05.713278Z","iopub.execute_input":"2025-11-28T08:45:05.713554Z","iopub.status.idle":"2025-11-28T08:45:05.721366Z","shell.execute_reply.started":"2025-11-28T08:45:05.71353Z","shell.execute_reply":"2025-11-28T08:45:05.720526Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# the number of malignant cancer cases from biopsy\nlen(df_train[(df_train['biopsy'] == 1) & (df_train['cancer'] == 1)])","metadata":{"papermill":{"duration":0.026787,"end_time":"2023-02-13T09:01:25.056814","exception":false,"start_time":"2023-02-13T09:01:25.030027","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T08:45:05.722428Z","iopub.execute_input":"2025-11-28T08:45:05.722676Z","iopub.status.idle":"2025-11-28T08:45:05.734571Z","shell.execute_reply.started":"2025-11-28T08:45:05.722652Z","shell.execute_reply":"2025-11-28T08:45:05.733674Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# 60% of biopsy resulted in not-malignanct cancer.\ndata = pd.DataFrame(np.concatenate([['Biopsy but Not Malignant'] * len(df_train[(df_train['biopsy'] == 1) & (df_train['cancer'] == 0)]) , ['Malignant Cancer'] *  len(df_train[df_train['cancer'] == 1]), ['Invasive Cancer'] *  len(df_train[(df_train['cancer'] == 1) & (df_train['invasive'] == 1)])]), columns = [\"class\"])\n\nsns.countplot(x = 'class', data = data)","metadata":{"papermill":{"duration":0.199887,"end_time":"2023-02-13T09:01:25.27297","exception":false,"start_time":"2023-02-13T09:01:25.073083","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T08:45:05.735687Z","iopub.execute_input":"2025-11-28T08:45:05.736588Z","iopub.status.idle":"2025-11-28T08:45:05.914871Z","shell.execute_reply.started":"2025-11-28T08:45:05.73656Z","shell.execute_reply":"2025-11-28T08:45:05.913888Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Attention\n\nThis is a very difficult question whether **\"not-malignant cancer\" cases should be selected from (1)\"the whole patients without malignant cancer,\" (2)\"the whole patients without taking biopsy,\" or (3)\"patients with biopsy but without malignant cancer.\" The first choice is optimal for initial screening for general people**, most of whom are healthy.\n\nPeople generally take biopsies because not only of suspected mammography images of breast cancer, but also of their clinical findings or family history. Thus, images from **\"patients with biopsy but without malignant cancer\" may be totally healthy or may include benign cancer or other diseases, such as inflammation**. Therefore, **the second choice may be optimal, if the initial screening is only conducted for detection of malignant breast cancer**, and the patients do not suffer from any other diseases.\n\n**The third choice may be optimal** for special screening for suspected cases of malignant cancer. This is particularly useful **when it is suspected that a patient might suffer malignant cancer**. This AI would be used before biopsy is conducted.\n\nIn this competition, the purpose of screening is not sufficiently clear, because the test data do not include information as to biopsy. This time we took the third choice, but another choice might be better for the purpose of the competition.","metadata":{}},{"cell_type":"code","source":"# The not-malignant cancer cases were limited into biopsy cases. \nDF_train = df_train[df_train['biopsy'] == 1].reset_index(drop = True)\nDF_train.head()","metadata":{"papermill":{"duration":0.038103,"end_time":"2023-02-13T09:01:25.328245","exception":false,"start_time":"2023-02-13T09:01:25.290142","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T08:45:05.916064Z","iopub.execute_input":"2025-11-28T08:45:05.916385Z","iopub.status.idle":"2025-11-28T08:45:05.933209Z","shell.execute_reply.started":"2025-11-28T08:45:05.916357Z","shell.execute_reply":"2025-11-28T08:45:05.932274Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# The number of positive (malignant) and negative (not-malignat) cases should be the same\n# to create a balanced dataset.\nDF_train = DF_train.groupby(['cancer']).apply(lambda x: x.sample(1158, replace = True)\n                                                      ).reset_index(drop = True)\nprint('New Data Size:', DF_train.shape[0])","metadata":{"papermill":{"duration":0.034133,"end_time":"2023-02-13T09:01:25.380147","exception":false,"start_time":"2023-02-13T09:01:25.346014","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T08:45:05.934488Z","iopub.execute_input":"2025-11-28T08:45:05.93514Z","iopub.status.idle":"2025-11-28T08:45:05.951578Z","shell.execute_reply.started":"2025-11-28T08:45:05.935101Z","shell.execute_reply":"2025-11-28T08:45:05.950768Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Generally, invasive cancer is confirmed by biopsy, not by mammography.\n# Maybe it is also extremely difficult for AI to detect invasive cancer from mammography.\ndata = pd.DataFrame(np.concatenate([['Biopsy but Not Malignant'] * len(DF_train[(DF_train['biopsy'] == 1) & (DF_train['cancer'] == 0)]) , ['Malignant Cancer'] *  len(DF_train[DF_train['cancer'] == 1]), ['Invasive Cancer'] *  len(DF_train[(DF_train['cancer'] == 1) & (DF_train['invasive'] == 1)])]), columns = [\"class\"])\n\nsns.countplot(x = 'class', data = data)","metadata":{"papermill":{"duration":0.19691,"end_time":"2023-02-13T09:01:25.593941","exception":false,"start_time":"2023-02-13T09:01:25.397031","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T08:45:05.952553Z","iopub.execute_input":"2025-11-28T08:45:05.952797Z","iopub.status.idle":"2025-11-28T08:45:06.060163Z","shell.execute_reply.started":"2025-11-28T08:45:05.952774Z","shell.execute_reply":"2025-11-28T08:45:06.059319Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Medical Image Data in DICOM\n\nThen, we try to pick up a sample image stored in the DICOM format. In fact, the data include various information about the patient.","metadata":{"papermill":{"duration":0.016938,"end_time":"2023-02-13T09:01:25.628253","exception":false,"start_time":"2023-02-13T09:01:25.611315","status":"completed"},"tags":[]}},{"cell_type":"code","source":"dcmfnm = '/kaggle/input/rsna-breast-cancer-detection/train_images/10006/1459541791.dcm'\n\nds = pydicom.dcmread(dcmfnm, force = True)\nprint(\"Display Meta Information\\n\", ds)\n\n# Get information with keyword.\np_id = ds.PatientID\nprint(\"\\n>Patient ID=\", p_id, type(p_id))","metadata":{"papermill":{"duration":0.109334,"end_time":"2023-02-13T09:01:25.756228","exception":false,"start_time":"2023-02-13T09:01:25.646894","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T08:45:06.061313Z","iopub.execute_input":"2025-11-28T08:45:06.061684Z","iopub.status.idle":"2025-11-28T08:45:06.170387Z","shell.execute_reply.started":"2025-11-28T08:45:06.061632Z","shell.execute_reply":"2025-11-28T08:45:06.169466Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# This is the way to visualize a medical image.\nimg = ds.pixel_array\nplt.imshow(img, cmap = 'gray')\nplt.show()","metadata":{"papermill":{"duration":2.50666,"end_time":"2023-02-13T09:01:28.280395","exception":false,"start_time":"2023-02-13T09:01:25.773735","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T08:45:06.175322Z","iopub.execute_input":"2025-11-28T08:45:06.175588Z","iopub.status.idle":"2025-11-28T08:45:08.392444Z","shell.execute_reply.started":"2025-11-28T08:45:06.175564Z","shell.execute_reply":"2025-11-28T08:45:08.391538Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"img","metadata":{"papermill":{"duration":0.027203,"end_time":"2023-02-13T09:01:28.325651","exception":false,"start_time":"2023-02-13T09:01:28.298448","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T08:45:08.393555Z","iopub.execute_input":"2025-11-28T08:45:08.393816Z","iopub.status.idle":"2025-11-28T08:45:08.40043Z","shell.execute_reply.started":"2025-11-28T08:45:08.393794Z","shell.execute_reply":"2025-11-28T08:45:08.39955Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# The size of medical image is extremely huge.\nimg.shape","metadata":{"papermill":{"duration":0.026516,"end_time":"2023-02-13T09:01:28.369805","exception":false,"start_time":"2023-02-13T09:01:28.343289","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T08:45:08.401714Z","iopub.execute_input":"2025-11-28T08:45:08.401933Z","iopub.status.idle":"2025-11-28T08:45:08.412984Z","shell.execute_reply.started":"2025-11-28T08:45:08.401913Z","shell.execute_reply":"2025-11-28T08:45:08.412152Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Take another image.\nds = pydicom.dcmread(os.path.join(RSNA_2022_path + '/' + str(DF_train.loc[0, 'patient_id']) + '/' + str(DF_train.loc[0, 'image_id']) + '.dcm'), force = True)\nimg = ds.pixel_array.astype(np.float32)\nimg","metadata":{"papermill":{"duration":0.523099,"end_time":"2023-02-13T09:01:28.910574","exception":false,"start_time":"2023-02-13T09:01:28.387475","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T08:45:08.414744Z","iopub.execute_input":"2025-11-28T08:45:08.415007Z","iopub.status.idle":"2025-11-28T08:45:09.060778Z","shell.execute_reply.started":"2025-11-28T08:45:08.414985Z","shell.execute_reply":"2025-11-28T08:45:09.059901Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"img.shape","metadata":{"papermill":{"duration":0.031077,"end_time":"2023-02-13T09:01:28.960236","exception":false,"start_time":"2023-02-13T09:01:28.929159","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T08:45:09.062045Z","iopub.execute_input":"2025-11-28T08:45:09.062416Z","iopub.status.idle":"2025-11-28T08:45:09.068166Z","shell.execute_reply.started":"2025-11-28T08:45:09.062381Z","shell.execute_reply":"2025-11-28T08:45:09.067267Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plt.imshow(img, cmap = 'gray')\nplt.show()","metadata":{"papermill":{"duration":0.704853,"end_time":"2023-02-13T09:01:29.713707","exception":false,"start_time":"2023-02-13T09:01:29.008854","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T08:45:09.069465Z","iopub.execute_input":"2025-11-28T08:45:09.070147Z","iopub.status.idle":"2025-11-28T08:45:09.667845Z","shell.execute_reply.started":"2025-11-28T08:45:09.070108Z","shell.execute_reply":"2025-11-28T08:45:09.666958Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# The size of image must be resized to reduce computational costs for machine learning.\nimg = np.resize(img, (1024, 1024))","metadata":{"papermill":{"duration":0.046846,"end_time":"2023-02-13T09:01:29.779323","exception":false,"start_time":"2023-02-13T09:01:29.732477","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T08:45:09.669119Z","iopub.execute_input":"2025-11-28T08:45:09.669487Z","iopub.status.idle":"2025-11-28T08:45:09.693924Z","shell.execute_reply.started":"2025-11-28T08:45:09.669453Z","shell.execute_reply":"2025-11-28T08:45:09.693268Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# But nothing can be seen by human eyes after the resizing.\nplt.imshow(img, cmap = 'gray')\nplt.show()","metadata":{"papermill":{"duration":0.253293,"end_time":"2023-02-13T09:01:30.050805","exception":false,"start_time":"2023-02-13T09:01:29.797512","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T08:45:09.694906Z","iopub.execute_input":"2025-11-28T08:45:09.695163Z","iopub.status.idle":"2025-11-28T08:45:09.863707Z","shell.execute_reply.started":"2025-11-28T08:45:09.695139Z","shell.execute_reply":"2025-11-28T08:45:09.862815Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Creation of Dataset and Data Loader for Training and Validation\n\nHere, we create a dataset and data loader for machine learning with PyTorch. This is **a typical programming process for image classification with PyTorch**, **except for 'pixel_array,' which is required for DICOM**. It may be better to **create the validation data set without data augmentation separately**.\n\n### In addition, **downsizing the medical images is inevitable here to carry out the machine learning on the Kaggle platform, which limits computational loads for machine learning. Relevant information is lost for the image classification, and the score will be low. In fact, it is indispensable to carry out machine learning without downsizing using your special computer to get a high score in this competition.**","metadata":{"papermill":{"duration":0.01829,"end_time":"2023-02-13T09:01:30.087828","exception":false,"start_time":"2023-02-13T09:01:30.069538","status":"completed"},"tags":[]}},{"cell_type":"code","source":"transform = transforms.Compose(\n    [transforms.ToTensor(),\n     transforms.Normalize((0.5), (0.5))])","metadata":{"papermill":{"duration":0.025884,"end_time":"2023-02-13T09:01:30.132433","exception":false,"start_time":"2023-02-13T09:01:30.106549","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T08:45:09.864773Z","iopub.execute_input":"2025-11-28T08:45:09.865031Z","iopub.status.idle":"2025-11-28T08:45:09.869382Z","shell.execute_reply.started":"2025-11-28T08:45:09.865007Z","shell.execute_reply":"2025-11-28T08:45:09.868304Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"class RSNA_Dataset(Dataset):\n    def __init__(self, img_data, img_path, transform = None):\n        self.img_path = img_path\n        self.transform = transform\n        self.img_data = img_data\n        \n    def __len__(self):\n        return len(self.img_data)\n    \n    def __getitem__(self, index):\n        img_name = os.path.join(RSNA_2022_path + '/' + str(self.img_data.loc[index, 'patient_id']) + '/' + str(self.img_data.loc[index, 'image_id']) + '.dcm')\n        ds = pydicom.dcmread(img_name, force = True) # for DICOM data\n        image = ds.pixel_array.astype(np.float32)\n        # The data augmentation process begins.\n        # Convert to PIL image.\n        image = Image.fromarray(image)\n        # random crop\n        i, j, h, w = transforms.RandomCrop.get_params(image, output_size = (256, 256))\n        image = transforms.functional.crop(image, i, j, h, w)\n        # random horizontal flipping\n        if random.random() > 0.5:\n            image = transforms.functional.hflip(image)\n        # Convert back to NumPy array\n        image = np.array(image)\n        # The data augmentation process ends. It can be cut for validation data set.\n        # resize\n        image = np.resize(image, (300, 300))\n        #image = cv2.cvtColor(image, cv2.COLOR_GRAY2RGB)\n        label = torch.tensor(self.img_data.loc[index, 'cancer'])\n        if self.transform is not None:\n            image = self.transform(image)\n        return image, label","metadata":{"execution":{"iopub.status.busy":"2025-11-28T08:45:09.870763Z","iopub.execute_input":"2025-11-28T08:45:09.87109Z","iopub.status.idle":"2025-11-28T08:45:09.880033Z","shell.execute_reply.started":"2025-11-28T08:45:09.871064Z","shell.execute_reply":"2025-11-28T08:45:09.879235Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"dataset = RSNA_Dataset(DF_train, RSNA_2022_path, transform)","metadata":{"papermill":{"duration":0.025802,"end_time":"2023-02-13T09:01:30.222871","exception":false,"start_time":"2023-02-13T09:01:30.197069","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T08:45:09.881195Z","iopub.execute_input":"2025-11-28T08:45:09.881462Z","iopub.status.idle":"2025-11-28T08:45:09.893755Z","shell.execute_reply.started":"2025-11-28T08:45:09.881434Z","shell.execute_reply":"2025-11-28T08:45:09.89302Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"len(dataset)","metadata":{"papermill":{"duration":0.026849,"end_time":"2023-02-13T09:01:30.267929","exception":false,"start_time":"2023-02-13T09:01:30.24108","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T08:45:09.894657Z","iopub.execute_input":"2025-11-28T08:45:09.89487Z","iopub.status.idle":"2025-11-28T08:45:09.907418Z","shell.execute_reply.started":"2025-11-28T08:45:09.89485Z","shell.execute_reply":"2025-11-28T08:45:09.906467Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"x, t = dataset[0]","metadata":{"papermill":{"duration":0.488539,"end_time":"2023-02-13T09:01:30.774638","exception":false,"start_time":"2023-02-13T09:01:30.286099","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T08:45:09.90852Z","iopub.execute_input":"2025-11-28T08:45:09.908804Z","iopub.status.idle":"2025-11-28T08:45:10.436007Z","shell.execute_reply.started":"2025-11-28T08:45:09.90878Z","shell.execute_reply":"2025-11-28T08:45:10.435222Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"x","metadata":{"papermill":{"duration":0.035786,"end_time":"2023-02-13T09:01:30.833129","exception":false,"start_time":"2023-02-13T09:01:30.797343","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T08:45:10.437222Z","iopub.execute_input":"2025-11-28T08:45:10.43757Z","iopub.status.idle":"2025-11-28T08:45:10.446089Z","shell.execute_reply.started":"2025-11-28T08:45:10.437534Z","shell.execute_reply":"2025-11-28T08:45:10.445244Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"type(x), x.dtype, x.shape","metadata":{"papermill":{"duration":0.028839,"end_time":"2023-02-13T09:01:30.88166","exception":false,"start_time":"2023-02-13T09:01:30.852821","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T08:45:10.447285Z","iopub.execute_input":"2025-11-28T08:45:10.447526Z","iopub.status.idle":"2025-11-28T08:45:10.455046Z","shell.execute_reply.started":"2025-11-28T08:45:10.447504Z","shell.execute_reply":"2025-11-28T08:45:10.454215Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"t","metadata":{"papermill":{"duration":0.028762,"end_time":"2023-02-13T09:01:30.928714","exception":false,"start_time":"2023-02-13T09:01:30.899952","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T08:45:10.455745Z","iopub.execute_input":"2025-11-28T08:45:10.456063Z","iopub.status.idle":"2025-11-28T08:45:10.464302Z","shell.execute_reply.started":"2025-11-28T08:45:10.45604Z","shell.execute_reply":"2025-11-28T08:45:10.463437Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train, Val = torch.utils.data.random_split(dataset = dataset, lengths = [1800, 516], generator = torch.Generator().manual_seed(42))","metadata":{"papermill":{"duration":0.026665,"end_time":"2023-02-13T09:01:30.973799","exception":false,"start_time":"2023-02-13T09:01:30.947134","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T08:45:10.46529Z","iopub.execute_input":"2025-11-28T08:45:10.465631Z","iopub.status.idle":"2025-11-28T08:45:10.474046Z","shell.execute_reply.started":"2025-11-28T08:45:10.465604Z","shell.execute_reply":"2025-11-28T08:45:10.473363Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"len(train), len(Val)","metadata":{"papermill":{"duration":0.028316,"end_time":"2023-02-13T09:01:31.020946","exception":false,"start_time":"2023-02-13T09:01:30.99263","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T08:45:10.475102Z","iopub.execute_input":"2025-11-28T08:45:10.475353Z","iopub.status.idle":"2025-11-28T08:45:10.484629Z","shell.execute_reply.started":"2025-11-28T08:45:10.47533Z","shell.execute_reply":"2025-11-28T08:45:10.483907Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"class RSNA_Valset(Dataset):\n    def __init__(self, img_data, img_path, transform = None):\n        self.img_path = img_path\n        self.transform = transform\n        self.img_data = img_data\n        \n    def __len__(self):\n        return len(self.img_data)\n    \n    def __getitem__(self, index):\n        img_name = os.path.join(RSNA_2022_path + '/' + str(self.img_data.loc[index, 'patient_id']) + '/' + str(self.img_data.loc[index, 'image_id']) + '.dcm')\n        ds = pydicom.dcmread(img_name, force = True) # for DICOM data\n        image = ds.pixel_array.astype(np.float32)\n        # resize\n        image = np.resize(image, (300, 300))\n        #image = cv2.cvtColor(image, cv2.COLOR_GRAY2RGB)\n        label = torch.tensor(self.img_data.loc[index, 'cancer'])\n        if self.transform is not None:\n            image = self.transform(image)\n        return image, label","metadata":{"execution":{"iopub.status.busy":"2025-11-28T08:45:10.485699Z","iopub.execute_input":"2025-11-28T08:45:10.485973Z","iopub.status.idle":"2025-11-28T08:45:10.494111Z","shell.execute_reply.started":"2025-11-28T08:45:10.48595Z","shell.execute_reply":"2025-11-28T08:45:10.493325Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"dataset_val = RSNA_Valset(DF_train, RSNA_2022_path, transform)","metadata":{"execution":{"iopub.status.busy":"2025-11-28T08:45:10.495071Z","iopub.execute_input":"2025-11-28T08:45:10.49532Z","iopub.status.idle":"2025-11-28T08:45:10.507397Z","shell.execute_reply.started":"2025-11-28T08:45:10.495298Z","shell.execute_reply":"2025-11-28T08:45:10.506588Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"len(dataset_val)","metadata":{"execution":{"iopub.status.busy":"2025-11-28T08:45:10.508188Z","iopub.execute_input":"2025-11-28T08:45:10.508398Z","iopub.status.idle":"2025-11-28T08:45:10.52006Z","shell.execute_reply.started":"2025-11-28T08:45:10.508379Z","shell.execute_reply":"2025-11-28T08:45:10.519286Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"Train, val = torch.utils.data.random_split(dataset = dataset_val, lengths = [1800, 516], generator = torch.Generator().manual_seed(42))","metadata":{"execution":{"iopub.status.busy":"2025-11-28T08:45:10.521135Z","iopub.execute_input":"2025-11-28T08:45:10.521489Z","iopub.status.idle":"2025-11-28T08:45:10.528984Z","shell.execute_reply.started":"2025-11-28T08:45:10.521455Z","shell.execute_reply":"2025-11-28T08:45:10.528277Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"len(Train), len(val)","metadata":{"execution":{"iopub.status.busy":"2025-11-28T08:45:10.529848Z","iopub.execute_input":"2025-11-28T08:45:10.530087Z","iopub.status.idle":"2025-11-28T08:45:10.542462Z","shell.execute_reply.started":"2025-11-28T08:45:10.530063Z","shell.execute_reply":"2025-11-28T08:45:10.541545Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**They are what we exactly need!**","metadata":{}},{"cell_type":"code","source":"# They are what we exactly need!\nlen(train), len(val)","metadata":{"execution":{"iopub.status.busy":"2025-11-28T08:45:10.543522Z","iopub.execute_input":"2025-11-28T08:45:10.544438Z","iopub.status.idle":"2025-11-28T08:45:10.55179Z","shell.execute_reply.started":"2025-11-28T08:45:10.544402Z","shell.execute_reply":"2025-11-28T08:45:10.550657Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Check that the dataset is balanced with not-malignant and malignant cases.\ncan = 0\nfor i in range(len(val)):\n    x, t = val[i]\n    if t == 1:\n        can += 1\nprint(can / len(val))        ","metadata":{"papermill":{"duration":368.124207,"end_time":"2023-02-13T09:07:39.163921","exception":false,"start_time":"2023-02-13T09:01:31.039714","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T08:45:10.553069Z","iopub.execute_input":"2025-11-28T08:45:10.553386Z","iopub.status.idle":"2025-11-28T08:50:32.314001Z","shell.execute_reply.started":"2025-11-28T08:45:10.553363Z","shell.execute_reply":"2025-11-28T08:50:32.313071Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_loader = torch.utils.data.DataLoader(train, batch_size = 128, shuffle = False, drop_last = True)\nval_loader = torch.utils.data.DataLoader(val, batch_size = 128)","metadata":{"papermill":{"duration":0.026605,"end_time":"2023-02-13T09:07:39.210197","exception":false,"start_time":"2023-02-13T09:07:39.183592","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T08:50:32.315333Z","iopub.execute_input":"2025-11-28T08:50:32.315753Z","iopub.status.idle":"2025-11-28T08:50:32.320631Z","shell.execute_reply.started":"2025-11-28T08:50:32.315717Z","shell.execute_reply":"2025-11-28T08:50:32.319796Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Creation of Dataset and Visualization (No Resize)\n\nWe also have to create a dataset for visualization, because **the resizing has caused images that cannot be observed by human eyes**.","metadata":{"papermill":{"duration":0.018942,"end_time":"2023-02-13T09:07:39.24802","exception":false,"start_time":"2023-02-13T09:07:39.229078","status":"completed"},"tags":[]}},{"cell_type":"code","source":"class RSNA_Dataset_Visual(Dataset):\n    def __init__(self, img_data, img_path, transform = None):\n        self.img_path = img_path\n        self.transform = transform\n        self.img_data = img_data\n        \n    def __len__(self):\n        return len(self.img_data)\n    \n    def __getitem__(self, index):\n        img_name = os.path.join(RSNA_2022_path + '/' + str(self.img_data.loc[index, 'patient_id']) + '/' + str(self.img_data.loc[index, 'image_id']) + '.dcm')\n        ds = pydicom.dcmread(img_name, force = True)\n        image = ds.pixel_array.astype(np.float32)\n        #image = np.resize(image, (300, 300)) # no risize\n        #image = cv2.cvtColor(image, cv2.COLOR_GRAY2RGB)\n        label = torch.tensor(self.img_data.loc[index, 'cancer'])\n        if self.transform is not None:\n            image = self.transform(image)\n        return image, label","metadata":{"papermill":{"duration":0.029029,"end_time":"2023-02-13T09:07:39.295699","exception":false,"start_time":"2023-02-13T09:07:39.26667","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T08:50:32.321839Z","iopub.execute_input":"2025-11-28T08:50:32.322154Z","iopub.status.idle":"2025-11-28T08:50:32.33472Z","shell.execute_reply.started":"2025-11-28T08:50:32.322121Z","shell.execute_reply":"2025-11-28T08:50:32.333834Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"dataset_visual = RSNA_Dataset_Visual(DF_train, RSNA_2022_path, transform)","metadata":{"papermill":{"duration":0.02565,"end_time":"2023-02-13T09:07:39.339867","exception":false,"start_time":"2023-02-13T09:07:39.314217","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T08:50:32.335767Z","iopub.execute_input":"2025-11-28T08:50:32.336017Z","iopub.status.idle":"2025-11-28T08:50:32.349029Z","shell.execute_reply.started":"2025-11-28T08:50:32.335996Z","shell.execute_reply":"2025-11-28T08:50:32.348122Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Use the same random seed to create the same dataset as before except for the resizing. \ntrain_visual, val_visual = torch.utils.data.random_split(dataset = dataset_visual, lengths = [1800, 516], generator = torch.Generator().manual_seed(42))","metadata":{"papermill":{"duration":0.026751,"end_time":"2023-02-13T09:07:39.385564","exception":false,"start_time":"2023-02-13T09:07:39.358813","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T08:50:32.35011Z","iopub.execute_input":"2025-11-28T08:50:32.350447Z","iopub.status.idle":"2025-11-28T08:50:32.359717Z","shell.execute_reply.started":"2025-11-28T08:50:32.350419Z","shell.execute_reply":"2025-11-28T08:50:32.358822Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def img_display(img):\n    img = img / 2 + 0.5 # unnormalize\n    npimg = img.numpy()\n    #npimg = np.transpose(npimg, (1, 2, 0))\n    return npimg","metadata":{"papermill":{"duration":0.025764,"end_time":"2023-02-13T09:07:39.429904","exception":false,"start_time":"2023-02-13T09:07:39.40414","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T08:50:32.360786Z","iopub.execute_input":"2025-11-28T08:50:32.36168Z","iopub.status.idle":"2025-11-28T08:50:32.372978Z","shell.execute_reply.started":"2025-11-28T08:50:32.361645Z","shell.execute_reply":"2025-11-28T08:50:32.372239Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Get some training images.\narthopod_types = {0: 'Not Malignant', 1: 'Malignant Cancer'}\n# viewing data examples used for training\nfig, axis = plt.subplots(2, 5, figsize = (15, 10))\nfor i, ax in enumerate(axis.flat):\n    with torch.no_grad():\n        image, label = train_visual[i]\n        ax.imshow(img_display(image).squeeze(0)) # add image\n        ax.set(title = f\"{arthopod_types[label.item()]}\") # add label","metadata":{"papermill":{"duration":9.015798,"end_time":"2023-02-13T09:07:48.464326","exception":false,"start_time":"2023-02-13T09:07:39.448528","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T08:50:32.373902Z","iopub.execute_input":"2025-11-28T08:50:32.37412Z","iopub.status.idle":"2025-11-28T08:50:43.93215Z","shell.execute_reply.started":"2025-11-28T08:50:32.3741Z","shell.execute_reply":"2025-11-28T08:50:43.931162Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Creation of Model\n\nThis time we created **a simple AI model in PyTorch** as an example. This model might not be suitable for medical image data having huge data sizes. However, **it is extremely difficult to carry out machine learning using a new deep learning model with medical data on a publicly available platform like here, because the data sizes of medical data are extremely huge. A huge computational capacity would be required.** You can **replace** it with a more advanced model. **But, you may need a special computer that can satisfy the huge computational loads.**","metadata":{"papermill":{"duration":0.022509,"end_time":"2023-02-13T09:07:48.510019","exception":false,"start_time":"2023-02-13T09:07:48.48751","status":"completed"},"tags":[]}},{"cell_type":"code","source":"class Net(nn.Module):\n    def __init__(self):\n        super(Net, self).__init__()\n        # 1 input image channel, 16 output channels, 3x3 square convolution kernel\n        self.conv1 = nn.Conv2d(1, 16, kernel_size = 3,stride = 2,padding = 1)\n        self.conv2 = nn.Conv2d(16, 32, kernel_size = 3,stride = 2, padding = 1)\n        self.conv3 = nn.Conv2d(32, 64, kernel_size = 3,stride = 2, padding = 1)\n        self.conv4 = nn.Conv2d(64, 64, kernel_size = 3,stride = 2, padding = 1)\n        self.pool = nn.MaxPool2d(2, 2)\n        self.dropout = nn.Dropout2d(0.4)\n        self.batchnorm1 = nn.BatchNorm2d(16)\n        self.batchnorm2 = nn.BatchNorm2d(32)\n        self.batchnorm3 = nn.BatchNorm2d(64)\n        self.fc1 = nn.Linear(64 * 5 * 5, 512 )\n        self.fc2 = nn.Linear(512, 256)\n        self.fc3 = nn.Linear(256, 1)\n        \n    def forward(self, x):\n        x = self.batchnorm1(F.relu(self.conv1(x)))\n        x = self.batchnorm2(F.relu(self.conv2(x)))\n        x = self.dropout(self.batchnorm2(self.pool(x)))\n        x = self.batchnorm3(self.pool(F.relu(self.conv3(x))))\n        x = self.dropout(self.conv4(x))\n        x = x.view(-1, 64 * 5 * 5) # Flatten layer\n        x = self.dropout(self.fc1(x))\n        x = self.dropout(self.fc2(x))\n        x = F.logsigmoid(self.fc3(x))\n        return x","metadata":{"papermill":{"duration":0.056163,"end_time":"2023-02-13T09:07:48.58822","exception":false,"start_time":"2023-02-13T09:07:48.532057","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T08:50:43.933339Z","iopub.execute_input":"2025-11-28T08:50:43.93367Z","iopub.status.idle":"2025-11-28T08:50:43.943154Z","shell.execute_reply.started":"2025-11-28T08:50:43.933645Z","shell.execute_reply":"2025-11-28T08:50:43.942349Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#model = Net() # On CPU\ndevice = torch.device(\"cuda:0\" if torch.cuda.is_available() else \"cpu\")\nmodel = Net().to(device) # On GPU\nprint(model)","metadata":{"papermill":{"duration":0.107934,"end_time":"2023-02-13T09:07:48.726673","exception":false,"start_time":"2023-02-13T09:07:48.618739","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T08:50:43.944347Z","iopub.execute_input":"2025-11-28T08:50:43.945164Z","iopub.status.idle":"2025-11-28T08:50:46.509611Z","shell.execute_reply.started":"2025-11-28T08:50:43.945137Z","shell.execute_reply":"2025-11-28T08:50:46.508725Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Training and Validation\n\nThis is the training and validation process for binary image classification. The following code is **a typical sample for binary image classification with PyTorch**.","metadata":{"papermill":{"duration":0.021887,"end_time":"2023-02-13T09:07:48.771543","exception":false,"start_time":"2023-02-13T09:07:48.749656","status":"completed"},"tags":[]}},{"cell_type":"code","source":"criterion = nn.BCELoss()\noptimizer = optim.Adam(model.parameters(), lr = 0.001)\nscheduler = torch.optim.lr_scheduler.LinearLR(optimizer, start_factor = 1, end_factor = 0.1, total_iters = 8)","metadata":{"papermill":{"duration":0.030542,"end_time":"2023-02-13T09:07:48.824331","exception":false,"start_time":"2023-02-13T09:07:48.793789","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T08:50:46.510896Z","iopub.execute_input":"2025-11-28T08:50:46.511274Z","iopub.status.idle":"2025-11-28T08:50:46.516817Z","shell.execute_reply.started":"2025-11-28T08:50:46.511237Z","shell.execute_reply":"2025-11-28T08:50:46.515922Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"n_epochs = 6\nvalid_loss_min = np.Inf\nval_loss = []\nval_acc = []\ntrain_loss = []\ntrain_acc = []\ntotal_step = len(train_loader)\n# training\nfor epoch in range(1, n_epochs + 1):\n    running_loss = 0.0\n    scheduler.step(epoch)\n    correct = 0\n    total = 0\n    print(f'Epoch {epoch}\\n')\n    for batch_idx, (data_, target_) in enumerate(train_loader):\n        data_, target_ = data_.to(device), target_.to(device) # on GPU\n        # zero the parameter gradients\n        optimizer.zero_grad()\n        # forward + backward + optimize\n        outputs = model(data_)\n        pred = torch.sigmoid(outputs)\n        target = target_.unsqueeze(1).float()\n        loss = criterion(pred, target)\n        loss.backward()\n        optimizer.step()\n        # print statistics\n        running_loss += loss.item()\n        pred = pred > 0.40 # normally 0.5\n        accuracy = (target == pred).sum().item() / target.size(0)\n\n        if (batch_idx) % 3 == 0:\n            print ('Epoch [{}/{}], Step [{}/{}], Loss: {:.4f}' \n                   .format(epoch, n_epochs, batch_idx, total_step, loss.item()))\n    train_acc.append(100 * accuracy)\n    train_loss.append(running_loss / total_step)\n    print(f'\\ntrain loss: {np.mean(train_loss):.4f}, train acc: {(100 * accuracy):.4f}')\n    batch_loss = 0\n    total_t = 0\n    correct_t = 0\n# validation\n    with torch.no_grad():\n        model.eval()\n        for data_t, target_t in (val_loader):\n            data_t, target_t = data_t.to(device), target_t.to(device) # on GPU\n            outputs_t = model(data_t)\n            pred_t = torch.sigmoid(outputs_t)\n            target_t = target_t.unsqueeze(1).float()\n            loss_t = criterion(pred_t, target_t)\n            batch_loss += loss_t.item()\n            pred_t = pred_t > 0.40 # normally 0.5\n            accuracy_t = (target_t == pred_t).sum().item() / target_t.size(0)\n        val_acc.append(100 * accuracy_t)\n        val_loss.append(batch_loss / len(val_loader))\n        network_learned = batch_loss < valid_loss_min\n        print(f'validation loss: {np.mean(val_loss):.4f}, validation acc: {(100 * accuracy_t):.4f}\\n')\n        # Saving the best weight. \n        if network_learned:\n            valid_loss_min = batch_loss\n            torch.save(model.state_dict(), 'cancer_classification.pt')\n            print('Detected network improvement, saving current model')\n    scheduler.step()\n    model.train()","metadata":{"papermill":{"duration":8611.131065,"end_time":"2023-02-13T11:31:19.97735","exception":false,"start_time":"2023-02-13T09:07:48.846285","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T08:50:46.518104Z","iopub.execute_input":"2025-11-28T08:50:46.518452Z","iopub.status.idle":"2025-11-28T10:55:16.590972Z","shell.execute_reply.started":"2025-11-28T08:50:46.518411Z","shell.execute_reply":"2025-11-28T10:55:16.590011Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Save and Visualize the Results\n\nWe see **the accuracy and loss in training and validation** as well as **confusion matrix**, and save the model's parameters, which can also be used in another notebook.","metadata":{"papermill":{"duration":0.02426,"end_time":"2023-02-13T11:31:20.026099","exception":false,"start_time":"2023-02-13T11:31:20.001839","status":"completed"},"tags":[]}},{"cell_type":"code","source":"fig = plt.figure(figsize = (20, 10))\nplt.title(\"Train - Validation Accuracy\")\nplt.plot(train_acc, label = 'train')\nplt.plot(val_acc, label = 'validation')\nplt.xlabel('num_epochs', fontsize = 12)\nplt.ylabel('accuracy', fontsize = 12)\nplt.legend(loc = 'best')","metadata":{"papermill":{"duration":0.28155,"end_time":"2023-02-13T11:31:20.331883","exception":false,"start_time":"2023-02-13T11:31:20.050333","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T10:55:16.592377Z","iopub.execute_input":"2025-11-28T10:55:16.592965Z","iopub.status.idle":"2025-11-28T10:55:16.840445Z","shell.execute_reply.started":"2025-11-28T10:55:16.592925Z","shell.execute_reply":"2025-11-28T10:55:16.839632Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"fig = plt.figure(figsize = (20, 10))\nplt.title(\"Train - Validation Loss\")\nplt.plot(train_loss, label = 'train')\nplt.plot(val_loss, label = 'validation')\nplt.xlabel('num_epochs', fontsize = 12)\nplt.ylabel('loss', fontsize = 12)\nplt.legend(loc = 'best')","metadata":{"papermill":{"duration":0.26598,"end_time":"2023-02-13T11:31:20.623086","exception":false,"start_time":"2023-02-13T11:31:20.357106","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T10:55:16.841725Z","iopub.execute_input":"2025-11-28T10:55:16.842328Z","iopub.status.idle":"2025-11-28T10:55:17.079591Z","shell.execute_reply.started":"2025-11-28T10:55:16.842289Z","shell.execute_reply":"2025-11-28T10:55:17.078591Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# importing trained network with better loss of validation\nmodel.load_state_dict(torch.load('cancer_classification.pt'))","metadata":{"papermill":{"duration":0.042298,"end_time":"2023-02-13T11:31:20.69137","exception":false,"start_time":"2023-02-13T11:31:20.649072","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T10:55:17.080989Z","iopub.execute_input":"2025-11-28T10:55:17.081798Z","iopub.status.idle":"2025-11-28T10:55:17.097007Z","shell.execute_reply.started":"2025-11-28T10:55:17.081762Z","shell.execute_reply":"2025-11-28T10:55:17.096191Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# \"True\" and \"False\" mean that the prediction by the model is correct or wrong, respectively.\nmodel = model.cpu() # CPU\ndataiter = iter(val_loader)\nimages, labels = dataiter.next()\narthopod_types = {0: 'Not Malignant', 1: 'Malignant Cancer'}\n# viewing data examples used for training\nfig, axis = plt.subplots(2, 5, figsize = (15, 10))\nfor i, ax in enumerate(axis.flat):\n    with torch.no_grad():\n        model.eval()\n        image_visual, label_visual = val_visual[i]\n        image, label = images[i], labels[i]\n        ax.imshow(img_display(image_visual.squeeze(0))) # add image\n        image_tensor = image.unsqueeze_(0)\n        output_ = model(image_tensor)\n        output_ = torch.sigmoid(output_)\n        if output_.item() > 0.40: # normally 0.5\n            k = (label.item() == 1)\n        else:\n            k = (label.item() == 0)\n        ax.set_title(str(arthopod_types[label.item()]) + \":\" + str(k)) # add label","metadata":{"papermill":{"duration":85.178115,"end_time":"2023-02-13T11:32:45.895386","exception":false,"start_time":"2023-02-13T11:31:20.717271","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T10:55:17.098043Z","iopub.execute_input":"2025-11-28T10:55:17.098302Z","iopub.status.idle":"2025-11-28T10:56:41.162687Z","shell.execute_reply.started":"2025-11-28T10:55:17.098279Z","shell.execute_reply":"2025-11-28T10:56:41.161741Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"model = model.cpu() # CPU\nmodel.eval()\ny_pred = []\ny_true = []\n\n# Iterate over test data.\nwith torch.no_grad():\n    for inputs, labels in val_loader:\n            outputs = model(inputs) # Feed Network.\n            outputs = torch.sigmoid(outputs)\n\n            for i in range(len(outputs)):\n                output = outputs[i]            \n                if output.item() > 0.40: # normally 0.5\n                    y_pred.append(int(1)) # Save prediction.\n                else:\n                    y_pred.append(int(0)) # Save prediction.\n\n            labels = labels.data.cpu().numpy()\n            y_true.extend(labels) # Save truth.\n\n# constant for classes\nclasses = ('Not Malignant', 'Malignant Cancer')","metadata":{"papermill":{"duration":319.676138,"end_time":"2023-02-13T11:38:05.601677","exception":false,"start_time":"2023-02-13T11:32:45.925539","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T10:56:41.163785Z","iopub.execute_input":"2025-11-28T10:56:41.164051Z","iopub.status.idle":"2025-11-28T11:01:08.304924Z","shell.execute_reply.started":"2025-11-28T10:56:41.164026Z","shell.execute_reply":"2025-11-28T11:01:08.304041Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Build confusion matrix.\ncm = confusion_matrix(y_true, y_pred)","metadata":{"papermill":{"duration":0.040908,"end_time":"2023-02-13T11:38:05.67339","exception":false,"start_time":"2023-02-13T11:38:05.632482","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T11:01:08.306175Z","iopub.execute_input":"2025-11-28T11:01:08.306715Z","iopub.status.idle":"2025-11-28T11:01:08.313124Z","shell.execute_reply.started":"2025-11-28T11:01:08.306673Z","shell.execute_reply":"2025-11-28T11:01:08.312308Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"target_name = {'Not Malignant': 0, 'Malignant Cancer': 1}\nprint('Confusion Matrix')\nplt.figure(figsize = (10, 10))\n_ = sns.heatmap(cm.T, annot = True, fmt = 'd', cbar = True, square = True, xticklabels = target_name.keys(),\n             yticklabels = target_name.keys())\nplt.xlabel('Truth')\nplt.ylabel('Predicted')","metadata":{"papermill":{"duration":0.273102,"end_time":"2023-02-13T11:38:05.975466","exception":false,"start_time":"2023-02-13T11:38:05.702364","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T11:01:08.314303Z","iopub.execute_input":"2025-11-28T11:01:08.314587Z","iopub.status.idle":"2025-11-28T11:01:08.502438Z","shell.execute_reply.started":"2025-11-28T11:01:08.314564Z","shell.execute_reply":"2025-11-28T11:01:08.501642Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Prediction of Test Data\n\nFinally, we **read test data and predict the results** using the trained model and then **submit the answer in a CSV file**.","metadata":{"papermill":{"duration":0.028885,"end_time":"2023-02-13T11:38:06.034851","exception":false,"start_time":"2023-02-13T11:38:06.005966","status":"completed"},"tags":[]}},{"cell_type":"code","source":"RSNA_test_path = '/kaggle/input/rsna-breast-cancer-detection/test_images'","metadata":{"papermill":{"duration":0.036539,"end_time":"2023-02-13T11:38:06.100689","exception":false,"start_time":"2023-02-13T11:38:06.06415","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T11:01:08.50341Z","iopub.execute_input":"2025-11-28T11:01:08.503739Z","iopub.status.idle":"2025-11-28T11:01:08.508084Z","shell.execute_reply.started":"2025-11-28T11:01:08.503705Z","shell.execute_reply":"2025-11-28T11:01:08.507226Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df_test = pd.read_csv('/kaggle/input/rsna-breast-cancer-detection/test.csv')\ndf_test.head()","metadata":{"papermill":{"duration":0.061541,"end_time":"2023-02-13T11:38:06.191505","exception":false,"start_time":"2023-02-13T11:38:06.129964","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T11:01:08.509305Z","iopub.execute_input":"2025-11-28T11:01:08.509635Z","iopub.status.idle":"2025-11-28T11:01:08.530468Z","shell.execute_reply.started":"2025-11-28T11:01:08.5096Z","shell.execute_reply":"2025-11-28T11:01:08.529603Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"class RSNA_Testset(Dataset):\n    def __init__(self, img_data, img_path, transform = None):\n        self.img_path = img_path\n        self.transform = transform\n        self.img_data = img_data\n        \n    def __len__(self):\n        return len(self.img_data)\n    \n    def __getitem__(self, index):\n        img_name = os.path.join(RSNA_test_path + '/' + str(self.img_data.loc[index, 'patient_id']) + '/' + str(self.img_data.loc[index, 'image_id']) + '.dcm')\n        ds = pydicom.dcmread(img_name, force = True)\n        image = ds.pixel_array.astype(np.float32)\n        image = np.resize(image, (300, 300))\n        #image = cv2.cvtColor(image, cv2.COLOR_GRAY2RGB)\n        #label = torch.tensor(self.img_data.loc[index, 'cancer'])\n        if self.transform is not None:\n            image = self.transform(image)\n        return image#, label","metadata":{"papermill":{"duration":0.040559,"end_time":"2023-02-13T11:38:06.261478","exception":false,"start_time":"2023-02-13T11:38:06.220919","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T11:01:08.531558Z","iopub.execute_input":"2025-11-28T11:01:08.532415Z","iopub.status.idle":"2025-11-28T11:01:08.538764Z","shell.execute_reply.started":"2025-11-28T11:01:08.53239Z","shell.execute_reply":"2025-11-28T11:01:08.537955Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"testset = RSNA_Testset(df_test, RSNA_test_path, transform)","metadata":{"papermill":{"duration":0.036619,"end_time":"2023-02-13T11:38:06.327251","exception":false,"start_time":"2023-02-13T11:38:06.290632","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T11:01:08.545876Z","iopub.execute_input":"2025-11-28T11:01:08.546096Z","iopub.status.idle":"2025-11-28T11:01:08.551121Z","shell.execute_reply.started":"2025-11-28T11:01:08.546074Z","shell.execute_reply":"2025-11-28T11:01:08.550323Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Without data loader","metadata":{}},{"cell_type":"code","source":"model = model.to(device) # On GPU\nmodel.eval()\ncancer = []\nwith torch.no_grad():\n    for i in range(len(testset)):\n        out = testset[i].unsqueeze_(0)\n        out = out.to(device)\n        out = model(out)\n        out = torch.sigmoid(out)\n        cancer.append(out.item())\ncancer","metadata":{"papermill":{"duration":2.345762,"end_time":"2023-02-13T11:38:08.702254","exception":false,"start_time":"2023-02-13T11:38:06.356492","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T11:01:08.552334Z","iopub.execute_input":"2025-11-28T11:01:08.552656Z","iopub.status.idle":"2025-11-28T11:01:09.979485Z","shell.execute_reply.started":"2025-11-28T11:01:08.552625Z","shell.execute_reply":"2025-11-28T11:01:09.978533Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## With data loader","metadata":{}},{"cell_type":"code","source":"# Load the test set.\ntestset = RSNA_Testset(df_test, RSNA_test_path, transform)\n\n# Create a DataLoader for the test set.\ntestloader = torch.utils.data.DataLoader(testset, batch_size = 32, shuffle = False)\n\n# Move the model to the GPU (if available).\ndevice = torch.device(\"cuda:0\" if torch.cuda.is_available() else \"cpu\")\nmodel = model.to(device)\n\n# Put the model in evaluation mode.\nmodel.eval()\n\n# Make predictions on the test set.\ncancer = []\nwith torch.no_grad():\n    for images in testloader:\n        # Move the batch of images to the GPU (if available).\n        images = images.to(device)\n\n        # Pass the batch of images through the model to get the predicted cancer probabilities.\n        out = model(images)\n\n        # Apply a sigmoid function to get a probability between 0 and 1.\n        out = torch.sigmoid(out)\n\n        # Add the predicted probabilities to the list\n        cancer.extend(out.cpu().numpy())\n        \n        cancer_probs = np.concatenate(cancer)\n        \n        cancer = [x[0] for x in cancer]\n\n# Print the predicted probabilities.\nprint(cancer)","metadata":{"execution":{"iopub.status.busy":"2025-11-28T11:01:09.980585Z","iopub.execute_input":"2025-11-28T11:01:09.980855Z","iopub.status.idle":"2025-11-28T11:01:11.270996Z","shell.execute_reply.started":"2025-11-28T11:01:09.980831Z","shell.execute_reply":"2025-11-28T11:01:11.270035Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df_test['cancer'] = cancer\ndf_test","metadata":{"papermill":{"duration":0.045937,"end_time":"2023-02-13T11:38:08.778934","exception":false,"start_time":"2023-02-13T11:38:08.732997","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T11:01:11.272376Z","iopub.execute_input":"2025-11-28T11:01:11.272831Z","iopub.status.idle":"2025-11-28T11:01:11.284656Z","shell.execute_reply.started":"2025-11-28T11:01:11.272785Z","shell.execute_reply":"2025-11-28T11:01:11.283667Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"submission = df_test.loc[:, 'prediction_id':'cancer']\nsubmission","metadata":{"papermill":{"duration":0.051298,"end_time":"2023-02-13T11:38:08.860074","exception":false,"start_time":"2023-02-13T11:38:08.808776","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T11:01:11.28584Z","iopub.execute_input":"2025-11-28T11:01:11.286086Z","iopub.status.idle":"2025-11-28T11:01:11.302451Z","shell.execute_reply.started":"2025-11-28T11:01:11.286062Z","shell.execute_reply":"2025-11-28T11:01:11.30164Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Choose mean, max, min, or something else.\nsubmission = submission.groupby('prediction_id').mean().reset_index()\nsubmission","metadata":{"papermill":{"duration":0.04781,"end_time":"2023-02-13T11:38:08.937687","exception":false,"start_time":"2023-02-13T11:38:08.889877","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T11:01:11.303607Z","iopub.execute_input":"2025-11-28T11:01:11.303934Z","iopub.status.idle":"2025-11-28T11:01:11.320969Z","shell.execute_reply.started":"2025-11-28T11:01:11.303902Z","shell.execute_reply":"2025-11-28T11:01:11.320155Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"submission = pd.DataFrame(data = {'prediction_id': df_test['prediction_id'], 'cancer': cancer})\nsubmission","metadata":{"papermill":{"duration":0.043929,"end_time":"2023-02-13T11:38:09.012057","exception":false,"start_time":"2023-02-13T11:38:08.968128","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T11:01:11.322359Z","iopub.execute_input":"2025-11-28T11:01:11.322695Z","iopub.status.idle":"2025-11-28T11:01:11.334364Z","shell.execute_reply.started":"2025-11-28T11:01:11.322663Z","shell.execute_reply":"2025-11-28T11:01:11.333571Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Attention\nI prefer max() to mean() or min(), because it is possible that **cancer can be observed from one view (high score) and cannot be observed from the other view (low score)**. If cancer is positive from one view, this case should be positive regardless of the other view.","metadata":{}},{"cell_type":"code","source":"submission = submission.groupby('prediction_id').max().reset_index()\nsubmission","metadata":{"papermill":{"duration":0.048208,"end_time":"2023-02-13T11:38:09.091172","exception":false,"start_time":"2023-02-13T11:38:09.042964","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T11:01:11.335253Z","iopub.execute_input":"2025-11-28T11:01:11.335531Z","iopub.status.idle":"2025-11-28T11:01:11.351019Z","shell.execute_reply.started":"2025-11-28T11:01:11.335496Z","shell.execute_reply":"2025-11-28T11:01:11.350028Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"submission.to_csv('submission.csv', index = False)","metadata":{"papermill":{"duration":0.043742,"end_time":"2023-02-13T11:38:09.164927","exception":false,"start_time":"2023-02-13T11:38:09.121185","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-11-28T11:01:11.351916Z","iopub.execute_input":"2025-11-28T11:01:11.352156Z","iopub.status.idle":"2025-11-28T11:01:11.364852Z","shell.execute_reply.started":"2025-11-28T11:01:11.352133Z","shell.execute_reply":"2025-11-28T11:01:11.363922Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Conclusion\n\nThis is the entire process for the competition. However, **the used simple model appeared to be insufficient for medical image classification tasks**.\n\nIn addition, **downsizing the medical images was inevitable here to carry out the machine learning on the Kaggle platform**, which limits computational loads for machine learning. Therefore, **you need a special computer that can satisfy the huge computational loads to get a high score in this competition**. ","metadata":{}},{"cell_type":"markdown","source":"## Credits\n\nPlease go and visit these below mentioned notebooks and support their works and advice. Without the help of the below mentioned notebooks and advice, it would have been much difficult for me to approach the solution.\n1. [how to read .dcm (DICOM) data](https://www.kaggle.com/code/micheldc55/how-to-read-dcm-dicom-data)\n2. [PyTorch-Tutorial (The Classification)](https://www.kaggle.com/code/basu369victor/pytorch-tutorial-the-classification)\n3. [rsna-2022-whl](https://www.kaggle.com/code/vslaykovsky/rsna-2022-whl)\n4. [Thomas Konstantin](https://www.kaggle.com/code/gokifujiya/medical-image-classification-baseline-with-pytorch/comments?scriptVersionId=133629908)","metadata":{"papermill":{"duration":0.030211,"end_time":"2023-02-13T11:38:09.225726","exception":false,"start_time":"2023-02-13T11:38:09.195515","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"I am a medical doctor working on **artificial intelligence (AI) for medicine**. At present AI is also widely used in the medical field. Particularly, AI performs in the healthcare sector following tasks: **image classification, object detection, semantic segmentation, GANs, text classification, etc**. **If you are interested in AI for medicine, please see my other notebooks.**","metadata":{}}]}