{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# RSNA EDA + PCA + Logistic Regression\n\n# Introduction\n\nHello, everyone! This is my first time to do EDA. Hope you could have fun as I did from this notebook.\n\nThis competition aims to detect breast cancer with mammograms and other information in train.csv file including patients age, image view, etc. I will explore some intrinsic information of the table and relationships between patients and machine info with cancer. \n\nAt last, I will do PCA and a logitic regression to predict cancer or not. ","metadata":{}},{"cell_type":"markdown","source":"# Install and imports","metadata":{}},{"cell_type":"code","source":"# !pip install -qU python-gdcm pydicom pylibjpeg\n!pip install /kaggle/input/rsna-2022-whl/{pydicom-2.3.0-py3-none-any.whl,pylibjpeg-1.4.0-py3-none-any.whl,python_gdcm-3.0.15-cp37-cp37m-manylinux_2_17_x86_64.manylinux2014_x86_64.whl}","metadata":{"execution":{"iopub.status.busy":"2023-01-30T01:14:35.97376Z","iopub.execute_input":"2023-01-30T01:14:35.974244Z","iopub.status.idle":"2023-01-30T01:15:11.194372Z","shell.execute_reply.started":"2023-01-30T01:14:35.974151Z","shell.execute_reply":"2023-01-30T01:15:11.192537Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pydicom\nimport gdcm\nimport matplotlib.pyplot as plt\nimport cv2\nfrom PIL import Image\nimport pandas as pd\nimport seaborn as sns\nfrom pandas.plotting import scatter_matrix\nsns.set_style('darkgrid')\nsns.set_color_codes('bright')\nimport os\nfrom sklearn.decomposition import PCA\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.cluster import KMeans\nimport cv2\nfrom tqdm import tqdm\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.linear_model import LogisticRegression\nimport numpy as np","metadata":{"execution":{"iopub.status.busy":"2023-01-30T01:15:11.199246Z","iopub.execute_input":"2023-01-30T01:15:11.199616Z","iopub.status.idle":"2023-01-30T01:15:12.417625Z","shell.execute_reply.started":"2023-01-30T01:15:11.199581Z","shell.execute_reply":"2023-01-30T01:15:12.41653Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# EDA","metadata":{}},{"cell_type":"code","source":"base_img_dir = \"/kaggle/input/rsna-breast-cancer-detection/train_images\"\ntrain_df = pd.read_csv(\"/kaggle/input/rsna-breast-cancer-detection/train.csv\")\ntrain_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-01-30T01:15:12.418903Z","iopub.execute_input":"2023-01-30T01:15:12.419295Z","iopub.status.idle":"2023-01-30T01:15:12.592481Z","shell.execute_reply.started":"2023-01-30T01:15:12.419252Z","shell.execute_reply":"2023-01-30T01:15:12.591403Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"total_images = len(train_df)\nprint(f\"There are {total_images} images in all\")\npatient_num = len(train_df[\"patient_id\"].unique())\nprint(f\"There are {patient_num} patients in all.\")","metadata":{"execution":{"iopub.status.busy":"2023-01-30T01:15:12.59466Z","iopub.execute_input":"2023-01-30T01:15:12.595045Z","iopub.status.idle":"2023-01-30T01:15:12.609049Z","shell.execute_reply.started":"2023-01-30T01:15:12.595013Z","shell.execute_reply":"2023-01-30T01:15:12.607789Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We saw there are much more mammograms than patients which means each patient took more than 1 images. Let's see how many mammograms each patient took.","metadata":{}},{"cell_type":"code","source":"patient_image = train_df.groupby(\"patient_id\").image_id.count()\nsns.histplot(patient_image.values, bins=20)\nplt.title(\"Number of mammograms for each patients\")\nplt.xlabel(\"Number\")","metadata":{"execution":{"iopub.status.busy":"2023-01-30T01:15:12.610503Z","iopub.execute_input":"2023-01-30T01:15:12.611597Z","iopub.status.idle":"2023-01-30T01:15:13.026215Z","shell.execute_reply.started":"2023-01-30T01:15:12.611562Z","shell.execute_reply":"2023-01-30T01:15:13.025095Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we can see, most patients took 4 ~ 6 mammograms. Few people took more than 6 images.","metadata":{}},{"cell_type":"markdown","source":"Let's see how many patients got cancer. In these 11913 patients, around 96% have no cancer while 4% (486 patients) having cancer. So it is quite unbalanced.","metadata":{}},{"cell_type":"code","source":"def has_cancer(c):\n    if c > 0:\n        return True\n    else:\n        return False\n\ncancer_per_patient = train_df.groupby(\"patient_id\").cancer.sum().apply(lambda c: has_cancer(c)).values\nax = sns.countplot(cancer_per_patient)\nax.bar_label(ax.containers[0])\nplt.xlabel(\"Cancer\")\nplt.ylabel(\"Count\")\nplt.title(\"Number of patients with cancer\")","metadata":{"execution":{"iopub.status.busy":"2023-01-30T01:15:13.027605Z","iopub.execute_input":"2023-01-30T01:15:13.02797Z","iopub.status.idle":"2023-01-30T01:15:13.244846Z","shell.execute_reply.started":"2023-01-30T01:15:13.027936Z","shell.execute_reply":"2023-01-30T01:15:13.243667Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From following age distribution, we could find ages of most patients are between 40 to 70. Particularlly, the number of patients surged at the age of 40. More people tend to do diagnosis in hosptial when they reach 40 years old.","metadata":{}},{"cell_type":"code","source":"def get_age(a):\n    return a[0]\n    \npatient_age = train_df[train_df['age'].isnull()==False].groupby(\"patient_id\").age.unique().apply(lambda a: get_age(a))\nsns.histplot(patient_age.values, bins=60)\nplt.xlabel(\"Age\")\nplt.title(\"Patient ages\")","metadata":{"execution":{"iopub.status.busy":"2023-01-30T01:15:13.246333Z","iopub.execute_input":"2023-01-30T01:15:13.246693Z","iopub.status.idle":"2023-01-30T01:15:14.281984Z","shell.execute_reply.started":"2023-01-30T01:15:13.246659Z","shell.execute_reply":"2023-01-30T01:15:14.280762Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's take a look at correlations to gain some insights.","metadata":{}},{"cell_type":"code","source":"corr_matrix = train_df.corr()\ncorr_matrix['cancer'].sort_values(ascending=False)","metadata":{"execution":{"iopub.status.busy":"2023-01-30T01:15:14.283467Z","iopub.execute_input":"2023-01-30T01:15:14.283853Z","iopub.status.idle":"2023-01-30T01:15:14.321362Z","shell.execute_reply.started":"2023-01-30T01:15:14.283819Z","shell.execute_reply":"2023-01-30T01:15:14.320097Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It seems \"invasive\" and \"biopsy\" have strong positive correlation with \"cancer\". Let's first recap what they mean. \n* \"biopsy\" - Whether or not a follow-up biopsy was performed on the breast. Only provided for train.\n* \"invasive\" - If the breast is positive for cancer, whether or not the cancer proved to be invasive. Only provided for train.\n\nIt is understandable that the doctor may require to perform a follow-up biopsy if he found some unusual things in the mammograms. But there are no information about \"biopsy\" as well as \"invasive\" in testing dataset. \n\nWe also notice that there is slight positive correlation (not strong) between \"age\" and \"cancer\". We could explore a little bit.","metadata":{}},{"cell_type":"code","source":"attribs = [\"age\", \"cancer\"]\nscatter_matrix(train_df[attribs])","metadata":{"execution":{"iopub.status.busy":"2023-01-30T01:15:14.32289Z","iopub.execute_input":"2023-01-30T01:15:14.323323Z","iopub.status.idle":"2023-01-30T01:15:15.700423Z","shell.execute_reply.started":"2023-01-30T01:15:14.323287Z","shell.execute_reply":"2023-01-30T01:15:15.699394Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's zoom in on the correlation scatterplot between age and cancer.","metadata":{}},{"cell_type":"code","source":"train_df.plot(kind='scatter', x=\"age\", y=\"cancer\")","metadata":{"execution":{"iopub.status.busy":"2023-01-29T05:12:40.859613Z","iopub.execute_input":"2023-01-29T05:12:40.860069Z","iopub.status.idle":"2023-01-29T05:12:41.233388Z","shell.execute_reply.started":"2023-01-29T05:12:40.860035Z","shell.execute_reply":"2023-01-29T05:12:41.231957Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I draw following graph to see whether there is relationship between age and cancer.","metadata":{}},{"cell_type":"code","source":"age_cancer = train_df.groupby([\"patient_id\", \"cancer\"]).age.unique().apply(lambda a: get_age(a))\nage_cancer_df = pd.DataFrame(age_cancer).reset_index()\nfig, ax = plt.subplots(figsize=(10, 5))\nsns.histplot(age_cancer_df[age_cancer_df[\"cancer\"]==0][\"age\"], color=\"b\", bins=60, kde=True, ax=ax)\nax.set_ylabel(\"No cancer count\")\nax2 = ax.twinx()\nax2.set_ylabel(\"Cancer count\")\nsns.histplot(age_cancer_df[age_cancer_df[\"cancer\"]==1][\"age\"], color=\"r\", bins=60, kde=True, ax=ax2)","metadata":{"execution":{"iopub.status.busy":"2023-01-30T02:12:09.836159Z","iopub.execute_input":"2023-01-30T02:12:09.836586Z","iopub.status.idle":"2023-01-30T02:12:11.349161Z","shell.execute_reply.started":"2023-01-30T02:12:09.83655Z","shell.execute_reply":"2023-01-30T02:12:11.347944Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In the above figure, red represent \"Patients with cancer\" while blue is \"Patients with no cancer\". Looking at the distribution of age, we could see the mean patients' age with cancer is greater than patients without cancer. In another word, **it is more likely that people get cancer when they get older.**","metadata":{}},{"cell_type":"markdown","source":"From the following violin figure, it is also obvious that patients with cancer is more concentrated at elder ones (age between 60 - 70).","metadata":{}},{"cell_type":"code","source":"colors = [\"light blue\", \"light red\"]\nsns.violinplot(x=\"cancer\", y=\"age\", data=train_df, palette=sns.xkcd_palette(colors))","metadata":{"execution":{"iopub.status.busy":"2023-01-30T01:15:25.761231Z","iopub.execute_input":"2023-01-30T01:15:25.761822Z","iopub.status.idle":"2023-01-30T01:15:26.172117Z","shell.execute_reply.started":"2023-01-30T01:15:25.76177Z","shell.execute_reply":"2023-01-30T01:15:26.170836Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"How about implant affects cancer? Are women having implants more easily to get cancer?\n\nLet's first see how many patients having implants.","metadata":{}},{"cell_type":"code","source":"patient_implant = train_df.groupby(\"patient_id\").implant.max()\nax = sns.histplot(patient_implant.values, bins=2, discrete=True)\nax.bar_label(ax.containers[0])\nplt.title(\"Patients with and without implant\")\nplt.xlabel(\"Implant\")\nax.set_xticks(range(0, 2))\nax.set_xticklabels([\"0\", \"1\"])\npatient_with_implant = len(patient_implant[patient_implant.values>0])\nprint(f\"There are {patient_with_implant} ({round(100*patient_with_implant/patient_num, 1)}%) patients having implant.\")","metadata":{"execution":{"iopub.status.busy":"2023-01-30T01:15:29.191636Z","iopub.execute_input":"2023-01-30T01:15:29.192101Z","iopub.status.idle":"2023-01-30T01:15:29.437424Z","shell.execute_reply.started":"2023-01-30T01:15:29.192062Z","shell.execute_reply":"2023-01-30T01:15:29.436324Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"How many percentage of cancer if the patient having implant.","metadata":{}},{"cell_type":"code","source":"patient_df = train_df.groupby([\"patient_id\", \"cancer\"]).implant.max().reset_index()\nax = sns.histplot(x=\"cancer\", hue=\"implant\", data=patient_df, stat=\"percent\", multiple=\"stack\", bins=2, discrete=True)\nax.bar_label(ax.containers[0])\nax.set_xticks(range(2))\nax.set_xticklabels([\"0\", \"1\"])\nplt.title(\"Patients implant percentage\")","metadata":{"execution":{"iopub.status.busy":"2023-01-30T01:15:32.703437Z","iopub.execute_input":"2023-01-30T01:15:32.704002Z","iopub.status.idle":"2023-01-30T01:15:32.997137Z","shell.execute_reply.started":"2023-01-30T01:15:32.703955Z","shell.execute_reply":"2023-01-30T01:15:32.996281Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Because there are only few patients (171 patients, 1.4% out of 11913 patients) have implant, it is hard to say whether it would cause cancer. Remember there are 486 cancer cases in all. Only 3 patients who having implant got cancer. Other 483 patients got cancer but have no implant.","metadata":{}},{"cell_type":"code","source":"patient_df[patient_df[\"cancer\"]==1].groupby(\"implant\").count()","metadata":{"execution":{"iopub.status.busy":"2023-01-30T01:15:50.792009Z","iopub.execute_input":"2023-01-30T01:15:50.792458Z","iopub.status.idle":"2023-01-30T01:15:50.807369Z","shell.execute_reply.started":"2023-01-30T01:15:50.792423Z","shell.execute_reply":"2023-01-30T01:15:50.806295Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We know patients doing diagnosis at different machines. Are there some bias for one particular machine detecting more cancers like the movie \"COMA\"(1978)? Let's first take a look at numbers of patients at different machines. ","metadata":{}},{"cell_type":"code","source":"machine_df = train_df.groupby(\"machine_id\").count()[\"patient_id\"].reset_index()\nplt.figure(figsize=(10,5))\nsns.barplot(x=\"machine_id\", y=\"patient_id\", data=machine_df)\nplt.ylabel(\"Count\")","metadata":{"execution":{"iopub.status.busy":"2023-01-30T02:10:23.85861Z","iopub.execute_input":"2023-01-30T02:10:23.859425Z","iopub.status.idle":"2023-01-30T02:10:24.10956Z","shell.execute_reply.started":"2023-01-30T02:10:23.859376Z","shell.execute_reply":"2023-01-30T02:10:24.108517Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Most patients are diagnosed on Machine 49, so it would be more cancer detected on this machine. Usage of Machine 21, 29 and 48 are almost equal. Let's see how many cancer cases found on those machines.","metadata":{}},{"cell_type":"code","source":"machine_cancer_df = train_df.groupby([\"machine_id\", \"cancer\"]).count()[\"patient_id\"].reset_index()\nplt.figure(figsize=(10,5))\nsns.barplot(x=\"machine_id\", y=\"patient_id\", data=machine_cancer_df, hue=\"cancer\")\nplt.ylabel(\"Count\")","metadata":{"execution":{"iopub.status.busy":"2023-01-30T02:06:14.272044Z","iopub.execute_input":"2023-01-30T02:06:14.272467Z","iopub.status.idle":"2023-01-30T02:06:14.657759Z","shell.execute_reply.started":"2023-01-30T02:06:14.27243Z","shell.execute_reply":"2023-01-30T02:06:14.656696Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There are more cancer cases detected on Machine 49 which is normal. Machine 21, 29, 48 serviced almost same number of patients and detected almost equal cancer cases. For better visualization, let's draw the cancer percentage over all these machines.","metadata":{}},{"cell_type":"code","source":"total_machine_df = machine_cancer_df.groupby(\"machine_id\").sum().reset_index()\nmachine_cancer_percent = machine_cancer_df.merge(total_machine_df, on=\"machine_id\")\nmachine_cancer_percent[\"cancer_percent\"] = 100.0*machine_cancer_percent[\"patient_id_x\"]/machine_cancer_percent[\"patient_id_y\"]\ncancer_percent = machine_cancer_percent[machine_cancer_percent[\"cancer_x\"]==1]\nplt.figure(figsize=(10,5))\nsns.barplot(x=\"machine_id\", y=\"cancer_percent\", data=cancer_percent)\nplt.ylabel(\"Cancer percent(%)\")\nplt.title(\"Percentage of cancer detected\")","metadata":{"execution":{"iopub.status.busy":"2023-01-30T02:09:37.534844Z","iopub.execute_input":"2023-01-30T02:09:37.536254Z","iopub.status.idle":"2023-01-30T02:09:37.786784Z","shell.execute_reply.started":"2023-01-30T02:09:37.5362Z","shell.execute_reply":"2023-01-30T02:09:37.785649Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"machine_num = train_df.groupby(\"machine_id\").count()[\"patient_id\"].reset_index()\nmost_machine = machine_num[machine_num[\"machine_id\"]< 50][\"patient_id\"].sum()/machine_num[\"patient_id\"].sum()\nprint(f\"{round(most_machine * 100.0, 1)} % patients are token mammograms on Machine 21, 29, 48, 49.\")","metadata":{"execution":{"iopub.status.busy":"2023-01-30T02:44:32.283596Z","iopub.execute_input":"2023-01-30T02:44:32.284018Z","iopub.status.idle":"2023-01-30T02:44:32.309806Z","shell.execute_reply.started":"2023-01-30T02:44:32.283984Z","shell.execute_reply":"2023-01-30T02:44:32.308721Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Percentage of cancer detected on machine 21, 29, 48, 49 are around 2% ~ 2.5% which is quite close. Those 4 machines served most of patients, occupying 89%. Other machines are much less used.","metadata":{}},{"cell_type":"markdown","source":"Let's take a look at images with and without cancer.","metadata":{}},{"cell_type":"code","source":"image_no_cancer = train_df[train_df[\"cancer\"]==0]\nimage_cancer = train_df[train_df[\"cancer\"]==1]\ndef get_image_path(image_df, n):\n    ids = image_df[[\"patient_id\", \"image_id\"]].sample(n, random_state=42)\n    paths = []\n    for i in range(len(ids)):\n        path = os.path.join(base_img_dir, str(ids[\"patient_id\"].values[i]), str(ids[\"image_id\"].values[i])+'.dcm')\n        paths.append(path)\n    return paths\n\nn = 5\nno_cancer_paths = get_image_path(image_no_cancer, n)\ncancer_paths = get_image_path(image_cancer, n)\n\nplt.figure(figsize=(15, 8))\n\nfor j in range(n):\n    plt.subplot(1, n, j+1)\n    im = pydicom.dcmread(no_cancer_paths[j])\n    plt.imshow(im.pixel_array, cmap='bone')\n    plt.grid(False)\n    plt.title(f\"No cancer\")\n    plt.axis(\"off\")\nplt.show()\nplt.figure(figsize=(15, 8))\nfor j in range(n):\n    plt.subplot(1, n, j + 1)\n    im = pydicom.dcmread(cancer_paths[j])\n    plt.imshow(im.pixel_array, cmap='bone')\n    plt.grid(False)\n    plt.title(f\"Cancer\")\n    plt.axis(\"off\")\nplt.show()\n# fig.tight_layout()","metadata":{"execution":{"iopub.status.busy":"2023-01-30T02:44:46.728181Z","iopub.execute_input":"2023-01-30T02:44:46.728579Z","iopub.status.idle":"2023-01-30T02:45:03.840781Z","shell.execute_reply.started":"2023-01-30T02:44:46.728545Z","shell.execute_reply":"2023-01-30T02:45:03.839949Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It seems there are some concentration part in images with cancer while it is more balanced distribution in image without cancer. But it is hard to say. Let's try to use PCA to reduce dimensions and do a logistic regression to predict whether it is with cancer.","metadata":{}},{"cell_type":"markdown","source":"# PCA\n\nI will use the resized images with 256 x 256 from [this notebook](https://www.kaggle.com/code/theoviel/dicom-resized-png-jpg).","metadata":{}},{"cell_type":"code","source":"def get_image_label(img_path, df):\n    X = []\n    y = []\n    for file in tqdm(os.listdir(img_path)):\n        if file.endswith('.png'):\n            im = cv2.imread(os.path.join(img_path, file), 0)\n            im = cv2.resize(im, (64, 64))  # reduce to 64 x 64 to do pca\n            im = im.reshape((-1))\n            X.append(im)\n            fname = file.split('.')[0]\n            p_id, im_id = fname.split(\"_\")\n            p_id = int(p_id)\n            im_id = int(im_id)\n            data = df.loc[(df[\"patient_id\"]==p_id) & (df[\"image_id\"]==im_id)]\n            label = data[\"cancer\"].values[0]\n            y.append(label)\n    return X, y\n            \nimg_path = \"/kaggle/input/rsna-breast-cancer-256-pngs\"\nX, y = get_image_label(img_path, train_df)\nX_ = np.array(X)\ny_ = np.array(y)\nX_train, X_val, y_train, y_val = train_test_split(X_, y_, test_size=0.2, stratify=y_)","metadata":{"execution":{"iopub.status.busy":"2023-01-30T02:45:35.408692Z","iopub.execute_input":"2023-01-30T02:45:35.40941Z","iopub.status.idle":"2023-01-30T02:55:51.117222Z","shell.execute_reply.started":"2023-01-30T02:45:35.409369Z","shell.execute_reply":"2023-01-30T02:55:51.116209Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pca = PCA(n_components=0.99)\nX_pca = pca.fit_transform(X_train)\nprint(f\"Before PCA, image dimensions are {X_train.shape[1:]}\")\nprint(f\"After PCA (keep 0.99 variances), image dimensions reduced to {X_pca.shape[1:]}\")","metadata":{"execution":{"iopub.status.busy":"2023-01-30T03:18:03.156627Z","iopub.execute_input":"2023-01-30T03:18:03.157664Z","iopub.status.idle":"2023-01-30T03:20:11.978374Z","shell.execute_reply.started":"2023-01-30T03:18:03.15761Z","shell.execute_reply":"2023-01-30T03:20:11.97693Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"x_reconstructed = pca.inverse_transform(X_pca[:10])\nplt.figure(figsize=(15,8))\nfor i in range(len(x_reconstructed)):\n    x = x_reconstructed[i].reshape((64, 64))\n    plt.subplot(2, 5, i+1)\n    plt.imshow(x, cmap='gray')\n    plt.title(y_train[i])\n    plt.axis(\"off\")\nplt.tight_layout()","metadata":{"execution":{"iopub.status.busy":"2023-01-30T03:17:31.716841Z","iopub.execute_input":"2023-01-30T03:17:31.717947Z","iopub.status.idle":"2023-01-30T03:17:33.36703Z","shell.execute_reply.started":"2023-01-30T03:17:31.717903Z","shell.execute_reply":"2023-01-30T03:17:33.366153Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It is hard to do image classification without deep neural network given such a big image. But here we just do a simple demonstration with pca and logistic regression.\n\nYou could notice the reconstructed images are quite vague. ","metadata":{}},{"cell_type":"markdown","source":"# Logistic Regression to classify breast cancer","metadata":{}},{"cell_type":"code","source":"lr_clf = LogisticRegression()\nlr_clf.fit(X_pca, y_train)\nlr_clf.score(X_pca, y_train)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_pca_val = pca.transform(X_val)\nlr_clf.score(X_pca_val, y_val)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Submission","metadata":{}},{"cell_type":"code","source":"test_path = \"/kaggle/input/rsna-breast-cancer-detection/test.csv\"\ntest_img_dir = \"/kaggle/input/rsna-breast-cancer-detection/test_images\"\ntest_df = pd.read_csv(test_path)\nprediction_id = []\ncancer = []\nfor i in range(len(test_df)):\n    line = test_df.iloc[i]\n    p_id = str(line[\"patient_id\"])\n    im_id = str(line[\"image_id\"]) + '.dcm'\n    prediction_id.append(line[\"prediction_id\"])\n    path = os.path.join(test_img_dir, p_id, im_id)\n    im = pydicom.dcmread(path).pixel_array\n    im = np.array(im)\n    im = cv2.resize(im, (64, 64))\n    im = im.reshape((-1))\n    im_pca = pca.transform([im])\n    prob = lr_clf.predict_proba(im_pca)[0][1]\n    cancer.append(prob)\n    \nsubmission = pd.DataFrame(data={\"prediction_id\": prediction_id, \"cancer\": cancer})\nsubmission = submission.sort_values(\"cancer\", ascending=False)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission = submission.drop_duplicates(subset=\"prediction_id\", keep=\"first\")\nsubmission.to_csv(\"submission.csv\", index=False)\nsubmission","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"That's it! Thank you for your reading. If you feel this notebook is helpful, don't forget to **upvote**. Have a good one!","metadata":{}}]}