{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Introduction\n\nIn The Mythical Man Month, Fred Brooks, the father of the IBM 360, opined \"you should always \"plan to throw one away ... you will anyway\"\n\nThis notebook:\n\n* Is derived from:\n  * [My initial competition entry](https://www.kaggle.com/code/julianmacnamara/rsna-screening-mammography-breast-cancer-detection)\n  * [RSNA Breast Cancer Detection (PTTOA) v1.1](https://www.kaggle.com/code/julianmacnamara/rsna-breast-cancer-detection-pttoa-v1-1)\n* Uses:\n  * fp16\n  * The [RSNA Breast Cancer Detection - 512x512 pngs](https://www.kaggle.com/datasets/theoviel/rsna-breast-cancer-512-pngs) dataset which was created by Theo Viel\n  * A batch size of 32\n  * ConvNeXt (convnext_small_22k_224)\n  * An ensemble approach to consuder the influence of different attributes such as machine_id on submission scores","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\n\nimport timm\n\nimport cv2\n\nimport shutil\nimport random\nimport warnings\n\nwarnings.filterwarnings(\"ignore\", category=UserWarning) \n\nimport pydicom\nfrom pydicom.pixel_data_handlers.util import apply_voi_lut\n\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport os\nfrom pathlib import Path\nimport glob\nimport random\n\nfrom fastai.data.all import *\nfrom fastai.vision.all import *\nfrom fastai.callback.fp16 import *\nfrom fastai.vision.widgets import *","metadata":{"execution":{"iopub.status.busy":"2023-02-18T13:39:03.206739Z","iopub.execute_input":"2023-02-18T13:39:03.207156Z","iopub.status.idle":"2023-02-18T13:39:07.645376Z","shell.execute_reply.started":"2023-02-18T13:39:03.207076Z","shell.execute_reply":"2023-02-18T13:39:07.644343Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!mkdir -p /root/.cache/torch/hub/checkpoints\n\nsrc = '/kaggle/input/convnext/convnext_small_22k_224.pth'\ndst = '/root/.cache/torch/hub/checkpoints/convnext_small_22k_224.pth'\n\nshutil.copy(src, dst)","metadata":{"execution":{"iopub.status.busy":"2023-02-18T13:39:15.86918Z","iopub.execute_input":"2023-02-18T13:39:15.869948Z","iopub.status.idle":"2023-02-18T13:39:20.420851Z","shell.execute_reply.started":"2023-02-18T13:39:15.869897Z","shell.execute_reply":"2023-02-18T13:39:20.419659Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"path='/kaggle/input/rsna-difficult-images-20230206/images_difficult_cases/'","metadata":{"execution":{"iopub.status.busy":"2023-02-18T13:39:24.655783Z","iopub.execute_input":"2023-02-18T13:39:24.656299Z","iopub.status.idle":"2023-02-18T13:39:24.6813Z","shell.execute_reply.started":"2023-02-18T13:39:24.65625Z","shell.execute_reply":"2023-02-18T13:39:24.677063Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Checking image file\n\nimg = PILImage.create(path+'64056_326013080.png')\nimg.to_thumb(128)","metadata":{"execution":{"iopub.status.busy":"2023-02-18T13:39:29.716312Z","iopub.execute_input":"2023-02-18T13:39:29.716759Z","iopub.status.idle":"2023-02-18T13:39:29.767684Z","shell.execute_reply.started":"2023-02-18T13:39:29.716721Z","shell.execute_reply":"2023-02-18T13:39:29.766666Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_csv = pd.read_csv('/kaggle/input/rsna-breast-cancer-detection/test.csv')\ntest_csv","metadata":{"execution":{"iopub.status.busy":"2023-02-18T13:39:33.024043Z","iopub.execute_input":"2023-02-18T13:39:33.024501Z","iopub.status.idle":"2023-02-18T13:39:33.065808Z","shell.execute_reply.started":"2023-02-18T13:39:33.024462Z","shell.execute_reply":"2023-02-18T13:39:33.064778Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### As noted in [A Brief Intro to Mammography](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369262) \n\n* For this task, we are focused on screening mammograms, which again means that only scores of BI-RADS 0, 1, or 2 are possible. One can essentially think of 0 as \"abnormal\" and 1 and 2 as \"normal.\" There can be subjectivity in assigning BI-RADS 1 or 2. For example, if there are stable findings that are almost certainly benign breast cysts, the mammogram may be assigned 1 or 2 depending on the radiologist.\n\n* Our task in the challenge is to predict cancer or no cancer, a binary value. This may be obvious, but not all BI-RADS 0 cases have cancer. The final cancer label will depend on the outcome of the diagnostic mammogram and the biopsy results, if obtained.\n\nSince the test set only includes MLO and CC views only rows that have these views are included\n\n[revised_01_20230206.csv](/kaggle/input/rsna-csv-files/revised_01_20230206.csv) contains:\n\n* 6,946 original and augmented images where cancer was diagnosed regardless of their BIRADS score<br>\nMammograms diagnosed with cancer were augmented using [this notebook](https://www.kaggle.com/code/julianmacnamara/rsna-augmentation). Essentially each image was flipped, inverted, mirrored, autocontrasted with a cutoff of 0.5 and equalised\n\n* 18,021 images where the BIRADS score was either 1 or 2 but cancer had not been diagnosed\n\nbut this scored poorly on submission\n\n[revised_02_20230206](/kaggle/input/rsna-csv-files/revised_01_20230206.csv) contains:\n\n* The same 6,946 original and augmented images where cancer was diagnosed<br>\n\n* 7,697 images where \"difficult_negative_case\" was true\n\n#### 7. February\n\n* Ran with randomized version of revised_02_20230206.csv. Have yet to submit\n* Changed resnet18 to resnet50. Submission score of 0.02 suggests over-fitting\n\n#### 12 February\n\n* Changed architecture to ConvNeXt\n* Created revised_03_20230212.csv which is the same as revised_02_20230206 but with an additional 8,000 records where \"difficult_negative_case\" was false\n\n#### 14 February\n\n* Using Mixup. mixup_01.csv contains:\n  * 1,158 images where cancer was diagnosed\n  * 1,185 random images where cancer was not diagnosed but \"difficult_negative_case\" was true\n  * 1,184 random images where cancer was not diagnosed but \"difficult_negative_case\" was false\n  Mixup was a disappointment\n* Experimenting with looking at different machines\n\n#### 16 February\n\n* Submitted approach based on machine_id to the competition to see if it scores and, if so, how well. Scored poorly\n\n#### 18 February\n\n* Submitted model based on machine_id 49 only","metadata":{}},{"cell_type":"code","source":"train_csv = pd.read_csv('/kaggle/input/rsna-csv-files/revised_04_m_0215.csv')\ntrain_csv['cancer'] = train_csv['cancer'].astype(str)\ntrain_csv['lbl'] = train_csv['lbl'].astype(str)\ntrain_csv['lbl1'] = train_csv['lbl1'].astype(str)\n\ntrain_csv.head()","metadata":{"execution":{"iopub.status.busy":"2023-02-18T13:39:49.70713Z","iopub.execute_input":"2023-02-18T13:39:49.707586Z","iopub.status.idle":"2023-02-18T13:39:50.069472Z","shell.execute_reply.started":"2023-02-18T13:39:49.707547Z","shell.execute_reply":"2023-02-18T13:39:50.068388Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(20,6))\nplt.subplot(1,2,1)\nax1 = sns.countplot(data=train_csv, x='lbl')\nplt.xticks(rotation = -45)\nfor container in ax1.containers:\n    ax1.bar_label(container)\nplt.title('Labels');","metadata":{"execution":{"iopub.status.busy":"2023-02-18T13:40:16.323948Z","iopub.execute_input":"2023-02-18T13:40:16.324427Z","iopub.status.idle":"2023-02-18T13:40:16.893648Z","shell.execute_reply.started":"2023-02-18T13:40:16.324386Z","shell.execute_reply":"2023-02-18T13:40:16.892597Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_49=train_csv[train_csv['machine_id']==49]\ndf_49=df_49[[\"img_path\",\"cancer\", \"machine_id\"]]\ndf_49.head()","metadata":{"execution":{"iopub.status.busy":"2023-02-18T13:40:41.484247Z","iopub.execute_input":"2023-02-18T13:40:41.48469Z","iopub.status.idle":"2023-02-18T13:40:41.516444Z","shell.execute_reply.started":"2023-02-18T13:40:41.484653Z","shell.execute_reply":"2023-02-18T13:40:41.515521Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# df_48=train_csv[train_csv['machine_id']==48]\n# df_48=df_48[[\"img_path\",\"cancer\", \"machine_id\"]]\n# df_48.head()","metadata":{"execution":{"iopub.status.busy":"2023-02-18T13:29:11.732543Z","iopub.execute_input":"2023-02-18T13:29:11.732962Z","iopub.status.idle":"2023-02-18T13:29:11.737718Z","shell.execute_reply.started":"2023-02-18T13:29:11.732919Z","shell.execute_reply":"2023-02-18T13:29:11.736367Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# df_29=train_csv[train_csv['machine_id']==29]\n# df_29=df_29[[\"img_path\",\"cancer\", \"machine_id\"]]\n# df_29.head()","metadata":{"execution":{"iopub.status.busy":"2023-02-18T13:29:11.739696Z","iopub.execute_input":"2023-02-18T13:29:11.740176Z","iopub.status.idle":"2023-02-18T13:29:11.74916Z","shell.execute_reply.started":"2023-02-18T13:29:11.740142Z","shell.execute_reply":"2023-02-18T13:29:11.747916Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# df_21=train_csv[train_csv['machine_id']==21]\n# df_21=df_21[[\"img_path\",\"cancer\", \"machine_id\"]]\n# df_21.head()","metadata":{"execution":{"iopub.status.busy":"2023-02-18T13:29:11.75063Z","iopub.execute_input":"2023-02-18T13:29:11.751198Z","iopub.status.idle":"2023-02-18T13:29:11.759462Z","shell.execute_reply.started":"2023-02-18T13:29:11.751152Z","shell.execute_reply":"2023-02-18T13:29:11.758605Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Creating the learner with a batch size of 32 and training","metadata":{}},{"cell_type":"code","source":"ens_model=[]","metadata":{"execution":{"iopub.status.busy":"2023-02-18T13:40:50.268643Z","iopub.execute_input":"2023-02-18T13:40:50.269115Z","iopub.status.idle":"2023-02-18T13:40:50.2742Z","shell.execute_reply.started":"2023-02-18T13:40:50.269075Z","shell.execute_reply":"2023-02-18T13:40:50.273249Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# df_49\n\ndef get_x(r): return r['img_path']\ndef get_y(r): return r['cancer']\n\ndblock = DataBlock(blocks=(ImageBlock, CategoryBlock),\n                   get_x = get_x,\n                   get_y = get_y,\n                   splitter=RandomSplitter (valid_pct=0.2),                   \n                   item_tfms=Resize(512),                   \n                   batch_tfms=[Normalize.from_stats(*imagenet_stats)])\n\ndsets = dblock.datasets(df_49)\n\ndls = dblock.dataloaders(df_49, bs=32)\n\nlearn_49 = vision_learner(dls,\n                       'convnext_small_in22k',\n                       metrics=accuracy).to_fp16()\n\nlrs = learn_49.lr_find()","metadata":{"execution":{"iopub.status.busy":"2023-02-18T13:40:57.315777Z","iopub.execute_input":"2023-02-18T13:40:57.316276Z","iopub.status.idle":"2023-02-18T13:42:57.488716Z","shell.execute_reply.started":"2023-02-18T13:40:57.316237Z","shell.execute_reply":"2023-02-18T13:42:57.487536Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"learn_49.fit_one_cycle(5, lrs)\n\nens_model.append(learn_49)\n\ninterp = ClassificationInterpretation.from_learner(learn_49)\ninterp.plot_confusion_matrix()","metadata":{"execution":{"iopub.status.busy":"2023-02-18T13:43:09.531044Z","iopub.execute_input":"2023-02-18T13:43:09.53151Z","iopub.status.idle":"2023-02-18T13:50:17.718851Z","shell.execute_reply.started":"2023-02-18T13:43:09.531464Z","shell.execute_reply":"2023-02-18T13:50:17.717138Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# df_48\n\n# def get_x(r): return r['img_path']\n# def get_y(r): return r['cancer']\n\n# dblock = DataBlock(blocks=(ImageBlock, CategoryBlock),\n#                    get_x = get_x,\n#                    get_y = get_y,\n#                    splitter=RandomSplitter (valid_pct=0.2),                   \n#                    item_tfms=Resize(512),                   \n#                    batch_tfms=[Normalize.from_stats(*imagenet_stats)])\n\n# dsets = dblock.datasets(df_48)\n\n# dls = dblock.dataloaders(df_48, bs=32)\n\n# learn_48 = vision_learner(dls,\n#                        'convnext_small_in22k',\n#                        metrics=accuracy).to_fp16()\n\n# lrs = learn_48.lr_find()\n\n# learn_48.fit_one_cycle(5, lrs)\n\n# ens_model.append(learn_48)\n\n# interp = ClassificationInterpretation.from_learner(learn_48)\n# interp.plot_confusion_matrix()","metadata":{"execution":{"iopub.status.busy":"2023-02-18T13:32:03.433593Z","iopub.execute_input":"2023-02-18T13:32:03.434937Z","iopub.status.idle":"2023-02-18T13:32:03.443931Z","shell.execute_reply.started":"2023-02-18T13:32:03.434891Z","shell.execute_reply":"2023-02-18T13:32:03.440686Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# df_29\n\n# def get_x(r): return r['img_path']\n# def get_y(r): return r['cancer']\n\n# dblock = DataBlock(blocks=(ImageBlock, CategoryBlock),\n#                    get_x = get_x,\n#                    get_y = get_y,\n#                    splitter=RandomSplitter (valid_pct=0.2),                   \n#                    item_tfms=Resize(512),                   \n#                    batch_tfms=[Normalize.from_stats(*imagenet_stats)])\n\n# dsets = dblock.datasets(df_29)\n\n# dls = dblock.dataloaders(df_29, bs=32)\n\n# learn_29 = vision_learner(dls,\n#                        'convnext_small_in22k',\n#                        metrics=accuracy).to_fp16()\n\n# lrs = learn_29.lr_find()\n\n# learn_29.fit_one_cycle(5, lrs)\n\n# ens_model.append(learn_29)\n\n# interp = ClassificationInterpretation.from_learner(learn_29)\n# interp.plot_confusion_matrix()","metadata":{"execution":{"iopub.status.busy":"2023-02-18T13:32:03.445812Z","iopub.execute_input":"2023-02-18T13:32:03.446589Z","iopub.status.idle":"2023-02-18T13:32:03.457445Z","shell.execute_reply.started":"2023-02-18T13:32:03.446544Z","shell.execute_reply":"2023-02-18T13:32:03.455394Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# df_21\n\n# def get_x(r): return r['img_path']\n# def get_y(r): return r['cancer']\n\n# dblock = DataBlock(blocks=(ImageBlock, CategoryBlock),\n#                    get_x = get_x,\n#                    get_y = get_y,\n#                    splitter=RandomSplitter (valid_pct=0.2),                   \n#                    item_tfms=Resize(512),                   \n#                    batch_tfms=[Normalize.from_stats(*imagenet_stats)])\n\n# dsets = dblock.datasets(df_21)\n\n# dls = dblock.dataloaders(df_21, bs=32)\n\n# learn_21 = vision_learner(dls,\n#                        'convnext_small_in22k',\n#                        metrics=accuracy).to_fp16()\n\n# lrs = learn_21.lr_find()\n\n# learn_21.fit_one_cycle(5, lrs)\n\n# ens_model.append(learn_29)\n\n# interp = ClassificationInterpretation.from_learner(learn_29)\n# interp.plot_confusion_matrix()","metadata":{"execution":{"iopub.status.busy":"2023-02-18T13:32:03.459395Z","iopub.execute_input":"2023-02-18T13:32:03.460377Z","iopub.status.idle":"2023-02-18T13:32:03.473617Z","shell.execute_reply.started":"2023-02-18T13:32:03.460329Z","shell.execute_reply":"2023-02-18T13:32:03.472154Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# ens_model","metadata":{"execution":{"iopub.status.busy":"2023-02-18T13:52:36.658076Z","iopub.execute_input":"2023-02-18T13:52:36.658735Z","iopub.status.idle":"2023-02-18T13:52:36.664439Z","shell.execute_reply.started":"2023-02-18T13:52:36.658675Z","shell.execute_reply":"2023-02-18T13:52:36.663424Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Prediction and submission","metadata":{}},{"cell_type":"code","source":"path='/kaggle/input/rsna-test-images-as-pngs/'","metadata":{"execution":{"iopub.status.busy":"2023-02-18T13:52:36.666856Z","iopub.execute_input":"2023-02-18T13:52:36.667602Z","iopub.status.idle":"2023-02-18T13:52:36.682331Z","shell.execute_reply.started":"2023-02-18T13:52:36.667558Z","shell.execute_reply":"2023-02-18T13:52:36.681389Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Source: https://www.kaggle.com/code/beezus666/dicom-to-png-to-predict\n\n# Get images\ntest_images = sorted([os.path.join(path, file) for file in os.listdir(path)])\nimages = [cv2.imread(file) for file in test_images]","metadata":{"execution":{"iopub.status.busy":"2023-02-18T13:52:36.683559Z","iopub.execute_input":"2023-02-18T13:52:36.684804Z","iopub.status.idle":"2023-02-18T13:52:36.762584Z","shell.execute_reply.started":"2023-02-18T13:52:36.684765Z","shell.execute_reply":"2023-02-18T13:52:36.760899Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 49\n# pass images to fast.ai learner and get predictions\ntest_dl = learn_49.dls.test_dl(images)\npreds_batch, _ = learn_49.get_preds(dl=test_dl)\npredsdec, _, decoded = learn_49.get_preds(dl=test_dl, with_decoded=True)\npredsdec[:10]\n\n# extract the prediction, which is the 2nd value in the tensor above\n\narray = preds_batch.numpy()\nlist_of_lists = array.tolist()\nvalues_49 = [lst[1] for lst in list_of_lists]","metadata":{"execution":{"iopub.status.busy":"2023-02-18T13:52:36.766862Z","iopub.execute_input":"2023-02-18T13:52:36.767618Z","iopub.status.idle":"2023-02-18T13:52:37.5374Z","shell.execute_reply.started":"2023-02-18T13:52:36.767576Z","shell.execute_reply":"2023-02-18T13:52:37.536235Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# # 48\n# # pass images to fast.ai learner and get predictions\n# test_dl = learn_48.dls.test_dl(images)\n# preds_batch, _ = learn_48.get_preds(dl=test_dl)\n# predsdec, _, decoded = learn_48.get_preds(dl=test_dl, with_decoded=True)\n# predsdec[:10]\n\n# # extract the prediction, which is the 2nd value in the tensor above\n\n# array = preds_batch.numpy()\n# list_of_lists = array.tolist()\n# values_48 = [lst[1] for lst in list_of_lists]","metadata":{"execution":{"iopub.status.busy":"2023-02-18T13:52:37.539554Z","iopub.execute_input":"2023-02-18T13:52:37.540277Z","iopub.status.idle":"2023-02-18T13:52:37.546305Z","shell.execute_reply.started":"2023-02-18T13:52:37.540231Z","shell.execute_reply":"2023-02-18T13:52:37.544878Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# # 29\n# # pass images to fast.ai learner and get predictions\n# test_dl = learn_29.dls.test_dl(images)\n# preds_batch, _ = learn_29.get_preds(dl=test_dl)\n# predsdec, _, decoded = learn_29.get_preds(dl=test_dl, with_decoded=True)\n# predsdec[:10]\n\n# # extract the prediction, which is the 2nd value in the tensor above\n\n# array = preds_batch.numpy()\n# list_of_lists = array.tolist()\n# values_29 = [lst[1] for lst in list_of_lists]","metadata":{"execution":{"iopub.status.busy":"2023-02-18T13:52:37.548269Z","iopub.execute_input":"2023-02-18T13:52:37.549121Z","iopub.status.idle":"2023-02-18T13:52:37.559519Z","shell.execute_reply.started":"2023-02-18T13:52:37.54908Z","shell.execute_reply":"2023-02-18T13:52:37.558383Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# # 21\n# # pass images to fast.ai learner and get predictions\n# test_dl = learn_21.dls.test_dl(images)\n# preds_batch, _ = learn_21.get_preds(dl=test_dl)\n# predsdec, _, decoded = learn_21.get_preds(dl=test_dl, with_decoded=True)\n# predsdec[:10]\n\n# # extract the prediction, which is the 2nd value in the tensor above\n\n# array = preds_batch.numpy()\n# list_of_lists = array.tolist()\n# values_21 = [lst[1] for lst in list_of_lists]","metadata":{"execution":{"iopub.status.busy":"2023-02-18T13:52:37.561569Z","iopub.execute_input":"2023-02-18T13:52:37.562789Z","iopub.status.idle":"2023-02-18T13:52:37.574646Z","shell.execute_reply.started":"2023-02-18T13:52:37.562748Z","shell.execute_reply":"2023-02-18T13:52:37.573515Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# columns=['values_49','values_48','values_29','values_21']\n# df_pred=pd.DataFrame(list(zip(values_49,values_48,values_29,values_21)),\n#                     columns=columns)\n# df_pred['pred']=df_pred.mean(axis=1)","metadata":{"execution":{"iopub.status.busy":"2023-02-18T13:52:37.576724Z","iopub.execute_input":"2023-02-18T13:52:37.577595Z","iopub.status.idle":"2023-02-18T13:52:37.586375Z","shell.execute_reply.started":"2023-02-18T13:52:37.577553Z","shell.execute_reply":"2023-02-18T13:52:37.585134Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# print(df_pred)","metadata":{"execution":{"iopub.status.busy":"2023-02-18T13:52:37.594016Z","iopub.execute_input":"2023-02-18T13:52:37.595101Z","iopub.status.idle":"2023-02-18T13:52:37.60715Z","shell.execute_reply.started":"2023-02-18T13:52:37.595055Z","shell.execute_reply":"2023-02-18T13:52:37.605996Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# col_list=df_pred.pred.values.tolist()\n# print(col_list)","metadata":{"execution":{"iopub.status.busy":"2023-02-18T13:52:37.608606Z","iopub.execute_input":"2023-02-18T13:52:37.60988Z","iopub.status.idle":"2023-02-18T13:52:37.617582Z","shell.execute_reply.started":"2023-02-18T13:52:37.609836Z","shell.execute_reply":"2023-02-18T13:52:37.616387Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sorted_files = sorted(os.listdir(path))\nno_extensions = [os.path.splitext(name)[0] for name in sorted_files]","metadata":{"execution":{"iopub.status.busy":"2023-02-18T13:52:37.618874Z","iopub.execute_input":"2023-02-18T13:52:37.619241Z","iopub.status.idle":"2023-02-18T13:52:37.629432Z","shell.execute_reply.started":"2023-02-18T13:52:37.619204Z","shell.execute_reply":"2023-02-18T13:52:37.628175Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# preds_df = pd.DataFrame(data = {'concat_id':no_extensions, 'cancer':col_list})\npreds_df = pd.DataFrame(data = {'concat_id':no_extensions, 'cancer':values_49})\npreds_df\n","metadata":{"execution":{"iopub.status.busy":"2023-02-18T13:52:37.631464Z","iopub.execute_input":"2023-02-18T13:52:37.632251Z","iopub.status.idle":"2023-02-18T13:52:37.654177Z","shell.execute_reply.started":"2023-02-18T13:52:37.632211Z","shell.execute_reply":"2023-02-18T13:52:37.652886Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"info_df=pd.read_csv('/kaggle/input/rsna-breast-cancer-detection/test.csv')\ninfo_df['concat_id'] = info_df['image_id'].astype(str)\ninfo_df=info_df[[\"patient_id\",\"image_id\",\"laterality\",\"concat_id\"]]\n\nmerged_df = pd.merge(right=info_df, left=preds_df, on='concat_id')\nmerged_df['prediction_id'] = merged_df['patient_id'].astype(str)+'_'+merged_df['laterality'].astype(str)","metadata":{"execution":{"iopub.status.busy":"2023-02-18T13:52:37.655764Z","iopub.execute_input":"2023-02-18T13:52:37.656457Z","iopub.status.idle":"2023-02-18T13:52:37.688148Z","shell.execute_reply.started":"2023-02-18T13:52:37.656407Z","shell.execute_reply":"2023-02-18T13:52:37.68703Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Select only the 'prediction_id' and 'cancer' columns\nresulting_df = merged_df[['prediction_id', 'cancer']]\nresulting_df = resulting_df.groupby('prediction_id', as_index=False).mean()\nresulting_df = resulting_df.sort_index()","metadata":{"execution":{"iopub.status.busy":"2023-02-18T13:52:37.689927Z","iopub.execute_input":"2023-02-18T13:52:37.690623Z","iopub.status.idle":"2023-02-18T13:52:37.705035Z","shell.execute_reply.started":"2023-02-18T13:52:37.690583Z","shell.execute_reply":"2023-02-18T13:52:37.703643Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"resulting_df","metadata":{"execution":{"iopub.status.busy":"2023-02-18T13:52:37.710879Z","iopub.execute_input":"2023-02-18T13:52:37.713544Z","iopub.status.idle":"2023-02-18T13:52:37.730549Z","shell.execute_reply.started":"2023-02-18T13:52:37.713489Z","shell.execute_reply":"2023-02-18T13:52:37.729153Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"i=resulting_df['cancer'][0]\nj=resulting_df['cancer'][1]\n\nfinal_df=pd.read_csv('/kaggle/input/rsna-breast-cancer-detection/sample_submission.csv')\nfinal_df.loc[0:,\"cancer\"]=i\nfinal_df.loc[1:,\"cancer\"]=j\nfinal_df","metadata":{"execution":{"iopub.status.busy":"2023-02-18T13:52:37.735704Z","iopub.execute_input":"2023-02-18T13:52:37.739106Z","iopub.status.idle":"2023-02-18T13:52:37.769267Z","shell.execute_reply.started":"2023-02-18T13:52:37.739062Z","shell.execute_reply":"2023-02-18T13:52:37.762532Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"shutil.rmtree(\"/kaggle/working/models\")","metadata":{"execution":{"iopub.status.busy":"2023-02-18T13:52:37.771047Z","iopub.execute_input":"2023-02-18T13:52:37.771745Z","iopub.status.idle":"2023-02-18T13:52:37.779614Z","shell.execute_reply.started":"2023-02-18T13:52:37.771704Z","shell.execute_reply":"2023-02-18T13:52:37.777059Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"final_df.to_csv('submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2023-02-18T13:52:37.78913Z","iopub.execute_input":"2023-02-18T13:52:37.789905Z","iopub.status.idle":"2023-02-18T13:52:37.804878Z","shell.execute_reply.started":"2023-02-18T13:52:37.78985Z","shell.execute_reply":"2023-02-18T13:52:37.803864Z"},"trusted":true},"execution_count":null,"outputs":[]}]}