{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        #print(os.path.join(dirname, filename))\n        pass\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-01-19T13:30:55.760921Z","iopub.execute_input":"2023-01-19T13:30:55.761336Z","iopub.status.idle":"2023-01-19T13:31:40.293412Z","shell.execute_reply.started":"2023-01-19T13:30:55.761251Z","shell.execute_reply":"2023-01-19T13:31:40.290365Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%load_ext autoreload\n%autoreload 2","metadata":{"execution":{"iopub.status.busy":"2023-01-19T13:31:40.299631Z","iopub.execute_input":"2023-01-19T13:31:40.302628Z","iopub.status.idle":"2023-01-19T13:31:40.351214Z","shell.execute_reply.started":"2023-01-19T13:31:40.302589Z","shell.execute_reply":"2023-01-19T13:31:40.350301Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nfrom pathlib import Path","metadata":{"execution":{"iopub.status.busy":"2023-01-19T13:31:40.355427Z","iopub.execute_input":"2023-01-19T13:31:40.357686Z","iopub.status.idle":"2023-01-19T13:31:40.394891Z","shell.execute_reply.started":"2023-01-19T13:31:40.357638Z","shell.execute_reply":"2023-01-19T13:31:40.393882Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- First I have tried to work on preprocessed dicom images, which is published [here](https://www.kaggle.com/datasets/remekkinas/rsna-breast-cancer-detection-poi-images)\n- In my local laptop which has 4 GB GPU it took around 40 minutes to run one epoch. I thought to create a smaller dataset with smaller image size and smaller number of images.\n_ Those smaller data can be found [here](https://www.kaggle.com/datasets/hasangoni/rsna-small-for-faster-experimentation).\n- With those data it took me around 1minutes for one epoch. Now I can create a quick pipeline, to see every thing works or not. \n","metadata":{}},{"cell_type":"code","source":"!pip install -Uq fastai\n!pip install timm","metadata":{"execution":{"iopub.status.busy":"2023-01-19T13:31:40.40089Z","iopub.execute_input":"2023-01-19T13:31:40.403447Z","iopub.status.idle":"2023-01-19T13:32:14.819647Z","shell.execute_reply.started":"2023-01-19T13:31:40.40341Z","shell.execute_reply":"2023-01-19T13:32:14.818355Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Path\nhttps://docs.python.org/3/library/pathlib.html","metadata":{}},{"cell_type":"code","source":"from pathlib import Path\nfrom fastai.vision.all import *","metadata":{"execution":{"iopub.status.busy":"2023-01-19T13:32:14.824825Z","iopub.execute_input":"2023-01-19T13:32:14.827351Z","iopub.status.idle":"2023-01-19T13:32:19.156944Z","shell.execute_reply.started":"2023-01-19T13:32:14.827306Z","shell.execute_reply":"2023-01-19T13:32:19.155679Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Defining paths","metadata":{}},{"cell_type":"code","source":"root_path = Path(r'/kaggle/input/')\n\nactual_path = Path(fr'{root_path}/rsna-breast-cancer-detection')\nsmall_image_path = Path(fr'{root_path}/rsna-small-for-faster-experimentation/data_small')\n                 \n# checking how much files are avilable    \nsmall_image_path.ls()\n","metadata":{"execution":{"iopub.status.busy":"2023-01-19T13:32:19.162134Z","iopub.execute_input":"2023-01-19T13:32:19.164731Z","iopub.status.idle":"2023-01-19T13:32:19.260531Z","shell.execute_reply.started":"2023-01-19T13:32:19.164684Z","shell.execute_reply":"2023-01-19T13:32:19.259617Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This is taken from fast ai https://docs.fast.ai/data.transforms.html\n","metadata":{}},{"cell_type":"code","source":"fn = get_image_files(small_image_path)\nlen(fn)","metadata":{"execution":{"iopub.status.busy":"2023-01-19T13:32:19.265552Z","iopub.execute_input":"2023-01-19T13:32:19.267948Z","iopub.status.idle":"2023-01-19T13:32:20.45505Z","shell.execute_reply.started":"2023-01-19T13:32:19.267909Z","shell.execute_reply":"2023-01-19T13:32:20.453636Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"type(fn)","metadata":{"execution":{"iopub.status.busy":"2023-01-19T13:32:36.273497Z","iopub.execute_input":"2023-01-19T13:32:36.274027Z","iopub.status.idle":"2023-01-19T13:32:36.339801Z","shell.execute_reply.started":"2023-01-19T13:32:36.273982Z","shell.execute_reply":"2023-01-19T13:32:36.338822Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fn[0]","metadata":{"execution":{"iopub.status.busy":"2023-01-19T13:32:39.748825Z","iopub.execute_input":"2023-01-19T13:32:39.749353Z","iopub.status.idle":"2023-01-19T13:32:39.813548Z","shell.execute_reply.started":"2023-01-19T13:32:39.749309Z","shell.execute_reply":"2023-01-19T13:32:39.812229Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"actual_path.ls()","metadata":{"execution":{"iopub.status.busy":"2023-01-19T13:32:43.546386Z","iopub.execute_input":"2023-01-19T13:32:43.546895Z","iopub.status.idle":"2023-01-19T13:32:43.620561Z","shell.execute_reply.started":"2023-01-19T13:32:43.546815Z","shell.execute_reply":"2023-01-19T13:32:43.619274Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Reading Training csv files and process little bit ","metadata":{}},{"cell_type":"code","source":"df_trn = pd.read_csv(f'{actual_path}/train.csv')\ndf_trn=(\n    df_trn\n    .assign(image_id = lambda df: df['image_id'].astype(str))\n)\ndf_trn.head()","metadata":{"execution":{"iopub.status.busy":"2023-01-19T13:32:49.037131Z","iopub.execute_input":"2023-01-19T13:32:49.037594Z","iopub.status.idle":"2023-01-19T13:32:49.300976Z","shell.execute_reply.started":"2023-01-19T13:32:49.037556Z","shell.execute_reply":"2023-01-19T13:32:49.299893Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"type(df_trn.iloc[0]['image_id'])","metadata":{"execution":{"iopub.status.busy":"2023-01-19T13:32:52.536488Z","iopub.execute_input":"2023-01-19T13:32:52.536995Z","iopub.status.idle":"2023-01-19T13:32:52.60082Z","shell.execute_reply.started":"2023-01-19T13:32:52.53695Z","shell.execute_reply":"2023-01-19T13:32:52.599642Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Creating a function to get label name from the file path and training dataframe","metadata":{}},{"cell_type":"code","source":"fn_id = [ i.stem.split('_')[1] for i in fn]\nlabel_df = [df_trn.loc[df_trn['image_id'] == i, 'cancer'].values[0] for i in fn_id]\ndef get_y_label(x):\n    return df_trn.loc[df_trn['image_id'] == x.stem.split('_')[1], 'cancer'].values[0]","metadata":{"execution":{"iopub.status.busy":"2023-01-19T13:32:57.681693Z","iopub.execute_input":"2023-01-19T13:32:57.682173Z","iopub.status.idle":"2023-01-19T13:33:31.536765Z","shell.execute_reply.started":"2023-01-19T13:32:57.682123Z","shell.execute_reply":"2023-01-19T13:33:31.535656Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"type(fn_id)","metadata":{"execution":{"iopub.status.busy":"2023-01-19T13:34:17.516496Z","iopub.execute_input":"2023-01-19T13:34:17.517126Z","iopub.status.idle":"2023-01-19T13:34:17.587352Z","shell.execute_reply.started":"2023-01-19T13:34:17.517075Z","shell.execute_reply":"2023-01-19T13:34:17.586055Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(len(label_df))\nprint(df_trn.shape)","metadata":{"execution":{"iopub.status.busy":"2023-01-19T13:34:20.481032Z","iopub.execute_input":"2023-01-19T13:34:20.481492Z","iopub.status.idle":"2023-01-19T13:34:20.543638Z","shell.execute_reply.started":"2023-01-19T13:34:20.481453Z","shell.execute_reply":"2023-01-19T13:34:20.542414Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"get_y_label(fn[0])","metadata":{"execution":{"iopub.status.busy":"2023-01-19T13:34:22.465977Z","iopub.execute_input":"2023-01-19T13:34:22.466464Z","iopub.status.idle":"2023-01-19T13:34:22.536377Z","shell.execute_reply.started":"2023-01-19T13:34:22.466424Z","shell.execute_reply":"2023-01-19T13:34:22.535351Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(fn[0])\nprint(fn_id[0])","metadata":{"execution":{"iopub.status.busy":"2023-01-19T13:34:24.602227Z","iopub.execute_input":"2023-01-19T13:34:24.602675Z","iopub.status.idle":"2023-01-19T13:34:24.665809Z","shell.execute_reply.started":"2023-01-19T13:34:24.602635Z","shell.execute_reply":"2023-01-19T13:34:24.664361Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"label_df[:20]","metadata":{"execution":{"iopub.status.busy":"2023-01-19T13:34:26.257082Z","iopub.execute_input":"2023-01-19T13:34:26.257539Z","iopub.status.idle":"2023-01-19T13:34:26.320726Z","shell.execute_reply.started":"2023-01-19T13:34:26.257501Z","shell.execute_reply":"2023-01-19T13:34:26.319602Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Creating a dataloader to train ","metadata":{}},{"cell_type":"markdown","source":"resize , aug_transforms\nhttps://docs.fast.ai/vision.augment.html","metadata":{}},{"cell_type":"markdown","source":"https://docs.fast.ai/vision.data.html","metadata":{}},{"cell_type":"code","source":"dls = ImageDataLoaders.from_path_func(\n                                small_image_path,# path of the image files\n                                fnames=fn,# file names\n                                label_func=get_y_label,# function to get the label\n                                valid_pct=0.2,# percentage of validation data\n                                seed=42,# now seed to get repetitive result\n                                item_tfms=Resize(166, method='squish'),# resize image using squash method\n                                batch_tfms=aug_transforms(size=128, min_scale=0.75) # in Gpu batch augmenation\n                                )\ndls.device = default_device() # I guess it is not required, it automatically select the default_device, but just making sure.","metadata":{"execution":{"iopub.status.busy":"2023-01-19T13:42:40.622058Z","iopub.execute_input":"2023-01-19T13:42:40.622528Z","iopub.status.idle":"2023-01-19T13:43:09.601431Z","shell.execute_reply.started":"2023-01-19T13:42:40.622486Z","shell.execute_reply":"2023-01-19T13:43:09.598881Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"default_device()","metadata":{"execution":{"iopub.status.busy":"2023-01-19T13:48:16.30412Z","iopub.execute_input":"2023-01-19T13:48:16.304595Z","iopub.status.idle":"2023-01-19T13:48:16.371234Z","shell.execute_reply.started":"2023-01-19T13:48:16.304553Z","shell.execute_reply":"2023-01-19T13:48:16.370127Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Seeing whether our batch works","metadata":{}},{"cell_type":"code","source":"dls.show_batch(max_n=30)","metadata":{"execution":{"iopub.status.busy":"2023-01-19T13:48:25.146952Z","iopub.execute_input":"2023-01-19T13:48:25.147401Z","iopub.status.idle":"2023-01-19T13:48:29.334079Z","shell.execute_reply.started":"2023-01-19T13:48:25.147362Z","shell.execute_reply":"2023-01-19T13:48:29.33311Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Creating a dataloader telling structure, metrics, and use floating point 16","metadata":{}},{"cell_type":"markdown","source":"https://docs.fast.ai/vision.learner.html#vision_learner","metadata":{}},{"cell_type":"code","source":"learn = vision_learner(dls, 'resnet26d', metrics=error_rate, path='.').to_fp16()","metadata":{"execution":{"iopub.status.busy":"2023-01-19T13:54:43.755952Z","iopub.execute_input":"2023-01-19T13:54:43.756418Z","iopub.status.idle":"2023-01-19T13:54:50.077257Z","shell.execute_reply.started":"2023-01-19T13:54:43.756378Z","shell.execute_reply":"2023-01-19T13:54:50.076164Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Searching for learing rate. Nice and easy function developed by fastai","metadata":{}},{"cell_type":"markdown","source":"https://fastai1.fast.ai/callbacks.lr_finder.html","metadata":{}},{"cell_type":"markdown","source":"https://sgugger.github.io/how-do-you-find-a-good-learning-rate.html","metadata":{}},{"cell_type":"code","source":"learn.lr_find(suggest_funcs=(valley, slide))","metadata":{"execution":{"iopub.status.busy":"2023-01-19T13:57:54.044942Z","iopub.execute_input":"2023-01-19T13:57:54.045442Z","iopub.status.idle":"2023-01-19T13:59:30.25228Z","shell.execute_reply.started":"2023-01-19T13:57:54.045399Z","shell.execute_reply":"2023-01-19T13:59:30.246878Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Training our learner","metadata":{}},{"cell_type":"markdown","source":"https://docs.fast.ai/callback.schedule.html","metadata":{}},{"cell_type":"code","source":"learn.fine_tune(10, 0.012)#3","metadata":{"execution":{"iopub.status.busy":"2023-01-19T14:12:44.09853Z","iopub.execute_input":"2023-01-19T14:12:44.099025Z","iopub.status.idle":"2023-01-19T14:25:14.594758Z","shell.execute_reply.started":"2023-01-19T14:12:44.098975Z","shell.execute_reply":"2023-01-19T14:25:14.593555Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Saving our learner to make offline submission","metadata":{}},{"cell_type":"code","source":"learn.export()","metadata":{"execution":{"iopub.status.busy":"2023-01-19T14:26:16.348284Z","iopub.execute_input":"2023-01-19T14:26:16.348732Z","iopub.status.idle":"2023-01-19T14:26:16.707065Z","shell.execute_reply.started":"2023-01-19T14:26:16.348694Z","shell.execute_reply":"2023-01-19T14:26:16.705924Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- with learn.export() a pickle file is generated and we need to load it during inference","metadata":{}},{"cell_type":"code","source":"ss = pd.read_csv(f'{actual_path}/sample_submission.csv')\ndf_tst = pd.read_csv(f'{actual_path}/test.csv')\ndf_tst.head()","metadata":{"execution":{"iopub.status.busy":"2023-01-19T14:26:39.437951Z","iopub.execute_input":"2023-01-19T14:26:39.438404Z","iopub.status.idle":"2023-01-19T14:26:39.528564Z","shell.execute_reply.started":"2023-01-19T14:26:39.438366Z","shell.execute_reply":"2023-01-19T14:26:39.527383Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- Now we can do first experimentation, because each epoch requires approximate 1 minute","metadata":{}},{"cell_type":"markdown","source":"# Just seeing test images and reading submission files","metadata":{}},{"cell_type":"code","source":"ss.head()","metadata":{"execution":{"iopub.status.busy":"2023-01-19T14:27:45.01484Z","iopub.execute_input":"2023-01-19T14:27:45.015369Z","iopub.status.idle":"2023-01-19T14:27:45.091217Z","shell.execute_reply.started":"2023-01-19T14:27:45.015327Z","shell.execute_reply":"2023-01-19T14:27:45.089896Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ls ../input/rsna-breast-cancer-detection/test_images/10008","metadata":{"execution":{"iopub.status.busy":"2023-01-19T14:28:45.98427Z","iopub.execute_input":"2023-01-19T14:28:45.984748Z","iopub.status.idle":"2023-01-19T14:28:47.237601Z","shell.execute_reply.started":"2023-01-19T14:28:45.984701Z","shell.execute_reply":"2023-01-19T14:28:47.235936Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_tst.shape[0]","metadata":{"execution":{"iopub.status.busy":"2023-01-19T14:29:31.21868Z","iopub.execute_input":"2023-01-19T14:29:31.219166Z","iopub.status.idle":"2023-01-19T14:29:31.292837Z","shell.execute_reply.started":"2023-01-19T14:29:31.219125Z","shell.execute_reply":"2023-01-19T14:29:31.291821Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission = pd.DataFrame(\n                         data={\n                             'prediction_id': df_tst['prediction_id'], \n                             'cancer': np.random.rand(df_tst.shape[0])}\n                        ).drop_duplicates(subset='prediction_id')\nsubmission.head()","metadata":{"execution":{"iopub.status.busy":"2023-01-19T14:29:05.734229Z","iopub.execute_input":"2023-01-19T14:29:05.734765Z","iopub.status.idle":"2023-01-19T14:29:05.838824Z","shell.execute_reply.started":"2023-01-19T14:29:05.734714Z","shell.execute_reply":"2023-01-19T14:29:05.837893Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission.to_csv('submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2023-01-19T14:30:34.831286Z","iopub.execute_input":"2023-01-19T14:30:34.831749Z","iopub.status.idle":"2023-01-19T14:30:34.900386Z","shell.execute_reply.started":"2023-01-19T14:30:34.831711Z","shell.execute_reply":"2023-01-19T14:30:34.899179Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"learn.save(r'/kaggle/working/learner_first')","metadata":{"execution":{"iopub.status.busy":"2023-01-19T14:31:15.87229Z","iopub.execute_input":"2023-01-19T14:31:15.872734Z","iopub.status.idle":"2023-01-19T14:31:17.169553Z","shell.execute_reply.started":"2023-01-19T14:31:15.872695Z","shell.execute_reply":"2023-01-19T14:31:17.168533Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- One problem will be during submission, is that we need offline inference.\n- So we need to save timm library.\n- I have used fastkaggle library to save the timm library. So we use first with internet connection pip install fastkaggle library, then save the timm libray\n- [Here](https://www.kaggle.com/code/hasangoni/timm-library-for-offline-submission) one can find the timm library.\n- Then one just need to add those data in the offline notebook and use pip install.\n- For offline submission, I have published [here](https://www.kaggle.com/code/hasangoni/offline-inference/edit/run/115595817), I am getting ```Submission Scoring Error``. Don't know what is happening here. -> solution is actually to process the test images from dicom files to files for the learner. Right now the offline submission works. Thanks to @radek1 for his useful tips.\n","metadata":{"execution":{"iopub.status.busy":"2023-01-05T19:55:12.63809Z","iopub.execute_input":"2023-01-05T19:55:12.638596Z","iopub.status.idle":"2023-01-05T19:55:12.659331Z","shell.execute_reply.started":"2023-01-05T19:55:12.638553Z","shell.execute_reply":"2023-01-05T19:55:12.656188Z"}}},{"cell_type":"markdown","source":"- IF this notebook is helpful, please try to upvote it. This will help me to keep me motivated and publish more notebook like this.","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}