{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"Resources used: \nhttps://www.kaggle.com/code/masatakaitakura/eda-for-beginner-rsna-mammography-breast-cancer\nhttps://www.kaggle.com/code/radek1/fast-ai-starter-pack-train-inference\nhttps://www.kaggle.com/code/asimple/pytorch-dataloader-pattern-rsna\n\nhttps://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/370587\n\n","metadata":{}},{"cell_type":"markdown","source":"# 1. Help your model see the data better\n\nA couple of ideas here. First of all, normalization. Should we normalize each image to between 0 and 1? That is not great for neural networks in general (they learn better on zero-centered data). Equally importantly, this way we lose the relative pixel intensities between images.\n\nMaybe one scan being lighter than the other is important?\n\nA related idea here is putting different window in each of the 3 channels (though I haven't had much luck with this approach in the past).","metadata":{}},{"cell_type":"code","source":"! pip install einops\n! pip install pylibjpeg\n! pip install python-gdcm","metadata":{"execution":{"iopub.status.busy":"2022-12-08T22:56:42.078113Z","iopub.execute_input":"2022-12-08T22:56:42.078998Z","iopub.status.idle":"2022-12-08T22:57:04.26791Z","shell.execute_reply.started":"2022-12-08T22:56:42.078927Z","shell.execute_reply":"2022-12-08T22:57:04.266731Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import os\nimport cv2\nimport torch\nimport pylibjpeg\nimport numpy as np\nimport pandas as pd\nimport seaborn as sns\nfrom glob import glob\nimport torch.nn as nn\nfrom tqdm import tqdm\nimport pydicom as dicom\nfrom pydicom import dcmread\nfrom einops import rearrange\nfrom torchvision import transforms\nfrom matplotlib import pyplot as plt\nfrom mpl_toolkits.axes_grid1 import ImageGrid\nfrom torch.utils.data import Dataset, DataLoader","metadata":{"execution":{"iopub.status.busy":"2022-12-08T22:57:04.270286Z","iopub.execute_input":"2022-12-08T22:57:04.270983Z","iopub.status.idle":"2022-12-08T22:57:06.897492Z","shell.execute_reply.started":"2022-12-08T22:57:04.270941Z","shell.execute_reply":"2022-12-08T22:57:06.896483Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 2. Training with bigger batches\n\nThe bigger the batch, the higher the chance the loss will be stable.","metadata":{}},{"cell_type":"markdown","source":"\n","metadata":{}},{"cell_type":"code","source":"class CFG:\n    class data:\n        fold=0\n        batch_size=32\n        image_size=(224, 224)\n        path_to_train=\"../input/split-folds-rsna/5_folds_data.csv\"\n        path_to_dcms=\"../input/rsna-breast-cancer-detection/train_images\"\n        path_to_train_images=\"/kaggle/input/rsna-mammography-images-as-pngs/images_as_pngs_512/train_images_processed_512/\"","metadata":{"execution":{"iopub.status.busy":"2022-12-08T22:57:06.899072Z","iopub.execute_input":"2022-12-08T22:57:06.89973Z","iopub.status.idle":"2022-12-08T22:57:06.906032Z","shell.execute_reply.started":"2022-12-08T22:57:06.89969Z","shell.execute_reply":"2022-12-08T22:57:06.904978Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"code","source":"train_df = pd.read_csv(CFG.data.path_to_train)\ntrain_df['img_name'] =train_df['patient_id'].astype(str) + \"/\" + train_df['image_id'].astype(str) + \".png\"\nprint(f\"train.shape = {train_df.shape}\")\n\ntrain_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-12-08T22:57:06.90848Z","iopub.execute_input":"2022-12-08T22:57:06.90909Z","iopub.status.idle":"2022-12-08T22:57:07.112013Z","shell.execute_reply.started":"2022-12-08T22:57:06.909053Z","shell.execute_reply":"2022-12-08T22:57:07.111032Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"code","source":"train = train_df.query(f'fold != {CFG.data.fold}').reset_index(drop=True)\nvalid = train_df.query(f'fold == {CFG.data.fold}').reset_index(drop=True)\n\ntrain.shape, valid.shape","metadata":{"execution":{"iopub.status.busy":"2022-12-08T22:57:07.113404Z","iopub.execute_input":"2022-12-08T22:57:07.114245Z","iopub.status.idle":"2022-12-08T22:57:07.150746Z","shell.execute_reply.started":"2022-12-08T22:57:07.114208Z","shell.execute_reply":"2022-12-08T22:57:07.149213Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"assert not round(233 / 10746 - 925 / 42802)\n\ntrain.cancer.value_counts(), valid.cancer.value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-12-08T22:57:07.152682Z","iopub.execute_input":"2022-12-08T22:57:07.153441Z","iopub.status.idle":"2022-12-08T22:57:07.171166Z","shell.execute_reply.started":"2022-12-08T22:57:07.153401Z","shell.execute_reply":"2022-12-08T22:57:07.170044Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"code","source":"transform = transforms.Compose([\n    transforms.ToTensor()\n])","metadata":{"execution":{"iopub.status.busy":"2022-12-08T22:57:07.17524Z","iopub.execute_input":"2022-12-08T22:57:07.176043Z","iopub.status.idle":"2022-12-08T22:57:07.185457Z","shell.execute_reply.started":"2022-12-08T22:57:07.175998Z","shell.execute_reply":"2022-12-08T22:57:07.182013Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 3. Training with smaller LR\n\nSimilar argument to the above -- if any one step that we make is less likely to take us too far in the wrong direction, our training might be more stable.","metadata":{}},{"cell_type":"code","source":"class RSNAData(Dataset):\n    def __init__(self, df, img_folder, image_size, transform=None, is_test=False):\n        self.df = df\n        self.is_test = is_test\n        self.transform = transform\n        self.image_size = image_size\n        self.img_folder = img_folder\n        \n    def __getitem__(self, idx):\n        img_path = os.path.join(self.img_folder, self.df['img_name'][idx])\n        img = cv2.imread(img_path)\n        img = cv2.resize(img, self.image_size)\n        if self.transform:\n            img = self.transform(image=img)['image']\n        img = torch.tensor(img, dtype=torch.float)\n        \n        # Rearrange the image dimensions so that channels are first in format\n        # This is because VIT Model requires Channels (c) to come first\n        \n        img = rearrange(img, 'h w c -> c h w')\n        \n        if not self.is_test:\n            target = self.df['cancer'][idx]\n            target = torch.tensor(target, dtype=torch.float)\n            return {\n                \"X\": img,\n                \"y\": target,\n            }\n        return {\"X\": img,}\n    \n    def __len__(self):\n        return len(self.df)","metadata":{"execution":{"iopub.status.busy":"2022-12-08T22:57:20.252536Z","iopub.execute_input":"2022-12-08T22:57:20.252929Z","iopub.status.idle":"2022-12-08T22:57:20.263253Z","shell.execute_reply.started":"2022-12-08T22:57:20.252894Z","shell.execute_reply":"2022-12-08T22:57:20.262194Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_dataset = RSNAData(df=train, img_folder=CFG.data.path_to_train_images, image_size=CFG.data.image_size)\nvalid_dataset = RSNAData(df=valid, img_folder=CFG.data.path_to_train_images, image_size=CFG.data.image_size)\n\ntrain_loader = DataLoader(train_dataset, batch_size=CFG.data.batch_size, shuffle=True)\nvalid_loader = DataLoader(valid_dataset, batch_size=CFG.data.batch_size, shuffle=False)","metadata":{"execution":{"iopub.status.busy":"2022-12-08T22:57:21.971281Z","iopub.execute_input":"2022-12-08T22:57:21.971653Z","iopub.status.idle":"2022-12-08T22:57:21.977898Z","shell.execute_reply.started":"2022-12-08T22:57:21.97162Z","shell.execute_reply":"2022-12-08T22:57:21.976833Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"markdown","source":"# 4. Cleaning the data\n\nRemoving images that are malformed or mislabeled can go a very long way to stabilizing (and improving) results.","metadata":{}},{"cell_type":"code","source":"batch_sample_images = next(iter(train_loader))[\"X\"]\nbatch_sample_images.size()","metadata":{"execution":{"iopub.status.busy":"2022-12-08T23:01:05.745122Z","iopub.execute_input":"2022-12-08T23:01:05.745483Z","iopub.status.idle":"2022-12-08T23:01:06.152876Z","shell.execute_reply.started":"2022-12-08T23:01:05.745452Z","shell.execute_reply":"2022-12-08T23:01:06.151833Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = plt.figure(figsize=(20, 15))\ngrid = ImageGrid(fig, 111,\n                 nrows_ncols=(8, 4),\n                 axes_pad=0.25\n)\n\nfor ax, img in zip(grid, batch_sample_images):\n    ax.imshow(img.permute(1, 2, 0))\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-12-08T23:05:55.729095Z","iopub.execute_input":"2022-12-08T23:05:55.729458Z","iopub.status.idle":"2022-12-08T23:05:59.56787Z","shell.execute_reply.started":"2022-12-08T23:05:55.729425Z","shell.execute_reply":"2022-12-08T23:05:59.566847Z"},"trusted":true},"execution_count":null,"outputs":[]}]}