{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Introduction\n\nThis notebook was developed for analyzing the main data delivered on the RSNA competition https://www.kaggle.com/competitions/rsna-breast-cancer-detection, focused on the main CSV table and the images.","metadata":{}},{"cell_type":"code","source":"","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data overview\n\nThe description of the data is shown on the following list, where is specified that some of them are exclusive for the training set and also they give an introduction to the data form:\n\n- site_id - ID code for the source hospital.\n- patient_id - ID code for the patient.\n- image_id - ID code for the image.\n- laterality - Whether the image is of the left or right breast.\n- view - The orientation of the image. The default for a screening exam is to capture two views per breast.\n- age - The patient's age in years.\n- implant - Whether or not the patient had breast implants. Site 1 only provides breast implant information at the patient level, not at the breast level.\n- density - A rating for how dense the breast tissue is, with A being the least dense and D being the most dense. Extremely dense tissue can make diagnosis more difficult. **Only provided for train**.\n- machine_id - An ID code for the imaging device.\n- cancer - Whether or not the breast was positive for malignant cancer. The target value. **Only provided for train**.\n- biopsy - Whether or not a follow-up biopsy was performed on the breast. **Only provided for train**.\n- invasive - If the breast is positive for cancer, whether or not the cancer proved to be invasive. **Only provided for train**.\n- BIRADS - 0 if the breast required follow-up, 1 if the breast was rated as negative for cancer, and 2 if the breast was rated as normal. Only provided for train.\n- prediction_id - The ID for the matching submission row. Multiple images will share the same prediction ID. **Test only**.\n- difficult_negative_case - True if the case was unusually difficult. **Only provided for train.**\n\nFrom this description, the main EDA was built as a reference for the future Cancer Detection.","metadata":{}},{"cell_type":"markdown","source":"First, the main libraries are loaded, focused on plotting and statistical analysis.","metadata":{}},{"cell_type":"code","source":"# !pip install pylibjpeg pylibjpeg-libjpeg pylibjpeg-openjpeg ","metadata":{"execution":{"iopub.status.busy":"2023-02-15T10:21:24.938215Z","iopub.execute_input":"2023-02-15T10:21:24.938828Z","iopub.status.idle":"2023-02-15T10:21:24.96566Z","shell.execute_reply.started":"2023-02-15T10:21:24.938689Z","shell.execute_reply":"2023-02-15T10:21:24.964533Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing\nimport seaborn as sns #plotting\nimport matplotlib.pyplot as plt #plotting\nimport os\n\nfrom pathlib import Path\nimport cv2\nfrom joblib import Parallel, delayed\n\nimport torch\nimport torch.nn as nn\nfrom torch.utils.data import Dataset, DataLoader\nfrom sklearn.model_selection import train_test_split\nimport torch.optim as optim\nfrom torchvision import models\n\nfrom PIL import Image\nfrom torchvision import transforms\n\n# import pylibjpeg\nimport pydicom\nimport pydicom.data\n\ntry:\n    import dicomsdl\nexcept:\n    !pip install /kaggle/input/rsna-dicomsd1-lib/dicomsdl-0.109.1-cp37-cp37m-manylinux_2_12_x86_64.manylinux2010_x86_64.whl\n    import dicomsdl\n","metadata":{"_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","execution":{"iopub.status.busy":"2023-02-15T10:21:24.967248Z","iopub.execute_input":"2023-02-15T10:21:24.967584Z","iopub.status.idle":"2023-02-15T10:21:58.750632Z","shell.execute_reply.started":"2023-02-15T10:21:24.967547Z","shell.execute_reply":"2023-02-15T10:21:58.749332Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Later, the train table was saved on a pandas DataFrame. At first glace, it shows that each patient has more than two images and that there are some missing elements on the **BIRADS** and **density** columns.","metadata":{}},{"cell_type":"code","source":"sample_table=pd.read_csv(\"/kaggle/input/rsna-breast-cancer-detection/train.csv\")\nsample_table.head()","metadata":{"execution":{"iopub.status.busy":"2023-02-15T10:21:58.752932Z","iopub.execute_input":"2023-02-15T10:21:58.75429Z","iopub.status.idle":"2023-02-15T10:21:58.885505Z","shell.execute_reply.started":"2023-02-15T10:21:58.75425Z","shell.execute_reply":"2023-02-15T10:21:58.884405Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_table.cancer.where(sample_table.image_id==1459541791).dropna()\n","metadata":{"execution":{"iopub.status.busy":"2023-02-15T10:21:58.888685Z","iopub.execute_input":"2023-02-15T10:21:58.889087Z","iopub.status.idle":"2023-02-15T10:21:58.90807Z","shell.execute_reply.started":"2023-02-15T10:21:58.889046Z","shell.execute_reply":"2023-02-15T10:21:58.907009Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\n# dicom = pydicom.dcmread('test_images/10008/1591370361.dcm')\n# img = dicom.pixel_array\n# img = (img - img.min()) / (img.max() - img.min())\n# if dicom.PhotometricInterpretation == \"MONOCHROME1\":\n#     img = 1 - img\n# # # plt.imshow(img, cmap=plt.cm.bone)  # set the color map to bone\n# # # plt.show()","metadata":{"execution":{"iopub.status.busy":"2023-02-15T10:21:58.909241Z","iopub.execute_input":"2023-02-15T10:21:58.909508Z","iopub.status.idle":"2023-02-15T10:21:58.914717Z","shell.execute_reply.started":"2023-02-15T10:21:58.909482Z","shell.execute_reply":"2023-02-15T10:21:58.913585Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# разные режимы датасета \nDATA_MODES = ['train', 'val', 'test']\n# все изображения будут масштабированы к размеру 224x224 px\nRESCALE_SIZE = 224\n# работаем на видеокарте\nDEVICE = torch.device(\"cuda\")\n\n\nclass RSNADataset(Dataset):\n    \"\"\"\n    Датасет с картинками, который паралельно подгружает их из папок\n    производит скалирование и превращение в торчевые тензоры\n    \"\"\"\n    def __init__(self, files, mode):\n        super().__init__()\n        # список файлов для загрузки\n        self.files = sorted(files)\n        # режим работы\n        self.mode = mode\n\n        if self.mode not in DATA_MODES:\n            print(f\"{self.mode} is not correct; correct modes: {DATA_MODES}\")\n            raise NameError\n\n        self.len_ = len(self.files)\n     \n#         self.label_encoder = LabelEncoder()\n#         if self.mode != 'test':\n#             self.labels = [int(pydicom.dcmread(path).InstanceNumber) for path in self.files]\n#             self.label_encoder.fit(self.labels)\n\n                      \n    def __len__(self):\n        return self.len_\n      \n    def load_sample(self, file):\n        image = Image.open(file)\n\n        return image\n  \n    def __getitem__(self, index):\n        # для преобразования изображений в тензоры PyTorch и нормализации входа\n        if self.mode == \"train\":  #добавлено для оптимизации  # transforms.RandomRotation(degrees=(0, 180)), \n#             transforms.RandomHorizontalFlip(p=0.5),\n            augmentations = transforms.Compose([\n                                   transforms.RandomHorizontalFlip(p=0.5),\n                               ])\n#             augmentations = transforms.RandomChoice([\n#                                    transforms.Compose([\n#                                        transforms.Resize(size=313, max_size=333),\n#                                        transforms.CenterCrop(size=300),\n#                                        transforms.RandomCrop(250)\n#                                        ]),\n#                                    transforms.RandomRotation(degrees=(-33, 33)),\n#                                    transforms.RandomHorizontalFlip(p=1),\n#                                    ])\n            transform = transforms.Compose([\n#                                     augmentations,\n                                    transforms.Resize(size=(RESCALE_SIZE, RESCALE_SIZE)),\n                                    transforms.ToTensor(),\n                                    transforms.Normalize([0.485], [0.229])\n                                    ]) \n                \n        else:\n            transform = transforms.Compose([\n                        transforms.Resize(size=(RESCALE_SIZE, RESCALE_SIZE)),\n                        transforms.ToTensor(),\n                        transforms.Normalize([0.485, 0.456, 0.406], [0.229, 0.224, 0.225]) \n                    ])\n#         #transform = transforms.Compose([  #добавлено для оптимизации\n#         #    transforms.ToTensor(),  #добавлено для оптимизации\n#         #    transforms.Normalize([0.485, 0.456, 0.406], [0.229, 0.224, 0.225])  #добавлено для оптимизации\n#         #])  #добавлено для оптимизации\n        im = self.load_sample(self.files[index])\n\n        # img = im.pixel_array\n        # img = (img - img.min()) / (img.max() - img.min())\n        # if im.PhotometricInterpretation == \"MONOCHROME1\":\n        #     img = 1 - img\n        x =transform(im)\n        x = x.expand(3,*x.shape[1:])\n\n        if self.mode == 'test':\n            return x\n        else:\n            im_id = int(self.files[index].split('.')[0].split('_')[-1])\n            y_array = sample_table.cancer.where(sample_table.image_id==im_id).dropna().values\n            y = torch.tensor(y_array)\n            return x, y\n        \n    def _prepare_sample(self, image):\n        image = image.resize((RESCALE_SIZE, RESCALE_SIZE))\n        return np.array(image)","metadata":{"execution":{"iopub.status.busy":"2023-02-15T10:21:58.916408Z","iopub.execute_input":"2023-02-15T10:21:58.917098Z","iopub.status.idle":"2023-02-15T10:21:58.933172Z","shell.execute_reply.started":"2023-02-15T10:21:58.917062Z","shell.execute_reply":"2023-02-15T10:21:58.932072Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# #TRAIN_DIR = Path('train/dataset')\n# #TRAIN_DIR = Path('train/simpsons_dataset') #for colab\n# TRAIN_DIR = Path('/kaggle/input/rsna-png-dataset-512512/output/train') #for colab\n# #TEST_DIR = Path('test/testset')\n# #TEST_DIR = Path('testset/testset')  #for colab\n# # TEST_DIR = Path('output/test')  #for colab\n\n# train_val_files = sorted(list(TRAIN_DIR.rglob('*.png')))\n# # test_files = sorted(list(TEST_DIR.rglob('*.png')))\n\n# train_paths=[]\n# for path in range(0,len(train_val_files)):\n#     train_paths.append(train_val_files[path].__str__())\n# # test_paths=[]\n# # for path in range(0,len(test_files)):\n# #     test_paths.append(test_files[path].__str__())","metadata":{"execution":{"iopub.status.busy":"2023-02-15T10:21:58.934602Z","iopub.execute_input":"2023-02-15T10:21:58.935039Z","iopub.status.idle":"2023-02-15T10:21:58.944711Z","shell.execute_reply.started":"2023-02-15T10:21:58.935001Z","shell.execute_reply":"2023-02-15T10:21:58.943746Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# TRAIN_DIR\n# train_paths","metadata":{"execution":{"iopub.status.busy":"2023-02-15T10:21:58.94637Z","iopub.execute_input":"2023-02-15T10:21:58.946933Z","iopub.status.idle":"2023-02-15T10:21:58.957573Z","shell.execute_reply.started":"2023-02-15T10:21:58.946892Z","shell.execute_reply":"2023-02-15T10:21:58.956343Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# files_tr, files_val = train_test_split(train_paths, test_size=0.2)\n# data_tr=RSNADataset(files_tr,'train')\n# data_val=RSNADataset(files_val,'train')","metadata":{"execution":{"iopub.status.busy":"2023-02-15T10:21:58.960935Z","iopub.execute_input":"2023-02-15T10:21:58.961241Z","iopub.status.idle":"2023-02-15T10:21:58.966991Z","shell.execute_reply.started":"2023-02-15T10:21:58.961216Z","shell.execute_reply":"2023-02-15T10:21:58.965656Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"os.makedirs('output/train/', exist_ok=True)\nos.makedirs('output/test/', exist_ok=True)","metadata":{}},{"cell_type":"code","source":"# train_paths=[]\n# for path in range(0,len(train_val_files)):\n#     train_paths.append(train_val_files[path].__str__())\n# test_paths=[]\n# for path in range(0,len(test_files)):\n#     test_paths.append(test_files[path].__str__())\n\n# def convert(path, train=True, show=False, size=256):\n#     subfolder = 'train' if train else 'test'\n#     patient = path.split('\\\\')[-2]\n#     image_number = path.split('\\\\')[-1][:-4] # removes the .dcm\n    \n#     dicom = pydicom.dcmread(path)\n#     img = dicom.pixel_array\n    \n#     img = (img - img.min())/(img.max() - img.min()) # Normalisation\n    \n#     if dicom.PhotometricInterpretation == \"MONOCHROME1\":\n#         img = 1 - img # some images are inverted\n    \n#     img = cv2.resize(img, (size,size))\n#     img = (img * 255).astype(np.uint8)\n    \n# #     if show:\n# #         plt.figure(figsize=(5,5))\n# #         plt.imshow(img, cmap='gray')\n# #         plt.title(f'Patient {patient}, Image {image_number}')\n# #         plt.show()\n        \n#     cv2.imwrite(f'output/{subfolder}/{patient}_{image_number}.png', img)\n    \n#     return img\n\n# os.makedirs('output/train/', exist_ok=True)\n# os.makedirs('output/test/', exist_ok=True)\n\n# import glob\n# from tqdm import tqdm\n\n# _ = Parallel(n_jobs=10)(\n#     delayed(convert)(x, train=True)\n#     for x in tqdm(train_paths)\n# )\n# _ = Parallel(n_jobs=10)(\n#     delayed(convert)(x, train=False)\n#     for x in tqdm(test_paths)\n# )","metadata":{"execution":{"iopub.status.busy":"2023-02-15T10:21:58.971908Z","iopub.execute_input":"2023-02-15T10:21:58.972228Z","iopub.status.idle":"2023-02-15T10:21:58.978072Z","shell.execute_reply.started":"2023-02-15T10:21:58.972181Z","shell.execute_reply":"2023-02-15T10:21:58.977003Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# data_tr.__getitem__(5000)[0]","metadata":{"execution":{"iopub.status.busy":"2023-02-15T10:21:58.979676Z","iopub.execute_input":"2023-02-15T10:21:58.980326Z","iopub.status.idle":"2023-02-15T10:21:58.988847Z","shell.execute_reply.started":"2023-02-15T10:21:58.980292Z","shell.execute_reply":"2023-02-15T10:21:58.988088Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class SegNet(nn.Module):\n    def __init__(self):\n        super().__init__()\n\n        # encoder (downsampling)\n        # Each enc_conv/dec_conv block should look like this:\n        # nn.Sequential(\n        #     nn.Conv2d(...),\n        #     ... (2 or 3 conv layers with relu and batchnorm),\n        # )\n        self.enc_conv0 = nn.Sequential(\n                            nn.Conv2d(1, 8, kernel_size=3,padding=1),\n                            nn.BatchNorm2d(8),\n                            nn.ReLU(),\n                            nn.Conv2d(8,8, kernel_size=3,padding=1),\n                            nn.BatchNorm2d(8),\n                            nn.ReLU()\n                        )\n        self.pool0 = nn.MaxPool2d(kernel_size=2, stride=1,return_indices=True)\n        self.enc_conv1 = nn.Sequential(\n                            nn.Conv2d(8,16, kernel_size=3,padding=1),\n                            nn.BatchNorm2d(16),\n                            nn.ReLU(),\n                            nn.Conv2d(16, 16, kernel_size=3,padding=1),\n                            nn.BatchNorm2d(16),\n                            nn.ReLU()\n                        )\n        self.pool1 =  nn.MaxPool2d(kernel_size=2, stride=1,return_indices=True)\n        self.enc_conv2 = nn.Sequential(\n                            nn.Conv2d(16,32, kernel_size=3,padding=1),\n                            nn.BatchNorm2d(32),\n                            nn.ReLU(),\n                            nn.Conv2d(32, 32, kernel_size=3,padding=1),\n                            nn.BatchNorm2d(32),\n                            nn.ReLU(),\n                            nn.Conv2d(32, 32, kernel_size=3,padding=1),\n                            nn.BatchNorm2d(32),\n                            nn.ReLU()\n                        )\n        self.pool2 =  nn.MaxPool2d(kernel_size=2, stride=1,return_indices=True)\n        self.enc_conv3 =nn.Sequential(\n                            nn.Conv2d(32,64, kernel_size=3,padding=1),\n                            nn.BatchNorm2d(64),\n                            nn.ReLU(),\n                            nn.Conv2d(64, 64, kernel_size=3,padding=1),\n                            nn.BatchNorm2d(64),\n                            nn.ReLU(),\n                            nn.Conv2d(64, 64, kernel_size=3,padding=1),\n                            nn.BatchNorm2d(64),\n                            nn.ReLU()\n                        )\n        self.pool3 =  nn.MaxPool2d(kernel_size=2, stride=1,return_indices=True)\n\n        # bottleneck\n        self.bottleneck_conv = nn.Sequential(\n                            nn.Conv2d(64,64, kernel_size=3,padding=1),\n                            nn.BatchNorm2d(64),\n                            nn.ReLU(),\n                            nn.Conv2d(64, 64, kernel_size=3,padding=1),\n                            nn.BatchNorm2d(64),\n                            nn.ReLU()\n                        )\n\n#         # decoder (upsampling)\n#         self.upsample0 =  nn.MaxUnpool2d( kernel_size=2, stride=1)\n\n#         self.dec_conv0 =nn.Sequential(\n#                             nn.ConvTranspose2d(64,64, kernel_size=3,padding=1),\n#                            nn.BatchNorm2d(64),\n#                             nn.ReLU(),\n#                             nn.ConvTranspose2d(64,64, kernel_size=3,padding=1),\n#                            nn.BatchNorm2d(64),\n#                             nn.ReLU()\n#                         )\n#         self.upsample1 =  nn.MaxUnpool2d(kernel_size=2, stride=1)\n#         self.dec_conv1 = nn.Sequential(\n#                             nn.ConvTranspose2d(64,64, kernel_size=3,padding=1),\n#                            nn.BatchNorm2d(64),\n#                             nn.ReLU(),\n#                             nn.ConvTranspose2d(64,32, kernel_size=3,padding=1),\n#                            nn.BatchNorm2d(32),\n#                             nn.ReLU()\n#                         )\n#         self.upsample2 =  nn.MaxUnpool2d(kernel_size=2, stride=1)\n#         self.dec_conv2 = nn.Sequential(\n#                             nn.ConvTranspose2d(32,32, kernel_size=3,padding=1),\n#                            nn.BatchNorm2d(32),\n#                             nn.ReLU(),\n#                             nn.ConvTranspose2d(32,16, kernel_size=3,padding=1),\n#                            nn.BatchNorm2d(16),\n#                             nn.ReLU()\n#                         )\n#         self.upsample3 =  nn.MaxUnpool2d(kernel_size=2, stride=1)\n#         self.dec_conv3 =nn.Sequential(\n#                             nn.ConvTranspose2d(16,16, kernel_size=3,padding=1),\n#                            nn.BatchNorm2d(16),\n#                             nn.ReLU(),\n#                             nn.ConvTranspose2d(16,8, kernel_size=3,padding=1),\n#                            nn.BatchNorm2d(8),\n#                             nn.ReLU(),\n#                         )\n        self.final=nn.Sequential(nn.Flatten(),\n                            nn.Linear(3097600, 1 ),\n                        )\n     \n\n    def forward(self, x):\n        # encoder\n        e0,indices0  =self.pool0(self.enc_conv0(x))\n        e1,indices1  =self.pool1(self.enc_conv1(e0)) \n        e2,indices2 =self.pool2(self.enc_conv2(e1))\n        e3,indices3  =self.pool3(self.enc_conv3(e2))\n\n        # bottleneck\n        b = self.bottleneck_conv(e3)\n\n        # decoder\n#         d0 = self.upsample0(self.dec_conv0(b),indices3)\n#         d1 = self.upsample1(self.dec_conv1(d0),indices2)\n#         d2 = self.upsample2(self.dec_conv2(d1),indices1)\n        d3 = self.final(b)\n\n        return torch.sigmoid(d3)","metadata":{"execution":{"iopub.status.busy":"2023-02-15T10:21:58.990394Z","iopub.execute_input":"2023-02-15T10:21:58.991099Z","iopub.status.idle":"2023-02-15T10:21:59.010259Z","shell.execute_reply.started":"2023-02-15T10:21:58.990996Z","shell.execute_reply":"2023-02-15T10:21:59.009334Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class CNN(nn.Module):\n    def __init__(self):\n        super(CNN, self).__init__()\n        self.network = models.resnet18(pretrained=False)\n        n_features = self.network.fc.out_features\n        print(n_features)\n        # add additional layer that maps 2048 extracted features from resnet to 1 feature determining the class\n        self.classifier_layer = nn.Sequential(\n            nn.Linear(n_features , 256),\n            nn.Dropout(0.3),\n            nn.Linear(256 , 1)\n        )\n    \n    def forward(self, xb):        \n        xb = self.network(xb)\n        xb = self.classifier_layer(xb)\n        return torch.sigmoid(xb)","metadata":{"execution":{"iopub.status.busy":"2023-02-15T10:21:59.011826Z","iopub.execute_input":"2023-02-15T10:21:59.012291Z","iopub.status.idle":"2023-02-15T10:21:59.022376Z","shell.execute_reply.started":"2023-02-15T10:21:59.012251Z","shell.execute_reply":"2023-02-15T10:21:59.021566Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from tqdm import tqdm","metadata":{"execution":{"iopub.status.busy":"2023-02-15T10:21:59.023657Z","iopub.execute_input":"2023-02-15T10:21:59.02447Z","iopub.status.idle":"2023-02-15T10:21:59.033314Z","shell.execute_reply.started":"2023-02-15T10:21:59.024435Z","shell.execute_reply":"2023-02-15T10:21:59.03247Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def train(model, opt, loss_fn, epochs, data_tr, data_val):\n    print('Started')\n    loss_history=[]\n    acc_history=[]\n    for epoch in range(epochs):\n        tic = time()\n        print('* Epoch %d/%d' % (epoch+1, epochs))\n        \n        avg_loss = 0\n        model.train()  # train mode\n        with tqdm(data_tr, unit=\"batch\") as tepoch:\n            for X_batch, Y_batch in tepoch:\n                # data to device\n                # print(X_batch)\n\n                X_batch=X_batch.type(torch.cuda.FloatTensor).to('cuda')\n                Y_batch=Y_batch.type(torch.cuda.FloatTensor).to('cuda')\n#                 print(Y_batch)\n                # set parameter gradients to zero\n                opt.zero_grad()\n                # forward\n                Y_pred =torch.tensor((model(X_batch)>0.5).double(),requires_grad=True)\n#                 print(Y_pred.grad_fn)\n                \n\n\n#                 print(Y_pred.size())\n\n\n                loss = loss_fn(Y_pred,Y_batch,weights)\n                loss.backward()\n#                 print(loss)\n                opt.step()  # update weights\n    \n                # calculate loss to show the user\n                avg_loss += loss / len(data_tr)\n            toc = time()\n            print('Train loss: %f' % avg_loss)\n\n            # show intermediate results\n        val_loss,val_acc = eval_epoch(model,data_val,loss_fn)\n        loss_history.append(val_loss)\n        acc_history.append(val_acc)\n        print(f'Valitadion Loss: {val_loss:.4f}, Validation Acc: {val_acc:.4f} ')\n    return loss_history,acc_history\n\n        \ndef eval_epoch(model, val_loader, criterion):\n    model.eval()\n    running_loss = 0.0\n    running_corrects = 0\n    processed_size = 0\n    with tqdm(val_loader, unit=\"batch\") as nepoch:\n        for inputs, labels in nepoch:\n            inputs = inputs.to(DEVICE)\n            labels = labels.to(DEVICE)\n\n            with torch.set_grad_enabled(False):\n                outputs =(model(inputs)>0.5).double()\n                loss = criterion(outputs, labels,weights)\n\n\n            running_loss += loss.item() \n            running_corrects += torch.sum(outputs == labels.data)\n            processed_size += inputs.size(0)\n    val_loss = running_loss / processed_size\n    val_acc = running_corrects.double() / processed_size\n    return val_loss, val_acc\n\n\n","metadata":{"execution":{"iopub.status.busy":"2023-02-15T10:21:59.034788Z","iopub.execute_input":"2023-02-15T10:21:59.03546Z","iopub.status.idle":"2023-02-15T10:21:59.048855Z","shell.execute_reply.started":"2023-02-15T10:21:59.035425Z","shell.execute_reply":"2023-02-15T10:21:59.047931Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# cross_entropy_loss = nn.CrossEntropyLoss()\n# from time import time","metadata":{"execution":{"iopub.status.busy":"2023-02-15T10:21:59.051642Z","iopub.execute_input":"2023-02-15T10:21:59.051962Z","iopub.status.idle":"2023-02-15T10:21:59.067569Z","shell.execute_reply.started":"2023-02-15T10:21:59.051914Z","shell.execute_reply":"2023-02-15T10:21:59.066657Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\n# weights=[i/len(sample_table) for i in sample_table.cancer.value_counts()]\n\n# sampler=torch.utils.data.WeightedRandomSampler(weights, num_samples=len(sample_table))","metadata":{"execution":{"iopub.status.busy":"2023-02-15T10:21:59.068981Z","iopub.execute_input":"2023-02-15T10:21:59.069676Z","iopub.status.idle":"2023-02-15T10:21:59.076711Z","shell.execute_reply.started":"2023-02-15T10:21:59.069636Z","shell.execute_reply":"2023-02-15T10:21:59.075817Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def BCELoss_class_weighted(y_pred,target,weights):\n    \"\"\"\n    weights[0] is weight for class 0 (negative class)\n    weights[1] is weight for class 1 (positive class)\n    \"\"\"\n\n    y_pred = torch.clamp(y_pred,min=1e-7,max=1-1e-7) # for numerical stability\n    bce = - weights[1] * target * torch.log(y_pred) - (1 - target) * weights[0] * torch.log(1 - y_pred)\n    return torch.mean(bce)","metadata":{"execution":{"iopub.status.busy":"2023-02-15T10:21:59.078095Z","iopub.execute_input":"2023-02-15T10:21:59.078612Z","iopub.status.idle":"2023-02-15T10:21:59.086314Z","shell.execute_reply.started":"2023-02-15T10:21:59.078577Z","shell.execute_reply":"2023-02-15T10:21:59.085038Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# train_loader = DataLoader(data_tr, batch_size=128, num_workers=2,sampler=sampler)\n# val_loader = DataLoader(data_val, batch_size=128, num_workers=2)\n","metadata":{"execution":{"iopub.status.busy":"2023-02-15T10:21:59.088005Z","iopub.execute_input":"2023-02-15T10:21:59.088909Z","iopub.status.idle":"2023-02-15T10:21:59.095153Z","shell.execute_reply.started":"2023-02-15T10:21:59.088873Z","shell.execute_reply":"2023-02-15T10:21:59.093834Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# torch.cuda.empty_cache()\n# model = CNN().type(torch.cuda.FloatTensor).to('cuda')\n# torch.cuda.memory_summary(device=None, abbreviated=False)","metadata":{"execution":{"iopub.status.busy":"2023-02-15T10:21:59.096741Z","iopub.execute_input":"2023-02-15T10:21:59.097818Z","iopub.status.idle":"2023-02-15T10:21:59.104055Z","shell.execute_reply.started":"2023-02-15T10:21:59.097779Z","shell.execute_reply":"2023-02-15T10:21:59.10293Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# max_epochs =1\n# opt= optim.Adam(model.parameters(),lr=0.00001)\n\n# loss_history,acc_history=train(model,opt , BCELoss_class_weighted, max_epochs, train_loader, val_loader)","metadata":{"execution":{"iopub.status.busy":"2023-02-15T10:21:59.105468Z","iopub.execute_input":"2023-02-15T10:21:59.106645Z","iopub.status.idle":"2023-02-15T10:21:59.113049Z","shell.execute_reply.started":"2023-02-15T10:21:59.106608Z","shell.execute_reply":"2023-02-15T10:21:59.111999Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def predict_convert(path, train=True, show=False, size=256):\n    \n    im=dicomsdl.open(test_paths[0])\n    img=im.pixelData()\n\n    img = (img - img.min())/(img.max() - img.min()) # Normalisation\n\n    if im.PhotometricInterpretation == \"MONOCHROME1\":\n        img = 1 - img # some images are inverted\n\n    img = cv2.resize(img, (size,size))\n    img = (img * 255).astype(np.uint8)\n\n#     if show:\n#         plt.figure(figsize=(5,5))\n#         plt.imshow(img, cmap='gray')\n#         plt.title(f'Patient {patient}, Image {image_number}')\n#         plt.show()\n\n#     cv2.imwrite(f'output/{subfolder}/{patient}_{image_number}.png', img)\n\n    return Image.fromarray(img)\n\nclass RSNADatasetTest(Dataset):\n    \"\"\"\n    Датасет с картинками, который паралельно подгружает их из папок\n    производит скалирование и превращение в торчевые тензоры\n    \"\"\"\n    def __init__(self, files, mode):\n        super().__init__()\n        # список файлов для загрузки\n        self.files = sorted(files)\n        # режим работы\n        self.mode = mode\n\n        if self.mode not in DATA_MODES:\n            print(f\"{self.mode} is not correct; correct modes: {DATA_MODES}\")\n            raise NameError\n\n        self.len_ = len(self.files)\n  \n                      \n    def __len__(self):\n        return self.len_\n      \n    def load_sample(self, file):\n        image = predict_convert(file)\n\n        return image\n  \n    def __getitem__(self, index):\n        if self.mode == \"train\":  \n            augmentations = transforms.Compose([\n                                   transforms.RandomHorizontalFlip(p=0.5),\n                               ])\n\n            transform = transforms.Compose([\n#                                     augmentations,\n                                    transforms.Resize(size=(RESCALE_SIZE, RESCALE_SIZE)),\n                                    transforms.ToTensor(),\n                                    transforms.Normalize([0.485], [0.229])\n                                    ]) \n                \n        else:\n            transform = transforms.Compose([\n                        transforms.Resize(size=(RESCALE_SIZE, RESCALE_SIZE)),\n                        transforms.ToTensor(),\n                        transforms.Normalize([0.485], [0.229]) \n                    ])\n\n        im = self.load_sample(self.files[index])\n        x =transform(im)\n        x = x.expand(3,*x.shape[1:])\n\n        im_id = self.files[index].split('.')[0].split('/')[-1]\n\n        return x, im_id\n        \n    def _prepare_sample(self, image):\n        image = image.resize((RESCALE_SIZE, RESCALE_SIZE))\n        return np.array(image)","metadata":{"execution":{"iopub.status.busy":"2023-02-15T10:21:59.11493Z","iopub.execute_input":"2023-02-15T10:21:59.115868Z","iopub.status.idle":"2023-02-15T10:21:59.131587Z","shell.execute_reply.started":"2023-02-15T10:21:59.115833Z","shell.execute_reply":"2023-02-15T10:21:59.130549Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"TEST_DIR = Path('/kaggle/input/rsna-breast-cancer-detection/test_images')  #for colab\n\ntest_files = sorted(list(TEST_DIR.rglob('*.dcm')))\n\ntest_paths=[]\nfor path in range(0,len(test_files)):\n    test_paths.append(test_files[path].__str__())","metadata":{"execution":{"iopub.status.busy":"2023-02-15T10:21:59.134612Z","iopub.execute_input":"2023-02-15T10:21:59.134963Z","iopub.status.idle":"2023-02-15T10:21:59.149323Z","shell.execute_reply.started":"2023-02-15T10:21:59.134929Z","shell.execute_reply":"2023-02-15T10:21:59.148278Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_paths","metadata":{"execution":{"iopub.status.busy":"2023-02-15T10:21:59.1513Z","iopub.execute_input":"2023-02-15T10:21:59.152194Z","iopub.status.idle":"2023-02-15T10:21:59.159314Z","shell.execute_reply.started":"2023-02-15T10:21:59.152154Z","shell.execute_reply":"2023-02-15T10:21:59.158093Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_test=RSNADatasetTest(test_paths,'test')\n\ndata_test.__getitem__(0)[0]","metadata":{"execution":{"iopub.status.busy":"2023-02-15T10:21:59.161173Z","iopub.execute_input":"2023-02-15T10:21:59.162194Z","iopub.status.idle":"2023-02-15T10:21:59.890382Z","shell.execute_reply.started":"2023-02-15T10:21:59.162152Z","shell.execute_reply":"2023-02-15T10:21:59.889399Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\ntestloader=DataLoader(data_test, batch_size=1,num_workers=2)\n\ndef predict(model,test_loader):\n    outputs={}\n    with tqdm(test_loader, unit=\"batch\") as nepoch:\n        for inputs, labels in nepoch:\n            inputs = inputs.to(DEVICE)\n\n            with torch.set_grad_enabled(False):\n                outputs[int(labels[0])] =model(inputs).to('cpu').numpy()[0][0]\n    return outputs\n                ","metadata":{"execution":{"iopub.status.busy":"2023-02-15T10:21:59.891868Z","iopub.execute_input":"2023-02-15T10:21:59.893017Z","iopub.status.idle":"2023-02-15T10:21:59.90044Z","shell.execute_reply.started":"2023-02-15T10:21:59.892979Z","shell.execute_reply":"2023-02-15T10:21:59.899411Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"torch.save(model.state_dict(), PATH)","metadata":{}},{"cell_type":"code","source":"# torch.save(model.state_dict(), 'output/model')\n\nmodel = CNN()\nmodel.load_state_dict(torch.load('/kaggle/input/base-resnet-model-for-rsna-compet/model'))\nmodel.type(torch.cuda.FloatTensor).to('cuda')","metadata":{"execution":{"iopub.status.busy":"2023-02-15T10:21:59.90349Z","iopub.execute_input":"2023-02-15T10:21:59.904763Z","iopub.status.idle":"2023-02-15T10:22:03.307642Z","shell.execute_reply.started":"2023-02-15T10:21:59.904705Z","shell.execute_reply":"2023-02-15T10:22:03.306595Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"preds=predict(model,testloader)\npreds","metadata":{"execution":{"iopub.status.busy":"2023-02-15T10:22:03.312666Z","iopub.execute_input":"2023-02-15T10:22:03.314958Z","iopub.status.idle":"2023-02-15T10:22:10.238595Z","shell.execute_reply.started":"2023-02-15T10:22:03.314924Z","shell.execute_reply":"2023-02-15T10:22:10.237465Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df=pd.read_csv('/kaggle/input/rsna-breast-cancer-detection/test.csv')\ntest_df['cancer']=0.1\ntest_df.index=test_df.image_id\ntest_df.loc[1591370361].cancer\nfor key in preds.keys():\n    test_df.cancer.loc[key]=preds[key]\n# test_df=test_df.iloc[:,-2:]\nsubmission=test_df.iloc[:,-2:].groupby('prediction_id').mean()\nsubmission.to_csv('submission.csv')\n","metadata":{"execution":{"iopub.status.busy":"2023-02-15T10:22:52.444256Z","iopub.execute_input":"2023-02-15T10:22:52.444643Z","iopub.status.idle":"2023-02-15T10:22:52.461439Z","shell.execute_reply.started":"2023-02-15T10:22:52.444609Z","shell.execute_reply":"2023-02-15T10:22:52.460443Z"},"trusted":true},"execution_count":null,"outputs":[]}]}