{"cells":[{"metadata":{},"cell_type":"markdown","source":"**1. Loading and visualizing the data**"},{"metadata":{"_uuid":"051d70d956493feee0c6d64651c6a088724dca2a","_execution_state":"idle","trusted":true},"cell_type":"code","source":"Sys.setenv(RETICULATE_PYTHON=\"/usr/local/share/.virtualenvs/r-reticulate/bin/\")\nlibrary(tensorflow) \nlibrary(tidyverse)\nlibrary(keras)\nlibrary(scales)\ntrain <- read.csv(\"../input/rsna-intracranial-hemorrhage-detection/rsna-intracranial-hemorrhage-detection/stage_2_train.csv\",\n                  stringsAsFactors=FALSE)\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"library(oro.dicom)\ndcmImage <- readDICOMFile(\"../input/rsna-intracranial-hemorrhage-detection/rsna-intracranial-hemorrhage-detection/stage_2_train/ID_000012eaf.dcm\")\ndim(dcmImage[[2]])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"image(t(dcmImage$img), col=grey(0:64/64), axes=FALSE, xlab=\"\", ylab=\"\",\n               main=\"Example image\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train <- train %>% distinct() # remove identical rows\ntrain<-train %>% separate(ID, sep = \"_\", remove = FALSE,into = c(NA,\"ID\",\"type\"))\ntrain$ID<-paste0(\"ID_\",train$ID)\nhead(train)\nbarplot(table(train  %>%  filter(type==\"any\") %>% select(Label)), main = \"distribution of 'any' variable\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"positive_cases <- train  %>%  filter(Label==1) \ntable(positive_cases$type)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"barplot(table(positive_cases$type),las=2)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_spread <- spread(train, key=type , value = Label)\nhead(train_spread)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**2. Writing a data generator function**\n\nSince the given training data set is huge, it is advantageous to load small batches of data from disk into memory and then feed them to a neural network.\nThe generator function randomly picks `batch_size` number of files from disk. Randomly choosing the files (instead of sequentially) allows us to\nrepeat the following cycle:\n\ntrain the model -> save the model/weights -> load the model -> resume training \n\nwithout having to keep track where we left off."},{"metadata":{"trusted":true},"cell_type":"code","source":"files <- list.files(\"../input/rsna-intracranial-hemorrhage-detection/rsna-intracranial-hemorrhage-detection/stage_2_train\")\n#files <- files[-264486] # remove corrupted file\n\npositive_ids <- train_spread %>% dplyr::filter(any==1) %>% select(ID)\npositive_ids <- as.character(positive_ids[,1])\nnegative_ids <- train_spread %>% dplyr::filter(any==0) %>% select(ID)\nnegative_ids <- as.character(negative_ids[,1])\n\n# use train set without \"any\" variable\ntrain_spread <- train_spread %>% dplyr::select(-any)\n\n# create train and validation set\nset.seed(246)\nindex <- sample(1:length(files),floor(0.8*length(files)),replace=FALSE) \ntrain_files <- files[index] \n\npos_files <- intersect(files, paste0(positive_ids,\".dcm\"))\nneg_files <- intersect(files, paste0(negative_ids,\".dcm\"))\n\nval_files <- files[-index]\n \n# generates random training batches from disk,\n# pos_rate specifies if batch is drawn randomly or \n# with custom rate of positive cases\ngenerator <- function(batch_size, file, pos_rate = NULL){\n \n function(){   \n   samples <- array(0, dim = c(batch_size, 512, 512, 1))\n   targets <- array(0, dim = c(batch_size, 5))\n  \n \n     if (is.null(pos_rate)) {\n     file_sample <- sample(file, batch_size, replace=FALSE) \n     } else {\n     num_pos_samples <- ceiling(batch_size*pos_rate)\n     pos <- sample(intersect(file, pos_files), size = num_pos_samples)\n     neg <- sample(intersect(file, neg_files), size = batch_size - num_pos_samples)\n     file_sample <- sample(c(pos, neg)) \n     }  \n     \n     \n   for (i in 1:batch_size){\n     file_name <- file_sample[i]\n     id <- str_remove(file_name,\".dcm\") \n     file_path <- paste0(\"../input/rsna-intracranial-hemorrhage-detection/rsna-intracranial-hemorrhage-detection/stage_2_train/\", file_name)  \n     pixel_matrix <- readDICOMFile(file_path)[[2]] \n     scaled_pixel_matrix <- rescale(matrix(pixel_matrix,ncol=1),range=c(-1,1))\n     samples[i,,,] <- array(scaled_pixel_matrix, dim=c(512,512,1)) \n     target_row <- train_spread %>% dplyr::filter(ID==id)\n     row <- as.numeric(target_row[1,2:6])  \n     targets[i,] <- as.array(row)  \n     }\n  list(samples, targets)\n }     \n}\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**3. Convolutional Neural Network**\n\nLooking at the distribution of the \"any\" variable, we can notice that most of the samples do not contain a hemorrhage. This should make it difficult for a CNN to\nlearn patterns in the images, since our network gets a lot of input where there is no pattern to pick up.   "},{"metadata":{"trusted":true},"cell_type":"code","source":"barplot(table(train  %>%  filter(type==\"any\") %>% select(Label)), main = \"distribution of 'any' variable\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"It seems like a reasonable idea to start feeding the network with images where mostly there is a hemorrhage, so the network is able to pick up patterns and \nthen feed more images without a hemorrhage. This can be done by adjusting the `pos_rate` variable in our generator function."},{"metadata":{"trusted":true},"cell_type":"code","source":"model <- keras_model_sequential() %>% \n  layer_conv_2d(filters = 32, kernel_size = c(3, 3), activation = \"relu\",\n                 input_shape = c(512, 512,1)) %>%  \n  layer_batch_normalization()  %>% \n  layer_max_pooling_2d(pool_size = c(2, 2)) %>% \n  layer_conv_2d(filters = 64, kernel_size = c(3, 3), activation = \"relu\") %>% \n  layer_batch_normalization()  %>% \n  layer_max_pooling_2d(pool_size = c(2, 2)) %>% \n  layer_conv_2d(filters = 128, kernel_size = c(3, 3), activation = \"relu\") %>% \n  layer_max_pooling_2d(pool_size = c(2, 2)) %>% \n  layer_conv_2d(filters = 128, kernel_size = c(3, 3), activation = \"relu\") %>% \n  layer_max_pooling_2d(pool_size = c(2, 2)) %>%\n  layer_batch_normalization()  %>% \n  layer_conv_2d(filters = 256, kernel_size = c(3, 3), activation = \"relu\") %>% \n  layer_batch_normalization()  %>% \n  layer_max_pooling_2d(pool_size = c(2, 2)) %>% \n  layer_flatten() %>% \n  layer_dense(units = 128, activation = \"relu\") %>% \n  layer_dense(units = 5, activation = \"sigmoid\")\n\n#model ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"model %>% compile(\n  loss = \"categorical_crossentropy\",\n  optimizer = optimizer_rmsprop(lr = 1e-3)\n)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# batch_size should not be too big, otherwise the resulting tensors allocate too much RAM for\n# a kaggle kernel to handle; rather enlarge the speps_per_epoche variable in the fit_generator\n# function to feed the model more data\ntrain_generator <- generator(batch_size = 16, file = train_files, pos_rate = 0.9) \nval_generator <- generator(batch_size = 16, file = val_files, pos_rate = 0.9)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"history <- model %>% fit_generator(\n   generator=train_generator,\n   steps_per_epoch = 100,\n   epochs = 10,\n   validation_data = val_generator,\n   validation_steps = 10\n )\n\nplot(history)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"I have not done any hyperparameter tuning yet, thus the performance is pretty bad. \nWith the current settings, i'm not able to run training on gpus (this was possible a few month ago with the same code)."},{"metadata":{"trusted":true},"cell_type":"code","source":"save_model_hdf5(model, \"rsna_model.h5\")","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"R","language":"R","name":"ir"},"language_info":{"mimetype":"text/x-r-source","name":"R","pygments_lexer":"r","version":"3.4.2","file_extension":".r","codemirror_mode":"r"}},"nbformat":4,"nbformat_minor":1}