{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# RSNA - EfficientNet Baseline\n_____\n\n### EfficientNet\n\nEfficientNet is one of the most solid baselines known today. It is a family of convolutional neural networks that have achieved state-of-the-art accuracy on ImageNet while also being smaller and faster than other models.\nThe main idea of EfficientNet is scaling up CNNs in a principled way. It uses a scalable architecture, named compound scaling, which balances network depth, width, and resolution to achieve superior performance.\n\n\n#### Model Size (B5)\n\nAs written above, the main claim of efficientnet is providing a method for scaling up neural networks. \nUsing this approach, the authors of the paper released the official efficientnet architecture scaled to various sizes [B0, B1, B2.. B7. And two monstrosities L1 and L2].\nAs it is usually empirically the case (Also hinted by the [Scaling Laws for Neural Language Models](https://arxiv.org/abs/2001.08361)): Performance increase with model size. \nOn this notebook we use the **B5** variant of efficientnet. \nFeel free to experiment with larger sizes.","metadata":{"execution":{"iopub.status.busy":"2022-07-29T07:11:03.457277Z"},"papermill":{"duration":0.006207,"end_time":"2022-08-14T14:38:50.804121","exception":false,"start_time":"2022-08-14T14:38:50.797914","status":"completed"},"pycharm":{"name":"#%% md\n"},"tags":[]}},{"cell_type":"markdown","source":"**Installations**","metadata":{"papermill":{"duration":0.004763,"end_time":"2022-08-14T14:38:50.814116","exception":false,"start_time":"2022-08-14T14:38:50.809353","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# import os\n# import shutil\n\n# # Copy Efficientnet source code to a directory with write access\n# # shutil.copytree('../input/efficientnet-keras-source-code', '/kaggle/efficientnet-keras-source-code/')\n# # Pip install required packages\n# os.system(\"pip install -q ../input/keras-applications/Keras_Applications-1.0.8-py3-none-any.whl\")\n# os.system(\"pip install -q /kaggle/efficientnet-keras-source-code\")","metadata":{"execution":{"iopub.status.busy":"2022-10-02T08:08:25.61494Z","iopub.execute_input":"2022-10-02T08:08:25.615303Z","iopub.status.idle":"2022-10-02T08:08:25.61945Z","shell.execute_reply.started":"2022-10-02T08:08:25.615273Z","shell.execute_reply":"2022-10-02T08:08:25.618512Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import sys\nsys.path.append('../input/kerasapplications')\n# sys.path.append('../input/efficientnet-keras-source-code/')\nimport keras_applications\n# import efficientnet.tfkeras as efficientnet","metadata":{"execution":{"iopub.status.busy":"2022-10-02T08:09:56.687812Z","iopub.execute_input":"2022-10-02T08:09:56.688188Z","iopub.status.idle":"2022-10-02T08:09:56.693206Z","shell.execute_reply.started":"2022-10-02T08:09:56.688158Z","shell.execute_reply":"2022-10-02T08:09:56.692082Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Imports**","metadata":{"papermill":{"duration":0.007534,"end_time":"2022-08-14T14:38:57.389676","exception":false,"start_time":"2022-08-14T14:38:57.382142","status":"completed"},"tags":[]}},{"cell_type":"code","source":"import os\nimport cv2\nimport glob\nimport traceback\nimport cv2 as cv\nimport numpy as np\nimport pandas as pd\nfrom path import Path\nfrom tqdm import tqdm\nimport nibabel as nib\nimport pydicom as dicom\nimport tensorflow as tf\nfrom keras import layers\nfrom pydicom import dcmread\nfrom tensorflow import keras\nimport tensorflow_hub as hub\nimport matplotlib.pyplot as plt\nfrom tensorflow.keras import backend as K\nfrom pydicom.data import get_testdata_files\nfrom tensorflow.keras.utils import to_categorical\nfrom sklearn.model_selection import StratifiedKFold\nfrom pydicom.pixel_data_handlers.util import apply_voi_lut\nfrom tensorflow.keras.layers import Input, Dense, Flatten, Conv2D\nfrom tensorflow.keras.preprocessing.image import load_img, img_to_array\n\nimport tensorflow as tf\nimport keras_applications\n# import efficientnet.tfkeras  as efficientnet\n#import efficientnet.keras as efn \n# from keras_efficientnets import *\nfrom tensorflow.keras.applications import EfficientNetB5\nimport tensorflow as tf","metadata":{"_kg_hide-input":true,"papermill":{"duration":0.922527,"end_time":"2022-08-14T14:38:58.320365","exception":false,"start_time":"2022-08-14T14:38:57.397838","status":"completed"},"pycharm":{"name":"#%%\n"},"tags":[],"execution":{"iopub.status.busy":"2022-10-02T08:10:16.860492Z","iopub.execute_input":"2022-10-02T08:10:16.860872Z","iopub.status.idle":"2022-10-02T08:10:16.870172Z","shell.execute_reply.started":"2022-10-02T08:10:16.860841Z","shell.execute_reply":"2022-10-02T08:10:16.869163Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Data Loading\n","metadata":{"papermill":{"duration":0.004855,"end_time":"2022-08-14T14:38:58.331127","exception":false,"start_time":"2022-08-14T14:38:58.326272","status":"completed"},"tags":[]}},{"cell_type":"code","source":"bad = np.array([['1.2.826.0.1.3680043.10197_C1', '1.2.826.0.1.3680043.10197','C1'],['1.2.826.0.1.3680043.10454_C1', '1.2.826.0.1.3680043.10454','C1'],['1.2.826.0.1.3680043.10690_C1', '1.2.826.0.1.3680043.10690','C1']], dtype=np.object)\n\ntrain_df = pd.read_csv(\"../input/rsna-2022-cervical-spine-fracture-detection/train.csv\")\ntest_df = pd.read_csv(\"../input/rsna-2022-cervical-spine-fracture-detection/test.csv\")\n\ntrain_dir = '../input/rsna-2022-cervical-spine-fracture-detection/train_images'\ntest_dir = '../input/rsna-2022-cervical-spine-fracture-detection/test_images'\nfirst_image = os.path.join(test_dir, test_df['StudyInstanceUID'].iloc[0])\n\nnew_submission = []\nmeans = train_df.median(numeric_only=True).to_dict()\nmeans = dict(zip(train_df.columns[1:], np.average(train_df[train_df.columns[1:]], axis=0, weights=train_df[\"patient_overall\"] + 1)))\nprediction_type = test_df['prediction_type'].tolist()\nsubmission = pd.read_csv('../input/rsna-2022-cervical-spine-fracture-detection/sample_submission.csv')\nfor i in range(len(submission)):        \n    new_submission.append(means[prediction_type[i]])\nsubmission['fractured'] = new_submission\n\n\nif(test_df.values[0][0] == bad[0][0]): test_df = pd.DataFrame({\"row_id\": ['1.2.826.0.1.3680043.22327_C1', '1.2.826.0.1.3680043.25399_C1', '1.2.826.0.1.3680043.5876_C1'], \"StudyInstanceUID\": ['1.2.826.0.1.3680043.22327', '1.2.826.0.1.3680043.25399', '1.2.826.0.1.3680043.5876'], \"prediction_type\": [\"C1\", \"C1\", \"C1\"]})  \nprediction_type_mapping = test_df['prediction_type'].map({'C1': 0, 'C2': 1, 'C3': 2, 'C4': 3, 'C5': 4, 'C6': 5, 'C7': 6}).values","metadata":{"_kg_hide-output":true,"papermill":{"duration":0.054187,"end_time":"2022-08-14T14:38:58.390389","exception":false,"start_time":"2022-08-14T14:38:58.336202","status":"completed"},"pycharm":{"name":"#%%\n"},"tags":[],"execution":{"iopub.status.busy":"2022-10-02T08:10:22.138889Z","iopub.execute_input":"2022-10-02T08:10:22.139259Z","iopub.status.idle":"2022-10-02T08:10:22.168569Z","shell.execute_reply.started":"2022-10-02T08:10:22.139229Z","shell.execute_reply":"2022-10-02T08:10:22.167464Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Utility Functions & Variables\n\n- **load_dicom:** For loading (and converting) a single dicom image.\n- **lisdirs:** For listing the data directory.\n- **train_dir / test_dir:** used later on for fast switching between train & validation paths","metadata":{"papermill":{"duration":0.005006,"end_time":"2022-08-14T14:38:58.400738","exception":false,"start_time":"2022-08-14T14:38:58.395732","status":"completed"},"tags":[]}},{"cell_type":"code","source":"def load_dicom(path, size = 64):\n    try:\n        img=dicom.dcmread(path)\n        img.PhotometricInterpretation = 'YBR_FULL'\n        data=img.pixel_array\n        data=data-np.min(data)\n        if np.max(data) != 0:\n            data=data/np.max(data)\n        data=(data*255).astype(np.uint8)        \n        return cv2.cvtColor(data.reshape(512, 512), cv2.COLOR_GRAY2RGB)\n    except:        \n        return np.zeros((512, 512, 3))\n\ndef listdirs(folder):\n    return [d for d in os.listdir(folder) if os.path.isdir(os.path.join(folder, d))]    \n\ntrain_dir = '../input/rsna-2022-cervical-spine-fracture-detection/train_images'\ntest_dir = '../input/rsna-2022-cervical-spine-fracture-detection/test_images'\npatients = sorted(os.listdir(train_dir))","metadata":{"_kg_hide-output":true,"papermill":{"duration":0.128065,"end_time":"2022-08-14T14:38:58.533924","exception":false,"start_time":"2022-08-14T14:38:58.405859","status":"completed"},"pycharm":{"name":"#%%\n"},"tags":[],"execution":{"iopub.status.busy":"2022-10-02T08:10:31.540217Z","iopub.execute_input":"2022-10-02T08:10:31.540583Z","iopub.status.idle":"2022-10-02T08:10:31.550699Z","shell.execute_reply.started":"2022-10-02T08:10:31.540551Z","shell.execute_reply":"2022-10-02T08:10:31.549454Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Training Samples**","metadata":{"papermill":{"duration":0.005778,"end_time":"2022-08-14T14:38:58.555289","exception":false,"start_time":"2022-08-14T14:38:58.549511","status":"completed"},"tags":[]}},{"cell_type":"code","source":"image_file = glob.glob(\"../input/rsna-2022-cervical-spine-fracture-detection/train_images/1.2.826.0.1.3680043.10001/*.dcm\")\nplt.figure(figsize=(20, 20))\n\nfor i in range(28):\n    ax = plt.subplot(7, 7, i + 1)\n    image_path = image_file[i]\n    image = load_dicom(image_path)\n    plt.axis('off')   \n    plt.imshow(image)","metadata":{"papermill":{"duration":2.45368,"end_time":"2022-08-14T14:39:01.013957","exception":false,"start_time":"2022-08-14T14:38:58.560277","status":"completed"},"pycharm":{"name":"#%%\n"},"tags":[],"execution":{"iopub.status.busy":"2022-10-02T08:10:33.502868Z","iopub.execute_input":"2022-10-02T08:10:33.503876Z","iopub.status.idle":"2022-10-02T08:10:35.830458Z","shell.execute_reply.started":"2022-10-02T08:10:33.503839Z","shell.execute_reply":"2022-10-02T08:10:35.82963Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Targets**","metadata":{"papermill":{"duration":0.011376,"end_time":"2022-08-14T14:39:01.03665","exception":false,"start_time":"2022-08-14T14:39:01.025274","status":"completed"},"tags":[]}},{"cell_type":"code","source":"image_file = glob.glob(\"../input/rsna-2022-cervical-spine-fracture-detection/segmentations/*.nii\")\nplt.figure(figsize=(20, 20))\n\nfor i in range(28):\n    ax = plt.subplot(7, 7, i + 1)\n    image_path = image_file[i]\n    nii_img = nib.load(image_path).get_fdata()\n    nib_image = nii_img[:,:,59]\n    plt.axis('off')\n    plt.imshow(nib_image)","metadata":{"papermill":{"duration":31.046813,"end_time":"2022-08-14T14:39:32.095029","exception":false,"start_time":"2022-08-14T14:39:01.048216","status":"completed"},"pycharm":{"name":"#%%\n"},"tags":[],"execution":{"iopub.status.busy":"2022-10-02T08:10:35.832566Z","iopub.execute_input":"2022-10-02T08:10:35.833137Z","iopub.status.idle":"2022-10-02T08:10:43.36753Z","shell.execute_reply.started":"2022-10-02T08:10:35.833102Z","shell.execute_reply":"2022-10-02T08:10:43.366519Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Data Generators\n\nSince the training data can be a bit large, we load it batch by batch when training & inference.\nThis is done through simple keras generators. ","metadata":{"papermill":{"duration":0.012574,"end_time":"2022-08-14T14:39:32.120329","exception":false,"start_time":"2022-08-14T14:39:32.107755","status":"completed"},"pycharm":{"name":"#%% md\n"},"tags":[]}},{"cell_type":"markdown","source":"#### Train Generator \n\n- Also yields labels","metadata":{"papermill":{"duration":0.01221,"end_time":"2022-08-14T14:39:32.144944","exception":false,"start_time":"2022-08-14T14:39:32.132734","status":"completed"},"pycharm":{"name":"#%% md\n"},"tags":[]}},{"cell_type":"code","source":"def RSNATrainGenerator(train_df, batch_size, infinite = True, base_path = train_dir):\n    while True:\n        trainset = []\n        trainidt = []\n        trainlabel = []\n        for i in (range(len(train_df))):\n            idt = train_df.loc[i, 'StudyInstanceUID']\n            path = os.path.join(train_dir, idt)\n            for im in os.listdir(path):\n                dc = dicom.read_file(os.path.join(path,im))\n                if dc.file_meta.TransferSyntaxUID.name =='JPEG Lossless, Non-Hierarchical, First-Order Prediction (Process 14 [Selection Value 1])':\n                    continue\n                img = load_dicom(os.path.join(path , im))\n                img = cv.resize(img, (128 , 128))\n                image = img_to_array(img)\n                image = image / 255.0\n                trainset += [image]\n                cur_label = []\n                cur_label.append(train_df.loc[i,'C1'])\n                cur_label.append(train_df.loc[i,'C2'])\n                cur_label.append(train_df.loc[i,'C3'])\n                cur_label.append(train_df.loc[i,'C4'])\n                cur_label.append(train_df.loc[i,'C5'])\n                cur_label.append(train_df.loc[i,'C6'])\n                cur_label.append(train_df.loc[i,'C7'])\n                trainlabel += [cur_label]\n                trainidt += [idt]\n                if len(trainidt) == batch_size:                    \n                    yield np.array(trainset), np.array(trainlabel)\n                    trainset, trainlabel, trainidt = [], [], []\n            i+=1","metadata":{"_kg_hide-output":true,"papermill":{"duration":0.027346,"end_time":"2022-08-14T14:39:32.184769","exception":false,"start_time":"2022-08-14T14:39:32.157423","status":"completed"},"pycharm":{"name":"#%%\n"},"tags":[],"execution":{"iopub.status.busy":"2022-10-02T08:10:43.372245Z","iopub.execute_input":"2022-10-02T08:10:43.374447Z","iopub.status.idle":"2022-10-02T08:10:43.392956Z","shell.execute_reply.started":"2022-10-02T08:10:43.37441Z","shell.execute_reply":"2022-10-02T08:10:43.391925Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Test Generator\n\n- Yields only image samples","metadata":{"papermill":{"duration":0.012384,"end_time":"2022-08-14T14:39:32.209528","exception":false,"start_time":"2022-08-14T14:39:32.197144","status":"completed"},"tags":[]}},{"cell_type":"code","source":"def RSNATestGenerator(test_df, batch_size, infinite = True, base_path = test_dir):\n    while 1:        \n        testset=[]\n        testidt=[]\n        for i in (range(len(test_df))):        \n            if type(test_df) is list: idt = test_df[i]\n            else: idt = test_df['StudyInstanceUID'].iloc[i]\n            path = os.path.join(base_path, idt)\n            if os.path.exists(path):\n                for im in os.listdir(path):\n                    dc = dicom.read_file(os.path.join(path,im))\n                    if dc.file_meta.TransferSyntaxUID.name =='JPEG Lossless, Non-Hierarchical, First-Order Prediction (Process 14 [Selection Value 1])':\n                        continue\n                    img=load_dicom(os.path.join(path,im))\n                    img=cv.resize(img,(128, 128))\n                    image=img_to_array(img)\n                    image=image/255.0\n                    testset+=[image]\n                    testidt+=[idt]\n                    if len(testset) == batch_size:                        \n                        yield np.array(testset)\n                        testset = []\n        if len(testset) > 0: yield np.array(testset)\n        if not infinite: break","metadata":{"_kg_hide-output":true,"papermill":{"duration":0.024957,"end_time":"2022-08-14T14:39:32.247245","exception":false,"start_time":"2022-08-14T14:39:32.222288","status":"completed"},"pycharm":{"name":"#%%\n"},"tags":[],"execution":{"iopub.status.busy":"2022-10-02T08:10:43.394727Z","iopub.execute_input":"2022-10-02T08:10:43.395438Z","iopub.status.idle":"2022-10-02T08:10:43.410223Z","shell.execute_reply.started":"2022-10-02T08:10:43.395403Z","shell.execute_reply":"2022-10-02T08:10:43.409339Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### The Model\n","metadata":{"papermill":{"duration":0.01245,"end_time":"2022-08-14T14:39:32.272481","exception":false,"start_time":"2022-08-14T14:39:32.260031","status":"completed"},"tags":[]}},{"cell_type":"code","source":"def get_model():\n    inp = keras.layers.Input((None, None ,1))\n    x = Conv2D(3, 3, padding = 'SAME')(inp)\n    x = EfficientNetB5(include_top=False)(x)\n    x = keras.layers.GlobalAveragePooling2D()(x)\n    out = keras.layers.Dense(7, 'sigmoid')(x)\n    model = keras.models.Model(inp, out)\n    model.summary()\n    model.compile(loss=\"binary_crossentropy\", optimizer = tf.keras.optimizers.Adam(learning_rate = 0.0001))\n    return model","metadata":{"_kg_hide-output":true,"papermill":{"duration":0.0221,"end_time":"2022-08-14T14:39:32.307003","exception":false,"start_time":"2022-08-14T14:39:32.284903","status":"completed"},"pycharm":{"name":"#%%\n"},"tags":[],"execution":{"iopub.status.busy":"2022-10-02T08:11:05.688625Z","iopub.execute_input":"2022-10-02T08:11:05.689547Z","iopub.status.idle":"2022-10-02T08:11:05.697963Z","shell.execute_reply.started":"2022-10-02T08:11:05.68951Z","shell.execute_reply":"2022-10-02T08:11:05.697016Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Main Training Cell\n\nWe use a stratified (based on patient_overall) KFold split and train 5 different models. ","metadata":{"papermill":{"duration":0.012455,"end_time":"2022-08-14T14:39:32.331757","exception":false,"start_time":"2022-08-14T14:39:32.319302","status":"completed"},"tags":[]}},{"cell_type":"code","source":"for train_idx, val_idx in StratifiedKFold(5).split(train_df, train_df['patient_overall']):    \n    K.clear_session()\n    x_train = train_df.iloc[train_idx].reset_index()\n    x_val = train_df.iloc[val_idx].reset_index()\n    model = get_model()\n    hist = model.fit(                            \n                                    RSNATrainGenerator(x_train, min(len(x_train), 1), infinite = False, base_path = train_dir),\n                                    epochs = 50,\n                                    verbose = 1,\n                                    callbacks = [keras.callbacks.EarlyStopping(monitor = 'val_loss', patience = 2, restore_best_weights = True)],\n                                    validation_steps = max((len(x_val) // 1), 1),\n                                    steps_per_epoch = max((len(x_train) // 1), 1),\n                                    validation_data = RSNATrainGenerator(x_val, min(len(x_val), 1), infinite = False, base_path = train_dir),\n                              )\n    val_pred = model.predict(RSNATestGenerator(x_val, min(len(test_df), 1), infinite = False, base_path = train_dir), steps = max((len(test_df) // 64), 1))    \n    try: # the best we can do at the moment..\n        preds = model.predict(RSNATestGenerator(test_df, min(len(test_df), 1), infinite = False, base_path = test_dir), steps = max((len(test_df) // 64), 1))\n        \n        new_preds = []\n        for pred_idx in range(len(preds)):\n            new_preds.append(preds[pred_idx][prediction_type_mapping[pred_idx]])\n        # submission['fractured'] += preds[:, prediction_type_mapping] / 5\n        submission['fractured'] += np.array(new_preds) / 5\n        \n    except: traceback.print_exc()    ","metadata":{"papermill":{"duration":926.451402,"end_time":"2022-08-14T14:54:58.795715","exception":false,"start_time":"2022-08-14T14:39:32.344313","status":"completed"},"pycharm":{"name":"#%%\n"},"tags":[],"execution":{"iopub.status.busy":"2022-10-02T08:11:53.066125Z","iopub.execute_input":"2022-10-02T08:11:53.066494Z","iopub.status.idle":"2022-10-02T08:26:58.619092Z","shell.execute_reply.started":"2022-10-02T08:11:53.066463Z","shell.execute_reply":"2022-10-02T08:26:58.618094Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission","metadata":{"papermill":{"duration":0.081306,"end_time":"2022-08-14T14:54:58.927506","exception":false,"start_time":"2022-08-14T14:54:58.8462","status":"completed"},"pycharm":{"name":"#%%\n"},"tags":[],"execution":{"iopub.status.busy":"2022-10-02T08:26:58.623846Z","iopub.execute_input":"2022-10-02T08:26:58.624805Z","iopub.status.idle":"2022-10-02T08:26:58.654114Z","shell.execute_reply.started":"2022-10-02T08:26:58.624764Z","shell.execute_reply":"2022-10-02T08:26:58.653046Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Submission","metadata":{"papermill":{"duration":0.053456,"end_time":"2022-08-14T14:54:59.031614","exception":false,"start_time":"2022-08-14T14:54:58.978158","status":"completed"},"tags":[]}},{"cell_type":"code","source":"submission.to_csv('submission.csv', index = 0)","metadata":{"papermill":{"duration":0.069591,"end_time":"2022-08-14T14:54:59.152542","exception":false,"start_time":"2022-08-14T14:54:59.082951","status":"completed"},"pycharm":{"name":"#%%\n"},"tags":[],"execution":{"iopub.status.busy":"2022-10-02T08:26:58.687605Z","iopub.execute_input":"2022-10-02T08:26:58.688537Z","iopub.status.idle":"2022-10-02T08:26:58.695728Z","shell.execute_reply.started":"2022-10-02T08:26:58.688502Z","shell.execute_reply":"2022-10-02T08:26:58.694772Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}