{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"#### Yolov5 Cervical Spine (Neck) Fracture Detection\nby @vbookshelf<br>\n10 August 2022\n","metadata":{"papermill":{"duration":0.096947,"end_time":"2021-07-18T05:37:05.065223","exception":false,"start_time":"2021-07-18T05:37:04.968276","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-08-06T08:00:38.103499Z","iopub.execute_input":"2022-08-06T08:00:38.103969Z","iopub.status.idle":"2022-08-06T08:00:38.109681Z","shell.execute_reply.started":"2022-08-06T08:00:38.103929Z","shell.execute_reply":"2022-08-06T08:00:38.108025Z"}}},{"cell_type":"markdown","source":"## Task\n\nTrain a Yolov5 computer vision model to automatically detect fractures on cervical spine (neck) CT scans.\n\n**Input:** CT scans of cervical vetebrae in png format.<br>\n**Output:** Bounding box coordinates, and a score that indicates the probability that a fracture is present.\n\nThis model does not identify the vertebrae where the fracture is present i.e. predictions are only for the patient_overall class.\n\n## Data\n\nThe competition training data consists of 2019 patient scans. Each scan contains multiple cervical spine CT images (slices) in dicom format. There are 7217 images that have bounding boxes that identify the location of fractures.\n\nTo create the training set I chose all images that have bounding boxes. I also randomly selected 1000 images that do not have bounding boxes i.e. no fracture is present. I converted all dicom images to png format without resizing - Yolov5 automatically resizes images during training. I set aside 20% of all images for validation, stratified by class (0 = normal, 1 = fracture).\n\nIn summary, in this notebook there are 6573 train images and 1644 validation images. Taining and validation images are in png format.\n\nI converted the training images from dicom to png format and stored them in a Kaggle dataset (spine-fracture-comp-data). It's attached to this notebook.\n\n\n## Approach\n\n- Use Yolov5l\n- Use png images with the resize parameter set to 512\n- Use image augmentation to reduce overfitting and improve the model's ability to generalize to unseen data.\n- Train for 80 epochs\n\n## Results\n\nFracture detection using computer vision is possible. A comparison of true and predicted bounding boxes is displayed in this notebook. A confusion matrix and classification report is also provided. Click this link to go straight to the training and validation review: <a href='#validation_review'>Training and Validation Review</a><br>\n\n## Learning Resources\n\n- If you are new to Yolov5 then you may find the following tutorial helpful. It explains step by step how Yolov5 works.<br>\nhttps://www.kaggle.com/code/vbookshelf/basics-of-yolo-v5-balloon-detection\n\n- If you want to build a web app to perform cervical spine fracture detection, this example may be helpful:<br>\nhttps://github.com/vbookshelf/COVID-19-CXR-Analyzer","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport os\n\n\nimport cv2\nimport time\nimport shutil\n\nimport torch\nimport torch.nn as nn\nimport torch.optim as optim\nimport torch.nn.functional as F\n\nimport torchvision\nimport torchvision.transforms as transforms\n\nfrom torch.utils.data import Dataset, DataLoader\nfrom torchvision import transforms, utils\n\nfrom torch.utils.data import Dataset, DataLoader\n\n#from tqdm import tqdm\n\n# tqdm doesn't work well in colab.\n# This is the solution:\n# https://stackoverflow.com/questions/41707229/tqdm-printing-to-newline\nimport tqdm.notebook as tq\n#for i in tq.tqdm(...):\n\n\nimport gc\n\n\nimport albumentations as albu\nfrom albumentations import Compose\n\n\nfrom sklearn import model_selection\nfrom sklearn.utils import shuffle\nfrom sklearn import metrics\nfrom sklearn.metrics import confusion_matrix\nfrom sklearn.metrics import classification_report, jaccard_score\nimport itertools\n\n\n# load image with Pillow\nfrom PIL import Image\n\nfrom numpy import asarray\n\nfrom skimage.transform import resize\n\n\nimport matplotlib.pyplot as plt\n\n# Don't Show Warning Messages\nimport warnings\nwarnings.filterwarnings('ignore')\n\n# Note: Pytorch uses a channels-first format:\n# [batch_size, num_channels, height, width]\n\nprint(torch.__version__)\nprint(torchvision.__version__)","metadata":{"papermill":{"duration":11.55616,"end_time":"2021-07-18T05:38:38.821538","exception":false,"start_time":"2021-07-18T05:38:27.265378","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-09-28T09:21:11.351262Z","iopub.execute_input":"2022-09-28T09:21:11.351823Z","iopub.status.idle":"2022-09-28T09:21:14.901861Z","shell.execute_reply.started":"2022-09-28T09:21:11.351701Z","shell.execute_reply":"2022-09-28T09:21:14.899761Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Set the seed values\n\nimport random\n\n# Set the seed value all over the place to make this reproducible.\nseed_val = 101\n\nos.environ['PYTHONHASHSEED'] = str(seed_val)\nrandom.seed(seed_val)\nnp.random.seed(seed_val)\ntorch.manual_seed(seed_val)\ntorch.cuda.manual_seed_all(seed_val)\ntorch.backends.cudnn.deterministic = True","metadata":{"papermill":{"duration":0.227737,"end_time":"2021-07-18T05:38:39.265366","exception":false,"start_time":"2021-07-18T05:38:39.037629","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-09-28T09:21:14.903729Z","iopub.execute_input":"2022-09-28T09:21:14.904191Z","iopub.status.idle":"2022-09-28T09:21:14.915992Z","shell.execute_reply.started":"2022-09-28T09:21:14.904161Z","shell.execute_reply":"2022-09-28T09:21:14.914887Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"os.listdir('../input/')","metadata":{"papermill":{"duration":0.22636,"end_time":"2021-07-18T05:38:39.705411","exception":false,"start_time":"2021-07-18T05:38:39.479051","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-09-28T09:21:14.918585Z","iopub.execute_input":"2022-09-28T09:21:14.918856Z","iopub.status.idle":"2022-09-28T09:21:14.928028Z","shell.execute_reply.started":"2022-09-28T09:21:14.91883Z","shell.execute_reply":"2022-09-28T09:21:14.927077Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"base_path = '../input/rsna-2022-cervical-spine-fracture-detection/'\n\nprep_data_path = '../input/spine-fracture-comp-data/'\n","metadata":{"papermill":{"duration":0.220316,"end_time":"2021-07-18T05:38:40.140516","exception":false,"start_time":"2021-07-18T05:38:39.9202","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-09-28T09:21:14.930124Z","iopub.execute_input":"2022-09-28T09:21:14.931194Z","iopub.status.idle":"2022-09-28T09:21:14.935612Z","shell.execute_reply.started":"2022-09-28T09:21:14.931156Z","shell.execute_reply":"2022-09-28T09:21:14.934609Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Config","metadata":{}},{"cell_type":"code","source":"# Yolo setup:\nNUM_EPOCHS = 80\nBATCH_SIZE = 32\nIMAGE_SIZE = 512 # Yolo will automatically resize the input images to this size.\n\n# This is the fold that Yolo is trained on\nCHOSEN_FOLD = 0\n\nNUM_FOLDS = 5\n\nNUM_CORES = os.cpu_count()\nNUM_CORES","metadata":{"papermill":{"duration":0.225834,"end_time":"2021-07-18T05:38:40.577433","exception":false,"start_time":"2021-07-18T05:38:40.351599","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-09-28T09:21:14.983826Z","iopub.execute_input":"2022-09-28T09:21:14.984162Z","iopub.status.idle":"2022-09-28T09:21:14.991024Z","shell.execute_reply.started":"2022-09-28T09:21:14.984133Z","shell.execute_reply":"2022-09-28T09:21:14.989941Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Set up Yolov5 - For offline use\n\nThe Yolov5 model being used here needs to have the internet turned for training to work. However, it does not need to have the internet on during inference.","metadata":{"papermill":{"duration":0.215149,"end_time":"2021-07-18T05:38:41.008781","exception":false,"start_time":"2021-07-18T05:38:40.793632","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Copy the yolov5 folder from the notebook to  /kaggle/working/\n\nshutil.copytree('../input/v2-balloon-detection-dataset/yolov5', '/kaggle/working/yolov5')\n\n# shutil.copytree('../input/my-yolov5-for-offline-use/yolov5', '/kaggle/working/yolov5')","metadata":{"papermill":{"duration":1.303094,"end_time":"2021-07-18T05:38:42.525305","exception":false,"start_time":"2021-07-18T05:38:41.222211","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-09-28T09:26:09.578781Z","iopub.execute_input":"2022-09-28T09:26:09.579187Z","iopub.status.idle":"2022-09-28T09:26:10.78272Z","shell.execute_reply.started":"2022-09-28T09:26:09.579155Z","shell.execute_reply":"2022-09-28T09:26:10.781848Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!ls","metadata":{"papermill":{"duration":1.013171,"end_time":"2021-07-18T05:38:43.756671","exception":false,"start_time":"2021-07-18T05:38:42.7435","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-09-28T09:26:14.338905Z","iopub.execute_input":"2022-09-28T09:26:14.340317Z","iopub.status.idle":"2022-09-28T09:26:15.297257Z","shell.execute_reply.started":"2022-09-28T09:26:14.340268Z","shell.execute_reply":"2022-09-28T09:26:15.296081Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Load the train data","metadata":{"papermill":{"duration":0.215286,"end_time":"2021-07-18T05:38:44.189361","exception":false,"start_time":"2021-07-18T05:38:43.974075","status":"completed"},"tags":[]}},{"cell_type":"code","source":"path = prep_data_path + 'df_data.csv'\n\ndf_data = pd.read_csv(path)\n\nprint(df_data.shape)\n\ndf_data.head()","metadata":{"papermill":{"duration":0.415195,"end_time":"2021-07-18T05:38:44.823811","exception":false,"start_time":"2021-07-18T05:38:44.408616","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-09-28T09:26:25.751964Z","iopub.execute_input":"2022-09-28T09:26:25.753285Z","iopub.status.idle":"2022-09-28T09:26:25.825857Z","shell.execute_reply.started":"2022-09-28T09:26:25.753225Z","shell.execute_reply":"2022-09-28T09:26:25.824937Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_data['label'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-09-28T09:26:28.074847Z","iopub.execute_input":"2022-09-28T09:26:28.075711Z","iopub.status.idle":"2022-09-28T09:26:28.090955Z","shell.execute_reply.started":"2022-09-28T09:26:28.075669Z","shell.execute_reply":"2022-09-28T09:26:28.089816Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Process the train data","metadata":{"papermill":{"duration":0.217084,"end_time":"2021-07-18T05:38:45.726327","exception":false,"start_time":"2021-07-18T05:38:45.509243","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Create a column called 'target'\n\ndf_data['target'] = list(df_data['label'])\n\n#df_data.head()","metadata":{"execution":{"iopub.status.busy":"2022-09-28T09:26:29.993558Z","iopub.execute_input":"2022-09-28T09:26:29.993947Z","iopub.status.idle":"2022-09-28T09:26:30.003038Z","shell.execute_reply.started":"2022-09-28T09:26:29.993912Z","shell.execute_reply":"2022-09-28T09:26:30.002081Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Targets:\n# 0 = normal\n# 1 = fracture\n\ndf_data['target'].value_counts()","metadata":{"papermill":{"duration":0.231376,"end_time":"2021-07-18T05:38:47.246869","exception":false,"start_time":"2021-07-18T05:38:47.015493","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-09-28T09:26:30.934691Z","iopub.execute_input":"2022-09-28T09:26:30.93509Z","iopub.status.idle":"2022-09-28T09:26:30.943447Z","shell.execute_reply.started":"2022-09-28T09:26:30.935054Z","shell.execute_reply":"2022-09-28T09:26:30.942463Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Create a column for the bbox info","metadata":{}},{"cell_type":"code","source":"# Load the bbox data\n\n# This dataframe lists all slides that have fractures\n\npath = base_path + 'train_bounding_boxes.csv'\ndf_bbox = pd.read_csv(path)\n\nprint(df_bbox.shape)\n\ndf_bbox.head()","metadata":{"execution":{"iopub.status.busy":"2022-09-28T09:26:32.930188Z","iopub.execute_input":"2022-09-28T09:26:32.930911Z","iopub.status.idle":"2022-09-28T09:26:32.962362Z","shell.execute_reply.started":"2022-09-28T09:26:32.930869Z","shell.execute_reply":"2022-09-28T09:26:32.961353Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def create_study_slice(row):\n    \n    study_id = str(row['StudyInstanceUID'])\n    slice_num = str(row['slice_number'])\n    \n    study_slice = study_id + '_' + slice_num\n    \n    return study_slice\n\ndf_bbox['study_slice'] = df_bbox.apply(create_study_slice, axis=1)\n\ndf_bbox = df_bbox.set_index('study_slice')\n\nprint(df_bbox.shape)\n\ndf_bbox.head()","metadata":{"execution":{"iopub.status.busy":"2022-09-28T09:26:34.183503Z","iopub.execute_input":"2022-09-28T09:26:34.185332Z","iopub.status.idle":"2022-09-28T09:26:34.29961Z","shell.execute_reply.started":"2022-09-28T09:26:34.185282Z","shell.execute_reply":"2022-09-28T09:26:34.298521Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Put the bbox info for each image into a list of bbox dicts\n\n# Note: According to df_bbox we only have one bbox per image (slice)\n\nstudy_slice_list = list(df_data['study_slice'])\n\nbbox_list = []\n\nfor i in range(0, len(df_data)):\n    \n    target = df_data.loc[i, 'target']\n    study_slice = df_data.loc[i, 'study_slice']\n    \n    if target == 1:\n    \n        x = df_bbox.loc[study_slice, 'x']\n        y = df_bbox.loc[study_slice, 'y']\n        width = df_bbox.loc[study_slice, 'width']\n        height = df_bbox.loc[study_slice, 'height']\n\n        bbox_dict ={\n            'x': x,\n            'y': y,\n            'width': width,\n            'height': height\n        }\n\n        bbox_list.append(bbox_dict)\n        \n    else:\n        bbox_list.append('none')\n        \n        \n# Add the bbox_list to df_data\n\ndf_data['boxes'] = bbox_list\n\n#print(df_data.shape)\n\n#df_data.head()","metadata":{"execution":{"iopub.status.busy":"2022-09-28T09:26:35.514561Z","iopub.execute_input":"2022-09-28T09:26:35.51492Z","iopub.status.idle":"2022-09-28T09:26:35.936033Z","shell.execute_reply.started":"2022-09-28T09:26:35.514889Z","shell.execute_reply":"2022-09-28T09:26:35.935062Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_data.head(2)","metadata":{"execution":{"iopub.status.busy":"2022-09-28T09:26:37.22247Z","iopub.execute_input":"2022-09-28T09:26:37.223165Z","iopub.status.idle":"2022-09-28T09:26:37.236531Z","shell.execute_reply.started":"2022-09-28T09:26:37.223127Z","shell.execute_reply":"2022-09-28T09:26:37.235553Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Display one entry\n\ndf_data.loc[0, 'boxes']","metadata":{"execution":{"iopub.status.busy":"2022-09-28T09:26:37.645956Z","iopub.execute_input":"2022-09-28T09:26:37.647059Z","iopub.status.idle":"2022-09-28T09:26:37.654552Z","shell.execute_reply.started":"2022-09-28T09:26:37.64699Z","shell.execute_reply":"2022-09-28T09:26:37.653419Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Helper functions","metadata":{"papermill":{"duration":0.401781,"end_time":"2021-07-18T05:38:47.878685","exception":false,"start_time":"2021-07-18T05:38:47.476904","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Confusion matrix where the size of the plot and the text size can be changed\n\n# Source: Scikit Learn website\n# http://scikit-learn.org/stable/auto_examples/\n# model_selection/plot_confusion_matrix.html#sphx-glr-auto-examples-model-\n# selection-plot-confusion-matrix-py\n\n\ndef plot_confusion_matrix(cm, classes,\n                          normalize=False,\n                          title='Confusion matrix',\n                          cmap=plt.cm.Blues,\n                         text_size=12):\n    \"\"\"\n    This function prints and plots the confusion matrix.\n    Normalization can be applied by setting `normalize=True`.\n    \"\"\"\n    if normalize:\n        cm = cm.astype('float') / cm.sum(axis=1)[:, np.newaxis]\n        print(\"Normalized confusion matrix\")\n    else:\n        print('Confusion matrix, without normalization')\n\n    print(cm)\n\n    plt.imshow(cm, interpolation='nearest', cmap=cmap)\n    plt.title(title)\n    plt.colorbar()\n    tick_marks = np.arange(len(classes))\n    plt.xticks(tick_marks, classes, rotation=45, fontsize=text_size)\n    plt.yticks(tick_marks, classes, fontsize=text_size)\n\n    fmt = '.2f' if normalize else 'd'\n    thresh = cm.max() / 2.\n    for i, j in itertools.product(range(cm.shape[0]), range(cm.shape[1])):\n        plt.text(j, i, format(cm[i, j], fmt),\n                 horizontalalignment=\"center\",\n                 color=\"white\" if cm[i, j] > thresh else \"black\", fontsize=text_size)\n\n    plt.ylabel('True label', fontsize=text_size)\n    plt.xlabel('Predicted label', fontsize=text_size)\n    plt.tight_layout()","metadata":{"papermill":{"duration":0.227809,"end_time":"2021-07-18T05:38:48.408535","exception":false,"start_time":"2021-07-18T05:38:48.180726","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-09-28T09:26:40.240092Z","iopub.execute_input":"2022-09-28T09:26:40.240782Z","iopub.status.idle":"2022-09-28T09:26:40.2521Z","shell.execute_reply.started":"2022-09-28T09:26:40.240746Z","shell.execute_reply":"2022-09-28T09:26:40.250179Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Source: https://www.kaggle.com/xhlulu/siim-covid-19-convert-to-jpg-256px\n\nfrom PIL import Image\nimport pydicom\nfrom pydicom.pixel_data_handlers.util import apply_voi_lut\n\n\ndef read_xray(path, voi_lut = True, fix_monochrome = True):\n    # Original from: https://www.kaggle.com/raddar/convert-dicom-to-np-array-the-correct-way\n    dicom = pydicom.read_file(path)\n    \n    # VOI LUT (if available by DICOM device) is used to transform raw DICOM data to \n    # \"human-friendly\" view\n    if voi_lut:\n        data = apply_voi_lut(dicom.pixel_array, dicom)\n    else:\n        data = dicom.pixel_array\n               \n    # depending on this value, X-ray may look inverted - fix that:\n    if fix_monochrome and dicom.PhotometricInterpretation == \"MONOCHROME1\":\n        data = np.amax(data) - data\n        \n    data = data - np.min(data)\n    data = data / np.max(data)\n    data = (data * 255).astype(np.uint8)\n        \n    return data\n\n\n\n\ndef resize(array, size, keep_ratio=False, resample=Image.LANCZOS):\n    # Original from: https://www.kaggle.com/xhlulu/vinbigdata-process-and-resize-to-image\n    im = Image.fromarray(array)\n    \n    if keep_ratio:\n        im.thumbnail((size, size), resample)\n    else:\n        im = im.resize((size, size), resample)\n    \n    return im","metadata":{"papermill":{"duration":0.411458,"end_time":"2021-07-18T05:38:50.75773","exception":false,"start_time":"2021-07-18T05:38:50.346272","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-09-28T09:26:41.420121Z","iopub.execute_input":"2022-09-28T09:26:41.420788Z","iopub.status.idle":"2022-09-28T09:26:41.52882Z","shell.execute_reply.started":"2022-09-28T09:26:41.420752Z","shell.execute_reply":"2022-09-28T09:26:41.527845Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Create the folds","metadata":{"papermill":{"duration":0.215665,"end_time":"2021-07-18T05:38:57.522433","exception":false,"start_time":"2021-07-18T05:38:57.306768","status":"completed"},"tags":[]}},{"cell_type":"code","source":"from sklearn.model_selection import KFold, StratifiedKFold\n\nskf = StratifiedKFold(n_splits=NUM_FOLDS, shuffle=True, random_state=101)\n\nfor fold, ( _, val_) in enumerate(skf.split(X=df_data, y=df_data.target)):\n      df_data.loc[val_ , \"fold\"] = fold\n        \ndf_data['fold'].value_counts()","metadata":{"papermill":{"duration":1.022182,"end_time":"2021-07-18T05:38:58.751994","exception":false,"start_time":"2021-07-18T05:38:57.729812","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-09-28T09:26:43.398332Z","iopub.execute_input":"2022-09-28T09:26:43.39891Z","iopub.status.idle":"2022-09-28T09:26:43.420873Z","shell.execute_reply.started":"2022-09-28T09:26:43.398874Z","shell.execute_reply":"2022-09-28T09:26:43.419832Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Display the target distribution in each fold\n\nfor fold_index in range(0, NUM_FOLDS):\n    \n    df_train = df_data[df_data['fold'] != fold_index]\n    df_val = df_data[df_data['fold'] == fold_index]\n\n    print(f'\\nFold {fold_index}')\n    print('.........')\n    print()\n    print('Train shape:',df_train.shape)\n    print('Val shape:',df_val.shape)\n    print()\n    print('Train target distribution')\n    print(df_train['target'].value_counts())\n    print()\n    print('Val target distribution')\n    print(df_val['target'].value_counts())","metadata":{"papermill":{"duration":0.254228,"end_time":"2021-07-18T05:38:59.825215","exception":false,"start_time":"2021-07-18T05:38:59.570987","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-09-28T09:26:44.409296Z","iopub.execute_input":"2022-09-28T09:26:44.409665Z","iopub.status.idle":"2022-09-28T09:26:44.438056Z","shell.execute_reply.started":"2022-09-28T09:26:44.409636Z","shell.execute_reply":"2022-09-28T09:26:44.437049Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Create the Yolo directory structure\n\nWe need to create a directory structure inside the yolov5 folder. This is where the training and validation data will need to be stored","metadata":{"papermill":{"duration":0.214448,"end_time":"2021-07-18T05:39:00.717004","exception":false,"start_time":"2021-07-18T05:39:00.502556","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Note the the following folder structure must be\n# located inside the yolov5 folder\n\n# base_dir\n    # images\n        # train (contains image files)\n        # validation (contains image files)\n    # labels \n        # train (contains .txt files)\n        # validation (contains .txt files)\n        \n# Yolo expects the bounding box dimensions to be\n# normalized to have values between 0 and 1.\n        \n# Label format in .txt file\n# class x-center y-center width height\n# E.g. 0 0.1 0.2 200 300\n\n# Each label is on a new line, in the .txt file:\n# 0 0.1 0.2 200 300\n# 0 0.1 0.2 200 300","metadata":{"papermill":{"duration":0.241002,"end_time":"2021-07-18T05:39:01.171558","exception":false,"start_time":"2021-07-18T05:39:00.930556","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-09-28T09:26:48.06635Z","iopub.execute_input":"2022-09-28T09:26:48.066716Z","iopub.status.idle":"2022-09-28T09:26:48.07172Z","shell.execute_reply.started":"2022-09-28T09:26:48.066685Z","shell.execute_reply":"2022-09-28T09:26:48.07067Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# change the working directory to yolov5\nos.chdir('/kaggle/working/yolov5')\n\n# Create a new directory (this is happening inside the yolov5 directory)\n\nbase_dir = 'base_dir'\nos.mkdir(base_dir)\n\n\n# Now we create folders inside 'base_dir':\n\n# base_dir\n\n    # images\n        # train\n        # validation\n\n    # labels\n        # train\n        # validation\n\n# images\nimages = os.path.join(base_dir, 'images')\nos.mkdir(images)\n\n# labels\nlabels = os.path.join(base_dir, 'labels')\nos.mkdir(labels)\n\n\n\n# Inside each folder we create seperate folders for each class\n\n# create new folders inside images\ntrain = os.path.join(images, 'train')\nos.mkdir(train)\nvalidation = os.path.join(images, 'validation')\nos.mkdir(validation)\n\n\n# create new folders inside labels\ntrain = os.path.join(labels, 'train')\nos.mkdir(train)\nvalidation = os.path.join(labels, 'validation')\nos.mkdir(validation)\n\n# Display the folder structure\n!tree base_dir","metadata":{"papermill":{"duration":1.026715,"end_time":"2021-07-18T05:39:02.410223","exception":false,"start_time":"2021-07-18T05:39:01.383508","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-09-28T09:26:49.182003Z","iopub.execute_input":"2022-09-28T09:26:49.182889Z","iopub.status.idle":"2022-09-28T09:26:50.161915Z","shell.execute_reply.started":"2022-09-28T09:26:49.182853Z","shell.execute_reply":"2022-09-28T09:26:50.160686Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Process the data\n\nHere we will write a function to process the training and validation data. \n\nWe need to create a separate txt file for each image that contains the details of all the bounding boxes on that image.  This function will also move the training and val data into the directory structure that we created above. We won't need to do any image resizing for Yolo. It will do that automatically during training.","metadata":{"papermill":{"duration":0.376731,"end_time":"2021-07-18T05:39:03.198014","exception":false,"start_time":"2021-07-18T05:39:02.821283","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Change the working directory\nos.chdir('/kaggle/working/')","metadata":{"papermill":{"duration":0.227543,"end_time":"2021-07-18T05:39:03.850619","exception":false,"start_time":"2021-07-18T05:39:03.623076","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-09-28T09:26:51.678362Z","iopub.execute_input":"2022-09-28T09:26:51.678751Z","iopub.status.idle":"2022-09-28T09:26:51.684477Z","shell.execute_reply.started":"2022-09-28T09:26:51.678716Z","shell.execute_reply":"2022-09-28T09:26:51.683068Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Choose the fold to train on.\n\nfold_index = CHOSEN_FOLD\n\ndf_train = df_data[df_data['fold'] != fold_index]\ndf_val = df_data[df_data['fold'] == fold_index]\n\nprint(df_train['target'].value_counts())\nprint(df_val['target'].value_counts())","metadata":{"papermill":{"duration":0.259454,"end_time":"2021-07-18T05:39:04.321393","exception":false,"start_time":"2021-07-18T05:39:04.061939","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-09-28T09:26:52.834332Z","iopub.execute_input":"2022-09-28T09:26:52.835297Z","iopub.status.idle":"2022-09-28T09:26:52.850691Z","shell.execute_reply.started":"2022-09-28T09:26:52.835252Z","shell.execute_reply":"2022-09-28T09:26:52.849617Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Iterate through each row in the dataframe\n\n# We run the function below separately for\n# the train and val sets.\n# Remember that each image gets it's own text file\n# containing the info for all bboxes on that image.\n\n# For each image:\n# 1- get the info for each bounding box\n# 2- write the bounding box info to a txt file\n# 3- save the txt file in the correct folder\n# 4- copy the image to the correct folder\n\n# Note on bboxes:\n# For each image we have a list of dictionaries. Each dict \n# contains the coords one one bbox on that image.\n# We don't need to do anything if the image does not have any bboxes.\n\n\ndef process_data_for_yolo(df, data_type='train'):\n\n    for _, row in tq.tqdm(df.iterrows(), total=len(df)):\n        \n        # Get the target\n        target = row['target']\n        \n        # Create the image file name\n        study_slice = row['study_slice']\n        fname = study_slice + '.png'\n        \n        \n        # Only create txt files for class 1 images\n        if target == 1:\n        \n            # Get the list of bboxes on the image.\n            # Each item in the list is a dict containing the image coords.\n            bbox_dict = row['boxes']\n            \n            # put the coords into a list\n            bbox_list = [bbox_dict]\n            \n            # These are the original image sizes.\n            # If we have resized the images then this must be changed to\n            # the new sizes. We will then also be using resized bbox coords.\n            image_width = row['w']\n            image_height = row['h']\n\n\n            # Convert into the Yolo input format\n            # ...................................\n            \n            yolo_data = []\n\n            # row by row\n            for coord_dict in bbox_list:\n\n                xmin = int(coord_dict['x'])\n                ymin = int(coord_dict['y'])\n                bbox_w = int(coord_dict['width'])\n                bbox_h = int(coord_dict['height'])\n\n                # We only have one class i.e. opacity\n                # We will set the class_id to 0 for all images.\n                # Class numbers must start from 0.\n                class_id = target\n\n                x_center = xmin + (bbox_w/2)\n                y_center = ymin + (bbox_h/2)\n\n\n                # Normalize\n                # Yolo expects the dimensions to be normalized i.e.\n                # all values between 0 and 1.\n\n                x_center = x_center/image_width\n                y_center = y_center/image_height\n                bbox_w = bbox_w/image_width\n                bbox_h = bbox_h/image_height\n\n                # [class_id, x-center, y-center, width, height]\n                yolo_list = [class_id, x_center, y_center, bbox_w, bbox_h]\n\n                yolo_data.append(yolo_list)\n\n            # convert to nump array\n            yolo_data = np.array(yolo_data)\n\n\n            # Write the image bbox info to a txt file\n            #image_id = image_name.split('.')[0]\n            np.savetxt(os.path.join('yolov5/base_dir', \n                        f\"labels/{data_type}/{study_slice}.txt\"),\n                        yolo_data, \n                        fmt=[\"%d\", \"%f\", \"%f\", \"%f\", \"%f\"]\n                        ) # fmt means format the columns\n\n\n\n        # Copy the image to images\n        # Set the path to the images here.\n        shutil.copyfile(\n            f\"{prep_data_path}/images_dir/{fname}\",\n            os.path.join('yolov5/base_dir', f\"images/{data_type}/{fname}\")\n        )\n        \n        \n\n# Call the function    \nprocess_data_for_yolo(df_train, data_type='train')\nprocess_data_for_yolo(df_val, data_type='validation')","metadata":{"papermill":{"duration":3.555252,"end_time":"2021-07-18T05:39:11.818383","exception":false,"start_time":"2021-07-18T05:39:08.263131","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-09-28T09:26:54.747858Z","iopub.execute_input":"2022-09-28T09:26:54.748544Z","iopub.status.idle":"2022-09-28T09:28:19.939993Z","shell.execute_reply.started":"2022-09-28T09:26:54.74851Z","shell.execute_reply":"2022-09-28T09:28:19.938861Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!ls","metadata":{"execution":{"iopub.status.busy":"2022-09-28T09:28:19.942008Z","iopub.execute_input":"2022-09-28T09:28:19.942613Z","iopub.status.idle":"2022-09-28T09:28:21.0023Z","shell.execute_reply.started":"2022-09-28T09:28:19.942575Z","shell.execute_reply":"2022-09-28T09:28:21.001095Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Check that the files have been created\n\nprint(len(os.listdir('yolov5/base_dir/images/train')))\nprint(len(os.listdir('yolov5/base_dir/images/validation')))\n\nprint(len(os.listdir('yolov5/base_dir/labels/train')))\nprint(len(os.listdir('yolov5/base_dir/labels/validation')))","metadata":{"papermill":{"duration":0.237175,"end_time":"2021-07-18T05:39:12.278949","exception":false,"start_time":"2021-07-18T05:39:12.041774","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-09-28T09:28:21.004513Z","iopub.execute_input":"2022-09-28T09:28:21.004961Z","iopub.status.idle":"2022-09-28T09:28:21.023307Z","shell.execute_reply.started":"2022-09-28T09:28:21.004915Z","shell.execute_reply":"2022-09-28T09:28:21.022172Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"text_file_list = os.listdir('yolov5/base_dir/labels/train')\n\ntext_file = text_file_list[0]\n\ntext_file","metadata":{"papermill":{"duration":0.234228,"end_time":"2021-07-18T05:39:12.743254","exception":false,"start_time":"2021-07-18T05:39:12.509026","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-09-28T09:28:21.026723Z","iopub.execute_input":"2022-09-28T09:28:21.027113Z","iopub.status.idle":"2022-09-28T09:28:21.03893Z","shell.execute_reply.started":"2022-09-28T09:28:21.027077Z","shell.execute_reply":"2022-09-28T09:28:21.037909Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Display the contents of a text file\n\n! cat 'yolov5/base_dir/labels/train/1.2.826.0.1.3680043.26979_172.txt'","metadata":{"papermill":{"duration":1.008666,"end_time":"2021-07-18T05:39:13.974188","exception":false,"start_time":"2021-07-18T05:39:12.965522","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-09-28T09:28:21.040475Z","iopub.execute_input":"2022-09-28T09:28:21.041108Z","iopub.status.idle":"2022-09-28T09:28:22.134542Z","shell.execute_reply.started":"2022-09-28T09:28:21.041073Z","shell.execute_reply":"2022-09-28T09:28:22.133263Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Create the yaml file\nYolo requires that we also create a yaml file inside the yolov5 folder.","metadata":{"papermill":{"duration":0.222359,"end_time":"2021-07-18T05:39:14.419466","exception":false,"start_time":"2021-07-18T05:39:14.197107","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Ref:\n# Reading and Writing YAML to a File in Python\n# https://stackabuse.com/reading-and-writing-yaml-to-a-file-in-python\n\nyaml_dict = {'train': 'base_dir/images/train',   # path to the train folder\n            'val': 'base_dir/images/validation', # path to the val folder\n            'nc': 2,                             # number of classes\n            'names': ['0', '1']}                # list of label names\n\n\n\n# Create the yaml file called my_data.yaml\n# We will save this file inside the yolov5 folder.\n\nimport yaml\n\nwith open(r'yolov5/my_data.yaml', 'w') as file:\n    documents = yaml.dump(yaml_dict, file)","metadata":{"papermill":{"duration":0.245553,"end_time":"2021-07-18T05:39:14.886293","exception":false,"start_time":"2021-07-18T05:39:14.64074","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-09-28T09:28:22.137387Z","iopub.execute_input":"2022-09-28T09:28:22.137913Z","iopub.status.idle":"2022-09-28T09:28:22.153546Z","shell.execute_reply.started":"2022-09-28T09:28:22.137873Z","shell.execute_reply":"2022-09-28T09:28:22.152595Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Check that the my_data.yaml file is in the yolov5 folder.\n# It should appear in the list of files.\n\nos.listdir('yolov5')","metadata":{"papermill":{"duration":0.233988,"end_time":"2021-07-18T05:39:15.342968","exception":false,"start_time":"2021-07-18T05:39:15.10898","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-09-28T09:28:22.157337Z","iopub.execute_input":"2022-09-28T09:28:22.158185Z","iopub.status.idle":"2022-09-28T09:28:22.171909Z","shell.execute_reply.started":"2022-09-28T09:28:22.158147Z","shell.execute_reply":"2022-09-28T09:28:22.170863Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Display the contents of the yaml file\n\n! cat 'yolov5/my_data.yaml'","metadata":{"papermill":{"duration":1.026811,"end_time":"2021-07-18T05:39:16.593576","exception":false,"start_time":"2021-07-18T05:39:15.566765","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-09-28T09:28:22.173718Z","iopub.execute_input":"2022-09-28T09:28:22.176561Z","iopub.status.idle":"2022-09-28T09:28:23.177665Z","shell.execute_reply.started":"2022-09-28T09:28:22.17652Z","shell.execute_reply":"2022-09-28T09:28:23.176457Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Create a custom hyperameter/augmentation yaml file","metadata":{"papermill":{"duration":0.223307,"end_time":"2021-07-18T05:39:17.039924","exception":false,"start_time":"2021-07-18T05:39:16.816617","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Ref:\n# Reading and Writing YAML to a File in Python\n# https://stackabuse.com/reading-and-writing-yaml-to-a-file-in-python\n\n\nyaml_dict = {\n    \n'lr0': 0.01,  # initial learning rate (SGD=1E-2, Adam=1E-3)\n'lrf': 0.032,  # final OneCycleLR learning rate (lr0 * lrf)\n'momentum': 0.937,  # SGD momentum/Adam beta1\n'weight_decay': 0.0005,  # optimizer weight decay 5e-4\n'warmup_epochs': 3.0,  # warmup epochs (fractions ok)\n'warmup_momentum': 0.8,  # warmup initial momentum\n'warmup_bias_lr': 0.1,  # warmup initial bias lr\n'box': 0.1,  # box loss gain\n'cls': 1.0,  # cls loss gain\n'cls_pw': 0.5,  # cls BCELoss positive_weight\n'obj': 2.0,  # obj loss gain (scale with pixels)\n'obj_pw': 0.5,  # obj BCELoss positive_weight\n'iou_t': 0.20,  # IoU training threshold\n'anchor_t': 4.0,  # anchor-multiple threshold\n'anchors': 0,  # anchors per output layer (0 to ignore)\n'fl_gamma': 0.0,  # focal loss gamma (efficientDet default gamma=1.5)\n'hsv_h': 0,  # image HSV-Hue augmentation (fraction)\n'hsv_s': 0,  # image HSV-Saturation augmentation (fraction)\n'hsv_v': 0,  # image HSV-Value augmentation (fraction)\n'degrees': 30.0,  # image rotation (+/- deg)\n'translate': 0.2,  # image translation (+/- fraction)\n'scale': 0.3,  # image scale (+/- gain)\n'shear': 0.0,  # image shear (+/- deg)\n'perspective': 0.0,  # image perspective (+/- fraction), range 0-0.001\n'flipud': 0.2,  # image flip up-down (probability)\n'fliplr': 0.5,  # image flip left-right (probability)\n'mosaic': 0.8,  # image mosaic (probability)\n'mixup': 0.0  # image mixup (probability)\n    \n}\n\n\n# Create the yaml file called my_hyp.yaml\n# We will save this file inside the yolov5 folder.\n\nimport yaml\n\nwith open(r'yolov5/my_hyp.yaml', 'w') as file:\n    documents = yaml.dump(yaml_dict, file)","metadata":{"papermill":{"duration":0.239227,"end_time":"2021-07-18T05:39:17.506221","exception":false,"start_time":"2021-07-18T05:39:17.266994","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-09-28T09:28:23.179672Z","iopub.execute_input":"2022-09-28T09:28:23.181304Z","iopub.status.idle":"2022-09-28T09:28:23.194093Z","shell.execute_reply.started":"2022-09-28T09:28:23.181259Z","shell.execute_reply":"2022-09-28T09:28:23.193156Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Check that the my_data.yaml file is in the yolov5 folder.\n# It should appear in the list of files.\n\nos.listdir('yolov5')","metadata":{"papermill":{"duration":0.234891,"end_time":"2021-07-18T05:39:17.96441","exception":false,"start_time":"2021-07-18T05:39:17.729519","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-09-28T09:28:23.198507Z","iopub.execute_input":"2022-09-28T09:28:23.198905Z","iopub.status.idle":"2022-09-28T09:28:23.210057Z","shell.execute_reply.started":"2022-09-28T09:28:23.198872Z","shell.execute_reply":"2022-09-28T09:28:23.209141Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Display the contents of the  my_hyp.yaml file\n\n! cat 'yolov5/my_hyp.yaml'","metadata":{"papermill":{"duration":1.010899,"end_time":"2021-07-18T05:39:19.194646","exception":false,"start_time":"2021-07-18T05:39:18.183747","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-09-28T09:28:23.211597Z","iopub.execute_input":"2022-09-28T09:28:23.212426Z","iopub.status.idle":"2022-09-28T09:28:24.176518Z","shell.execute_reply.started":"2022-09-28T09:28:23.212388Z","shell.execute_reply":"2022-09-28T09:28:24.1753Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Train the model","metadata":{"papermill":{"duration":0.239099,"end_time":"2021-07-18T05:39:19.65831","exception":false,"start_time":"2021-07-18T05:39:19.419211","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# change the working directory to yolov5\nos.chdir('/kaggle/working/yolov5')\n\n!pwd","metadata":{"papermill":{"duration":1.119913,"end_time":"2021-07-18T05:39:21.005032","exception":false,"start_time":"2021-07-18T05:39:19.885119","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-09-28T09:28:24.178678Z","iopub.execute_input":"2022-09-28T09:28:24.179029Z","iopub.status.idle":"2022-09-28T09:28:25.129885Z","shell.execute_reply.started":"2022-09-28T09:28:24.178982Z","shell.execute_reply":"2022-09-28T09:28:25.12868Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# What some of the parameters mean:\n\n# --weights => the pre-trained model that we are using.\n# We are using weights that have been downloaded and stored in a Kaggle dataset.\n\n# --save-txt => The predicted bbox coordinates get saved to a txt file. One txt file per image.\n# --save-conf => The conf score gets included in the above txt file.\n# --img => The image will be resized to this size before creating the mosaic.\n# --conf => The confidence threshold\n# --rect => Means don't use mosaic augmentation during training\n# --name => Give a model a name e.g. --name my_model\n# --batch => batch size\n# --epochs => number of training epochs\n# --data => the yaml file path\n# --exist-ok => do not increment the project names with each run i.e. don't change exp to epx2, exp3 etc.\n# --nosave => do not save the images/videos (helpful when deploying to a server)\n\n# It's helpful to review the source code in detect.py to know what the above parameters mean.\n# detect.py is located inside the yolov5 folder.","metadata":{"papermill":{"duration":0.234227,"end_time":"2021-07-18T05:39:21.512445","exception":false,"start_time":"2021-07-18T05:39:21.278218","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-09-28T09:28:25.131815Z","iopub.execute_input":"2022-09-28T09:28:25.132448Z","iopub.status.idle":"2022-09-28T09:28:25.138043Z","shell.execute_reply.started":"2022-09-28T09:28:25.132413Z","shell.execute_reply":"2022-09-28T09:28:25.136902Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Path to the weights stored in the dataset\n# yolo_model_path = '/kaggle/input/my-yolov5-for-offline-use/yolov5l.pt'\nyolo_model_path = '/kaggle/input/yolo-wt/yolov5s.pt'\n\n# If you uncomment and run this line you will get a request to enter a \n# wandb password. To solve this problem we include WANDB_MODE=\"dryrun\" in\n# the next line.\n#! python train.py --img 1024 --batch 8 --epochs 2 --data my_data.yaml --cfg models/yolov5s.yaml --name my_model\n\n# Note that now hyp=my_hyp.yaml in the printout blow.\n!WANDB_MODE=\"dryrun\" python train.py --img $IMAGE_SIZE --batch $BATCH_SIZE --epochs 4 --data my_data.yaml --hyp my_hyp.yaml --weights $yolo_model_path","metadata":{"_kg_hide-output":true,"papermill":{"duration":4424.742752,"end_time":"2021-07-18T06:53:06.478556","exception":false,"start_time":"2021-07-18T05:39:21.735804","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-09-28T09:46:58.914903Z","iopub.execute_input":"2022-09-28T09:46:58.916007Z","iopub.status.idle":"2022-09-28T09:56:36.16466Z","shell.execute_reply.started":"2022-09-28T09:46:58.915955Z","shell.execute_reply":"2022-09-28T09:56:36.163426Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!pwd","metadata":{"papermill":{"duration":4.359494,"end_time":"2021-07-18T06:53:14.249096","exception":false,"start_time":"2021-07-18T06:53:09.889602","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-09-28T10:02:10.153182Z","iopub.execute_input":"2022-09-28T10:02:10.15413Z","iopub.status.idle":"2022-09-28T10:02:11.152492Z","shell.execute_reply.started":"2022-09-28T10:02:10.15409Z","shell.execute_reply":"2022-09-28T10:02:11.151361Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Copy the trained model\n\nWe will copy the trained model to the Kaggle working directory. This will make the model easier to access.","metadata":{"papermill":{"duration":3.382839,"end_time":"2021-07-18T06:53:20.977817","exception":false,"start_time":"2021-07-18T06:53:17.594978","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# This is where the trained model is stored\n\npath = 'yolov5/runs/train/exp2/weights/'\n\nos.listdir(path)","metadata":{"papermill":{"duration":3.680092,"end_time":"2021-07-18T06:53:27.975342","exception":false,"start_time":"2021-07-18T06:53:24.29525","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-09-28T10:02:23.158443Z","iopub.execute_input":"2022-09-28T10:02:23.158876Z","iopub.status.idle":"2022-09-28T10:02:23.179825Z","shell.execute_reply.started":"2022-09-28T10:02:23.158836Z","shell.execute_reply":"2022-09-28T10:02:23.178957Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# change the working directory\nos.chdir('/kaggle/working/')\n\n!pwd","metadata":{"papermill":{"duration":4.147132,"end_time":"2021-07-18T06:53:35.505998","exception":false,"start_time":"2021-07-18T06:53:31.358866","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-09-28T10:02:26.735807Z","iopub.execute_input":"2022-09-28T10:02:26.736194Z","iopub.status.idle":"2022-09-28T10:02:27.678212Z","shell.execute_reply.started":"2022-09-28T10:02:26.736162Z","shell.execute_reply":"2022-09-28T10:02:27.67699Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Copy the best model to the kaggle/working\n\nshutil.copyfile(\n    '/kaggle/working/yolov5/runs/train/exp2/weights/best.pt',\n    '/kaggle/working/best.pt')\n\n\n!ls","metadata":{"papermill":{"duration":4.267281,"end_time":"2021-07-18T06:53:43.461979","exception":false,"start_time":"2021-07-18T06:53:39.194698","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-09-28T10:02:35.323845Z","iopub.execute_input":"2022-09-28T10:02:35.324863Z","iopub.status.idle":"2022-09-28T10:02:36.307141Z","shell.execute_reply.started":"2022-09-28T10:02:35.32482Z","shell.execute_reply":"2022-09-28T10:02:36.305936Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Get the name of the last experiment\n\nYolov5 saves every training run as an experiment.","metadata":{"papermill":{"duration":3.532048,"end_time":"2021-07-18T06:53:51.514652","exception":false,"start_time":"2021-07-18T06:53:47.982604","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# change the working directory to yolov5\nos.chdir('/kaggle/working/yolov5')\n\n!pwd","metadata":{"papermill":{"duration":4.55418,"end_time":"2021-07-18T06:53:59.475607","exception":false,"start_time":"2021-07-18T06:53:54.921427","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-09-28T10:02:39.11106Z","iopub.execute_input":"2022-09-28T10:02:39.11186Z","iopub.status.idle":"2022-09-28T10:02:40.054968Z","shell.execute_reply.started":"2022-09-28T10:02:39.111819Z","shell.execute_reply":"2022-09-28T10:02:40.053773Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"os.listdir('runs/train/')","metadata":{"papermill":{"duration":3.405475,"end_time":"2021-07-18T06:54:06.248916","exception":false,"start_time":"2021-07-18T06:54:02.843441","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-09-28T10:02:41.089658Z","iopub.execute_input":"2022-09-28T10:02:41.090293Z","iopub.status.idle":"2022-09-28T10:02:41.099086Z","shell.execute_reply.started":"2022-09-28T10:02:41.090253Z","shell.execute_reply":"2022-09-28T10:02:41.097896Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# get a list of experiments\nexp_list = os.listdir('runs/train/')\n\nexp_list","metadata":{"papermill":{"duration":3.377704,"end_time":"2021-07-18T06:54:13.375338","exception":false,"start_time":"2021-07-18T06:54:09.997634","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-09-28T10:02:42.899829Z","iopub.execute_input":"2022-09-28T10:02:42.900211Z","iopub.status.idle":"2022-09-28T10:02:42.907482Z","shell.execute_reply.started":"2022-09-28T10:02:42.900182Z","shell.execute_reply":"2022-09-28T10:02:42.9065Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Get the latest exp.\n# I found that the first item in the list is the latest experiment. Not\n# the last item as one would normally expect.\nexp = exp_list[1]\n\nexp","metadata":{"papermill":{"duration":4.272773,"end_time":"2021-07-18T06:54:21.146894","exception":false,"start_time":"2021-07-18T06:54:16.874121","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-09-28T10:02:50.907326Z","iopub.execute_input":"2022-09-28T10:02:50.907691Z","iopub.status.idle":"2022-09-28T10:02:50.913743Z","shell.execute_reply.started":"2022-09-28T10:02:50.907661Z","shell.execute_reply":"2022-09-28T10:02:50.912809Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Display the contents of the \"exp\" folder\nos.listdir(f'runs/train/{exp}')","metadata":{"papermill":{"duration":3.352582,"end_time":"2021-07-18T06:54:27.864267","exception":false,"start_time":"2021-07-18T06:54:24.511685","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-09-28T10:02:54.764981Z","iopub.execute_input":"2022-09-28T10:02:54.765409Z","iopub.status.idle":"2022-09-28T10:02:54.773976Z","shell.execute_reply.started":"2022-09-28T10:02:54.765373Z","shell.execute_reply":"2022-09-28T10:02:54.773075Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"| <a id='validation_review'></a>","metadata":{}},{"cell_type":"markdown","source":"## Training and Validation Review\n\n**Please look at the list of files in the output of the above cell.**\n\n- Yolo stores all the training curves as one png file. To view the training curves we need to display the png file.\n\n- **IMPORTANT NOTE:** The summary displayed at the end of training shows the resuts for the LAST epoch. This is not the results for the BEST epoch. The results for each epoch are logged in a file called results.txt. We will load that file into a pandas dataframe and then get the best epoch and the best mAP score.\n\n- Yolo also stores png images showing the true and predicted labels for each val batch. In these images the true and predicted bounding boxes are drawn in. One batch is shown on one image.\n\nThere's more results info available. The yolov5 folder will appear in the output of this notebook. I suggest that you download it and look at the contents of the yolov5/runs/train/exp folder.","metadata":{"papermill":{"duration":3.34843,"end_time":"2021-07-18T06:54:34.867454","exception":false,"start_time":"2021-07-18T06:54:31.519024","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Display the contents of the \"exp\" folder\nos.listdir(f'runs/train/{exp}')","metadata":{"papermill":{"duration":3.490772,"end_time":"2021-07-18T06:54:41.697612","exception":false,"start_time":"2021-07-18T06:54:38.20684","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-09-28T10:03:05.110043Z","iopub.execute_input":"2022-09-28T10:03:05.110812Z","iopub.status.idle":"2022-09-28T10:03:05.119223Z","shell.execute_reply.started":"2022-09-28T10:03:05.110772Z","shell.execute_reply":"2022-09-28T10:03:05.118116Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# change the working directory to yolov5\n\nos.chdir('/kaggle/working/yolov5')\n\n!pwd","metadata":{"papermill":{"duration":4.17146,"end_time":"2021-07-18T06:54:49.377868","exception":false,"start_time":"2021-07-18T06:54:45.206408","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-09-28T10:03:08.793572Z","iopub.execute_input":"2022-09-28T10:03:08.794372Z","iopub.status.idle":"2022-09-28T10:03:09.775242Z","shell.execute_reply.started":"2022-09-28T10:03:08.79433Z","shell.execute_reply":"2022-09-28T10:03:09.774078Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Display the training curves","metadata":{"papermill":{"duration":3.348482,"end_time":"2021-07-18T06:54:56.99484","exception":false,"start_time":"2021-07-18T06:54:53.646358","status":"completed"},"tags":[]}},{"cell_type":"code","source":"plt.figure(figsize = (15, 15))\nplt.imshow(plt.imread(f'runs/train/{exp}/results.png'))","metadata":{"papermill":{"duration":4.331596,"end_time":"2021-07-18T06:55:04.665328","exception":false,"start_time":"2021-07-18T06:55:00.333732","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-09-28T10:03:12.433931Z","iopub.execute_input":"2022-09-28T10:03:12.434993Z","iopub.status.idle":"2022-09-28T10:03:13.226698Z","shell.execute_reply.started":"2022-09-28T10:03:12.434952Z","shell.execute_reply":"2022-09-28T10:03:13.225781Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Get the best mAP and best epoch","metadata":{"papermill":{"duration":3.325504,"end_time":"2021-07-18T06:55:11.376077","exception":false,"start_time":"2021-07-18T06:55:08.050573","status":"completed"},"tags":[]}},{"cell_type":"code","source":"!ls","metadata":{"papermill":{"duration":4.486321,"end_time":"2021-07-18T06:55:19.433214","exception":false,"start_time":"2021-07-18T06:55:14.946893","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-09-28T10:03:22.296958Z","iopub.execute_input":"2022-09-28T10:03:22.297918Z","iopub.status.idle":"2022-09-28T10:03:23.259981Z","shell.execute_reply.started":"2022-09-28T10:03:22.29788Z","shell.execute_reply":"2022-09-28T10:03:23.258882Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Display the contents of the results.txt file\n\npath = f'runs/train/{exp}/results.txt'\n\n!cat $path","metadata":{"papermill":{"duration":5.156095,"end_time":"2021-07-18T06:55:27.939715","exception":false,"start_time":"2021-07-18T06:55:22.78362","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-09-28T10:03:24.603137Z","iopub.execute_input":"2022-09-28T10:03:24.603901Z","iopub.status.idle":"2022-09-28T10:03:25.670864Z","shell.execute_reply.started":"2022-09-28T10:03:24.60386Z","shell.execute_reply":"2022-09-28T10:03:25.669285Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Read the results from the training log: results.txt\n\n# https://stackoverflow.com/questions/3277503/how-to-read-a-file-line-by-line-into-a-list\n# https://stackoverflow.com/questions/65381312/how-to-convert-a-yolo-darknet-format-into-csv-file\n\n\nfilename = f'runs/train/{exp}/results.txt'\n\nfile_list = []\n\nwith open(filename) as f:\n    # read a line into a list, format: ['item item item', 'item item item', ...]\n    file_line_list = f.readlines()\n    \n    \nfor i in range(0, len(file_line_list)):\n    \n    # Get the first item in the list and split on the spaces.\n    # This returns a list of all items in the line: ['item', 'item', 'item']\n    line_list = file_line_list[i].split()\n    \n    # remove whitespace characters like `\\n` at the end of each line\n    line_list = [x.strip() for x in line_list]\n    \n    # Save the list.\n    # all_lines_list is a list of lists\n    file_list.append(line_list)\n    \nlen(file_list)","metadata":{"papermill":{"duration":3.36078,"end_time":"2021-07-18T06:55:34.638954","exception":false,"start_time":"2021-07-18T06:55:31.278174","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-09-28T10:03:33.364628Z","iopub.execute_input":"2022-09-28T10:03:33.365033Z","iopub.status.idle":"2022-09-28T10:03:33.375762Z","shell.execute_reply.started":"2022-09-28T10:03:33.364981Z","shell.execute_reply":"2022-09-28T10:03:33.37465Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Put the file data into a dataframe\n\ndf = pd.DataFrame(file_list)\n\ndf.head(10)","metadata":{"papermill":{"duration":3.393293,"end_time":"2021-07-18T06:55:41.765059","exception":false,"start_time":"2021-07-18T06:55:38.371766","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-09-28T10:03:37.15928Z","iopub.execute_input":"2022-09-28T10:03:37.159635Z","iopub.status.idle":"2022-09-28T10:03:37.181268Z","shell.execute_reply.started":"2022-09-28T10:03:37.159605Z","shell.execute_reply":"2022-09-28T10:03:37.180173Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# choose only the columns we want\n\ncol_names = ['epoch', 'P', 'R', 'map0.5', 'map0.5:0.95']\n\n# filter out specific columns\ndf_results = df[[0, 8, 9, 10, 11]]\n\ndf_results.columns = col_names\n\n# change the column names\ndf_results.head(10)","metadata":{"papermill":{"duration":3.75076,"end_time":"2021-07-18T06:55:48.88542","exception":false,"start_time":"2021-07-18T06:55:45.13466","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-09-28T10:03:42.23334Z","iopub.execute_input":"2022-09-28T10:03:42.234253Z","iopub.status.idle":"2022-09-28T10:03:42.249328Z","shell.execute_reply.started":"2022-09-28T10:03:42.234203Z","shell.execute_reply":"2022-09-28T10:03:42.248348Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Get the best map0.5\n\nbest_map = df_results['map0.5'].max()\n\nprint('---------------------')\n\nprint('Best map0.5:', best_map)\nprint()\n\n# print the row that contains the best map0.5\ndf = df_results[df_results['map0.5'] == best_map]\n\nprint(df.head())\n\nprint('---------------------')","metadata":{"papermill":{"duration":3.438387,"end_time":"2021-07-18T06:55:55.694216","exception":false,"start_time":"2021-07-18T06:55:52.255829","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-09-28T10:03:50.023049Z","iopub.execute_input":"2022-09-28T10:03:50.023416Z","iopub.status.idle":"2022-09-28T10:03:50.035004Z","shell.execute_reply.started":"2022-09-28T10:03:50.023374Z","shell.execute_reply":"2022-09-28T10:03:50.03368Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Display one batch of train images","metadata":{"execution":{"iopub.execute_input":"2021-07-05T08:19:53.683161Z","iopub.status.busy":"2021-07-05T08:19:53.682523Z","iopub.status.idle":"2021-07-05T08:19:53.68887Z","shell.execute_reply":"2021-07-05T08:19:53.687579Z","shell.execute_reply.started":"2021-07-05T08:19:53.683105Z"},"papermill":{"duration":3.572099,"end_time":"2021-07-18T06:56:03.775207","exception":false,"start_time":"2021-07-18T06:56:00.203108","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Train\n# One mosaic batch of train images with labels\n\nplt.figure(figsize = (15, 15))\nplt.imshow(plt.imread(f'runs/train/{exp}/train_batch0.jpg'))","metadata":{"papermill":{"duration":4.375391,"end_time":"2021-07-18T06:56:11.494857","exception":false,"start_time":"2021-07-18T06:56:07.119466","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-09-28T10:03:55.130345Z","iopub.execute_input":"2022-09-28T10:03:55.130707Z","iopub.status.idle":"2022-09-28T10:03:55.792671Z","shell.execute_reply.started":"2022-09-28T10:03:55.130675Z","shell.execute_reply":"2022-09-28T10:03:55.791859Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Display true and predicted val set bboxes\n\nHere we will display the true and predicted bboxes for two val batches.","metadata":{"papermill":{"duration":3.378424,"end_time":"2021-07-18T06:56:18.287695","exception":false,"start_time":"2021-07-18T06:56:14.909271","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# BATCH 0 - TRUE BBOXES\n\nplt.figure(figsize = (15, 15))\nplt.imshow(plt.imread(f'runs/train/{exp}/test_batch0_labels.jpg'))","metadata":{"papermill":{"duration":3.956924,"end_time":"2021-07-18T06:56:25.962776","exception":false,"start_time":"2021-07-18T06:56:22.005852","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-09-28T10:04:01.591842Z","iopub.execute_input":"2022-09-28T10:04:01.592225Z","iopub.status.idle":"2022-09-28T10:04:02.228541Z","shell.execute_reply.started":"2022-09-28T10:04:01.592193Z","shell.execute_reply":"2022-09-28T10:04:02.224523Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# BATCH 0 - PREDICTED BBOXES\n\nplt.figure(figsize = (15, 15))\nplt.imshow(plt.imread(f'runs/train/{exp}/test_batch0_pred.jpg'))","metadata":{"papermill":{"duration":4.657357,"end_time":"2021-07-18T06:56:34.292266","exception":false,"start_time":"2021-07-18T06:56:29.634909","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-09-28T10:04:38.22919Z","iopub.execute_input":"2022-09-28T10:04:38.230489Z","iopub.status.idle":"2022-09-28T10:04:38.986669Z","shell.execute_reply.started":"2022-09-28T10:04:38.230437Z","shell.execute_reply":"2022-09-28T10:04:38.985827Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# BATCH 1 - TRUE BBOXES\n\nplt.figure(figsize = (15, 15))\nplt.imshow(plt.imread(f'runs/train/{exp}/test_batch1_labels.jpg'))","metadata":{"papermill":{"duration":4.052771,"end_time":"2021-07-18T06:56:41.741357","exception":false,"start_time":"2021-07-18T06:56:37.688586","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-09-28T10:04:44.593062Z","iopub.execute_input":"2022-09-28T10:04:44.594156Z","iopub.status.idle":"2022-09-28T10:04:45.259488Z","shell.execute_reply.started":"2022-09-28T10:04:44.594106Z","shell.execute_reply":"2022-09-28T10:04:45.256766Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# BATCH 1 - PREDICTED BBOXES\n\nplt.figure(figsize = (15, 15))\nplt.imshow(plt.imread(f'runs/train/{exp}/test_batch1_pred.jpg'))","metadata":{"papermill":{"duration":4.138354,"end_time":"2021-07-18T06:56:49.649222","exception":false,"start_time":"2021-07-18T06:56:45.510868","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-09-28T10:04:50.163153Z","iopub.execute_input":"2022-09-28T10:04:50.164004Z","iopub.status.idle":"2022-09-28T10:04:50.817693Z","shell.execute_reply.started":"2022-09-28T10:04:50.163969Z","shell.execute_reply":"2022-09-28T10:04:50.816853Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"papermill":{"duration":3.82951,"end_time":"2021-07-18T06:56:56.920429","exception":false,"start_time":"2021-07-18T06:56:53.090919","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Make a prediction on the val set","metadata":{"papermill":{"duration":4.04488,"end_time":"2021-07-18T06:57:04.446831","exception":false,"start_time":"2021-07-18T06:57:00.401951","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# change the working directory\nos.chdir('/kaggle/working')\n\n!pwd","metadata":{"papermill":{"duration":4.204321,"end_time":"2021-07-18T06:57:12.466427","exception":false,"start_time":"2021-07-18T06:57:08.262106","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-09-28T10:04:58.7837Z","iopub.execute_input":"2022-09-28T10:04:58.784083Z","iopub.status.idle":"2022-09-28T10:04:59.9772Z","shell.execute_reply.started":"2022-09-28T10:04:58.784049Z","shell.execute_reply":"2022-09-28T10:04:59.976063Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create a folder to store the extracted files\nif os.path.isdir('yolo_images_dir') == False:\n    yolo_images_dir = 'yolo_images_dir'\n    os.mkdir(yolo_images_dir)\n    \n    \n    \nval_fname_list = list(df_val['study_slice'])\n\nfor study_slice in val_fname_list:\n    \n    fname = study_slice + '.png'\n    \n    # Copy the image to images\n    # Set the path to the images here.\n    shutil.copyfile(\n        f\"{prep_data_path}/images_dir/{fname}\",\n        f\"yolo_images_dir/{fname}\")\n    \n    \nlen(os.listdir('yolo_images_dir'))","metadata":{"papermill":{"duration":5.227336,"end_time":"2021-07-18T06:57:28.383883","exception":false,"start_time":"2021-07-18T06:57:23.156547","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-09-28T10:05:11.138032Z","iopub.execute_input":"2022-09-28T10:05:11.138746Z","iopub.status.idle":"2022-09-28T10:05:14.587971Z","shell.execute_reply.started":"2022-09-28T10:05:11.138702Z","shell.execute_reply":"2022-09-28T10:05:14.587035Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!ls","metadata":{"papermill":{"duration":4.662857,"end_time":"2021-07-18T06:57:36.513146","exception":false,"start_time":"2021-07-18T06:57:31.850289","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-09-28T10:05:16.495291Z","iopub.execute_input":"2022-09-28T10:05:16.495664Z","iopub.status.idle":"2022-09-28T10:05:17.448283Z","shell.execute_reply.started":"2022-09-28T10:05:16.495633Z","shell.execute_reply":"2022-09-28T10:05:17.44693Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# change the working directory\nos.chdir('/kaggle/working/yolov5')\n\n!pwd","metadata":{"papermill":{"duration":4.277636,"end_time":"2021-07-18T06:57:44.58411","exception":false,"start_time":"2021-07-18T06:57:40.306474","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-09-28T10:05:17.450958Z","iopub.execute_input":"2022-09-28T10:05:17.45138Z","iopub.status.idle":"2022-09-28T10:05:18.434255Z","shell.execute_reply.started":"2022-09-28T10:05:17.451338Z","shell.execute_reply":"2022-09-28T10:05:18.433084Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Make a prediction on all images in images_dir\n\n# The model only creates a txt file if it finds objects on an image.\n\ntest_images_path = '/kaggle/working/yolo_images_dir'\nyolo_model_path = '/kaggle/working/best.pt'\n\n# Ensembling two Yolo models\n# How to ensemble Yolov5 models:\n# Ref: https://github.com/ultralytics/yolov5/issues/318\n\n!python detect.py --source $test_images_path --weights $yolo_model_path --img $IMAGE_SIZE --save-txt --save-conf --exist-ok","metadata":{"_kg_hide-output":true,"papermill":{"duration":41.782806,"end_time":"2021-07-18T06:58:29.775248","exception":false,"start_time":"2021-07-18T06:57:47.992442","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-09-28T10:05:20.347433Z","iopub.execute_input":"2022-09-28T10:05:20.347828Z","iopub.status.idle":"2022-09-28T10:06:27.671902Z","shell.execute_reply.started":"2022-09-28T10:05:20.347794Z","shell.execute_reply":"2022-09-28T10:06:27.670692Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Process the predictions","metadata":{"papermill":{"duration":3.642808,"end_time":"2021-07-18T06:58:37.370203","exception":false,"start_time":"2021-07-18T06:58:33.727395","status":"completed"},"tags":[]}},{"cell_type":"code","source":"txt_files_list = os.listdir('runs/detect/exp/labels')\n\nprint(len(txt_files_list))\nprint(txt_files_list[0])","metadata":{"papermill":{"duration":4.536142,"end_time":"2021-07-18T06:58:45.507384","exception":false,"start_time":"2021-07-18T06:58:40.971242","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-09-28T10:06:34.409203Z","iopub.execute_input":"2022-09-28T10:06:34.409599Z","iopub.status.idle":"2022-09-28T10:06:34.416812Z","shell.execute_reply.started":"2022-09-28T10:06:34.409564Z","shell.execute_reply":"2022-09-28T10:06:34.415768Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Put the info inside all the txt files into one dataframe.\n# Remember that if the image does not have any bounding boxes\n# then Yolo does not create a txt file for it.\n\ntxt_files_list = os.listdir('runs/detect/exp/labels')\n\nfor i, txt_file in enumerate(txt_files_list):\n    \n    # set the path\n    path = f'runs/detect/exp/labels/{txt_file}'\n    \n    # create a list of column names\n    cols = ['class', 'x-center', 'y-center', 'bbox_width', 'bbox_height', 'conf-score']\n\n    # put the file contents into a dataframe\n    df = pd.read_csv(path, sep=\" \", header=None)\n    \n    # add the column names to the datafrae\n    df.columns = cols\n    \n    # Split the txt fname on the full stop and choose the first item \n    # in the list. The add the .jpg extension.\n    # 87a0829f53c1.txt becomes 87a0829f53c1_image\n    #fname = txt_file.split('.')[0] + '.png'\n    fname = txt_file.replace(\"txt\", \"png\")\n    \n    # add a new column with the fname\n    df['id'] = fname\n \n    # stack the dataframes for each txt file\n    if i == 0:\n        \n        df_test_preds = df\n    else:\n        \n        df_test_preds = pd.concat([df_test_preds, df], axis=0)\n       \n    \n    \nprint(len(txt_files_list))\nprint(df_test_preds['id'].nunique())\nprint(df_test_preds.shape)\n\ndf_test_preds.head()","metadata":{"papermill":{"duration":5.758562,"end_time":"2021-07-18T06:58:54.922543","exception":false,"start_time":"2021-07-18T06:58:49.163981","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-09-28T10:06:36.095914Z","iopub.execute_input":"2022-09-28T10:06:36.097023Z","iopub.status.idle":"2022-09-28T10:06:36.126473Z","shell.execute_reply.started":"2022-09-28T10:06:36.096962Z","shell.execute_reply":"2022-09-28T10:06:36.125442Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Add the predictions to df_val\n\n# reset the index\ndf_val = df_val.reset_index(drop=True)\n\n# create a new column called 'id'\ndf_val['id'] =  df_val['study_slice'] + '.png'\n\nval_pred_list = []\n\npred_list = list(df_test_preds['id'])\n\nfor i in range(0, len(df_val)):\n    \n    #fname = df_val.loc[i, 'study_slice']\n    #fname = study_slice + '.png'\n    \n    # get the fname\n    fname = df_val.loc[i, 'id']\n    \n    # The fname will only be in the pred list if Yolo created a txt file for the val image.\n    # Yolo will only create a txt file if a fracture was detected on the image.\n    if fname in pred_list:\n        \n        val_pred_list.append(1)\n    else:\n        val_pred_list.append(0)\n    \n    \ndf_val['preds'] = val_pred_list\n\n# Check the distribution of the predicted classes\ndf_val['preds'].value_counts()","metadata":{"papermill":{"duration":3.732707,"end_time":"2021-07-18T06:59:02.462832","exception":false,"start_time":"2021-07-18T06:58:58.730125","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-09-28T10:06:41.053751Z","iopub.execute_input":"2022-09-28T10:06:41.054147Z","iopub.status.idle":"2022-09-28T10:06:41.087902Z","shell.execute_reply.started":"2022-09-28T10:06:41.054115Z","shell.execute_reply":"2022-09-28T10:06:41.086989Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Confusion Matrix","metadata":{"papermill":{"duration":3.706424,"end_time":"2021-07-18T06:59:10.094022","exception":false,"start_time":"2021-07-18T06:59:06.387598","status":"completed"},"tags":[]}},{"cell_type":"code","source":"from sklearn.metrics import confusion_matrix\n\nCLASS_LIST = ['Normal', 'Fracture']\n    \n# targets\ny_true = list(df_val['target'])\n\n# get the preds as integers\ny_pred = list(df_val['preds'])\n\n# argmax returns the index of the max value in each row.\ncm = confusion_matrix(y_true, y_pred)\n\n# Display the confusion matrix.\nprint()\nprint(cm)\nprint(CLASS_LIST)","metadata":{"papermill":{"duration":4.773196,"end_time":"2021-07-18T06:59:18.504453","exception":false,"start_time":"2021-07-18T06:59:13.731257","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-09-28T10:07:01.120535Z","iopub.execute_input":"2022-09-28T10:07:01.120937Z","iopub.status.idle":"2022-09-28T10:07:01.132629Z","shell.execute_reply.started":"2022-09-28T10:07:01.120904Z","shell.execute_reply":"2022-09-28T10:07:01.131457Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cm_plot_labels = ['normal', 'fracture']\n\n# Set the size of the plot.\n#plt.figure(figsize=(10,7))\n\n# Set the size of the text\ntext_size=12\n\nplot_confusion_matrix(cm, cm_plot_labels, title='Confusion Matrix', text_size=text_size)","metadata":{"papermill":{"duration":3.838768,"end_time":"2021-07-18T06:59:26.204345","exception":false,"start_time":"2021-07-18T06:59:22.365577","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-09-28T10:07:05.14786Z","iopub.execute_input":"2022-09-28T10:07:05.148845Z","iopub.status.idle":"2022-09-28T10:07:05.399167Z","shell.execute_reply.started":"2022-09-28T10:07:05.148808Z","shell.execute_reply":"2022-09-28T10:07:05.397974Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Classification Report","metadata":{"papermill":{"duration":3.649494,"end_time":"2021-07-18T06:59:33.878606","exception":false,"start_time":"2021-07-18T06:59:30.229112","status":"completed"},"tags":[]}},{"cell_type":"code","source":"from sklearn.metrics import classification_report\n    \nreport = classification_report(y_true, y_pred, target_names=cm_plot_labels)\n\nprint()\nprint(report)","metadata":{"papermill":{"duration":4.01076,"end_time":"2021-07-18T06:59:41.630028","exception":false,"start_time":"2021-07-18T06:59:37.619268","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-08-10T07:27:24.829435Z","iopub.execute_input":"2022-08-10T07:27:24.829834Z","iopub.status.idle":"2022-08-10T07:27:24.844658Z","shell.execute_reply.started":"2022-08-10T07:27:24.8298Z","shell.execute_reply":"2022-08-10T07:27:24.843536Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Precision\n\nGiven a prediction, what is the probability that the prediction is correct. In other words, how much can we trust a prediction made by the model?\n\n### Recall\n\nWhat percentage of the total number of objects did the model detect? Given an object, what is the probability that the model will detect it? Example: Did the computer vision model in the self driving car detect all the pedestrians?","metadata":{"papermill":{"duration":4.032807,"end_time":"2021-07-18T07:01:09.4492","exception":false,"start_time":"2021-07-18T07:01:05.416393","status":"completed"},"tags":[]}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Delete images to prevent notebook commit errors","metadata":{"papermill":{"duration":3.660432,"end_time":"2021-07-18T07:01:16.804923","exception":false,"start_time":"2021-07-18T07:01:13.144491","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# change the working directory to yolov5\n\nos.chdir('/kaggle/working/yolov5')\n\n!pwd","metadata":{"papermill":{"duration":4.440213,"end_time":"2021-07-18T07:01:25.213035","exception":false,"start_time":"2021-07-18T07:01:20.772822","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-08-10T07:15:03.191767Z","iopub.status.idle":"2022-08-10T07:15:03.192849Z","shell.execute_reply.started":"2022-08-10T07:15:03.192584Z","shell.execute_reply":"2022-08-10T07:15:03.192609Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Delete the folder to prevent a Kaggle error.\n\nif os.path.isdir('base_dir') == True:\n    shutil.rmtree('base_dir')","metadata":{"papermill":{"duration":3.868434,"end_time":"2021-07-18T07:01:33.690109","exception":false,"start_time":"2021-07-18T07:01:29.821675","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-08-10T07:15:03.194063Z","iopub.status.idle":"2022-08-10T07:15:03.194905Z","shell.execute_reply.started":"2022-08-10T07:15:03.194641Z","shell.execute_reply":"2022-08-10T07:15:03.194665Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"papermill":{"duration":3.962382,"end_time":"2021-07-18T07:01:41.32886","exception":false,"start_time":"2021-07-18T07:01:37.366478","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# change the working directory\n\nos.chdir('/kaggle/working/')\n\n!pwd","metadata":{"papermill":{"duration":4.532111,"end_time":"2021-07-18T07:01:49.531092","exception":false,"start_time":"2021-07-18T07:01:44.998981","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-08-10T07:15:03.196473Z","iopub.status.idle":"2022-08-10T07:15:03.197332Z","shell.execute_reply.started":"2022-08-10T07:15:03.197055Z","shell.execute_reply":"2022-08-10T07:15:03.197079Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Delete the folder to prevent a Kaggle error.\n\nif os.path.isdir('images_dir') == True:\n    shutil.rmtree('images_dir')\n    \nif os.path.isdir('val_images_dir') == True:\n    shutil.rmtree('val_images_dir')","metadata":{"papermill":{"duration":3.99797,"end_time":"2021-07-18T07:01:57.52992","exception":false,"start_time":"2021-07-18T07:01:53.53195","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-08-10T07:15:03.198931Z","iopub.status.idle":"2022-08-10T07:15:03.199779Z","shell.execute_reply.started":"2022-08-10T07:15:03.199505Z","shell.execute_reply":"2022-08-10T07:15:03.199531Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!ls","metadata":{"papermill":{"duration":4.459884,"end_time":"2021-07-18T07:02:06.52872","exception":false,"start_time":"2021-07-18T07:02:02.068836","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-08-10T07:15:03.201429Z","iopub.status.idle":"2022-08-10T07:15:03.202733Z","shell.execute_reply.started":"2022-08-10T07:15:03.202458Z","shell.execute_reply":"2022-08-10T07:15:03.202482Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"papermill":{"duration":3.959534,"end_time":"2021-07-18T07:02:14.134322","exception":false,"start_time":"2021-07-18T07:02:10.174788","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]}]}