{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<img src=\"https://i.imgur.com/3iLsS6i.png\">\n\n<center><h1> - Exploratory Phase - </h1></center>\n\n> 🦴 **Goal**: Detect and localize cervical spine fractures within CT scans.\n\n### Cervical Fractures (Broken Neck)\n\nThere are **7 bones that make up the cervical vertebrae** (neck). They support the head and connect it to the shoulders and body. A fracture, or break, in one of the cervical vertebrae is commonly called a broken neck ([src here](https://orthoinfo.aaos.org/en/diseases--conditions/cervical-fracture-broken-neck/#:~:text=A%20fracture%2C%20or%20break%2C%20in,Athletes%20are%20also%20at%20risk.)).\n\n<center><img src=\"https://i.imgur.com/hynMD8Z.png\" width=700></center>\n\n### ⬇ Libraries","metadata":{}},{"cell_type":"code","source":"!pip install -qU \"python-gdcm\" pydicom pylibjpeg \"opencv-python-headless\"","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-08-26T19:09:36.53018Z","iopub.execute_input":"2022-08-26T19:09:36.531698Z","iopub.status.idle":"2022-08-26T19:09:54.066451Z","shell.execute_reply.started":"2022-08-26T19:09:36.531569Z","shell.execute_reply":"2022-08-26T19:09:54.065075Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Libraries\nimport os\nimport re\nimport gc\nimport cv2\nimport wandb\nfrom PIL import Image\nimport random\nimport math\nimport shutil\nfrom glob import glob\nfrom tqdm import tqdm\nfrom pprint import pprint\nfrom time import time\nimport warnings\nimport itertools\nimport pandas as pd\nimport numpy as np\nimport seaborn as sns\nimport matplotlib as mpl\nfrom matplotlib import cm\nimport matplotlib.patches as patches\nimport matplotlib.pyplot as plt\nimport matplotlib.image as mpimg\nfrom matplotlib.offsetbox import AnnotationBbox, OffsetImage\nfrom matplotlib.colors import ListedColormap, LinearSegmentedColormap\nfrom matplotlib.patches import Rectangle\nfrom IPython.display import display_html\nplt.rcParams.update({'font.size': 16})\n\n# .dcm handling\nimport pydicom\nimport nibabel as nib\nfrom pydicom.pixel_data_handlers.util import apply_voi_lut\n\n# Environment check\nwarnings.filterwarnings(\"ignore\")\nos.environ[\"WANDB_SILENT\"] = \"true\"\nCONFIG = {'competition': 'RSNA_SpineFructure', '_wandb_kernel': 'aot'}\n\n# Custom colors\nclass clr:\n    S = '\\033[1m' + '\\033[94m'\n    E = '\\033[0m'\n    \nmy_colors = [\"#5EAFD9\", \"#449DD1\", \"#3977BB\", \n             \"#2D51A5\", \"#5C4C8F\", \"#8B4679\",\n             \"#C53D4C\", \"#E23836\", \"#FF4633\", \"#FF5746\"]\nCMAP1 = ListedColormap(my_colors)\n\nprint(clr.S+\"Notebook Color Schemes:\"+clr.E)\nsns.palplot(sns.color_palette(my_colors))\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-26T19:09:54.069663Z","iopub.execute_input":"2022-08-26T19:09:54.070182Z","iopub.status.idle":"2022-08-26T19:09:56.002573Z","shell.execute_reply.started":"2022-08-26T19:09:54.070129Z","shell.execute_reply":"2022-08-26T19:09:56.000751Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 🐝 W&B Fork & Run\n\nIn order to run this notebook you will need to input your own **secret API key** within the `! wandb login $secret_value_0` line. \n\n🐝**How do you get your own API key?**\n\nSuper simple! Go to **https://wandb.ai/site** -> Login -> Click on your profile in the top right corner -> Settings -> Scroll down to API keys -> copy your very own key (for more info check [this amazing notebook for ML Experiment Tracking on Kaggle](https://www.kaggle.com/ayuraj/experiment-tracking-with-weights-and-biases)).\n\n<center><img src=\"https://i.imgur.com/fFccmoS.png\" width=500></center>","metadata":{}},{"cell_type":"code","source":"# 🐝 Secrets\nfrom kaggle_secrets import UserSecretsClient\nuser_secrets = UserSecretsClient()\nsecret_value_0 = user_secrets.get_secret(\"wandb\")\n\n! wandb login $secret_value_0","metadata":{"execution":{"iopub.status.busy":"2022-08-26T19:09:56.005097Z","iopub.execute_input":"2022-08-26T19:09:56.006085Z","iopub.status.idle":"2022-08-26T19:09:59.04635Z","shell.execute_reply.started":"2022-08-26T19:09:56.006015Z","shell.execute_reply":"2022-08-26T19:09:59.045062Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### ⬇ Helper Functions","metadata":{}},{"cell_type":"code","source":"def show_values_on_bars(axs, h_v=\"v\", space=0.4):\n    '''Plots the value at the end of the a seaborn barplot.\n    axs: the ax of the plot\n    h_v: weather or not the barplot is vertical/ horizontal'''\n    \n    def _show_on_single_plot(ax):\n        if h_v == \"v\":\n            for p in ax.patches:\n                _x = p.get_x() + p.get_width() / 2\n                _y = p.get_y() + p.get_height()\n                value = int(p.get_height())\n                ax.text(_x, _y, format(value, ','), ha=\"center\") \n        elif h_v == \"h\":\n            for p in ax.patches:\n                _x = p.get_x() + p.get_width() + float(space)\n                _y = p.get_y() + p.get_height()\n                value = int(p.get_width())\n                ax.text(_x, _y, format(value, ','), ha=\"left\")\n\n    if isinstance(axs, np.ndarray):\n        for idx, ax in np.ndenumerate(axs):\n            _show_on_single_plot(ax)\n    else:\n        _show_on_single_plot(axs)\n        \n        \ndef atoi(text):\n    return int(text) if text.isdigit() else text\n\ndef natural_keys(text):\n    '''\n    alist.sort(key=natural_keys) sorts in human order\n    http://nedbatchelder.com/blog/200712/human_sorting.html\n    (See Toothy's implementation in the comments)\n    '''\n    return [ atoi(c) for c in re.split(r'(\\d+)', text) ]\n        \n        \n# === 🐝 W&B ===\ndef save_dataset_artifact(run_name, artifact_name, path):\n    '''Saves dataset to W&B Artifactory.\n    run_name: name of the experiment\n    artifact_name: under what name should the dataset be stored\n    path: path to the dataset'''\n    \n    run = wandb.init(project='RSNA_SpineFructure', \n                     name=run_name, \n                     config=CONFIG)\n    artifact = wandb.Artifact(name=artifact_name, \n                              type='dataset')\n    artifact.add_file(path)\n\n    wandb.log_artifact(artifact)\n    wandb.finish()\n    print(\"Artifact has been saved successfully.\")\n    \n    \ndef create_wandb_plot(x_data=None, y_data=None, x_name=None, y_name=None, title=None, log=None, plot=\"line\"):\n    '''Create and save lineplot/barplot in W&B Environment.\n    x_data & y_data: Pandas Series containing x & y data\n    x_name & y_name: strings containing axis names\n    title: title of the graph\n    log: string containing name of log'''\n    \n    data = [[label, val] for (label, val) in zip(x_data, y_data)]\n    table = wandb.Table(data=data, columns = [x_name, y_name])\n    \n    if plot == \"line\":\n        wandb.log({log : wandb.plot.line(table, x_name, y_name, title=title)})\n    elif plot == \"bar\":\n        wandb.log({log : wandb.plot.bar(table, x_name, y_name, title=title)})\n    elif plot == \"scatter\":\n        wandb.log({log : wandb.plot.scatter(table, x_name, y_name, title=title)})\n        \n        \ndef create_wandb_hist(x_data=None, x_name=None, title=None, log=None):\n    '''Create and save histogram in W&B Environment.\n    x_data: Pandas Series containing x values\n    x_name: strings containing axis name\n    title: title of the graph\n    log: string containing name of log'''\n    \n    data = [[x] for x in x_data]\n    table = wandb.Table(data=data, columns=[x_name])\n    wandb.log({log : wandb.plot.histogram(table, x_name, title=title)})\n    \n    \n# 🐝 Log Cover Photo\nrun = wandb.init(project='RSNA_SpineFructure', name='CoverPhoto', config=CONFIG)\ncover = plt.imread(\"../input/rsna-fracture-detection/Kaggle Covers (6).png\")\nwandb.log({\"example\": wandb.Image(cover)})\nwandb.finish()","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2022-08-26T19:09:59.050143Z","iopub.execute_input":"2022-08-26T19:09:59.050981Z","iopub.status.idle":"2022-08-26T19:10:09.116769Z","shell.execute_reply.started":"2022-08-26T19:09:59.050935Z","shell.execute_reply":"2022-08-26T19:10:09.115348Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 1. Tabular Data [.csv]\n\n#### 🦴 **To Note**:\n* **Train Dataframe**:\n    * `patient_overall` - whether or not there is at least 1 vertebrae that is fractured\n    * `C1 - C7` - which of the vertebraes (if any) are fractured\n    * 2019 unique IDs in total\n* **Train BBox Dataframe**:\n    * contains bbox coordinates with the exact location of the fractures\n* **Train Images folder**:\n    * 2019 folders containing the `.dcm` CT scans of the patient","metadata":{}},{"cell_type":"code","source":"run = wandb.init(project='RSNA_SpineFructure', name='tabular_explore', config=CONFIG)","metadata":{"execution":{"iopub.status.busy":"2022-08-26T19:10:09.118433Z","iopub.execute_input":"2022-08-26T19:10:09.118866Z","iopub.status.idle":"2022-08-26T19:10:12.477513Z","shell.execute_reply.started":"2022-08-26T19:10:09.118822Z","shell.execute_reply":"2022-08-26T19:10:12.475243Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"BASE = \"../input/rsna-2022-cervical-spine-fracture-detection\"","metadata":{"execution":{"iopub.status.busy":"2022-08-26T19:10:12.479635Z","iopub.execute_input":"2022-08-26T19:10:12.480853Z","iopub.status.idle":"2022-08-26T19:10:13.08691Z","shell.execute_reply.started":"2022-08-26T19:10:12.48079Z","shell.execute_reply":"2022-08-26T19:10:13.085566Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def read_data():\n    '''Reads in all .csv files.'''\n    \n    train = pd.read_csv(\"../input/rsna-2022-cervical-spine-fracture-detection/train.csv\")\n    train_bbox = pd.read_csv(\"../input/rsna-2022-cervical-spine-fracture-detection/train_bounding_boxes.csv\")\n    test = pd.read_csv(\"../input/rsna-2022-cervical-spine-fracture-detection/test.csv\")\n    ss = pd.read_csv(\"../input/rsna-2022-cervical-spine-fracture-detection/sample_submission.csv\")\n    \n    return train, train_bbox, test, ss\n\n\ndef get_csv_info(csv, name=\"Default\"):\n    '''Prints main information for the speciffied .csv file.'''\n    \n    print(clr.S+f\"=== {name} ===\"+clr.E)\n    print(clr.S+f\"Shape:\"+clr.E, csv.shape)\n    print(clr.S+f\"Missing Values:\"+clr.E, csv.isna().sum().sum(), \"total missing datapoints.\")\n    print(clr.S+\"Columns:\"+clr.E, list(csv.columns), \"\\n\")\n    \n    display_html(csv.head())\n    print(\"\\n\")","metadata":{"execution":{"iopub.status.busy":"2022-08-26T19:10:13.088582Z","iopub.execute_input":"2022-08-26T19:10:13.088981Z","iopub.status.idle":"2022-08-26T19:10:13.713766Z","shell.execute_reply.started":"2022-08-26T19:10:13.088944Z","shell.execute_reply":"2022-08-26T19:10:13.712688Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Read in the data\ntrain, train_bbox, test, ss = read_data()\n\n# Print useful information on it\nfor csv, name in zip([train, train_bbox, test, ss],\n                     [\"Train\", \"Train Bbox\", \"Test\", \"Sample Submission\"]):\n    get_csv_info(csv, name)","metadata":{"execution":{"iopub.status.busy":"2022-08-26T19:10:13.71554Z","iopub.execute_input":"2022-08-26T19:10:13.716543Z","iopub.status.idle":"2022-08-26T19:10:14.448141Z","shell.execute_reply.started":"2022-08-26T19:10:13.716495Z","shell.execute_reply":"2022-08-26T19:10:14.446056Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### I. Fractures Analysis\n\n🦴 **Main Takeaways**:\n* *class is balanced* - there is ~ 50%-50% ratio between patients with and without fractures\n* *C7* - this is the bone with the most frequent injuries\n* *C3* - this is the bone with the least number of injuries","metadata":{}},{"cell_type":"code","source":"dt = pd.melt(train, \n             id_vars=['StudyInstanceUID', 'patient_overall'],\n             var_name=\"Vertebra\",\n             value_name=\"Flag\")\n\n# Plot\nfig, (ax1, ax2) = plt.subplots(1, 2, figsize=(24, 12))\nfig.suptitle('Fractures - Main Analysis', \n             weight=\"bold\", size=25)\n\nsns.countplot(data=train, x=\"patient_overall\", ax=ax1, palette=[my_colors[0], my_colors[4]])\nshow_values_on_bars(ax1, h_v=\"v\", space=0.4)\nax1.set_title(\"Patient Overall Fracture Flag [frequency]\", weight=\"bold\", size=19)\nax1.set_xlabel(\"Fracture Flag\", size = 18, weight=\"bold\")\nax1.set_ylabel(\"\")\nax1.set_yticks([])\n\nhatches = itertools.cycle(['', '//'])\nfor i, bar in enumerate(ax1.patches):\n    hatch = next(hatches)\n    bar.set_hatch(hatch)\n\nsns.countplot(data=dt, x=\"Vertebra\", hue=\"Flag\", ax=ax2, palette=[my_colors[0], my_colors[4]])\nshow_values_on_bars(ax2, h_v=\"v\", space=0.4)\nax2.set_title(\"Fracture Flag on Vertebra Type [frequency]\", weight=\"bold\", size=19)\nax2.set_xlabel(\"Vertebra\", size = 18, weight=\"bold\")\nax2.set_ylabel(\"\")\nax2.set_yticks([])\n\nfor i, bar in enumerate(ax2.patches):\n    hatch = ''\n    if i in [7, 8, 9, 10, 11, 12, 13]:\n        hatch = '//'\n    bar.set_hatch(hatch)\n\nsns.despine(right=True, top=True, left=True);","metadata":{"execution":{"iopub.status.busy":"2022-08-26T19:10:14.450134Z","iopub.execute_input":"2022-08-26T19:10:14.451586Z","iopub.status.idle":"2022-08-26T19:10:15.651316Z","shell.execute_reply.started":"2022-08-26T19:10:14.451522Z","shell.execute_reply":"2022-08-26T19:10:15.650129Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 🐝 Log in barplot to Dashboard\ncreate_wandb_plot(x_data=train[\"patient_overall\"].value_counts().index,\n                  y_data=train[\"patient_overall\"].value_counts().values,\n                  x_name=\"Fracture Flag\", y_name=\"Frequency\",\n                  title=\"Patient Overall Fracture Flag\",\n                  log=\"bar_overall\", plot=\"bar\")","metadata":{"execution":{"iopub.status.busy":"2022-08-26T19:10:15.65598Z","iopub.execute_input":"2022-08-26T19:10:15.656356Z","iopub.status.idle":"2022-08-26T19:10:16.242122Z","shell.execute_reply.started":"2022-08-26T19:10:15.656324Z","shell.execute_reply":"2022-08-26T19:10:16.240998Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### II. Number of Injuries / Patient\n\n🦴 **Main Takeaways**:\n* *injuries in 1 bone* - out of the 961 cases with at least 1 fracture, more than half have an injury in only 1 bone\n* *injuries in multiple bones* - very few cases have injuries in 4 bones or more in the same time","metadata":{}},{"cell_type":"code","source":"train[\"total_fractures\"] = train.iloc[:, 2:].sum(axis=1)\n\n# Plot\nplt.figure(figsize=(24, 12))\naxs = sns.countplot(data=train, x=\"total_fractures\", palette=my_colors)\nshow_values_on_bars(axs, h_v=\"v\", space=0.4)\nplt.title(\"Number of Bones Fractured per Patient\", weight=\"bold\", size=19)\nplt.xlabel(\"Total Bones Fractured\", size = 18, weight=\"bold\")\nplt.ylabel(\"\")\nplt.yticks([])\n\n# Hatch\nfor i, bar in enumerate(axs.patches):\n    hatch = ''\n    if i==0:\n        hatch = '-'\n    bar.set_hatch(hatch)\n\n# Arrow\nstyle = \"Simple, tail_width=2, head_width=14, head_length=16\"\nkw = dict(arrowstyle=style, color=my_colors[2])\narrow = patches.FancyArrowPatch((1.8, 900), (1, 660),\n                             connectionstyle=\"arc3,rad=.10\", **kw)\nplt.gca().add_patch(arrow)\nplt.text(x=2, y=900, s=f\"{round(623/961*100,2)}% cases have\",\n         color=my_colors[2], size=17, weight=\"bold\")\nplt.text(x=2, y=860, s=f\"only 1 bone fractured\",\n         color=my_colors[2], size=17, weight=\"bold\")\nplt.text(x=2, y=820, s=f\"in the same time.\",\n         color=my_colors[2], size=17, weight=\"bold\")\n    \n# Line    \nplt.axvline(x=3.5, linestyle = '--', color=my_colors[5], lw=2)\nplt.text(x=3.6, y=250, s=f\"{round(35/961*100,2)}% cases have 4+ bones fractured in the same time.\",\n         color=my_colors[5], size=17, weight=\"bold\")\n\nsns.despine(right=True, top=True, left=True);","metadata":{"execution":{"iopub.status.busy":"2022-08-26T19:10:16.243789Z","iopub.execute_input":"2022-08-26T19:10:16.244143Z","iopub.status.idle":"2022-08-26T19:10:17.369351Z","shell.execute_reply.started":"2022-08-26T19:10:16.244111Z","shell.execute_reply":"2022-08-26T19:10:17.368053Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 🐝 Log in barplot to Dashboard\ncreate_wandb_plot(x_data=train[\"total_fractures\"].value_counts().index,\n                  y_data=train[\"total_fractures\"].value_counts().values,\n                  x_name=\"Fracture Bones Fractured\", y_name=\"Frequency\",\n                  title=\"Number of Bones Fractured per Patient\",\n                  log=\"bar_bones\", plot=\"bar\")","metadata":{"execution":{"iopub.status.busy":"2022-08-26T19:10:17.370707Z","iopub.execute_input":"2022-08-26T19:10:17.371078Z","iopub.status.idle":"2022-08-26T19:10:17.977795Z","shell.execute_reply.started":"2022-08-26T19:10:17.371044Z","shell.execute_reply":"2022-08-26T19:10:17.97657Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 🐝 Finish this Experiment\nwandb.finish()","metadata":{"execution":{"iopub.status.busy":"2022-08-26T19:10:17.979428Z","iopub.execute_input":"2022-08-26T19:10:17.980249Z","iopub.status.idle":"2022-08-26T19:10:27.925404Z","shell.execute_reply.started":"2022-08-26T19:10:17.980197Z","shell.execute_reply":"2022-08-26T19:10:27.924265Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 2. Image Data [.dcm]\n\n## 2.1 Images Understanding\n\n🦴 **Steps**: Let's first look at a few examples to understand how these images look. Afterwards, we can further extract some details from the `.dcm` datasets and explore general features (like image size, number of slices per patient etc.).","metadata":{}},{"cell_type":"code","source":"run = wandb.init(project='RSNA_SpineFructure', name='image_explore', config=CONFIG)","metadata":{"execution":{"iopub.status.busy":"2022-08-26T19:10:27.926925Z","iopub.execute_input":"2022-08-26T19:10:27.927392Z","iopub.status.idle":"2022-08-26T19:10:31.366923Z","shell.execute_reply.started":"2022-08-26T19:10:27.927333Z","shell.execute_reply":"2022-08-26T19:10:31.365531Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def show_dcm_images(patient_id, rows=4, cols=8, random=0):\n    '''Show .dcm images based on id.'''\n    \n    N = rows*cols\n    wandb_logs = []\n\n    fig, axes = plt.subplots(nrows=rows, ncols=cols, figsize=(24,15))\n    fig.suptitle(f'ID: {patient_id}', weight=\"bold\", size=20)\n\n    # Get .dcm paths\n    dcm_paths = glob(f\"{BASE}/train_images/{patient_id}/*\")\n    print(clr.S+\"Number of TOTAL Slices:\"+clr.E, len(dcm_paths))\n    dcm_paths.sort(key=natural_keys)\n    dcm_paths = dcm_paths[random:(random+N)]\n    # Get corresponging datasets and images\n    datasets = [pydicom.dcmread(path) for path in dcm_paths]\n    images = [apply_voi_lut(dataset.pixel_array, dataset) for dataset in datasets]\n\n    # Loop through the information\n    for data, img, i in zip(datasets, images, range(N)):\n        slice_no = data.SOPInstanceUID.split(\".\")[-1]\n\n        # Plot the image\n        x = i // cols\n        y = i % cols\n\n        axes[x, y].imshow(img, cmap=\"bone\")\n        axes[x, y].set_title(f\"Slice: {slice_no}\", \n                  fontsize=14, weight='bold')\n        axes[x, y].axis('off');\n        \n        # 🐝 Create Image\n        wandb_logs.append(wandb.Image(img, \n                                      caption=f\"Slice: {slice_no}\"))\n        \n    # 🐝 Log images into Dashboard\n    wandb.log({f\"{data.SOPInstanceUID}\": wandb_logs})","metadata":{"execution":{"iopub.status.busy":"2022-08-26T19:10:31.369202Z","iopub.execute_input":"2022-08-26T19:10:31.369596Z","iopub.status.idle":"2022-08-26T19:10:31.994813Z","shell.execute_reply.started":"2022-08-26T19:10:31.369539Z","shell.execute_reply":"2022-08-26T19:10:31.993331Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Example 1","metadata":{}},{"cell_type":"code","source":"show_dcm_images(\"1.2.826.0.1.3680043.10001\",\n                rows=4, cols=8)","metadata":{"execution":{"iopub.status.busy":"2022-08-26T19:10:31.996622Z","iopub.execute_input":"2022-08-26T19:10:31.996999Z","iopub.status.idle":"2022-08-26T19:10:37.149057Z","shell.execute_reply.started":"2022-08-26T19:10:31.996961Z","shell.execute_reply":"2022-08-26T19:10:37.14776Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Example 2","metadata":{}},{"cell_type":"code","source":"show_dcm_images(\"1.2.826.0.1.3680043.14994\", \n                rows=4, cols=8, random=23)","metadata":{"execution":{"iopub.status.busy":"2022-08-26T19:10:37.150934Z","iopub.execute_input":"2022-08-26T19:10:37.151321Z","iopub.status.idle":"2022-08-26T19:10:42.29251Z","shell.execute_reply.started":"2022-08-26T19:10:37.151282Z","shell.execute_reply":"2022-08-26T19:10:42.291267Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 🐝 Finish Experiment\nwandb.finish()","metadata":{"execution":{"iopub.status.busy":"2022-08-26T19:10:42.294218Z","iopub.execute_input":"2022-08-26T19:10:42.294821Z","iopub.status.idle":"2022-08-26T19:10:47.89089Z","shell.execute_reply.started":"2022-08-26T19:10:42.294784Z","shell.execute_reply":"2022-08-26T19:10:47.889961Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 2.2 Saving the images\n\nThe `get_images()` function reads all the `.dcm` files for each Study Instance and saves the corresponding images as `.png` files. \n\n🦴 **File Saving**: Each saved file is named as follows: `StudyInstanceUID/InstanceNumber.png`\n* e.g.: `1.2.826.0.1.3680043.53/4.png`\n    * `StudyInstanceUID` -> 1.2.826.0.1.3680043.53\n    * `InstanceNumber` -> 4 (slice number) in `.png` format\n    \n*! **Update**: I've made a mistake and kept only the images that contain a bbox and excluded all the .dcm files that didn't have any fractures - I've rectified this mistake in V11 of this notebook.*","metadata":{}},{"cell_type":"code","source":"def get_images():\n    '''\n    Retrieves and saves all .dcm images into a zip file.\n    '''\n\n    for k in tqdm(range(len(train))):\n        \n        dt = train.iloc[k, :]\n        os.mkdir(\"../working/png_images/\" + str(dt.StudyInstanceUID))\n\n        # Get all .dcm paths for this Instance\n        dcm_paths = glob(f\"{BASE}/train_images/{dt.StudyInstanceUID}/*\")\n\n        for path in dcm_paths:\n            # Get datasets\n            dataset = pydicom.dcmread(path)\n\n            # Get images\n            image = apply_voi_lut(dataset.pixel_array, dataset)\n            cv2.imwrite(f\"../working/png_images/{dataset.StudyInstanceUID}/{dataset.InstanceNumber}.png\", image)\n            \n    print(\"PNG Images Saved.\")","metadata":{"execution":{"iopub.status.busy":"2022-08-26T19:10:47.892737Z","iopub.execute_input":"2022-08-26T19:10:47.893285Z","iopub.status.idle":"2022-08-26T19:10:47.900616Z","shell.execute_reply.started":"2022-08-26T19:10:47.893244Z","shell.execute_reply":"2022-08-26T19:10:47.899377Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> 🦴**Note**: For the purpose of this notebook running faster, I am commenting the cell below (that creates, saves and archives the `.png` images). You can find the full dataset [here](https://www.kaggle.com/datasets/andradaolteanu/rsna-fracture-detection).","metadata":{}},{"cell_type":"code","source":"# # === Uncomment cell to run ===\n\n# # Create file to store the images\n# ! rm -rf png_images\n# ! mkdir png_images\n\n# # Save images\n# get_images()\n\n# # Zip the png folder and remove original\n# zip_name = './zip_png_images/'\n# directory_name = './png_images/'\n\n# shutil.make_archive(zip_name, 'zip', directory_name)\n# ! rm -rf png_images","metadata":{"execution":{"iopub.status.busy":"2022-08-26T19:10:47.90233Z","iopub.execute_input":"2022-08-26T19:10:47.902863Z","iopub.status.idle":"2022-08-26T19:10:47.918058Z","shell.execute_reply.started":"2022-08-26T19:10:47.902807Z","shell.execute_reply":"2022-08-26T19:10:47.917018Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"total_files = len(glob(\"../input/rsna-fracture-detection/zip_png_images/*\"))\nprint(clr.S+\"Total .png files downloaded and saved:\"+clr.E, total_files)","metadata":{"execution":{"iopub.status.busy":"2022-08-26T19:10:47.919754Z","iopub.execute_input":"2022-08-26T19:10:47.920418Z","iopub.status.idle":"2022-08-26T19:10:48.044385Z","shell.execute_reply.started":"2022-08-26T19:10:47.920381Z","shell.execute_reply":"2022-08-26T19:10:48.043509Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 3. DICOM Metadata\n\n## 3.1 Retrieving Metadata for 1 file\n\n🦴 At the moment, the only information I would like to retrieve is the following:\n* Rows -> the height of the CT scan/image\n* Columns -> the width of the CT scan/image\n* SOPInstanceUID -> Unique identifier containing the `StudyInstanceUID` + slice number\n* ContentDate -> the date the image pixel data creation started\n* SliceThickness -> gives the thickness of the imaged slice (*TODO: maybe pair with `Spacing Between Slices` - gives the distance between two adjacent slices*)\n* InstanceNumber -> slice number\n* ImagePositionPatient -> the x, y, and z coordinates of the upper left hand corner (center of the first voxel transmitted) of the image, in mm\n* ImageOrientationPatient -> the direction cosines of the first row and the first column with respect to the patient\n\n> 🦴 **Note**: all attribute explanations [are here](https://dicom.innolitics.com/ciods/rt-dose/image-plane/00200037).","metadata":{}},{"cell_type":"code","source":"def get_observation_data(path):\n    '''\n    Get information from the .dcm files\n    '''\n\n    dataset = pydicom.read_file(path)\n    \n    # Dictionary to store the information from the image\n    observation_data = {\n        \"Rows\" : dataset.get(\"Rows\"),\n        \"Columns\" : dataset.get(\"Columns\"),\n        \"SOPInstanceUID\" : dataset.get(\"SOPInstanceUID\"),\n        \"ContentDate\" : dataset.get(\"ContentDate\"),\n        \"SliceThickness\" : dataset.get(\"SliceThickness\"),\n        \"InstanceNumber\" : dataset.get(\"InstanceNumber\"),\n        \"ImagePositionPatient\" : dataset.get(\"ImagePositionPatient\"),\n        \"ImageOrientationPatient\" : dataset.get(\"ImageOrientationPatient\"),\n    }\n\n    # String columns\n    str_columns = [\"SOPInstanceUID\", \"ContentDate\", \n                   \"SliceThickness\", \"InstanceNumber\"]\n    for k in str_columns:\n        observation_data[k] = str(dataset.get(k)) if k in dataset else None\n\n    \n    return observation_data","metadata":{"execution":{"iopub.status.busy":"2022-08-26T19:10:48.046312Z","iopub.execute_input":"2022-08-26T19:10:48.047196Z","iopub.status.idle":"2022-08-26T19:10:48.056581Z","shell.execute_reply.started":"2022-08-26T19:10:48.047145Z","shell.execute_reply":"2022-08-26T19:10:48.05526Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# An example\npath = \"../input/rsna-2022-cervical-spine-fracture-detection/train_images/1.2.826.0.1.3680043.10001/109.dcm\"\nexample = get_observation_data(path)\npprint(example)","metadata":{"execution":{"iopub.status.busy":"2022-08-26T19:10:48.059061Z","iopub.execute_input":"2022-08-26T19:10:48.059458Z","iopub.status.idle":"2022-08-26T19:10:48.089323Z","shell.execute_reply.started":"2022-08-26T19:10:48.059412Z","shell.execute_reply":"2022-08-26T19:10:48.088531Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 3.2 Create Metadata file","metadata":{}},{"cell_type":"code","source":"def get_metadata():\n    '''\n    Retrieves the desired metadata from the .dcm files and saves it into dataframe.\n    '''\n    \n    exceptions = 0\n    dicts = []\n\n    for k in tqdm(range(len(train))):\n        dt = train.iloc[k, :]\n\n        if dt.patient_overall == 1:\n            # Get all .dcm paths for this Instance\n            dcm_paths = glob(f\"{BASE}/train_images/{dt.StudyInstanceUID}/*\")\n\n            for path in dcm_paths:\n                try:\n                    # Get datasets\n                    dataset = get_observation_data(path)\n                    dicts.append(dataset)\n                except Exception as e:\n                    exceptions += 1\n                    continue\n                    \n    # Convert into df\n    meta_train_data = pd.DataFrame(data=dicts, columns=example.keys())\n    # Export information\n    meta_train_data.to_csv(\"meta_train.csv\", index=False)\n            \n    print(f\"Metadata created. Number of total fails: {exceptions}.\")","metadata":{"execution":{"iopub.status.busy":"2022-08-26T19:10:48.090685Z","iopub.execute_input":"2022-08-26T19:10:48.091187Z","iopub.status.idle":"2022-08-26T19:10:48.099026Z","shell.execute_reply.started":"2022-08-26T19:10:48.091154Z","shell.execute_reply":"2022-08-26T19:10:48.098046Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> 🦴 **Note**: All data is available to download [here](https://www.kaggle.com/datasets/andradaolteanu/rsna-fracture-detection).","metadata":{}},{"cell_type":"code","source":"# Create and save the metadata\n# This cell takes ~ 1 hour to run\n# get_metadata()","metadata":{"execution":{"iopub.status.busy":"2022-08-26T19:10:48.100707Z","iopub.execute_input":"2022-08-26T19:10:48.101412Z","iopub.status.idle":"2022-08-26T19:10:48.114572Z","shell.execute_reply.started":"2022-08-26T19:10:48.101366Z","shell.execute_reply":"2022-08-26T19:10:48.11353Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 3.3 Metadata Explore\n\nLet's play a bit with the data we have just created 😁","metadata":{}},{"cell_type":"code","source":"run = wandb.init(project='RSNA_SpineFructure', name='meta_explore', config=CONFIG)","metadata":{"execution":{"iopub.status.busy":"2022-08-26T19:10:48.11597Z","iopub.execute_input":"2022-08-26T19:10:48.116577Z","iopub.status.idle":"2022-08-26T19:10:51.877222Z","shell.execute_reply.started":"2022-08-26T19:10:48.116543Z","shell.execute_reply":"2022-08-26T19:10:51.875528Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Read in saved metadata\nmeta_train = pd.read_csv(\"../input/rsna-fracture-detection/meta_train.csv\")\nmeta_train[\"StudyInstanceUID\"] = meta_train[\"SOPInstanceUID\"].apply(lambda x: \".\".join(x.split(\".\")[:-2]))\n\n# Information\nprint(clr.S+\"Total number of .dcm files: (CT scans)\"+clr.E, len(meta_train))\nprint(clr.S+\"[sanity check] Total Number of Study Instances\"+clr.E, meta_train[\"StudyInstanceUID\"].nunique())\n\n# 🐝 Log info into Dashboard\nwandb.log({\"total_dcm_files\": len(meta_train),\n           \"total_study_instances\": meta_train[\"StudyInstanceUID\"].nunique()})","metadata":{"execution":{"iopub.status.busy":"2022-08-26T19:10:51.87937Z","iopub.execute_input":"2022-08-26T19:10:51.881043Z","iopub.status.idle":"2022-08-26T19:10:54.030297Z","shell.execute_reply.started":"2022-08-26T19:10:51.88096Z","shell.execute_reply":"2022-08-26T19:10:54.028532Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### I. Distribution of .dcm files on Study Instance\n\n🦴 **To Note**:\n* The distribution is skewed to the right\n* 78% of all observations 200 to 400 slices\n* Only 6% of study instances have less thnan 200 slices\n* There are 50 Study Instances with more than 600 slices","metadata":{}},{"cell_type":"code","source":"# Get the data\ndata = meta_train[\"StudyInstanceUID\"].value_counts().reset_index()\ndata.columns = [\"StudyInstanceUID\", \"count\"]","metadata":{"execution":{"iopub.status.busy":"2022-08-26T19:10:54.038558Z","iopub.execute_input":"2022-08-26T19:10:54.039577Z","iopub.status.idle":"2022-08-26T19:10:54.810277Z","shell.execute_reply.started":"2022-08-26T19:10:54.039517Z","shell.execute_reply":"2022-08-26T19:10:54.809316Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Plot 1\nplt.figure(figsize=(24, 12))\nsns.distplot(data[\"count\"], rug=True, bins=10,\n             rug_kws={\"color\": my_colors[1]},\n             kde_kws={\"color\": my_colors[8], \"lw\": 5, \"alpha\": 0.7},\n             hist_kws={\"histtype\": \"step\", \"linewidth\": 3, \"alpha\": 1, \"color\": my_colors[1]}\n            )\n\nplt.title(\"Distribution of .dcm files on Study Instance\", weight=\"bold\", size=25)\nplt.xlabel(\"Number of Slices\", size = 18, weight=\"bold\")\nplt.ylabel(\"Frequency\")\n\n# Arrow\nstyle = \"Simple, tail_width=1, head_width=12, head_length=14\"\nkw = dict(arrowstyle=style, color=\"black\")\narrow = patches.FancyArrowPatch((600, 0.0033), (300, 0.0020),\n                             connectionstyle=\"arc3,rad=-.10\", **kw)\nplt.gca().add_patch(arrow)\n\nplt.text(x=400, y=0.0035, s=f\"78% of observations have 200 to 400 slices.\", \n         color=\"black\", size=17)\n\nsns.despine(right=True, top=True, left=True);","metadata":{"execution":{"iopub.status.busy":"2022-08-26T19:10:54.81173Z","iopub.execute_input":"2022-08-26T19:10:54.812363Z","iopub.status.idle":"2022-08-26T19:10:55.908854Z","shell.execute_reply.started":"2022-08-26T19:10:54.812321Z","shell.execute_reply":"2022-08-26T19:10:55.907448Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 🐝 Log in histogram to Dashboard\ncreate_wandb_hist(x_data=data[\"count\"], \n                  x_name=\"Number of Slices\",\n                  title=\"Distribution of .dcm files on Study Instance\",\n                  log=\"hist_slices\")","metadata":{"execution":{"iopub.status.busy":"2022-08-26T19:10:55.910667Z","iopub.execute_input":"2022-08-26T19:10:55.91139Z","iopub.status.idle":"2022-08-26T19:10:56.8038Z","shell.execute_reply.started":"2022-08-26T19:10:55.911319Z","shell.execute_reply":"2022-08-26T19:10:56.802529Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"random.seed(25)\n\n# Add a fictive y\n# this \"y\" doesn't mean anything, it's just for\n# showcasing purposes\ndata[\"y\"] = [random.randint(0, 100) for i in range(len(data))]\nperc = round(data[data[\"count\"]<=10].shape[0]/len(data), 3)*100\n\nplt.figure(figsize=(24, 12))\nsns.scatterplot(data=data, x=\"count\", y=\"y\", size=\"count\", alpha=0.65, sizes=(500, 2000),\n               hue=\"count\", palette=CMAP1)\n\nplt.title(\"Distribution of .dcm files on Study Instance\", weight=\"bold\", size=25)\nplt.xlabel(\"Number of Slices\", size = 18, weight=\"bold\")\nplt.ylabel(\"\")\nplt.yticks([])\n\nplt.axvline(x=200, linestyle = '--', color=\"black\", lw=4)\nplt.axvline(x=400, linestyle = '--', color=\"black\", lw=4)\nplt.text(x=225, y=50, s=f\"~80% of data is here\", color=\"black\", size=20, weight=\"bold\")\n\nplt.axvline(x=600, linestyle = '--', color=\"black\", lw=4)\nplt.text(x=610, y=7, s=f\"only 50 instances with slices >=600\", color=\"black\", size=14, weight=\"bold\")\nplt.arrow(x=600, y=5, dx=200, dy=0, color=\"black\", lw=4, \n          head_width=2, head_length=8)\n\nplt.legend('',frameon=False)\nsns.despine(right=True, top=True, left=True);","metadata":{"execution":{"iopub.status.busy":"2022-08-26T19:10:56.805347Z","iopub.execute_input":"2022-08-26T19:10:56.805731Z","iopub.status.idle":"2022-08-26T19:10:58.050162Z","shell.execute_reply.started":"2022-08-26T19:10:56.805694Z","shell.execute_reply":"2022-08-26T19:10:58.049049Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### II. Image sizes\n\n🦴 **Note**:\n* 99.6% of images have a fixed size of 512 by 512.\n* The rest of the slices (the other 0.3%) should be resized to 512x512 as well.","metadata":{}},{"cell_type":"code","source":"# Create column with full image size (height x width)\nmeta_train[\"ImageSize\"] = meta_train[\"Rows\"].astype(str) + \" x \" + meta_train[\"Columns\"].astype(str)\n\n\n# Plot\nplt.figure(figsize=(24, 10))\naxs = sns.countplot(data=meta_train, x=\"ImageSize\", palette=my_colors[5:])\nshow_values_on_bars(axs, h_v=\"v\", space=0.4)\nplt.title(\"Frequency of Image Sizes in .dcm data\", weight=\"bold\", size=19)\nplt.xlabel(\"Image Size (height x width)\", size = 18, weight=\"bold\")\nplt.ylabel(\"\")\nplt.yticks([])\n\n# Hatch\nfor i, bar in enumerate(axs.patches):\n    hatch = ''\n    if i==0:\n        hatch = '/'\n    bar.set_hatch(hatch)\n\n# Arrow\nstyle = \"Simple, tail_width=2, head_width=14, head_length=16\"\nkw = dict(arrowstyle=style, color=my_colors[4])\narrow = patches.FancyArrowPatch((0.8, 300000), (0.2, 200000),\n                             connectionstyle=\"arc3,rad=.10\", **kw)\nplt.gca().add_patch(arrow)\nplt.text(x=0.8, y=300000, s=f\"99.6% of images have size 512x512\",\n         color=my_colors[4], size=17, weight=\"bold\")\n\nplt.text(x=1, y=70000, s=f\"The rest 0.32% should be resized too as 512x512\",\n         color=my_colors[6], size=17, weight=\"bold\")\nplt.axhline(xmin=0.4, xmax=0.95, y=60000, linestyle = '--', color=my_colors[6], lw=2)\n\nsns.despine(right=True, top=True, left=True);","metadata":{"execution":{"iopub.status.busy":"2022-08-26T19:10:58.052126Z","iopub.execute_input":"2022-08-26T19:10:58.052533Z","iopub.status.idle":"2022-08-26T19:10:59.85315Z","shell.execute_reply.started":"2022-08-26T19:10:58.052462Z","shell.execute_reply":"2022-08-26T19:10:59.852149Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 🐝 Log in barplot to Dashboard\ncreate_wandb_plot(x_data=meta_train[\"ImageSize\"].value_counts().index,\n                  y_data=meta_train[\"ImageSize\"].value_counts().values,\n                  x_name=\"Image Size (height x width)\", y_name=\"Frequency\",\n                  title=\"Frequency of Image Sizes in .dcm data\",\n                  log=\"bar_imgsize\", plot=\"bar\")","metadata":{"execution":{"iopub.status.busy":"2022-08-26T19:10:59.855515Z","iopub.execute_input":"2022-08-26T19:10:59.856072Z","iopub.status.idle":"2022-08-26T19:11:00.742919Z","shell.execute_reply.started":"2022-08-26T19:10:59.856038Z","shell.execute_reply":"2022-08-26T19:11:00.741717Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### III. Content Date\n\n> 🦴 **Note**: All .dcm files have the same creation date (2022-07-27), so this column will be dropped.","metadata":{}},{"cell_type":"code","source":"print(clr.S+\"The *only* ContentDate is:\"+clr.E, meta_train[\"ContentDate\"].value_counts().index)\n\nmeta_train.drop(columns=\"ContentDate\", axis=1, inplace=True)\nprint(clr.S+\"Column ContentDate was dropped.\"+clr.E)","metadata":{"execution":{"iopub.status.busy":"2022-08-26T19:11:00.744341Z","iopub.execute_input":"2022-08-26T19:11:00.744729Z","iopub.status.idle":"2022-08-26T19:11:01.658435Z","shell.execute_reply.started":"2022-08-26T19:11:00.744694Z","shell.execute_reply":"2022-08-26T19:11:01.657229Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"wandb.finish()","metadata":{"execution":{"iopub.status.busy":"2022-08-26T19:11:01.660221Z","iopub.execute_input":"2022-08-26T19:11:01.660985Z","iopub.status.idle":"2022-08-26T19:11:12.017873Z","shell.execute_reply.started":"2022-08-26T19:11:01.660934Z","shell.execute_reply":"2022-08-26T19:11:12.016912Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 4. Bounding Boxes\n\nLet's now analyse **only the slices/images with bounding boxes** in them (aka the slices that contain 1 or more ruptures.","metadata":{}},{"cell_type":"code","source":"run = wandb.init(project='RSNA_SpineFructure', name='bbox_explore', config=CONFIG)","metadata":{"execution":{"iopub.status.busy":"2022-08-26T19:11:12.01912Z","iopub.execute_input":"2022-08-26T19:11:12.019445Z","iopub.status.idle":"2022-08-26T19:11:16.232355Z","shell.execute_reply.started":"2022-08-26T19:11:12.019414Z","shell.execute_reply":"2022-08-26T19:11:16.231306Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(clr.S+\"Number of Study Instances that contain bounding boxes:\"+clr.E,\n      train_bbox[\"StudyInstanceUID\"].nunique())","metadata":{"execution":{"iopub.status.busy":"2022-08-26T19:11:16.23411Z","iopub.execute_input":"2022-08-26T19:11:16.23556Z","iopub.status.idle":"2022-08-26T19:11:17.364775Z","shell.execute_reply.started":"2022-08-26T19:11:16.235512Z","shell.execute_reply":"2022-08-26T19:11:17.363435Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 4.1 Number of Bounding Boxes per Study Instance\n\n🦴 **Note**:\n* ~50% of all slices have about 25 or less bboxes per study.\n    * ! ~16% of slices have less than 10 bboxes per study instance.\n* there are 6 studies that are outliers - with more than 100 bboxes per study.\n\n*! Keep in mind: number of bboxes == number of fractures*.","metadata":{}},{"cell_type":"code","source":"# Plot 1\nplt.figure(figsize=(24, 12))\nsns.distplot(train_bbox[\"StudyInstanceUID\"].value_counts().values, rug=True, bins=20,\n             rug_kws={\"color\": my_colors[1]},\n             kde_kws={\"color\": my_colors[8], \"lw\": 5, \"alpha\": 0.7},\n             hist_kws={\"histtype\": \"step\", \"linewidth\": 3, \"alpha\": 1, \"color\": my_colors[1]}\n            )\n\nplt.title(\"Number of Bounding Boxes per Study Instance\", weight=\"bold\", size=25)\nplt.xlabel(\"Number of BBoxes / Study Instance\", size = 18, weight=\"bold\")\nplt.ylabel(\"Frequency\")\n\n# Arrow\nstyle = \"Simple, tail_width=1, head_width=12, head_length=14\"\nkw = dict(arrowstyle=style, color=\"black\")\narrow = patches.FancyArrowPatch((70, 0.022), (20, 0.018),\n                             connectionstyle=\"arc3,rad=.10\", **kw)\nplt.gca().add_patch(arrow)\nplt.text(x=70, y=0.022, s=f\"~50% of slices have less than 25 bboxes/study.\", \n         color=\"black\", size=17)\n\nstyle = \"Simple, tail_width=1, head_width=12, head_length=14\"\nkw = dict(arrowstyle=style, color=\"black\")\narrow = patches.FancyArrowPatch((160, 0.008), (150, 0.001),\n                             connectionstyle=\"arc3,rad=.10\", **kw)\nplt.gca().add_patch(arrow)\nplt.text(x=130, y=0.0085, s=f\"Only 6 slices have 100+ bboxes/study.\", \n         color=\"black\", size=17)\n\nsns.despine(right=True, top=True, left=True);","metadata":{"execution":{"iopub.status.busy":"2022-08-26T19:11:17.367411Z","iopub.execute_input":"2022-08-26T19:11:17.367916Z","iopub.status.idle":"2022-08-26T19:11:18.66593Z","shell.execute_reply.started":"2022-08-26T19:11:17.367863Z","shell.execute_reply":"2022-08-26T19:11:18.664721Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 🐝 Log in histogram to Dashboard\ncreate_wandb_hist(x_data=train_bbox[\"StudyInstanceUID\"].value_counts().values, \n                  x_name=\"Number of BBoxes / Study Instance\",\n                  title=\"Number of Bounding Boxes per Study Instance\",\n                  log=\"hist_bboxes\")","metadata":{"execution":{"iopub.status.busy":"2022-08-26T19:11:18.669583Z","iopub.execute_input":"2022-08-26T19:11:18.670884Z","iopub.status.idle":"2022-08-26T19:11:19.737626Z","shell.execute_reply.started":"2022-08-26T19:11:18.670827Z","shell.execute_reply":"2022-08-26T19:11:19.736727Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 4.2 BBoxes on Images\n\nLet's now plot some bounding boxes on out CT scans.","metadata":{}},{"cell_type":"code","source":"# Compute x2 and y2\ntrain_bbox[\"x2\"] = train_bbox[\"x\"] + train_bbox[\"width\"]\ntrain_bbox[\"y2\"] = train_bbox[\"y\"] + train_bbox[\"height\"]\n\n# Rename columns\ntrain_bbox.rename(columns={\"x\": \"x1\", \"y\": \"y1\"}, inplace=True)\n\n# Change to int\ntrain_bbox[\"x1\"] = train_bbox[\"x1\"].apply(lambda x: int(x))\ntrain_bbox[\"x2\"] = train_bbox[\"x2\"].apply(lambda x: int(x))\ntrain_bbox[\"y1\"] = train_bbox[\"y1\"].apply(lambda x: int(x))\ntrain_bbox[\"y2\"] = train_bbox[\"y2\"].apply(lambda x: int(x))\n\ntrain_bbox.head(3)","metadata":{"execution":{"iopub.status.busy":"2022-08-26T19:11:19.740604Z","iopub.execute_input":"2022-08-26T19:11:19.741575Z","iopub.status.idle":"2022-08-26T19:11:20.774236Z","shell.execute_reply.started":"2022-08-26T19:11:19.741528Z","shell.execute_reply":"2022-08-26T19:11:20.773356Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def show_dcm_bboxes(studyInstanceUID, rows, cols):\n    \n    # Set defaults\n    BASE_PATH = \"../input/rsna-2022-cervical-spine-fracture-detection/train_images\"\n    N = rows*cols\n    \n    fig, axes = plt.subplots(nrows=rows, ncols=cols, figsize=(24,25))\n\n    # Get dataframe & .dcm paths\n    bbox_df = train_bbox[train_bbox[\"StudyInstanceUID\"]==studyInstanceUID].reset_index(drop=True).head(N)\n    png_paths = [f\"{BASE_PATH}/{studyInstanceUID}/{slice_no}.dcm\" \n                 for slice_no in bbox_df[\"slice_number\"]]\n\n    for path, k in zip(png_paths, range(N)):\n        dataset = pydicom.dcmread(path)\n        img = apply_voi_lut(dataset.pixel_array, dataset)\n        slice_no = png_paths[k].split(\"/\")[-1].split(\".\")[0]\n\n        # BBoxes\n        bbox_list = (bbox_df.loc[k, \"x1\"], bbox_df.loc[k, \"y1\"],\n                     bbox_df.loc[k, \"x2\"], bbox_df.loc[k, \"y2\"])\n        x1, y1, x2, y2 = [ int(x) for x in bbox_list ]\n\n        # Plot the image\n        x_plot = k // cols\n        y_plot = k % cols\n        \n        cv2.rectangle(img, (x1, y1), (x2, y2), (226, 56, 54), 5)\n        axes[x_plot, y_plot].imshow(img, cmap=\"bone\")\n        axes[x_plot, y_plot].set_title(f\"Slice: {slice_no}\", \n                                       fontsize=14, weight='bold')\n        axes[x_plot, y_plot].axis('off');","metadata":{"execution":{"iopub.status.busy":"2022-08-26T19:11:20.775382Z","iopub.execute_input":"2022-08-26T19:11:20.775923Z","iopub.status.idle":"2022-08-26T19:11:21.760747Z","shell.execute_reply.started":"2022-08-26T19:11:20.775878Z","shell.execute_reply":"2022-08-26T19:11:21.759544Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Example 1","metadata":{}},{"cell_type":"code","source":"show_dcm_bboxes(studyInstanceUID=\"1.2.826.0.1.3680043.5783\",\n                rows=6, cols=6)","metadata":{"execution":{"iopub.status.busy":"2022-08-26T19:11:21.762132Z","iopub.execute_input":"2022-08-26T19:11:21.762539Z","iopub.status.idle":"2022-08-26T19:11:26.77309Z","shell.execute_reply.started":"2022-08-26T19:11:21.762481Z","shell.execute_reply":"2022-08-26T19:11:26.771916Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Example 2","metadata":{}},{"cell_type":"code","source":"show_dcm_bboxes(studyInstanceUID=\"1.2.826.0.1.3680043.19778\",\n                rows=6, cols=6)","metadata":{"execution":{"iopub.status.busy":"2022-08-26T19:11:26.775292Z","iopub.execute_input":"2022-08-26T19:11:26.776454Z","iopub.status.idle":"2022-08-26T19:11:32.20493Z","shell.execute_reply.started":"2022-08-26T19:11:26.7764Z","shell.execute_reply":"2022-08-26T19:11:32.203697Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Example 3","metadata":{}},{"cell_type":"code","source":"show_dcm_bboxes(studyInstanceUID=\"1.2.826.0.1.3680043.31077\",\n                rows=6, cols=6)","metadata":{"execution":{"iopub.status.busy":"2022-08-26T19:11:32.207176Z","iopub.execute_input":"2022-08-26T19:11:32.207659Z","iopub.status.idle":"2022-08-26T19:11:37.997347Z","shell.execute_reply.started":"2022-08-26T19:11:32.207613Z","shell.execute_reply":"2022-08-26T19:11:37.996003Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 🐝 Finish Experiment\nwandb.finish()","metadata":{"execution":{"iopub.status.busy":"2022-08-26T19:11:37.999375Z","iopub.execute_input":"2022-08-26T19:11:37.999844Z","iopub.status.idle":"2022-08-26T19:11:44.029362Z","shell.execute_reply.started":"2022-08-26T19:11:37.999797Z","shell.execute_reply":"2022-08-26T19:11:44.028339Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 5. 3D CT Scans\n\n> 🙏 This part is completely **inspired** by [Seungwon Song](https://www.kaggle.com/songseungwon)'s notebook: [Break Time☕️ Just Enjoy drawing 3D cervical spine](https://www.kaggle.com/code/songseungwon/break-time-just-enjoy-drawing-3d-cervical-spine/notebook) which I LOVED so I had to create something similar and learn ♥\n\n🦴 **nifti file format** -> NIfTI (`.nii`) is a type of file format for neuroimaging.","metadata":{}},{"cell_type":"code","source":"# Get .nii paths\nnii_paths = glob(\"../input/rsna-2022-cervical-spine-fracture-detection/segmentations/*\")\n\n\nprint(clr.S+\"Total paths in [segmentations] folder:\"+clr.E, len(nii_paths))","metadata":{"execution":{"iopub.status.busy":"2022-08-26T19:11:44.030555Z","iopub.execute_input":"2022-08-26T19:11:44.031006Z","iopub.status.idle":"2022-08-26T19:11:44.047781Z","shell.execute_reply.started":"2022-08-26T19:11:44.030964Z","shell.execute_reply":"2022-08-26T19:11:44.046423Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class DrawMaskSample():\n    \n    def __init__(self, nii_paths, i, ax):\n        self.i = i\n        self.ax = ax\n        self.nii_sample = nib.load(nii_paths[self.i]).get_fdata()\n        self.get_xyz()\n        self.draw_sample_3d()\n        \n    def get_xyz(self):\n        self.xyz_li = []\n        cnt = 0\n        max_cnt = self.nii_sample.shape[-1]\n        for iter_z, (iter_img) in enumerate(self.nii_sample.transpose(2,0,1)):\n            for iter_x, iter_arr_y in enumerate(iter_img):\n                iter_arr_y = np.where(iter_arr_y)[0]\n                if len(iter_arr_y) >= 1:\n                    iter_arr_y = list(set([iter_arr_y.max(), iter_arr_y.min()]))\n                    xyz = [(iter_x, iter_y, iter_z) for iter_y in iter_arr_y if np.any(iter_y)]\n                    self.xyz_li.append(xyz)\n            cnt += 1\n\n    def draw_sample_3d(self):\n        xyz_matrix = np.array(list(itertools.chain.from_iterable(self.xyz_li)))\n        X = xyz_matrix[:,0]\n        Y = xyz_matrix[:,1]\n        Z = xyz_matrix[:,2]\n\n        self.ax.scatter(X, Y, Z, s=1, alpha=0.05, color='#e3dac9')\n        xlim, ylim, zlim = self.nii_sample.shape\n        self.ax.set_xlim(0, xlim)\n        self.ax.set_ylim(0, ylim)\n        self.ax.set_zlim(0, zlim)\n        self.ax.set_title(f'sample - ({self.i+1})')\n        self.ax.set_yticklabels([])\n        self.ax.set_xticklabels([])\n        self.ax.set_zticklabels([])\n        self.ax.set_facecolor('#081921')","metadata":{"execution":{"iopub.status.busy":"2022-08-26T19:11:44.049337Z","iopub.execute_input":"2022-08-26T19:11:44.049749Z","iopub.status.idle":"2022-08-26T19:11:44.063482Z","shell.execute_reply.started":"2022-08-26T19:11:44.049713Z","shell.execute_reply":"2022-08-26T19:11:44.062435Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = plt.figure(figsize=(24,24))\nfor i in range(9):\n    ax = fig.add_subplot(int(f'33{i+1}'), projection='3d')\n    DrawMaskSample(nii_paths, i, ax)\n\nplt.tight_layout()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-26T19:11:44.064934Z","iopub.execute_input":"2022-08-26T19:11:44.065617Z","iopub.status.idle":"2022-08-26T19:12:27.310964Z","shell.execute_reply.started":"2022-08-26T19:11:44.065573Z","shell.execute_reply":"2022-08-26T19:12:27.309931Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 💽 Save Data","metadata":{}},{"cell_type":"code","source":"# Save datasets\ntrain.to_csv(\"train.csv\", index=False)\nmeta_train.to_csv(\"meta_train_clean.csv\", index=False)","metadata":{"execution":{"iopub.status.busy":"2022-08-26T19:12:27.312681Z","iopub.execute_input":"2022-08-26T19:12:27.313815Z","iopub.status.idle":"2022-08-26T19:12:29.123608Z","shell.execute_reply.started":"2022-08-26T19:12:27.313768Z","shell.execute_reply":"2022-08-26T19:12:29.122331Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 🐝 Make them artifacts and store in Dashboard\nsave_dataset_artifact(run_name=\"save_train\", artifact_name=\"train\",\n                      path=\"../input/rsna-fracture-detection/train.csv\")\nsave_dataset_artifact(run_name=\"save_meta\", artifact_name=\"train_meta\",\n                      path=\"../input/rsna-fracture-detection/meta_train_clean.csv\")","metadata":{"execution":{"iopub.status.busy":"2022-08-26T19:12:29.126369Z","iopub.execute_input":"2022-08-26T19:12:29.127141Z","iopub.status.idle":"2022-08-26T19:12:48.596251Z","shell.execute_reply.started":"2022-08-26T19:12:29.127089Z","shell.execute_reply":"2022-08-26T19:12:48.594786Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# sns.palplot(sns.color_palette(my_colors))","metadata":{"execution":{"iopub.status.busy":"2022-08-26T19:12:48.598379Z","iopub.execute_input":"2022-08-26T19:12:48.598896Z","iopub.status.idle":"2022-08-26T19:12:48.605664Z","shell.execute_reply.started":"2022-08-26T19:12:48.598837Z","shell.execute_reply":"2022-08-26T19:12:48.604429Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# === WIP ===","metadata":{"execution":{"iopub.status.busy":"2022-08-26T19:12:48.607364Z","iopub.execute_input":"2022-08-26T19:12:48.607709Z","iopub.status.idle":"2022-08-26T19:12:48.615947Z","shell.execute_reply.started":"2022-08-26T19:12:48.607676Z","shell.execute_reply":"2022-08-26T19:12:48.614868Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<center><img src=\"https://i.imgur.com/0cx4xXI.png\"></center>\n\n### 🐝 W&B Dashboard\n\n> My [W&B Dashboard](https://wandb.ai/andrada/RSNA_SpineFructure?workspace=user-andrada).\n\n<center><video src=\"https://i.imgur.com/FXnUE0O.mp4\" width=800 controls></center>\n\n<center><img src=\"https://i.imgur.com/knxTRkO.png\"></center>\n\n### My Specs\n\n* 🖥 Z8 G4 Workstation\n* 💾 2 CPUs & 96GB Memory\n* 🎮 2x NVIDIA A6000\n* 💻 Zbook Studio G7 on the go","metadata":{}}]}