{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Mayo Clinic STRIP Stroke Blood Clot Origin 1:8 Pixel Resolution Reduction Dataset\n\n## Compentition\n* https://www.kaggle.com/competitions/mayo-clinic-strip-ai\n\nThis notebook provides a data set transform for a 1 to 8 size (pixel resolution) reduction of the original data set.  Additionally images have been stored as PNG format for a smaller loss-less format.\n\n## Notes\nWe've added restart logic into this code base.  The Kaggle instances tend to OOM (Out of Memory) while processing the `other` data set type.\n\n| Data Type | Processing Times |\n| --------- | ---------------- |\n| other | 1.5 hours |\n| test | 2 minutes |\n| train | 4 hours |\n\nLeverages image utilities from https://www.kaggle.com/code/joshuacburt/utils-image/notebook\n\n\nThe below samples could not be preprocessed using this techique (without OOM errors):\n\n| Data Type | image_id |\n| --------- | -------- |\n| other | 2c3c06_0 |","metadata":{}},{"cell_type":"markdown","source":"# Setup Runtime Environment\n\nThis control allows for easily moving this solution out of Kaggle notebooks and into a different system.","metadata":{}},{"cell_type":"code","source":"from pathlib import Path\n\nBASE_DIRECTORY: Path = Path(\"/kaggle/input/\")\nBASE_OUTPUT_DIRECTORY: Path = Path(\"/kaggle/working/\")\n\nRAW_DATA_DIRECTORY: Path = BASE_DIRECTORY.joinpath(\"mayo-clinic-strip-ai\")\nDOWNSAMPLED_DATA_DIRECTORY: Path = BASE_OUTPUT_DIRECTORY.joinpath(\"downsampled_data\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Next we define our function for loading and storing the samples.  It will check to see if the sample has previously been generated.  This allows us to restart after an OOM scecnario.  Additionally internal variables storing the image are removed as soon as possible and we manually request a garbage collection.","metadata":{}},{"cell_type":"code","source":"import gc\nfrom typing import List\nfrom utils_image import load_tiff_image_as_downsampled, store_image\n\n# These samples can not be processed on Kaggle.  They simply throw OOM when attempting to use them.  They are excluded here.\nsample_black_list: List = [\"/kaggle/input/mayo-clinic-strip-ai/other/2c3c06_0.tif\"]\n\ndef store_downsampled_image(image_id: str, target_data_type: str, force: bool = False) -> None:\n    \"\"\"Stores the downsampled (plus minimally preprocesed) image as a PNG.\"\"\"\n\n    input_file: Path = RAW_DATA_DIRECTORY.joinpath(target_data_type, f\"{image_id}.tif\")\n    output_file: Path = DOWNSAMPLED_DATA_DIRECTORY.joinpath(target_data_type, f\"{image_id}.png\")\n\n    # print(f\"Processing {input_file}\")\n\n    if force or (not output_file.exists() and input_file.resolve().as_posix() not in sample_black_list):\n        try:\n            array: np.ndarray = load_tiff_image_as_downsampled(input_file.as_posix())\n            store_image(image_data=array, target_file=output_file)\n        except MemoryError as error:\n            # Note that since this is an OOM scenario the intepreter may not be able to recover.\n            # Reference: https://docs.python.org/2/library/exceptions.html#exceptions.MemoryError\n            print(str(error))\n            print(f\"Processing {input_file} failed with OOM (Out of Memory).  This sample will be skipped.\")\n        finally:\n            del array\n            gc.collect()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Load the details for all three types of data (test, train, and other).  We use these data to perform the preprocessing on the data set.","metadata":{}},{"cell_type":"code","source":"from typing import Dict, List\n\nimport pandas as pd\nfrom joblib import Parallel, delayed\nfrom tqdm import tqdm\n\nmetadata: Dict = {}\ndata_types: List[str] = [\"train\", \"test\", \"other\"]\nfor target_data_type in data_types:\n    metadata[target_data_type] = pd.read_csv(RAW_DATA_DIRECTORY.joinpath(f\"{target_data_type}.csv\"))[\n        \"image_id\"\n    ].tolist()","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Don't use parallel options in Kaggle instances they will run out of memory.  Leverage the parallelization on other platforms.\nJob counts have been set to `1` to help alleviant compute constraints within Kaggle.","metadata":{}},{"cell_type":"code","source":"Parallel(n_jobs=1, prefer=\"threads\")(\n    delayed(store_downsampled_image)(image_id, \"train\")\n    for image_id in tqdm(metadata[\"train\"], desc=\"Downsampling training images\") \n)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"Parallel(n_jobs=1, prefer=\"threads\")(\n    delayed(store_downsampled_image)(image_id, \"other\")\n    for image_id in tqdm(metadata[\"other\"], desc=\"Downsampling other images\")\n)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"Parallel(n_jobs=1, prefer=\"threads\")(\n    delayed(store_downsampled_image)(image_id, \"test\")\n    for image_id in tqdm(metadata[\"test\"], desc=\"Downsampling test images\")\n)","metadata":{"trusted":true},"execution_count":null,"outputs":[]}]}