{"cells":[{"metadata":{},"cell_type":"markdown","source":"Dataset URL: https://www.kaggle.com/guiferviz/rsna_stage1_png_128\n\n# Preparing dataset\n\nThe DICOM format is so cool, but I prefer normal images :)\n\nWith 156GB (compressed) it is very difficult to work with the resources of the vast majority of the mortals.\nThis notebook shows you how to scale down all the images and create a new dataset easier to deal with.\nEven with the best computing resources, I don't think it's necessary to use the original size to get good accuracy.\n\nIMPORTANT: In this notebook runs in a subset of the data, so don't use the generated output. That is because the Kaggle notebook runs out of space if you use all the examples. If you want to run this by yourself you should run it in a different machine or opening the next notebook https://colab.research.google.com/gist/guiferviz/50912a681776d5afe012b1a9259bd637/resize-dataset.ipynb in Google Colab. If you try to unzip the data in Google Colab you will also run out of space, so I've used the amazing tool *fuse-zip* to mount the zip and work with the files in it without extracting any of those.\n\nSome code taken from:\n* https://www.kaggle.com/omission/eda-view-dicom-images-with-correct-windowing\n* https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109649#latest-631701\n\n# Constants"},{"metadata":{"trusted":true},"cell_type":"code","source":"# Desired output size.\nRESIZED_WIDTH, RESIZED_HEIGHT = 224, 224\nOUTPUT_FORMAT = \"png\"\n\nOUTPUT_DIR = \"output\"","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Imports"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import glob\n\nimport joblib\n\nimport numpy as np\n\nimport PIL\n\nimport pydicom\n\nimport tqdm","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Get images paths"},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"data_dir = \"../input/rsna-intracranial-hemorrhage-detection/rsna-intracranial-hemorrhage-detection\"\n!ls {data_dir}","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_dir = \"stage_2_train\"\ntrain_paths = glob.glob(f\"{data_dir}/{train_dir}/*.dcm\")\nprint(len(train_paths))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def get_patient_id(img_path):\n    img_dicom = pydicom.read_file(img_path)\n    return img_dicom.PatientID","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"get_patient_id(train_paths[0])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"from collections import defaultdict \nfrom tqdm import tqdm\nimport random\ndict_id = defaultdict(list)\nrandom.shuffle(train_paths)\n#print(train_paths)\nfor img_path in tqdm(train_paths):\n    patient_id = get_patient_id(img_path)\n    path = img_path[101:114] + 'png'\n    dict_id[patient_id].append(path)\nlen(dict_id)\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"!apt-get install zip","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"import numpy as np\n\n!mkdir -p {OUTPUT_DIR}/{train_dir}\n\nnp.save('./output/stage_2_train/my_file.npy', dict_id) \n!zip resized.zip ./output/stage_2_train -r\n!rm ./output/stage_2_train -r ","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}