{"cells":[{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import os\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport openslide\nimport matplotlib.pyplot as plt\nfrom PIL import Image\nimport cv2\nfrom tqdm.notebook import tqdm\nimport skimage.io\nfrom skimage.transform import resize, rescale\nimport glob","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"As the organizer has already mentioned that \n\n**train_label_masks: Segmentation masks showing which parts of the image led to the ISUP grade. Not all training images have label masks.**\n\nTherefore, this simple notebook aims to check how many images in the training images folder that are not labelled with masks."},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"train = pd.read_csv('/kaggle/input/prostate-cancer-grade-assessment/train.csv')\ntrain_images_path = '/kaggle/input/prostate-cancer-grade-assessment/train_images/'\ntrain_label_mask_path = '/kaggle/input/prostate-cancer-grade-assessment/train_label_masks/'\nimg_type = \"*.tiff\"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def sanity_tally(train_images_path, train_label_mask_path, img_type):\n    total_img_list = [os.path.basename(img_name) for img_name\n                      in glob.glob(os.path.join(train_images_path, img_type))]\n    ## get the image_name\n    total_img_list = [x[:-5] for x in total_img_list]\n    \n    mask_img_list  = [os.path.basename(img_name) for img_name \n                      in glob.glob(os.path.join(train_label_mask_path, img_type))]\n    \n    # note that the image name in train_label_mask will always be in this format: abcdefg_mask.tiff; therefore I needed to\n    # remove the last 10 characters to tally with the images in train_images.\n    mask_img_list  = [x[:-10] for x in mask_img_list]\n    set_diff1      = set(total_img_list) - set(mask_img_list)\n    set_diff2      = set(mask_img_list)  - set(total_img_list)\n    \n    if set(total_img_list)  == set(mask_img_list):\n        print(\"Sanity Check Status: True\")\n    else:\n        print(\"Sanity Check Status: Failed. \\nThe elements in train_images_path but not in the train_label_mask_path is {} and the number is {}.\\n\\n\\nThe elements in train_label_mask_path but not in train_images_path is {} and the number is {}\".format(\n                set_diff1, len(set_diff1), set_diff2, len(set_diff2)))\n    \n    return set_diff1, set_diff2","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"set_diff1, set_diff2 = sanity_tally(train_images_path,train_label_mask_path, img_type)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"From the sanity check above, and assuming that each image's name in the train_images should necessarily match the ones in train_label_masks, we can deduce that the masked images in train_label_masks is a subset of the images in the train_images. That is to say, all masked images in train_label_masks has a corresponding image in the train_images, but there exists 100 images in train_images that do not have a mask. Whether we decide to keep these \"un-labelled\" images is up to you to decide."},{"metadata":{},"cell_type":"markdown","source":"For simplicity sake, I want to have a bijective relationship between the train_images folder and the train_label_masks folder; and since the set difference is small (100), I can make do to delete these 100 images that are not **annotated** by the pathologists."},{"metadata":{"trusted":true},"cell_type":"code","source":"remove_images = list(set_diff1)\nnew_train = train[~train.image_id.isin(remove_images)]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"new_train = new_train.reset_index(drop=True)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"new_train","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Next, I conveniently borrowed [xhlulu's panda resize and save train data kernel](https://www.kaggle.com/xhlulu/panda-resize-and-save-train-data) and save the image as png file, where all of them are resized to 512x512."},{"metadata":{"trusted":true},"cell_type":"code","source":"save_dir = \"/kaggle/train/\"\nos.makedirs(save_dir, exist_ok=True)\n\ntrain_images_path = '/kaggle/input/prostate-cancer-grade-assessment/train_images/'\ntrain_label_mask_path = '/kaggle/input/prostate-cancer-grade-assessment/train_label_masks/'\nimg_type = \"*.tiff\"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"for img_id in tqdm(new_train.image_id):\n    load_path = train_images_path + img_id + '.tiff'\n    save_path = save_dir + img_id + '.jpg'\n    \n    biopsy = skimage.io.MultiImage(load_path)\n    img = cv2.resize(biopsy[-1], (512, 512))\n    cv2.imwrite(save_path, img)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"for img_id in tqdm(new_train.image_id):\n    load_path = train_label_mask_path + img_id + '_mask' + '.tiff'\n    save_path = save_dir + img_id + '.jpg'\n    \n    biopsy = skimage.io.MultiImage(load_path)\n    img = cv2.resize(biopsy[-1], (512, 512))\n    cv2.imwrite(save_path, img)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"!tar -czf images.tar.gz ../train/*.png","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}