{"cells":[{"metadata":{},"cell_type":"markdown","source":"# PANDA 256x256 crops dataset\nAs the images are quite large in the dataset for this competition, here I show a fast way to crop them to form 256x256 tiles and discard \"white tiles\".\n\n**Version 1.** Provides code to create tiles from train images and saves tiles to `/kaggle/working/data256`\n\n**Some aspects that may need to be improved:** \n* The `crop_and_tile` function is slicing up to 255 pixels out at the end of each image dimension to make the image size multiple of 256. I didn't check if any \"good\" pixels are being removed in this step.\n* Some regions of interest may turn out to be in the margin or corner of the image. This may not be ideal, one possibility is to generate tiles with some overlap. \n\n**Note:** So far I have no idea if this approach will be any good, it's just my first guess on this competition."},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"from fastai.vision import *\nimport skimage.io\nimport zipfile","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def crop_and_tile(fn, tiff_layer=2, empty_thr=250, tile_thr=0):\n    im = skimage.io.MultiImage(str(fn))[tiff_layer]\n    crop = tile_size*(im.shape[0]//tile_size), tile_size*(im.shape[1]//tile_size)\n    im = im[:crop[0], :crop[1]]\n    imr = im.reshape(im.shape[0]//tile_size,tile_size,im.shape[1]//tile_size, tile_size, 3)\n    imr = imr.transpose(1,3,0,2,4)\n    imr = imr.reshape(imr.shape[0], imr.shape[1], imr.shape[2]*imr.shape[3], imr.shape[4])\n    imr = imr.transpose(2,0,1,3)\n    not_empty = np.array([(im[...,0]<empty_thr).sum() for im in imr])>tile_thr\n    return imr[not_empty]\n\ndef save_tiles(path:Path, filename:str, tiles):\n    for i, t in enumerate(tiles):\n        im = PIL.Image.fromarray(t)\n        im.save(path/f'{filename}_{i}.png')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"tile_size = 256\npath = Path('/kaggle/input/prostate-cancer-grade-assessment')\nsave_path = Path(f'/kaggle/working/data{tile_size}')\nsave_path.mkdir(exist_ok=True)\ntrain_folder = 'train_images'\nmasks_folder = 'train_label_masks'\npath.ls()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Load train.csv\ntrain_df = pd.read_csv(path/'train.csv')\ntrain_df.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"files = (path/train_folder).ls()\n\ndef do_one(fn, *args):\n    tiles = crop_and_tile(fn)\n    save_tiles(save_path, fn.stem, tiles)\n    \nparallel(do_one, files, max_workers=4)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Plot some samples\nsaved_files = np.random.permutation(save_path.ls())\nfig, axes = plt.subplots(ncols=16,nrows=16,figsize=(12,12),dpi=120,facecolor='gray')\nfor ax, fn in zip(axes.flat, saved_files):\n    im = PIL.Image.open(fn)\n    ax.imshow(np.array(im))\n    ax.axis('off')","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}