{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"This kernel is based on 2 open-source repositories, providing an easy and convenient way for dealing with datasets:\n- AMID (Awesome Medical Imaging Datasets) -- a curated collection of popular medical imaging datasets with unified interfaces\n- connectome -- a powerful framework for preprocessing, data augmentation and caching in your experiments\n\nI'll show you how to standardize the images, remove small artifacts, cache the results and apply data augmentation all in a few lines of code. At the end the dataset will look something like this:\n\n```python\ndataset = Chain(\n    RSNABreastCancer('/kaggle/input/rsna-breast-cancer-detection/'),\n    Normalize(),\n    Apply(image=lambda x: zoom(np.float32(x), 0.25, order=1)),\n    GreatestComponent(),\n    CropBackground(),\n    \n    CacheToDisk('image'),\n    CacheToRam(),\n    \n    RandomFlip(),\n)\n```","metadata":{}},{"cell_type":"markdown","source":"# Install AMID","metadata":{}},{"cell_type":"code","source":"# connectome will be installed automatically, because `amid` depends on it\n\n!pip install amid","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2023-01-04T14:21:44.955477Z","iopub.execute_input":"2023-01-04T14:21:44.956107Z","iopub.status.idle":"2023-01-04T14:22:32.613307Z","shell.execute_reply.started":"2023-01-04T14:21:44.955982Z","shell.execute_reply":"2023-01-04T14:22:32.612089Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# First look at the dataset","metadata":{}},{"cell_type":"code","source":"from amid.rsna_bc import RSNABreastCancer\n\nds = RSNABreastCancer('/kaggle/input/rsna-breast-cancer-detection/')","metadata":{"execution":{"iopub.status.busy":"2023-01-04T14:22:32.616855Z","iopub.execute_input":"2023-01-04T14:22:32.617397Z","iopub.status.idle":"2023-01-04T14:22:33.81278Z","shell.execute_reply.started":"2023-01-04T14:22:32.617342Z","shell.execute_reply":"2023-01-04T14:22:33.811308Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dir(ds)","metadata":{"execution":{"iopub.status.busy":"2023-01-04T14:22:33.814539Z","iopub.execute_input":"2023-01-04T14:22:33.815754Z","iopub.status.idle":"2023-01-04T14:22:33.827842Z","shell.execute_reply.started":"2023-01-04T14:22:33.815697Z","shell.execute_reply":"2023-01-04T14:22:33.825213Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(ds.ids)","metadata":{"execution":{"iopub.status.busy":"2023-01-04T14:22:33.83183Z","iopub.execute_input":"2023-01-04T14:22:33.832339Z","iopub.status.idle":"2023-01-04T14:22:38.929681Z","shell.execute_reply.started":"2023-01-04T14:22:33.832301Z","shell.execute_reply":"2023-01-04T14:22:38.928262Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"i = ds.ids[0]\ni","metadata":{"execution":{"iopub.status.busy":"2023-01-04T14:22:38.931387Z","iopub.execute_input":"2023-01-04T14:22:38.93187Z","iopub.status.idle":"2023-01-04T14:22:38.964143Z","shell.execute_reply.started":"2023-01-04T14:22:38.93183Z","shell.execute_reply":"2023-01-04T14:22:38.962666Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ds.age(i), ds.implant(i), ds.cancer(i), ds.view(i)","metadata":{"execution":{"iopub.status.busy":"2023-01-04T14:22:38.965734Z","iopub.execute_input":"2023-01-04T14:22:38.96654Z","iopub.status.idle":"2023-01-04T14:22:39.011279Z","shell.execute_reply.started":"2023-01-04T14:22:38.966501Z","shell.execute_reply":"2023-01-04T14:22:39.009825Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"But the most interesting part is the image:","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\n\ndef show(x):\n    plt.imshow(x, cmap='gray')","metadata":{"execution":{"iopub.status.busy":"2023-01-04T14:22:39.013005Z","iopub.execute_input":"2023-01-04T14:22:39.013441Z","iopub.status.idle":"2023-01-04T14:22:39.022152Z","shell.execute_reply.started":"2023-01-04T14:22:39.013358Z","shell.execute_reply":"2023-01-04T14:22:39.0212Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"show(ds.image(i))","metadata":{"execution":{"iopub.status.busy":"2023-01-04T14:22:39.023702Z","iopub.execute_input":"2023-01-04T14:22:39.024338Z","iopub.status.idle":"2023-01-04T14:22:41.811222Z","shell.execute_reply.started":"2023-01-04T14:22:39.024299Z","shell.execute_reply":"2023-01-04T14:22:41.809852Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"i2 = '2028281852'\nshow(ds.image(i2))","metadata":{"execution":{"iopub.status.busy":"2023-01-04T14:22:41.813312Z","iopub.execute_input":"2023-01-04T14:22:41.813757Z","iopub.status.idle":"2023-01-04T14:22:45.951191Z","shell.execute_reply.started":"2023-01-04T14:22:41.8137Z","shell.execute_reply":"2023-01-04T14:22:45.949424Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Normalizing the intensity\n\nAs we saw earlier, images have different background and tissue intensities. Most models would benefit from a normalization step, which would make the background equal 0.\n\nWe'll write a \"transform\" for that:","metadata":{}},{"cell_type":"code","source":"from connectome import Transform\n\n\nclass Normalize(Transform):\n    \"\"\"\n    Inverts the image intensity if necessary, so that the background is always zero\n    \"\"\"\n\n    __inherit__ = True # inherit all fields from the previous layer\n\n    def image(image, padding_value, intensity_sign):\n        if padding_value is not None:\n            if padding_value > 0:\n                return padding_value - image\n            return image\n        \n        # if no padding value is available, we can conpute it ourselves\n        if intensity_sign == 1:\n            return image.max() - image\n\n        return image\n","metadata":{"execution":{"iopub.status.busy":"2023-01-04T14:22:45.957592Z","iopub.execute_input":"2023-01-04T14:22:45.958211Z","iopub.status.idle":"2023-01-04T14:22:45.969165Z","shell.execute_reply.started":"2023-01-04T14:22:45.958165Z","shell.execute_reply":"2023-01-04T14:22:45.967925Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"And apply it to the base dataset:","metadata":{}},{"cell_type":"code","source":"dataset = ds >> Normalize()","metadata":{"execution":{"iopub.status.busy":"2023-01-04T14:22:45.970976Z","iopub.execute_input":"2023-01-04T14:22:45.97229Z","iopub.status.idle":"2023-01-04T14:22:46.154043Z","shell.execute_reply.started":"2023-01-04T14:22:45.97224Z","shell.execute_reply":"2023-01-04T14:22:46.152726Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# nothing changed here, the bg was already black\nshow(dataset.image(i))","metadata":{"execution":{"iopub.status.busy":"2023-01-04T14:22:46.155932Z","iopub.execute_input":"2023-01-04T14:22:46.156355Z","iopub.status.idle":"2023-01-04T14:22:48.62914Z","shell.execute_reply.started":"2023-01-04T14:22:46.156319Z","shell.execute_reply":"2023-01-04T14:22:48.627803Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# but this one got normalized\nshow(dataset.image(i2))","metadata":{"execution":{"iopub.status.busy":"2023-01-04T14:22:48.631054Z","iopub.execute_input":"2023-01-04T14:22:48.631442Z","iopub.status.idle":"2023-01-04T14:22:52.560621Z","shell.execute_reply.started":"2023-01-04T14:22:48.631407Z","shell.execute_reply":"2023-01-04T14:22:52.559339Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now what is going on here?\n\nThe \"image\" method describes how to transform an image given the arguments from the previous layer. `RSNABreastCancer` has all the needed fields: image, padding_value and intensity_sign; so it can compute the new \"image\". \n\nThis is a pretty powerful approach, called \"dependecy injection\". If you're familiar with `pytest`'s fixtures, then you have already used it.","metadata":{}},{"cell_type":"markdown","source":"# Cropping the background\n\nNow all the images have a lot of useless blackness. Let's remove it.","metadata":{}},{"cell_type":"code","source":"show(dataset.image(i) > 0)","metadata":{"execution":{"iopub.status.busy":"2023-01-04T14:22:52.562499Z","iopub.execute_input":"2023-01-04T14:22:52.563028Z","iopub.status.idle":"2023-01-04T14:22:54.895435Z","shell.execute_reply.started":"2023-01-04T14:22:52.562968Z","shell.execute_reply":"2023-01-04T14:22:54.894328Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Simply thresholding the image by 0 seems to almost do the trick. We'll get back to that small artifact in the top right corner later.","metadata":{}},{"cell_type":"code","source":"class CropBackground(Transform):\n    __inherit__ = True\n    \n    def image(image):\n        mask = image > 0\n        xs, = mask.any(0).nonzero()\n        ys, = mask.any(1).nonzero()\n        return image[ys.min():ys.max() + 1, xs.min():xs.max() + 1]","metadata":{"execution":{"iopub.status.busy":"2023-01-04T14:22:54.896802Z","iopub.execute_input":"2023-01-04T14:22:54.897212Z","iopub.status.idle":"2023-01-04T14:22:54.90442Z","shell.execute_reply.started":"2023-01-04T14:22:54.897179Z","shell.execute_reply":"2023-01-04T14:22:54.903378Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from connectome import Chain\n\n\n# Chain(a, b, c) is a more convenient way to write a >> b >> c\ndataset = Chain(\n    ds,\n    Normalize(),\n    CropBackground(),\n)","metadata":{"execution":{"iopub.status.busy":"2023-01-04T14:22:54.905697Z","iopub.execute_input":"2023-01-04T14:22:54.906488Z","iopub.status.idle":"2023-01-04T14:22:54.956725Z","shell.execute_reply.started":"2023-01-04T14:22:54.906455Z","shell.execute_reply":"2023-01-04T14:22:54.955606Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"show(dataset.image(i))","metadata":{"execution":{"iopub.status.busy":"2023-01-04T14:22:54.959057Z","iopub.execute_input":"2023-01-04T14:22:54.959957Z","iopub.status.idle":"2023-01-04T14:22:56.89314Z","shell.execute_reply.started":"2023-01-04T14:22:54.959905Z","shell.execute_reply":"2023-01-04T14:22:56.891691Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Removing small artifacts\n\nNow what about this small inscription in the top right corner?","metadata":{}},{"cell_type":"code","source":"show(dataset.image(i) > 0)","metadata":{"execution":{"iopub.status.busy":"2023-01-04T14:22:56.894647Z","iopub.execute_input":"2023-01-04T14:22:56.895628Z","iopub.status.idle":"2023-01-04T14:22:58.760529Z","shell.execute_reply.started":"2023-01-04T14:22:56.895588Z","shell.execute_reply":"2023-01-04T14:22:58.758912Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It seems like the breast is always the greatest connected component, and the rest can be simply discarded:","metadata":{}},{"cell_type":"code","source":"from skimage.morphology import label\nimport numpy as np\n\n\nclass GetGreatestComponent(Transform):\n    __inherit__ = True\n    \n    def image(image):\n        # find all the connected components\n        lbl = label(image > 0)\n        # caclulate their sizes\n        values, counts = np.unique(lbl, return_counts=True)\n        # 0 is always the background - we don't need it\n        foreground = values != 0\n        component = values[foreground][counts[foreground].argmax()]\n        # select all the components greater than the background\n        #  + the greatest foreground component\n        components = set(values[counts > counts[~foreground]]) | {component}\n        if len(components) > 1:\n            # if there are several components - pick the one with the greatest intensity\n            component = max(components, key=lambda c: image[lbl == c].mean())\n\n        return image * (lbl == component)","metadata":{"execution":{"iopub.status.busy":"2023-01-04T14:22:58.76232Z","iopub.execute_input":"2023-01-04T14:22:58.763214Z","iopub.status.idle":"2023-01-04T14:22:59.42054Z","shell.execute_reply.started":"2023-01-04T14:22:58.763174Z","shell.execute_reply":"2023-01-04T14:22:59.419107Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dataset = Chain(\n    ds,\n    Normalize(),\n    GetGreatestComponent(),\n    CropBackground(),\n)","metadata":{"execution":{"iopub.status.busy":"2023-01-04T14:22:59.422294Z","iopub.execute_input":"2023-01-04T14:22:59.423121Z","iopub.status.idle":"2023-01-04T14:22:59.440703Z","shell.execute_reply.started":"2023-01-04T14:22:59.423064Z","shell.execute_reply":"2023-01-04T14:22:59.439108Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"show(dataset.image(i))","metadata":{"execution":{"iopub.status.busy":"2023-01-04T14:22:59.442461Z","iopub.execute_input":"2023-01-04T14:22:59.443009Z","iopub.status.idle":"2023-01-04T14:23:01.75539Z","shell.execute_reply.started":"2023-01-04T14:22:59.442959Z","shell.execute_reply":"2023-01-04T14:23:01.754237Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Works like a charm! Note that we first remove the small components, and then crop the background.","metadata":{}},{"cell_type":"markdown","source":"# Downsampling\n\nThe images in this dataset are pretty big ~ 5000x2000. There is a principle in computer vision: if you _can_ downsample your data without losing in model accuracy, you _should_ downsample it.","metadata":{}},{"cell_type":"markdown","source":"We could write another transform for that, but `connectome` has a convenient shortcut for smaller functions:","metadata":{}},{"cell_type":"code","source":"from connectome import Apply\nfrom scipy.ndimage import zoom\n\ndataset = Chain(\n    ds,\n    Normalize(),\n    # 0.25 - is the downsample factor. It should probably be tuned via cross-validation\n    Apply(image=lambda x: zoom(np.float32(x), 0.25, order=1)),\n    GetGreatestComponent(),\n    CropBackground(),\n)","metadata":{"execution":{"iopub.status.busy":"2023-01-04T14:23:01.757411Z","iopub.execute_input":"2023-01-04T14:23:01.758294Z","iopub.status.idle":"2023-01-04T14:23:01.777958Z","shell.execute_reply.started":"2023-01-04T14:23:01.758245Z","shell.execute_reply":"2023-01-04T14:23:01.776683Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"show(dataset.image(i))","metadata":{"execution":{"iopub.status.busy":"2023-01-04T14:23:01.779496Z","iopub.execute_input":"2023-01-04T14:23:01.779984Z","iopub.status.idle":"2023-01-04T14:23:03.009758Z","shell.execute_reply.started":"2023-01-04T14:23:01.779939Z","shell.execute_reply":"2023-01-04T14:23:03.008473Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Caching","metadata":{}},{"cell_type":"markdown","source":"Both downsampling and connected components removal are pretty heavy operations. Let's see how much time they take:","metadata":{}},{"cell_type":"code","source":"%%timeit\n\ndataset.image(i);","metadata":{"execution":{"iopub.status.busy":"2023-01-04T14:23:03.011381Z","iopub.execute_input":"2023-01-04T14:23:03.011828Z","iopub.status.idle":"2023-01-04T14:23:11.07913Z","shell.execute_reply.started":"2023-01-04T14:23:03.011789Z","shell.execute_reply":"2023-01-04T14:23:11.078046Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"almost a second per image... and we have 54k of them! So it would take more than half a day to simply load them. Caching to rescue!","metadata":{}},{"cell_type":"markdown","source":"### Caching to RAM","metadata":{}},{"cell_type":"code","source":"from connectome import CacheToRam\n\n\ndataset = Chain(\n    ds,\n    Normalize(),\n    # 0.25 - is the downsample factor. It should probably be tuned via cross-validation\n    Apply(image=lambda x: zoom(np.float32(x), 0.25, order=1)),\n    GetGreatestComponent(),\n    CropBackground(),\n    \n    CacheToRam(),\n)","metadata":{"execution":{"iopub.status.busy":"2023-01-04T14:23:11.08105Z","iopub.execute_input":"2023-01-04T14:23:11.081734Z","iopub.status.idle":"2023-01-04T14:23:11.102729Z","shell.execute_reply.started":"2023-01-04T14:23:11.081695Z","shell.execute_reply":"2023-01-04T14:23:11.101222Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# the first call will fill the cache\ndataset.image(i);","metadata":{"execution":{"iopub.status.busy":"2023-01-04T14:23:11.104575Z","iopub.execute_input":"2023-01-04T14:23:11.10544Z","iopub.status.idle":"2023-01-04T14:23:12.058629Z","shell.execute_reply.started":"2023-01-04T14:23:11.105396Z","shell.execute_reply":"2023-01-04T14:23:12.057342Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%timeit\n\ndataset.image(i);","metadata":{"execution":{"iopub.status.busy":"2023-01-04T14:23:12.060149Z","iopub.execute_input":"2023-01-04T14:23:12.060539Z","iopub.status.idle":"2023-01-04T14:23:25.136484Z","shell.execute_reply.started":"2023-01-04T14:23:12.060501Z","shell.execute_reply":"2023-01-04T14:23:25.135058Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"And now it's almost 1000 times faster!\nHowewer, bevare that caching to ram the whole dataset will require around 40GB of memory. \nYou can lru_cache only a part of the dataset like so `CacheToRam(size=10000)` ","metadata":{}},{"cell_type":"markdown","source":"### Caching to disk","metadata":{}},{"cell_type":"markdown","source":"Caching to RAM has another downside - it's not persistent. This means that each time you reset your notebook you'll have to recompute all the transformations all over again. \n\nLet's add another layer of persistent caching to disk:","metadata":{}},{"cell_type":"markdown","source":"First, set up a cache storage. Create the file `~/.config/amid/.bev.yml` with the following content:\n\n```\nmain:\n  storage: /path/to/storage\n  cache: /path/to/cache\n```\n\nwhere `/path/to/storage` and `/path/to/cache` are some paths in your filesystem\n\nand run `amid init`","metadata":{}},{"cell_type":"markdown","source":"In my case this would look something like this:","metadata":{}},{"cell_type":"code","source":"%%bash\n\nmkdir -p ~/.config/amid\n\ncat >~/.config/amid/.bev.yml <<EOL\nmain:\n  storage: /tmp/storage\n  cache: /tmp/cache\nEOL\n\namid init","metadata":{"execution":{"iopub.status.busy":"2023-01-04T14:23:25.142332Z","iopub.execute_input":"2023-01-04T14:23:25.142785Z","iopub.status.idle":"2023-01-04T14:23:26.260653Z","shell.execute_reply.started":"2023-01-04T14:23:25.142733Z","shell.execute_reply":"2023-01-04T14:23:26.25918Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I chose /tmp/... here, because it's a Kaggle kernel, you should use a less temporary location.","metadata":{}},{"cell_type":"code","source":"from amid.internals import CacheToDisk\n\n\ndataset = Chain(\n    ds,\n    Normalize(),\n    # 0.25 - is the downsample factor. It should probably be tuned via cross-validation\n    Apply(image=lambda x: zoom(np.float32(x), 0.25, order=1)),\n    GetGreatestComponent(),\n    CropBackground(),\n    \n    CacheToDisk('image'), # we only want to cache this field\n    CacheToRam(),\n)","metadata":{"execution":{"iopub.status.busy":"2023-01-04T14:23:26.262736Z","iopub.execute_input":"2023-01-04T14:23:26.263193Z","iopub.status.idle":"2023-01-04T14:23:26.381622Z","shell.execute_reply.started":"2023-01-04T14:23:26.263149Z","shell.execute_reply":"2023-01-04T14:23:26.380279Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# the first call will fill the cache\ndataset.image(i);","metadata":{"execution":{"iopub.status.busy":"2023-01-04T14:23:26.383268Z","iopub.execute_input":"2023-01-04T14:23:26.383652Z","iopub.status.idle":"2023-01-04T14:23:27.46454Z","shell.execute_reply.started":"2023-01-04T14:23:26.383617Z","shell.execute_reply":"2023-01-04T14:23:27.463191Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"and now we have a persistent cache which will be alive even after you restart your notebook.\n\nAnd, what's even cooler, it's smart enough to detect changes in your pipeline and invalidate itself:","metadata":{}},{"cell_type":"code","source":"show(dataset.image(i))","metadata":{"execution":{"iopub.status.busy":"2023-01-04T14:23:27.466327Z","iopub.execute_input":"2023-01-04T14:23:27.466864Z","iopub.status.idle":"2023-01-04T14:23:27.773708Z","shell.execute_reply.started":"2023-01-04T14:23:27.466818Z","shell.execute_reply":"2023-01-04T14:23:27.77216Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Now we'll change the downsampling factor:\n\ndataset = Chain(\n    ds,\n    Normalize(),\n    Apply(image=lambda x: zoom(np.float32(x), 0.1, order=1)),\n    GetGreatestComponent(),\n    CropBackground(),\n    \n    CacheToDisk('image'),\n    CacheToRam(),\n)","metadata":{"execution":{"iopub.status.busy":"2023-01-04T14:23:27.77542Z","iopub.execute_input":"2023-01-04T14:23:27.776651Z","iopub.status.idle":"2023-01-04T14:23:27.83897Z","shell.execute_reply.started":"2023-01-04T14:23:27.776581Z","shell.execute_reply":"2023-01-04T14:23:27.837908Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"show(dataset.image(i))","metadata":{"execution":{"iopub.status.busy":"2023-01-04T14:23:27.840899Z","iopub.execute_input":"2023-01-04T14:23:27.842025Z","iopub.status.idle":"2023-01-04T14:23:29.014971Z","shell.execute_reply.started":"2023-01-04T14:23:27.841982Z","shell.execute_reply":"2023-01-04T14:23:29.013638Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Even though we didn't modify the cache, it managed to detect the change and update itself accordingly!","metadata":{}},{"cell_type":"markdown","source":"# Data augmentation\n\nBecause the breast can be either on the left or right side of the image, we may want to add some data augmentation to our pipeline. This is how it's done with `connectome`: ","metadata":{}},{"cell_type":"code","source":"from connectome import impure\n\n\nclass RandomFlip(Transform):\n    __inherit__ = True\n\n    @impure\n    def image(image):\n        # flip the image with a 50% chance\n        if np.random.binomial(1, 0.5):\n            return np.flip(image, axis=1)\n        return image","metadata":{"execution":{"iopub.status.busy":"2023-01-04T14:23:29.016963Z","iopub.execute_input":"2023-01-04T14:23:29.017672Z","iopub.status.idle":"2023-01-04T14:23:29.026026Z","shell.execute_reply.started":"2023-01-04T14:23:29.017634Z","shell.execute_reply":"2023-01-04T14:23:29.024559Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Note the `impure` decorator. Its purpose is to help the cache layers detect \"impure\" functions, i.e. those that don't only depend on their inputs.","metadata":{}},{"cell_type":"code","source":"dataset = Chain(\n    ds,\n    Normalize(),\n    Apply(image=lambda x: zoom(np.float32(x), 0.25, order=1)),\n    GetGreatestComponent(),\n    CropBackground(),\n    \n    CacheToDisk('image'),\n    CacheToRam(),\n\n    RandomFlip(),\n)","metadata":{"execution":{"iopub.status.busy":"2023-01-04T14:23:29.028007Z","iopub.execute_input":"2023-01-04T14:23:29.028546Z","iopub.status.idle":"2023-01-04T14:23:29.097245Z","shell.execute_reply.started":"2023-01-04T14:23:29.028495Z","shell.execute_reply":"2023-01-04T14:23:29.095699Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"show(dataset.image(i))","metadata":{"execution":{"iopub.status.busy":"2023-01-04T14:23:29.099868Z","iopub.execute_input":"2023-01-04T14:23:29.100702Z","iopub.status.idle":"2023-01-04T14:23:29.420426Z","shell.execute_reply.started":"2023-01-04T14:23:29.100643Z","shell.execute_reply":"2023-01-04T14:23:29.419002Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"show(dataset.image(i))","metadata":{"execution":{"iopub.status.busy":"2023-01-04T14:23:29.422371Z","iopub.execute_input":"2023-01-04T14:23:29.422803Z","iopub.status.idle":"2023-01-04T14:23:29.711115Z","shell.execute_reply.started":"2023-01-04T14:23:29.422738Z","shell.execute_reply":"2023-01-04T14:23:29.709479Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"show(dataset.image(i))","metadata":{"execution":{"iopub.status.busy":"2023-01-04T14:23:29.71273Z","iopub.execute_input":"2023-01-04T14:23:29.713176Z","iopub.status.idle":"2023-01-04T14:23:30.024076Z","shell.execute_reply.started":"2023-01-04T14:23:29.713139Z","shell.execute_reply":"2023-01-04T14:23:30.022734Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"show(dataset.image(i))","metadata":{"execution":{"iopub.status.busy":"2023-01-04T14:23:30.025651Z","iopub.execute_input":"2023-01-04T14:23:30.026234Z","iopub.status.idle":"2023-01-04T14:23:30.315687Z","shell.execute_reply.started":"2023-01-04T14:23:30.026192Z","shell.execute_reply":"2023-01-04T14:23:30.314412Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now each function call randomly flips the image.","metadata":{}},{"cell_type":"markdown","source":"Note that we use the augmentation after caching, because our memory usage would explode otherwise, and `connectome` is smart enough to protect us from this. The following code would fail with a ValueError exception:\n\n```python\ndataset = Chain(\n    ds,\n    Normalize(),\n    Apply(image=lambda x: zoom(np.float32(x), 0.25, order=1)),\n    GetGreatestComponent(),\n    CropBackground(),\n    \n    RandomFlip(),\n    \n    CacheToDisk('image'),\n    CacheToRam(),\n)\n```","metadata":{"execution":{"iopub.status.busy":"2023-01-04T11:51:33.373305Z","iopub.execute_input":"2023-01-04T11:51:33.373738Z","iopub.status.idle":"2023-01-04T11:51:33.382027Z","shell.execute_reply.started":"2023-01-04T11:51:33.373707Z","shell.execute_reply":"2023-01-04T11:51:33.380752Z"}}},{"cell_type":"markdown","source":"# Final dataset\n\nCombinig all the layers, our final dataset will look something like this:","metadata":{}},{"cell_type":"code","source":"dataset = Chain(\n    RSNABreastCancer('/kaggle/input/rsna-breast-cancer-detection/'),\n    Normalize(),\n    Apply(image=lambda x: zoom(np.float32(x), 0.25, order=1)),\n    GetGreatestComponent(),\n    CropBackground(),\n    \n    CacheToDisk('image'),\n    CacheToRam(),\n    \n    RandomFlip(),\n)","metadata":{"execution":{"iopub.status.busy":"2023-01-04T14:23:30.317693Z","iopub.execute_input":"2023-01-04T14:23:30.318704Z","iopub.status.idle":"2023-01-04T14:23:30.440312Z","shell.execute_reply.started":"2023-01-04T14:23:30.318647Z","shell.execute_reply":"2023-01-04T14:23:30.438644Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"and you can use it while training either like this:","metadata":{}},{"cell_type":"code","source":"x = dataset.image(i)\ny = dataset.cancer(i)\n\nx.shape, y","metadata":{"execution":{"iopub.status.busy":"2023-01-04T14:23:30.442248Z","iopub.execute_input":"2023-01-04T14:23:30.443319Z","iopub.status.idle":"2023-01-04T14:23:30.480494Z","shell.execute_reply.started":"2023-01-04T14:23:30.443237Z","shell.execute_reply":"2023-01-04T14:23:30.479149Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"or you can _compile_ a function that returns the pair:","metadata":{}},{"cell_type":"code","source":"loader = dataset._compile(['image', 'cancer'])\nx, y = loader(i)\n\nx.shape, y","metadata":{"execution":{"iopub.status.busy":"2023-01-04T14:23:30.482904Z","iopub.execute_input":"2023-01-04T14:23:30.484092Z","iopub.status.idle":"2023-01-04T14:23:30.498463Z","shell.execute_reply.started":"2023-01-04T14:23:30.484021Z","shell.execute_reply":"2023-01-04T14:23:30.496952Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"These are the main preprocessing steps I would use for my baseline model. If `connectome` got you interested, check out the [docs](https://neuro-ml.github.io/connectome/tutorials/00%20-%20Intro/) for a more detailed tutorial.","metadata":{}}]}