{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    print(dirname)\n#     for filename in filenames:\n#         print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_kg_hide-output":false,"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-30T09:00:44.244541Z","iopub.execute_input":"2022-07-30T09:00:44.245024Z","iopub.status.idle":"2022-07-30T09:00:44.373749Z","shell.execute_reply.started":"2022-07-30T09:00:44.244921Z","shell.execute_reply":"2022-07-30T09:00:44.372574Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# import all the dependencies at once\nimport matplotlib.pyplot as plt\n\nimport tifffile as tifi\n\nimport cv2 as cv","metadata":{"execution":{"iopub.status.busy":"2022-07-30T09:14:23.563383Z","iopub.execute_input":"2022-07-30T09:14:23.563852Z","iopub.status.idle":"2022-07-30T09:14:23.572135Z","shell.execute_reply.started":"2022-07-30T09:14:23.563811Z","shell.execute_reply":"2022-07-30T09:14:23.570391Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h2 style=\"color: #363636; font-family: Segoe UI; font-size: 1.9em; font-weight: 600;\">Exploring Metadata of Dataset</h2>\n\n<span style=\"color: #555555;\">Let's first have a look at our current working directory and then we can navigate to our datasets and some of the training files as well.<br/>\nSo, as it turns out cli commands doesn't work as I expected them to, so let's directly jump into the dataset exploration. We will start with the metadata for train dataset first.\n</span>","metadata":{}},{"cell_type":"code","source":"train_meta = pd.read_csv('../input/mayo-clinic-strip-ai/train.csv')\ntrain_meta.head().style.set_properties(**{'background-color': '#363636',\n                                          'color': '#F3F3F3',\n                                          'border': '1px solid #606060'})\n# !ls ../input/mayo-clinic-strip-ai","metadata":{"execution":{"iopub.status.busy":"2022-07-30T09:00:44.419298Z","iopub.execute_input":"2022-07-30T09:00:44.419804Z","iopub.status.idle":"2022-07-30T09:00:44.51695Z","shell.execute_reply.started":"2022-07-30T09:00:44.419771Z","shell.execute_reply":"2022-07-30T09:00:44.515728Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<span style=\"color: #555555\">Okay, cool. So, after few failed attempts we are able to see the ```train.csv```. Its a good news. From the ```head``` data we can make sense of all the labels in the csv. Like, ```image_num``` means that this image sample was taken from a single patient.</span><br/>\n<span style=\"color: #555555\">Now, lets have a look at the details of the csv data.</span>","metadata":{}},{"cell_type":"code","source":"train_meta.info()","metadata":{"execution":{"iopub.status.busy":"2022-07-30T09:00:44.518648Z","iopub.execute_input":"2022-07-30T09:00:44.51934Z","iopub.status.idle":"2022-07-30T09:00:44.540665Z","shell.execute_reply.started":"2022-07-30T09:00:44.519292Z","shell.execute_reply":"2022-07-30T09:00:44.539331Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<span style=\"color: #555555\">This is not interesting much but we can say that ```center_id``` and ```image_num``` are integers, and rest are some alpha-numeric ID columns. And most importantly we have 754 images in the train dataset.<br/> \nNow, I think it's important to check out the distribution of labels amongst these records. So, lets dive right into it.</span>","metadata":{}},{"cell_type":"code","source":"train_meta.label.value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-07-30T09:00:44.543804Z","iopub.execute_input":"2022-07-30T09:00:44.544246Z","iopub.status.idle":"2022-07-30T09:00:44.557221Z","shell.execute_reply.started":"2022-07-30T09:00:44.544211Z","shell.execute_reply":"2022-07-30T09:00:44.556125Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<span style=\"color: #555555\">This is something interesting for us. You can clearly figure out that there are only 2 types of labels in our train dataset and the ```CE``` label is the dominant one. Let's try to represent it using a graph to showcase this difference more clearly and visually asthetically. </span>","metadata":{}},{"cell_type":"code","source":"plt.gca().set_facecolor(\"#F3F3F3\")\ntrain_meta.label.value_counts().plot(kind='bar', color=\"#555555\")","metadata":{"execution":{"iopub.status.busy":"2022-07-30T09:00:44.558993Z","iopub.execute_input":"2022-07-30T09:00:44.560323Z","iopub.status.idle":"2022-07-30T09:00:44.766408Z","shell.execute_reply.started":"2022-07-30T09:00:44.560273Z","shell.execute_reply":"2022-07-30T09:00:44.765279Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<span style=\"color: #555555\">Okay, its visually clear that ```CE``` is more common in our train dataset then LAA.</span><br/>\n<span style=\"color: #555555\">Lets explore the metadata for **test dataset** now.</span>","metadata":{}},{"cell_type":"code","source":"test_meta = pd.read_csv('../input/mayo-clinic-strip-ai/test.csv')\ntest_meta.head().style.set_properties(**{\n    'background-color': '#363636',\n    'color': '#F3F3F3',\n    'border': '1px solid #606060'\n})","metadata":{"execution":{"iopub.status.busy":"2022-07-30T09:00:44.768014Z","iopub.execute_input":"2022-07-30T09:00:44.768484Z","iopub.status.idle":"2022-07-30T09:00:44.789626Z","shell.execute_reply.started":"2022-07-30T09:00:44.768434Z","shell.execute_reply":"2022-07-30T09:00:44.788286Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<span style=\"color: #555555\">Okay, this doesn't make much sense so lets dive into the details of our test dataset.</span>","metadata":{}},{"cell_type":"code","source":"test_meta.info()","metadata":{"execution":{"iopub.status.busy":"2022-07-30T09:00:44.791142Z","iopub.execute_input":"2022-07-30T09:00:44.791476Z","iopub.status.idle":"2022-07-30T09:00:44.804458Z","shell.execute_reply.started":"2022-07-30T09:00:44.791444Z","shell.execute_reply":"2022-07-30T09:00:44.80323Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<span style=\"color: #555555\">One very important thing to note here is that we only have data for **4** records in the test folder, so we can assume that the testing data for this competition is very limited. Or maybe we are missing something here. To be 100% sure we will look at the image data once we have gone through all the metadatas.</span>","metadata":{}},{"cell_type":"markdown","source":"<span style=\"color: #555555\">Let's explore the metadata for **others**.</span>","metadata":{}},{"cell_type":"code","source":"others_meta = pd.read_csv('../input/mayo-clinic-strip-ai/other.csv')\nothers_meta.head().style.set_properties(**{\n    'background-color': '#363636',\n    'color': '#F3F3F3',\n    'border': '1px solid #606060'\n})","metadata":{"execution":{"iopub.status.busy":"2022-07-30T09:00:44.806408Z","iopub.execute_input":"2022-07-30T09:00:44.806994Z","iopub.status.idle":"2022-07-30T09:00:44.835909Z","shell.execute_reply.started":"2022-07-30T09:00:44.806938Z","shell.execute_reply":"2022-07-30T09:00:44.833434Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<span style=\"color: #555555\">There is something interesting going on here. If you look closely you'll find a new col ```other_specified``` which wasn't there in other 2 datasets. And surprisingly it has a varied range of values. But I am not able to comprehend the requirement of including this dataset and how I can utilize it to improve the performance of my model. We will see to this at later point.</span><br/>\n<span style=\"color: #555555\">Lets have a look at the different values of ```other_specified```.</span>","metadata":{}},{"cell_type":"code","source":"others_meta.other_specified.value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-07-30T09:00:44.837343Z","iopub.execute_input":"2022-07-30T09:00:44.837813Z","iopub.status.idle":"2022-07-30T09:00:44.846399Z","shell.execute_reply.started":"2022-07-30T09:00:44.837778Z","shell.execute_reply":"2022-07-30T09:00:44.845204Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<span style=\"color: #555555\">We have quite a variety here. For the sake of simplicity and also being a fan of visual graphics lets plot this data and see how does it look.</span>","metadata":{}},{"cell_type":"code","source":"plt.gca().set_facecolor(\"#F3F3F3\")\nothers_meta.other_specified.value_counts().plot(kind='bar', color='#555555')","metadata":{"execution":{"iopub.status.busy":"2022-07-30T09:00:44.848387Z","iopub.execute_input":"2022-07-30T09:00:44.848816Z","iopub.status.idle":"2022-07-30T09:00:45.056428Z","shell.execute_reply.started":"2022-07-30T09:00:44.848784Z","shell.execute_reply":"2022-07-30T09:00:45.055143Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<span style=\"color: #555555\">You can see that ```Dissection``` is the most common label for the data. But again not sure what to do with this information. As it is famously known that information without means/context is of no use to anyone. So, we'll see how we can utilize this set of data in our model.</span>","metadata":{}},{"cell_type":"markdown","source":"<h2 style=\"color: #363636; font-family: Segoe UI; font-size: 1.9em; font-weight: 600;\">Exploring Images in the Dataset</h2>\n\n<span style=\"color: #555555\">Here comes the most important part of our data exploration. We are finally going to look at the images that we have to work upon. And I am honestly hoping that it will solve some of our doubts from the metadata exploration.</span><br/>\n\n<span style=\"color: #555555\">Okay, then. Let's dive into the colorful world of images!</span>","metadata":{}},{"cell_type":"markdown","source":"<span style=\"color: #555555\">But first, lets install ```imutils``` so that we can play around with the images more efficiently and a bit with ease of use.</span>","metadata":{}},{"cell_type":"code","source":"!pip install imutils","metadata":{"execution":{"iopub.status.busy":"2022-07-30T09:00:45.057896Z","iopub.execute_input":"2022-07-30T09:00:45.058246Z","iopub.status.idle":"2022-07-30T09:00:52.446266Z","shell.execute_reply.started":"2022-07-30T09:00:45.058204Z","shell.execute_reply":"2022-07-30T09:00:52.445142Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<span style=\"color: #555555\">So ```imutils``` is installed. Let's import some of its magic in our notebook.</span>","metadata":{}},{"cell_type":"code","source":"from imutils import paths","metadata":{"execution":{"iopub.status.busy":"2022-07-30T09:00:52.447876Z","iopub.execute_input":"2022-07-30T09:00:52.448803Z","iopub.status.idle":"2022-07-30T09:00:52.642175Z","shell.execute_reply.started":"2022-07-30T09:00:52.448762Z","shell.execute_reply":"2022-07-30T09:00:52.641224Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<span style=\"color: #555555\">**A question:** ```What is the right way to perform a set of operations multiple times with some different parameters?```<br/> You might be tempted to say **```loops```** but here I am referring to **```functions```** that can be called multiple times by simply varying the input params. So, lets create a function to read all the images for all three of the data types, and also to get the image id for each of the images.</span>","metadata":{}},{"cell_type":"code","source":"'''\n- function: get_images\n- parameters: folder_path = string (for the path to the images dataset); folder_name = string (for the name of the folder containing the images)\n- return: array of images\n- description: to get all the images inside the given folder\n'''\ndef get_images(folder_path, folder_name):\n    images = sorted(list(paths.list_images(folder_path)))\n    print(f'There are {len(images)} images in {folder_name} folder.')\n    return images","metadata":{"execution":{"iopub.status.busy":"2022-07-30T09:00:52.647625Z","iopub.execute_input":"2022-07-30T09:00:52.648113Z","iopub.status.idle":"2022-07-30T09:00:52.654501Z","shell.execute_reply.started":"2022-07-30T09:00:52.648065Z","shell.execute_reply":"2022-07-30T09:00:52.653573Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"'''\n- function: get_images_ids\n- parameters: images_array = array (for the images whose ids we need)\n- return: array of image ids\n- description: to get all the IDs for given set of images\n'''\ndef get_images_ids(images_array):\n    return [img.split('/')[-1].rstrip('.tif') for img in images_array]","metadata":{"execution":{"iopub.status.busy":"2022-07-30T09:00:52.655833Z","iopub.execute_input":"2022-07-30T09:00:52.656172Z","iopub.status.idle":"2022-07-30T09:00:52.666027Z","shell.execute_reply.started":"2022-07-30T09:00:52.656141Z","shell.execute_reply":"2022-07-30T09:00:52.664426Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<span style=\"color: #555555\">Okay, so our functions for getting the images and their respective image ids is completed. Now, lets get all the images we can get from our dataset. </span>","metadata":{}},{"cell_type":"markdown","source":"<h2 style=\"color: #363636; font-family: Segoe UI; font-size: 1.5em; font-weight: 600;\">Train Images Dataset </h2>","metadata":{}},{"cell_type":"code","source":"train_images = get_images('../input/mayo-clinic-strip-ai/train', 'train')","metadata":{"execution":{"iopub.status.busy":"2022-07-30T09:00:52.66789Z","iopub.execute_input":"2022-07-30T09:00:52.668832Z","iopub.status.idle":"2022-07-30T09:00:52.684145Z","shell.execute_reply.started":"2022-07-30T09:00:52.668741Z","shell.execute_reply":"2022-07-30T09:00:52.682872Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<span style=\"color: #555555\">Let's get the ids for train images as well.</span>","metadata":{}},{"cell_type":"code","source":"train_ids = get_images_ids(train_images)","metadata":{"execution":{"iopub.status.busy":"2022-07-30T09:00:52.686523Z","iopub.execute_input":"2022-07-30T09:00:52.687394Z","iopub.status.idle":"2022-07-30T09:00:52.693208Z","shell.execute_reply.started":"2022-07-30T09:00:52.687346Z","shell.execute_reply":"2022-07-30T09:00:52.692363Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h2 style=\"color: #363636; font-family: Segoe UI; font-size: 1.5em; font-weight: 600;\">Test Images Dataset </h2>","metadata":{}},{"cell_type":"code","source":"test_images = get_images('../input/mayo-clinic-strip-ai/test', 'test')","metadata":{"execution":{"iopub.status.busy":"2022-07-30T10:04:22.399561Z","iopub.execute_input":"2022-07-30T10:04:22.400544Z","iopub.status.idle":"2022-07-30T10:04:22.406372Z","shell.execute_reply.started":"2022-07-30T10:04:22.4005Z","shell.execute_reply":"2022-07-30T10:04:22.405253Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<span style=\"color: #555555\">**```Did you notice something here?```**<br/>No?<br/>Don't worry let me explain and show it to you. Do you remember when we were going over the test images metadata we encountered that there are only 4 records in the csv. And we had a strange feeling about it! Well, that strange feeling is correct, the test dataset contains only 4 images onto which we have to test our model. It is quite challenging as we might not be able to completely interpret the performance of our model. But again it's okay, we'll see what, when, and how things unfold. Let's move on for now!</span>","metadata":{}},{"cell_type":"code","source":"test_ids = get_images_ids(test_images)\ntest_ids","metadata":{"execution":{"iopub.status.busy":"2022-07-30T10:04:46.805921Z","iopub.execute_input":"2022-07-30T10:04:46.806646Z","iopub.status.idle":"2022-07-30T10:04:46.814739Z","shell.execute_reply.started":"2022-07-30T10:04:46.806591Z","shell.execute_reply":"2022-07-30T10:04:46.813417Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h2 style=\"color: #363636; font-family: Segoe UI; font-size: 1.5em; font-weight: 600;\">Other Images Dataset </h2>","metadata":{}},{"cell_type":"code","source":"other_images = get_images('../input/mayo-clinic-strip-ai/other', 'other')","metadata":{"execution":{"iopub.status.busy":"2022-07-30T09:00:52.717199Z","iopub.execute_input":"2022-07-30T09:00:52.717743Z","iopub.status.idle":"2022-07-30T09:00:52.73127Z","shell.execute_reply.started":"2022-07-30T09:00:52.717712Z","shell.execute_reply":"2022-07-30T09:00:52.729748Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"other_ids = get_images_ids(other_images)","metadata":{"execution":{"iopub.status.busy":"2022-07-30T09:00:52.732786Z","iopub.execute_input":"2022-07-30T09:00:52.733994Z","iopub.status.idle":"2022-07-30T09:00:52.739238Z","shell.execute_reply.started":"2022-07-30T09:00:52.73395Z","shell.execute_reply":"2022-07-30T09:00:52.738357Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<span style=\"color: #555555\">**Its time!** (in Vecna's voice)<br/>It's time to explore the images finally! So, let's dig into it.<br/>But again we are going to open a large number of images so you might have guessed by now that we are going to utilize a function to make this work a bit clean for us. So, let's implement the function to read an image.</span>","metadata":{}},{"cell_type":"code","source":"'''\n- function: read_image\n- parameters: img_path = string (absolute path of the image to be read)\n- return: image (image instance of the tif file), and string (name of the image read)\n- description: to read the image from given path\n'''\ndef read_image(img_path):\n    img = tifi.imread(img_path)\n    \n    # the file name of the image that we are reading\n    img_name = img_path.split('/')[-1].rstrip('tif')\n    print(f'Read image {img_name}')\n    \n    return img, img_name","metadata":{"execution":{"iopub.status.busy":"2022-07-30T09:00:52.740577Z","iopub.execute_input":"2022-07-30T09:00:52.741698Z","iopub.status.idle":"2022-07-30T09:00:52.74978Z","shell.execute_reply.started":"2022-07-30T09:00:52.74165Z","shell.execute_reply":"2022-07-30T09:00:52.748926Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<span style=\"color: #555555\">Let's read some random image from our train dataset.<br/>*When I was trying to plot a random image from our train dataset I observed that the **```CPU``` utilization was overshooting and soon we ```ran out of memory```**. So, I decided to pick an image that is already present in the notebook from the very beginning, so that CPU doesn't have to do much work to search for that image, and then read it out.*</span>","metadata":{}},{"cell_type":"code","source":"img, img_name = read_image(train_images[2])","metadata":{"execution":{"iopub.status.busy":"2022-07-30T09:03:30.082966Z","iopub.execute_input":"2022-07-30T09:03:30.083923Z","iopub.status.idle":"2022-07-30T09:03:50.333539Z","shell.execute_reply.started":"2022-07-30T09:03:30.08386Z","shell.execute_reply":"2022-07-30T09:03:50.332212Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<span style=\"color: #555555\">*Well it turns out the CPU utilization overshooting issue was not because of accessing some random image from dataset. And I'm still not able to find out the reason for it. Let's proceed..*</span>","metadata":{}},{"cell_type":"markdown","source":"<span style=\"color: #555555\">We have an instance of our train image. Let's deep dive into this image and see what we can understand about all the other images with the help of this random image sample.<br/>The first thing that comes to my mind is to check out the size of this image, and then plot the image on the notebook so that we can see what actually is it! So, lets do it!</span>","metadata":{}},{"cell_type":"code","source":"img.shape","metadata":{"execution":{"iopub.status.busy":"2022-07-30T09:00:54.478513Z","iopub.status.idle":"2022-07-30T09:00:54.479627Z","shell.execute_reply.started":"2022-07-30T09:00:54.479305Z","shell.execute_reply":"2022-07-30T09:00:54.479337Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<span style=\"color: #555555\">You can observe that the size of this image is quite large, and if the size of this random image in train set is so big, it is safe to assume that the size of other images will also be this large.<br/> Another thing to note here is the fact that this image uses 3 channels, i.e, it can be either RGB image, HSV image or another 3 channel image. Hence, it has to be a colored image.</span>","metadata":{}},{"cell_type":"markdown","source":"<span style=\"color: #555555\">Now, I don't know if this question struck you or not, but I wondered **```why is the shape/size of an image so important to investigate```** that before even looking at the image we studied its shape.<br/>Well, that's because we are supposed to classify these images. And what's better approach then **Neural Networks** classify images with high accuracy. ```But still where does the size of image comes in??```<br/>As soon as you talk about ```Neural Networks``` the size/shape/dimension of the object in question becomes important. Because Neural Networks play around with the dimension of its input breaking it down and then building it up in order to learn new features about it. Hence, we need to understand the shape of the input image before even looking at it.</span>","metadata":{}},{"cell_type":"markdown","source":"<span style=\"color: #555555\">I hope this clears your doubt but if it doesn't don't worry most people don't and it really doesn't come in between your progress as you will keep hitting at it until one day it will start making sense automatically.<br/>Without any further adieu, let's look at our image!<br/>But first let's create a function to display images as we are going to need it so much!</span>","metadata":{}},{"cell_type":"code","source":"'''\n- function: show_image\n- parameters: img = Image (image instance of the image file)\n- return: none\n- description: to plot the image using matplotlib\n'''\ndef show_image(img):\n    plt.figure(figsize=(10,10))\n    plt.imshow(img, cmap='Greys')\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-30T09:16:51.169752Z","iopub.execute_input":"2022-07-30T09:16:51.170233Z","iopub.status.idle":"2022-07-30T09:16:51.17724Z","shell.execute_reply.started":"2022-07-30T09:16:51.170196Z","shell.execute_reply":"2022-07-30T09:16:51.175621Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<span style=\"color: #555555\">*Okay so due to the excessive size of the image we were running out of memory, so we should first scale down the image, then have a look at it.*<br/>So, let's create a function to scale down the image or better let's create a centralized function to resize the image based on the params to either scale it up or scale it down. Okay enough talking lets code it out!</span>","metadata":{}},{"cell_type":"code","source":"'''\n- function: resize_image\n- parameters: img = Image (an Image instance); resize_op = string (string to whether scale up or down the image) (accepted values => UP: to scale up, scale down for everthing else)\n- return: image (resized image)\n- description: to resize the image based on the operation\n'''\ndef resize_image(img, resize_op='DOWN'):\n    if resize_op == 'UP':\n        return cv.resize(img, (int(img.shape[1]*33)), int(img.shape[0]*33), interpolation=cv.INTER_LINEAR)\n    return cv.resize(img,(int(img.shape[1]/33),int(img.shape[0]/33)),interpolation= cv.INTER_LINEAR)","metadata":{"execution":{"iopub.status.busy":"2022-07-30T09:23:05.902847Z","iopub.execute_input":"2022-07-30T09:23:05.903696Z","iopub.status.idle":"2022-07-30T09:23:05.911258Z","shell.execute_reply.started":"2022-07-30T09:23:05.903646Z","shell.execute_reply":"2022-07-30T09:23:05.910164Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"resz_img = resize_image(img, 'DOWN')","metadata":{"execution":{"iopub.status.busy":"2022-07-30T09:23:08.558914Z","iopub.execute_input":"2022-07-30T09:23:08.559383Z","iopub.status.idle":"2022-07-30T09:23:08.580251Z","shell.execute_reply.started":"2022-07-30T09:23:08.559339Z","shell.execute_reply":"2022-07-30T09:23:08.579336Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<span style=\"color: #555555\">Now I am hoping we will be able to look at the image, finally!</span>","metadata":{}},{"cell_type":"code","source":"show_image(resz_img)","metadata":{"execution":{"iopub.status.busy":"2022-07-30T09:24:13.848664Z","iopub.execute_input":"2022-07-30T09:24:13.849749Z","iopub.status.idle":"2022-07-30T09:24:14.196582Z","shell.execute_reply.started":"2022-07-30T09:24:13.849682Z","shell.execute_reply":"2022-07-30T09:24:14.195285Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<span style=\"color: #555555\">**Yayy!** So we are able to finally see our train image. But here's something interesting our train image doesn't have a single cell in it. Rather it has multiple cells and I think we have to segment these cells so that we can improve our system's performance. Before this I think it'll be helpful to have a look at one of the test image so that we can be sure that we need to segment our input images.<br/>So, let's explore our **test images**.</span>","metadata":{}},{"cell_type":"code","source":"test_img, test_img_name = read_image(test_images[1])\nresz_test_img = resize_image(test_img)","metadata":{"execution":{"iopub.status.busy":"2022-07-30T10:19:06.242829Z","iopub.execute_input":"2022-07-30T10:19:06.243298Z","iopub.status.idle":"2022-07-30T10:19:08.884492Z","shell.execute_reply.started":"2022-07-30T10:19:06.243259Z","shell.execute_reply":"2022-07-30T10:19:08.883195Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"show_image(resz_test_img)","metadata":{"execution":{"iopub.status.busy":"2022-07-30T10:19:11.351461Z","iopub.execute_input":"2022-07-30T10:19:11.351857Z","iopub.status.idle":"2022-07-30T10:19:11.683824Z","shell.execute_reply.started":"2022-07-30T10:19:11.351827Z","shell.execute_reply":"2022-07-30T10:19:11.683014Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<span style=\"color: #555555\">Okay, this is something we didn't expect. So, I think we don't have to segment our train dataset, and we can directly apply some classifier to our input dataset, and see how it does. But I still have some speculations about the performance of our classifier as it may infer some noise or irrelevant information as a feature. But I think it's something we should bother about later when we will implement the classifier.</span>","metadata":{}},{"cell_type":"markdown","source":"<h2 style=\"color: #363636; font-family: Segoe UI; font-size: 1.9em; font-weight: 600;\">Closing Thoughts</h2>\n\n* <span style=\"color: #555555\">We started by looking the metadata for our images dataset. And we found out some interesting facts about the images.</span>\n* <span style=\"color: #555555\">While exploring the metadata we encountered the others dataset and we did some search to find out whether this dataset makes sense or not. And till now we are in doubt.</span>\n* <span style=\"color: #555555\">Then we finally got a chance to explore the images. And we found out that in our train set as well as our test set we have a group of tissues as a single input. This lead to us speculating on the usage of segmentation to improve the performance of our classifier.</span>","metadata":{}},{"cell_type":"markdown","source":"<h2 style=\"color: #363636; font-family: Segoe UI; font-size: 1.9em; font-weight: 600;\">What's Next?</h2>\n\n<span style=\"color: #555555\">Well, I look forward to reading some research papers around the application of Deep Learning to classify tissue images. I hope that within the monsterously large library of these research papers I might be able to find something that could answer all my questions and allow me to build a decent classifier.<br/>Don't worry I am planning on explaining and showcasing everything that I find out.<br/>Until then, **```Happy Pondering```**...</span>","metadata":{}}]}