{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"All credits go to @vslaykovsky!\n\nNever knew Connected Components Analysis before! This notebook explains about **Connected Components Analysis** and the notebook code https://www.kaggle.com/code/vslaykovsky/rsna-cut-off-empty-space-from-images/notebook\n\n**Upvote if you think this is useful!** 👋\n\n","metadata":{}},{"cell_type":"markdown","source":"# RSNA: Cut Off Empty Space from Images\n\nThis notebook generates images with all empty space removed. \nThis magnifies important parts of the data and helps to train better models with the same compute. \n\nNote that extra artifacts such as labels, noise are also removed from images!\n\nIt uses pregenerated PNGs from https://www.kaggle.com/code/theoviel/dicom-resized-png-jpg as an input.\n\nSo if you want to use this preprocessing in inference:\n1. Run code from https://www.kaggle.com/code/theoviel/dicom-resized-png-jpg on test data first.\n2. Run the code from this notebook on the output of (1)\n\nSee https://www.kaggle.com/vslaykovsky/infer-effnetv2-aux-targets-weighted-loss-thres for an example use in inference. \n\n\n**UPD**\n\n\n\n**Upvote if you think this is awesome!**","metadata":{}},{"cell_type":"code","source":"from concurrent.futures import ProcessPoolExecutor\nimport cv2\nimport glob\nimport numpy as np\nfrom tqdm import tqdm\nimport os\nfrom PIL import Image\nimport matplotlib.pyplot as plt\nimport re\n\nDEBUG = False ","metadata":{"execution":{"iopub.status.busy":"2023-04-23T21:37:18.589637Z","iopub.execute_input":"2023-04-23T21:37:18.590241Z","iopub.status.idle":"2023-04-23T21:37:18.59831Z","shell.execute_reply.started":"2023-04-23T21:37:18.59019Z","shell.execute_reply":"2023-04-23T21:37:18.597003Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def fit_image(fname):\n    X = cv2.imread(fname)\n\n    # Some images have narrow exterior \"frames\" that complicate selection of the main data. Cutting off the frame\n    X = X[5:-5, 5:-5]\n    \n    # regions of non-empty pixels\n    output= cv2.connectedComponentsWithStats((X > 20).astype(np.uint8)[:, :, 0], 8, cv2.CV_32S)\n\n    # stats.shape == (N, 5), where N is the number of regions, 5 dimensions correspond to:\n    # left, top, width, height, area_size\n    stats = output[2]\n\n    # finding max area which always corresponds to the breast data. \n    idx = stats[1:, 4].argmax() + 1\n    x1, y1, w, h = stats[idx][:4]\n    x2 = x1 + w\n    y2 = y1 + h\n    \n    # cutting out the breast data\n    X_fit = X[y1: y2, x1: x2]\n    \n    patient_id, im_id = re.findall('(\\d+)_(\\d+).png', os.path.basename(fname))[0]\n    os.makedirs(patient_id, exist_ok=True)\n    cv2.imwrite(f'{patient_id}/{im_id}.png', X_fit[:, :, 0])\n\ndef fit_all_images(all_images):\n    with ProcessPoolExecutor(4) as p:\n        for i in tqdm(p.map(fit_image, all_images), total=len(all_images)):\n            pass\n\n\nfit_image('/kaggle/input/rsna-breast-cancer-1024-pngs/output/52316_201080742.png')\nnp.array(Image.open('52316/201080742.png'))","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-04-23T21:37:19.240933Z","iopub.execute_input":"2023-04-23T21:37:19.241375Z","iopub.status.idle":"2023-04-23T21:37:19.751277Z","shell.execute_reply.started":"2023-04-23T21:37:19.241341Z","shell.execute_reply":"2023-04-23T21:37:19.749829Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The code defines a function called fit_image that takes a file name as an input and does the following steps:\n- Line2: It reads the image file using cv2.imread() method, which returns a numpy array of pixels. The variable X stores this array.\n- Line5: It cuts off the exterior frame of the image by slicing the array X with a margin of 5 pixels on each side. This is done to remove any unwanted noise or background from the image.\n- Line8: It uses cv2.connectedComponentsWithStats() method, which finds and labels all the connected regions of non-empty pixels in the image. This method returns a tuple of four elements: the number of labels, the labeled image, the statistics of each label, and the centroids of each label. The variable output stores this tuple.\n- Line12: It extracts the statistics of each label from the output tuple. The statistics are a numpy array of shape (N, 5), where N is the number of labels, and 5 dimensions correspond to: left, top, width, height, and area size of each label. The variable stats stores this array.\n- Line15: It finds the index of the label with the maximum area size, which corresponds to the breast data in the image. This is done by skipping the first label (which is the background) and using numpy.argmax() method on the fourth column (which is the area size) of stats array. The variable idx stores this index.\n- Line16-18: It extracts the left, top, width, and height values of the label with the maximum area size from stats array using idx. It then calculates the right and bottom values by adding width and height to left and top respectively. The variables x1, y1, x2, y2 store these values.\n- Line21: It cuts out the breast data from the image by slicing the array X with x1, y1, x2, y2 values. The variable X_fit stores this sliced array.\n- Line23: It uses regular expressions to find and extract the patient ID and image ID from the file name using re.findall() method. The variables patient_id and im_id store these values as strings.\n- Line24: It creates a directory with patient ID as its name using os.makedirs() method if it does not exist already. This is done to organize and store the images by patient ID.\n- Line25: It writes and saves the sliced image as a png file using cv2.imwrite() method in the directory with patient ID as its name. The file name is composed of image ID and png extension.","metadata":{}},{"cell_type":"markdown","source":"## Line 5: Cutting off the frame","metadata":{}},{"cell_type":"code","source":"fname = '/kaggle/input/rsna-breast-cancer-1024-pngs/output/10006_462822612.png'\na = cv2.imread(fname)\na1 = a[5:-5, 5:-5]\n# Create a figure and two subplots\nfig, (ax1, ax2) = plt.subplots(nrows=1, ncols=2)\n\n# Display the images\nax1.imshow(a)\nax2.imshow(a1)\n\n# Show the plot\nplt.show()\nprint(a.shape)\nprint(a1.shape)","metadata":{"execution":{"iopub.status.busy":"2023-04-23T21:37:24.079172Z","iopub.execute_input":"2023-04-23T21:37:24.079577Z","iopub.status.idle":"2023-04-23T21:37:24.621113Z","shell.execute_reply.started":"2023-04-23T21:37:24.079544Z","shell.execute_reply":"2023-04-23T21:37:24.619731Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The expression `(X > 20).astype(np.uint8)[:, :, 0]` performs several operations on the `X` array:\n\n1. `X > 20` creates a boolean array with the same shape as `X`, where each element is `True` if the corresponding element in `X` is greater than 20 and `False` otherwise.\n2. `.astype(np.uint8)` converts the boolean array to an array of type `np.uint8`, where `True` values are converted to 1 and `False` values are converted to 0.\n3. `[:, :, 0]` selects the first channel of the resulting 3D array.\n\nThe overall effect of this expression is to create a **binary image** where pixels with values greater than 20 in the first channel of the `X` array are set to 1 and all other pixels are set to 0. This binary image is then passed to the `cv2.connectedComponentsWithStats` function to find connected regions of non-empty pixels.","metadata":{}},{"cell_type":"code","source":"binary_img = (a1>20).astype(np.uint8)[:, :, 0]\nplt.imshow(binary_img)\nbinary_img.shape","metadata":{"execution":{"iopub.status.busy":"2023-04-23T21:37:25.193994Z","iopub.execute_input":"2023-04-23T21:37:25.194459Z","iopub.status.idle":"2023-04-23T21:37:25.574578Z","shell.execute_reply.started":"2023-04-23T21:37:25.194408Z","shell.execute_reply":"2023-04-23T21:37:25.573697Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#np.set_printoptions(threshold=np.inf)\nnp.set_printoptions(threshold=10, edgeitems=3)\nprint(binary_img)","metadata":{"execution":{"iopub.status.busy":"2023-04-23T21:37:27.359727Z","iopub.execute_input":"2023-04-23T21:37:27.36097Z","iopub.status.idle":"2023-04-23T21:37:27.36784Z","shell.execute_reply.started":"2023-04-23T21:37:27.360913Z","shell.execute_reply":"2023-04-23T21:37:27.366476Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Line8: Connected Components","metadata":{}},{"cell_type":"markdown","source":"> Connected components labeling scans an image and groups its pixels into components based on **pixel connectivity**, i.e. all pixels in a connected component share similar pixel intensity values and are in some way connected with each other. Once all groups have been determined, each pixel is labeled with a graylevel or a color (color labeling) according to the component it was assigned to.\nhttps://homepages.inf.ed.ac.uk/rbf/HIPR2/label.htm\n\n`connectedComponentsWithStats()` is a function provided by the OpenCV library that performs connected component analysis on a binary image, which is an image where each pixel is either black or white (foreground or background). The function labels each foreground pixel with a unique integer label, depending on the connected component it belongs to.\n\nThe function takes the following inputs:\n\n- A **binary image** (usually represented as a 2D numpy array) to be analyzed\n- A connectivity value (either 4 or 8) that specifies how neighboring pixels are defined\n- An optional output data type for the labeled image, which defaults to `cv2.CV_32S`\n\nThe function returns a tuple of 4 elements:\n- The total number of connected components found in the image (**including the background**)\n- The labeled image, where each foreground pixel is labeled with a unique integer corresponding to its connected component.\n- The statistics matrix is a matrix that contains information about each connected component, such as its size (number of pixels) and its bounding box coordinates.\n  - The statistics are a numpy array of shape (N, 5), where N is the number of labels, and 5 dimensions correspond to: left, top, width, height, and area size of each label.\n- The centroids matrix is a matrix that contains the centroid coordinates of each connected component\n\nBy analyzing the labeled image and the statistics array, we can extract information about the connected components in the image, such as their location, size, and shape. This information can be useful for a variety of computer vision tasks, such as object detection, image segmentation, and feature extraction.\n\n","metadata":{}},{"cell_type":"markdown","source":"```python\noutput= cv2.connectedComponentsWithStats((X > 20).astype(np.uint8)[:, :, 0], 8, cv2.CV_32S)\n```\nWe need to choose CV_32S because it can handle signed values, which are needed for the cv2.connectedComponentsWithStats() function to label different regions in the image1. CV_32S can also store a larger range of possible values than CV_8U, which is limited to 0-255. However, CV_32S is not compatible with some other functions, such as cv2.medianBlur(), and it needs to be converted to 8 bits for displaying or saving1.","metadata":{}},{"cell_type":"code","source":"(totalLabels, labels, stats, centroid) = cv2.connectedComponentsWithStats((a1 > 10).astype(np.uint8)[:, :, 0], 8, cv2.CV_32S)\nprint(\"Number of connected components:\", totalLabels)\nprint(\"The labeled image： \\n\", labels)\nprint(\"The labeled image shape:\", labels.shape)\nprint(\"The centroid coordinates of each connected component: \\n\",centroid)","metadata":{"execution":{"iopub.status.busy":"2023-04-23T21:37:34.467338Z","iopub.execute_input":"2023-04-23T21:37:34.467823Z","iopub.status.idle":"2023-04-23T21:37:34.485451Z","shell.execute_reply.started":"2023-04-23T21:37:34.467779Z","shell.execute_reply":"2023-04-23T21:37:34.484012Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# If you want to check the whole matrix, uncomment line 2 code\n# np.set_printoptions(threshold=np.inf)\nprint(\"The labeled image： \\n\", labels)\nnp.set_printoptions(threshold=5)\nnp.unique(labels)","metadata":{"execution":{"iopub.status.busy":"2023-04-23T21:37:37.103853Z","iopub.execute_input":"2023-04-23T21:37:37.104359Z","iopub.status.idle":"2023-04-23T21:37:37.13954Z","shell.execute_reply.started":"2023-04-23T21:37:37.104315Z","shell.execute_reply":"2023-04-23T21:37:37.138117Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# segment areas - first is the background\nfig, axs = plt.subplots(1, 3)\nfor i in range(totalLabels):\n    componentMask = (labels == i).astype(\"uint8\") * 255\n    output = np.zeros(a1.astype(np.uint8)[:, :, 0].shape, dtype=\"uint8\")\n    output = cv2.bitwise_or(output, componentMask)\n    axs[i].imshow(output)\nplt.subplots_adjust(hspace=1)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-23T21:37:39.806243Z","iopub.execute_input":"2023-04-23T21:37:39.807031Z","iopub.status.idle":"2023-04-23T21:37:40.574292Z","shell.execute_reply.started":"2023-04-23T21:37:39.806972Z","shell.execute_reply":"2023-04-23T21:37:40.573113Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We got three connected components:\n- the background\n- noise\n- the breast","metadata":{}},{"cell_type":"markdown","source":"## Example","metadata":{}},{"cell_type":"markdown","source":"The image has two connected regions. We can use **Connected Component Labeling and Analysis**(cv2.connectedComponentsWithStars) to extract background and two components","metadata":{}},{"cell_type":"code","source":"imgR = np.array([[  0,   0,   0,   0,   0,   0,   0,   0,   0,   0],\n       [  0,   0, 200, 200,   0, 100,   0,   0,   0,   0],\n       [  0,   0, 200, 200,   0, 100, 100,   0,   0,   0],\n       [  0,   0, 200, 200,   0,   0, 100,   0,   0,   0],\n       [  0,   0,   0,   0,   0,   0, 100, 100,   0,   0]], dtype=np.uint8)\nplt.imshow(imgR)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-23T21:37:43.229791Z","iopub.execute_input":"2023-04-23T21:37:43.230193Z","iopub.status.idle":"2023-04-23T21:37:43.419646Z","shell.execute_reply.started":"2023-04-23T21:37:43.230161Z","shell.execute_reply":"2023-04-23T21:37:43.418513Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The imgR array already represent a grayscale image. Each element in the array represents the intensity of a pixel in the image, with 0 being black and 255 being white.\n\nIf you have a color image represented as a NumPy array with shape (height, width, 3) (where the last dimension represents the red, green, and blue color channels), you can use the cvtColor function from the cv2 library to convert it to grayscale. Here’s an example:\n```python\nimport cv2\n\n# Load a color image\nimg = cv2.imread('color_image.png')\n\n# Convert the image to grayscale\ngray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)\n```","metadata":{}},{"cell_type":"code","source":"# Connected Component Labeling and Analysis\n(totalLabels, label_ids, stats, centroid) = cv2.connectedComponentsWithStats((imgR > 20).astype(np.uint8), 4, cv2.CV_32S)\n\n# segment areas - first is the background\nfig, axs = plt.subplots(3, 1)\nfor i in range(totalLabels):\n    componentMask = (label_ids == i).astype(\"uint8\") * 255\n    output = np.zeros(imgR.shape, dtype=\"uint8\")\n    output = cv2.bitwise_or(output, componentMask)\n    axs[i].imshow(output)\nplt.subplots_adjust(hspace=0.5)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-23T21:37:46.031431Z","iopub.execute_input":"2023-04-23T21:37:46.032059Z","iopub.status.idle":"2023-04-23T21:37:46.361826Z","shell.execute_reply.started":"2023-04-23T21:37:46.032003Z","shell.execute_reply":"2023-04-23T21:37:46.360524Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's analyse the results: totalLabels, label_ids, values, centroid","metadata":{}},{"cell_type":"code","source":"print(totalLabels)","metadata":{"execution":{"iopub.status.busy":"2023-04-23T21:37:48.974761Z","iopub.execute_input":"2023-04-23T21:37:48.975641Z","iopub.status.idle":"2023-04-23T21:37:48.984688Z","shell.execute_reply.started":"2023-04-23T21:37:48.975564Z","shell.execute_reply":"2023-04-23T21:37:48.983159Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"label_ids","metadata":{"execution":{"iopub.status.busy":"2023-04-23T21:37:55.264471Z","iopub.execute_input":"2023-04-23T21:37:55.264893Z","iopub.status.idle":"2023-04-23T21:37:55.272314Z","shell.execute_reply.started":"2023-04-23T21:37:55.264856Z","shell.execute_reply":"2023-04-23T21:37:55.270896Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(stats)","metadata":{"execution":{"iopub.status.busy":"2023-04-23T21:37:57.915202Z","iopub.execute_input":"2023-04-23T21:37:57.915748Z","iopub.status.idle":"2023-04-23T21:37:57.923521Z","shell.execute_reply.started":"2023-04-23T21:37:57.915689Z","shell.execute_reply":"2023-04-23T21:37:57.921735Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(centroid)","metadata":{"execution":{"iopub.status.busy":"2023-04-23T21:38:00.625472Z","iopub.execute_input":"2023-04-23T21:38:00.626113Z","iopub.status.idle":"2023-04-23T21:38:00.633577Z","shell.execute_reply.started":"2023-04-23T21:38:00.626072Z","shell.execute_reply":"2023-04-23T21:38:00.632103Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"all_images = glob.glob('/kaggle/input/rsna-breast-cancer-1024-pngs/output/*')\nif DEBUG:\n    all_images = np.random.choice(all_images, size=100)\nfit_all_images(all_images)","metadata":{"execution":{"iopub.status.busy":"2023-04-23T20:19:18.822901Z","iopub.execute_input":"2023-04-23T20:19:18.823861Z","iopub.status.idle":"2023-04-23T20:34:48.477001Z","shell.execute_reply.started":"2023-04-23T20:19:18.823814Z","shell.execute_reply":"2023-04-23T20:34:48.474567Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"np.random.seed(123)\n\nfor fname in np.random.choice(glob.glob('*/*'), size=20):\n    plt.figure(figsize=(20, 10))\n    patient_id, im_id = re.findall('(\\d+)/(\\d+).png', fname)[0]\n    plt.suptitle(f'[{fname}]')\n    im1 = Image.open(fname).convert('F')\n    plt.subplot(121).imshow(im1)\n    plt.subplot(121).set_title(f'Output image {im1.size}')\n    im2 = Image.open(f'/kaggle/input/rsna-breast-cancer-1024-pngs/output/{patient_id}_{im_id}.png').convert('F')\n    plt.subplot(122).imshow(\n        im2\n    )\n    plt.subplot(122).set_title(f'Source image {im2.size}')","metadata":{"execution":{"iopub.status.busy":"2023-04-23T21:38:22.143668Z","iopub.execute_input":"2023-04-23T21:38:22.145087Z","iopub.status.idle":"2023-04-23T21:38:40.871552Z","shell.execute_reply.started":"2023-04-23T21:38:22.145028Z","shell.execute_reply":"2023-04-23T21:38:40.870452Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the 20 randomly choosed images above, we notice that noises like text are removed from the background and we cut off the breast part successfully.","metadata":{}},{"cell_type":"code","source":"import pandas as pd\ndf = pd.DataFrame(dict(zip(('y', 'x', 'c'), cv2.imread(i).shape)) for i in glob.glob('*/*.png'))","metadata":{"execution":{"iopub.status.busy":"2023-04-23T20:51:51.906381Z","iopub.execute_input":"2023-04-23T20:51:51.907039Z","iopub.status.idle":"2023-04-23T20:58:29.276913Z","shell.execute_reply.started":"2023-04-23T20:51:51.906976Z","shell.execute_reply":"2023-04-23T20:58:29.275875Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df[['x', 'y']].plot.hist(alpha=0.7, bins=100)","metadata":{"execution":{"iopub.status.busy":"2023-04-23T21:39:07.14245Z","iopub.execute_input":"2023-04-23T21:39:07.142907Z","iopub.status.idle":"2023-04-23T21:39:08.147322Z","shell.execute_reply.started":"2023-04-23T21:39:07.142868Z","shell.execute_reply":"2023-04-23T21:39:08.146102Z"},"trusted":true},"execution_count":null,"outputs":[]}]}