{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Prostate Cancer Notebook\n\nProstate cancer detection and EDA procedures\n","metadata":{}},{"cell_type":"markdown","source":"# References\n\n* https://www.kaggle.com/competitions/prostate-cancer-grade-assessment/data\n* https://www.kaggle.com/wouterbulten/getting-started-with-the-panda-dataset\n* https://www.kaggle.com/rohitsingh9990/panda-eda-better-visualization-simple-baseline","metadata":{}},{"cell_type":"markdown","source":"\n\n### What is Prostate Cancer?\nProstate cancer is cancer that occurs in the prostate ,a small walnut-shaped gland in men that produces the seminal fluid that nourishes and transports sperm.\n\nProstate cancer is one of the most common types of cancer in men. Usually prostate cancer grows slowly and is initially confined to the prostate gland, where it may not cause serious harm. However, while some types of prostate cancer grow slowly and may need minimal or even no treatment, other types are aggressive and can spread quickly.\n\n<img src=\"https://www.mayoclinic.org/-/media/kcms/gbs/patient-consumer/images/2013/11/15/17/38/ds00043_-my01633_im01561_prostca1thu_jpg.jpg\" height=\"100px\">\n\n### How it is tested and detected?\nProstate screening tests might include:\n\n* Digital rectal exam (DRE): During a DRE, your doctor inserts a gloved, lubricated finger into your rectum to examine your prostate, which is adjacent to the rectum. If your doctor finds any abnormalities in the texture, shape or size of the gland, you may need further tests.\n* Prostate-specific antigen (PSA) test: A blood sample is drawn from a vein in your arm and analyzed for PSA, a substance that's naturally produced by your prostate gland. It's normal for a small amount of PSA to be in your bloodstream. However, if a higher than normal level is found, it may indicate prostate infection, inflammation, enlargement or cancer.\n\nIf a DRE or PSA test detects an abnormality, your doctor may recommend further tests to determine whether you have prostate cancer, such as:\n\n* Ultrasound : If other tests raise concerns, your doctor may use transrectal ultrasound to further evaluate your prostate. A small probe, about the size and shape of a cigar, is inserted into your rectum. The probe uses sound waves to create a picture of your prostate gland.\n* Collecting a sample of prostate tissue : If initial test results suggest prostate cancer, your doctor may recommend a procedure to collect a sample of cells from your prostate (prostate biopsy). Prostate biopsy is often done using a thin needle that's inserted into the prostate to collect tissue. The tissue sample is analyzed in a lab to determine whether cancer cells are present.","metadata":{}},{"cell_type":"markdown","source":"### What is ISUP grade now?\nAccording to current guidelines by the International Society of Urological Pathology (ISUP), the Gleason scores are summarized into an ISUP grade on a scale from 1 to 5 according to the following rule:\n\n* Gleason score 6 = ISUP grade 1\u2028\n* Gleason score 7 (3 + 4) = ISUP grade 2\u2028\n* Gleason score 7 (4 + 3) = ISUP grade 3\u2028\n* Gleason score 8 = ISUP grade 4\u2028\n* Gleason score 9-10 = ISUP grade 5\u2028\n\nIf there is no cancer in the sample, we use the label ISUP grade 0 in this competition. \n\n<img src=\"https://storage.googleapis.com/kaggle-media/competitions/PANDA/Screen%20Shot%202020-04-08%20at%202.03.53%20PM.png\" height=\"100px\">\n\n### How has the Gleason scores been generated in the dataset?\nEach WSI in this challenge contains one, or in some cases two, thin tissue sections cut from a single biopsy sample. Prior to scanning, the tissue is stained with haematoxylin & eosin (H&E). This is a standard way of staining the originally transparent tissue to produce some contrast. The samples are made up of glandular tissue and connective tissue. The glands are hollow structures, which can be seen as white “holes” or branched cavities in the WSI. The appearance of the glands forms the basis of the Gleason grading system. The glandular structure characteristic of healthy prostate tissue is progressively lost with increasing grade. The grading system recognizes three categories: 3, 4, and 5. \n\n* [A]Benign prostate glands with folded epithelium :The cytoplasm is pale and the nuclei small and regular. The glands are grouped together.\n* [B]Prostatic adenocarcinoma : Gleason Pattern 3 has no loss of glandular differentiation. Small glands infiltrate between benign glands. The cytoplasm is often dark and the nuclei enlarged with dark chromatin and some prominent nucleoli. Each epithelial unit is separate and has a lumen.\n* [C]Prostatic adenocarcinoma : Gleason Pattern 4 has partial loss of glandular differentiation. There is an attempt to form lumina but the tumor fails to form complete, well-developed glands. This microphotograph shows irregular cribriform cancer, i.e. epithelial sheets with multiple lumina. There are also some poorly formed small glands and some fused glands. All of these are included in Gleason Pattern 4.\n* [D]Prostatic adenocarcinoma : Gleason Pattern 5 has an almost complete loss of glandular differentiation. Dispersed single cancer cells are seen in the stroma. Gleason Pattern 5 may also contain solid sheets or strands of cancer cells. All microphotographs show hematoxylin and eosin stains at 20x lens magnification.\n\n<img src=\"https://storage.googleapis.com/kaggle-media/competitions/PANDA/GleasonPattern_4squares%20copy500.png\" height=\"100px\">","metadata":{}},{"cell_type":"markdown","source":"# Import the Libraries ","metadata":{}},{"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport os\n\n# Visualizing data\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport PIL\nfrom IPython.display import Image, display\nfrom plotly import graph_objs as go # Plotly for the interactive viewer\nimport plotly.express as px\nimport plotly.figure_factory as ff\n\n#import openslide\n","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-03-20T17:33:51.465316Z","iopub.execute_input":"2023-03-20T17:33:51.46566Z","iopub.status.idle":"2023-03-20T17:33:51.471539Z","shell.execute_reply.started":"2023-03-20T17:33:51.465627Z","shell.execute_reply":"2023-03-20T17:33:51.470537Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load Data","metadata":{}},{"cell_type":"code","source":"\n# Input data files are available in the read-only \"../input/\" directory\ntrain_data = pd.read_csv(\"/kaggle/input/prostate-cancer-grade-assessment/train.csv\")\ntest_data = pd.read_csv(\"/kaggle/input/prostate-cancer-grade-assessment/test.csv\")\nsub_data = pd.read_csv(\"/kaggle/input/prostate-cancer-grade-assessment/sample_submission.csv\")","metadata":{"execution":{"iopub.status.busy":"2023-03-20T17:33:59.985482Z","iopub.execute_input":"2023-03-20T17:33:59.986406Z","iopub.status.idle":"2023-03-20T17:34:00.02245Z","shell.execute_reply.started":"2023-03-20T17:33:59.986349Z","shell.execute_reply":"2023-03-20T17:34:00.021409Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Explore Data","metadata":{}},{"cell_type":"code","source":"train_data.head()# view the first 5 rows of data","metadata":{"execution":{"iopub.status.busy":"2023-03-20T17:34:05.626457Z","iopub.execute_input":"2023-03-20T17:34:05.626886Z","iopub.status.idle":"2023-03-20T17:34:05.64083Z","shell.execute_reply.started":"2023-03-20T17:34:05.626842Z","shell.execute_reply":"2023-03-20T17:34:05.639782Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Definitions of the major columns and the labels from each institute\n[train/test].csv\n\n**image_id:** ID code for the image.\n\ndata_provider: The name of the institution that provided the data. Both the Karolinska Institute and Radboud University Medical Center contributed data. They used different scanners with slightly different maximum microscope resolutions and worked with different pathologists for labeling their images.\n\n**isup_grade:** Train only. The target variable. The severity of the cancer on a 0-5 scale.\n\n**gleason_score:** Train only. An alternate cancer severity rating system with more levels than the ISUP scale. \n\n[train/test]_images: The images. Each is a large multi-level tiff file. You can expect roughly 1,000 images in the hidden test set. Note that slightly different procedures were in place for the images used in the test set than the training set. Some of the training set images have stray pen marks on them, but the test set slides are free of pen marks.\n\n**train_label_masks:** Segmentation masks showing which parts of the image led to the ISUP grade. Not all training images have label masks, and there may be false positives or false negatives in the label masks for a variety of reasons. These masks are provided to assist with the development of strategies for selecting the most useful subsamples of the images. The mask values depend on the data provider:\n\n**Radboud: Prostate glands are individually labelled. Valid values are:**\n\n0: background (non tissue) or unknown\n\n1: stroma (connective tissue, non-epithelium tissue)\n\n2: healthy (benign) epithelium\n\n3: cancerous epithelium (Gleason 3)\n\n4: cancerous epithelium (Gleason 4)\n\n5: cancerous epithelium (Gleason 5)\n\n**Karolinska: Regions are labelled. Valid values are:**\n\n0: background (non tissue) or unknown\n\n1: benign tissue (stroma and epithelium combined)\n\n2: cancerous tissue (stroma and epithelium combined)","metadata":{}},{"cell_type":"code","source":"#number of rows and columns in the train data \nprint('Number of rows and columns:{}' .format(train_data.shape))\nprint('Data providers: {}'.format(train_data[\"data_provider\"].unique()))\nprint('ISUP Grade: {}'.format(len(train_data[\"isup_grade\"].unique())))\nprint('Gleason Score: {}'.format(len(train_data[\"gleason_score\"].unique())))\nprint('image ids : {}' .format(len(train_data[\"image_id\"].unique())))","metadata":{"execution":{"iopub.status.busy":"2023-03-20T17:34:12.097847Z","iopub.execute_input":"2023-03-20T17:34:12.098589Z","iopub.status.idle":"2023-03-20T17:34:12.114052Z","shell.execute_reply.started":"2023-03-20T17:34:12.098531Z","shell.execute_reply":"2023-03-20T17:34:12.112536Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_data.head()# view the first 5 rows of data","metadata":{"execution":{"iopub.status.busy":"2023-03-20T17:34:18.100926Z","iopub.execute_input":"2023-03-20T17:34:18.1013Z","iopub.status.idle":"2023-03-20T17:34:18.112777Z","shell.execute_reply.started":"2023-03-20T17:34:18.101265Z","shell.execute_reply":"2023-03-20T17:34:18.111652Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#number of rows and columns in the train data \nprint('Number of rows and columns:{}' .format(test_data.shape))\nprint('Data providers: {}'.format(test_data[\"data_provider\"].unique()))\nprint('image ids : {}' .format(len(test_data[\"image_id\"].unique())))","metadata":{"execution":{"iopub.status.busy":"2023-03-20T17:34:26.895826Z","iopub.execute_input":"2023-03-20T17:34:26.896183Z","iopub.status.idle":"2023-03-20T17:34:26.906186Z","shell.execute_reply.started":"2023-03-20T17:34:26.896151Z","shell.execute_reply":"2023-03-20T17:34:26.905042Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# get the data types, null values, and memory usage of each column\ntrain_data.info()","metadata":{"execution":{"iopub.status.busy":"2023-03-20T17:34:31.810105Z","iopub.execute_input":"2023-03-20T17:34:31.810693Z","iopub.status.idle":"2023-03-20T17:34:31.824762Z","shell.execute_reply.started":"2023-03-20T17:34:31.810655Z","shell.execute_reply":"2023-03-20T17:34:31.823591Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"***Handle Missing Data***","metadata":{}},{"cell_type":"code","source":"train_data.isna().sum()   # count the number of null values in each column","metadata":{"execution":{"iopub.status.busy":"2023-03-20T17:34:36.438195Z","iopub.execute_input":"2023-03-20T17:34:36.438762Z","iopub.status.idle":"2023-03-20T17:34:36.450763Z","shell.execute_reply.started":"2023-03-20T17:34:36.43871Z","shell.execute_reply":"2023-03-20T17:34:36.449596Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Describe Data","metadata":{}},{"cell_type":"code","source":"\"\"\"get the count, mean, standard deviation, minimum, \nand maximum values of each numerical column\"\"\"\ntrain_data.describe()   ","metadata":{"execution":{"iopub.status.busy":"2023-03-20T17:34:40.809881Z","iopub.execute_input":"2023-03-20T17:34:40.810215Z","iopub.status.idle":"2023-03-20T17:34:40.829579Z","shell.execute_reply.started":"2023-03-20T17:34:40.810183Z","shell.execute_reply":"2023-03-20T17:34:40.828298Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### What role does the GLEASON score play in this? \nFinding out the grade (aggressiveness) of the cancer cells is the next step after a biopsy confirms the presence of cancer. Your cancer is examined in a lab by a pathologist to discover how much the cancer cells differ from the normal cells. A malignancy with a higher grade is more likely to be aggressive and spread quickly. \n\nA Gleason score is the most popular scale for determining the grade of prostate cancer cells. Although the lower end of the range isn't utilised as frequently, Gleason scoring incorporates two scores and can range from 2 (nonaggressive disease) to 10 (highly aggressive cancer).","metadata":{}},{"cell_type":"markdown","source":"# ***Train Dataset***","metadata":{}},{"cell_type":"markdown","source":"The ***unique function*** in pandas is used to find the unique values from a series. A series is a single column of a data frame. ","metadata":{}},{"cell_type":"code","source":"train_data['gleason_score'].unique()","metadata":{"execution":{"iopub.status.busy":"2023-03-20T17:34:50.393783Z","iopub.execute_input":"2023-03-20T17:34:50.394395Z","iopub.status.idle":"2023-03-20T17:34:50.401038Z","shell.execute_reply.started":"2023-03-20T17:34:50.394356Z","shell.execute_reply":"2023-03-20T17:34:50.400121Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**if the Gleason score is written as 3+4=7, it means most of the tumor is grade 3 and less is grade 4, and they are added for a Gleason score of 7. Other ways that this Gleason score may be listed in your report are Gleason 7/10, Gleason 7 (3+4), or combined Gleason grade of 7.**\n\n***If a tumor is all the same grade (for example, grade 3), then the Gleason score is reported as 3+3=6.***","metadata":{}},{"cell_type":"code","source":"print(len(train_data[train_data['gleason_score']=='0+0']['isup_grade']))\nprint(len(train_data[train_data['gleason_score']=='negative']['isup_grade']))","metadata":{"execution":{"iopub.status.busy":"2023-03-20T17:34:55.042672Z","iopub.execute_input":"2023-03-20T17:34:55.042995Z","iopub.status.idle":"2023-03-20T17:34:55.056483Z","shell.execute_reply.started":"2023-03-20T17:34:55.042961Z","shell.execute_reply":"2023-03-20T17:34:55.054749Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**'Negative'** means when we already have 0+0 for no cancer","metadata":{}},{"cell_type":"markdown","source":"# Exploratory Data Analysis (EDA) \nis used to analyze and investigate data sets and summarize their main characteristics, often employing data visualization methods.","metadata":{}},{"cell_type":"markdown","source":"The train_data DataFrame is grouped by the **'isup_grade'** column using the **groupby()** function.\n\nThe count of **'image_id'** for each group is calculated using the **count()** function and is extracted using the indexing operator [].\n\nThe resulting DataFrame is then sorted in **descending order** based on the count of **'image_id'** using the **sort_values()** function.\n\nThe resulting DataFrame is assigned to a new variable named data.\n\nThe last line of the code **data.style.background_gradient(cmap='viridis')** applies a gradient color to the DataFrame data to highlight the values. \n\nSpecifically, it applies a range of colors gradient to the DataFrame using the background_gradient() method of pandas Styler object. \n\nhttps://jmsallan.netlify.app/blog/the-viridis-palettes/","metadata":{}},{"cell_type":"code","source":"\ndata = train_data.groupby('isup_grade').count()['image_id'].reset_index().sort_values(by='image_id',ascending=False)\ndata.style.background_gradient(cmap= 'viridis')","metadata":{"execution":{"iopub.status.busy":"2023-03-20T17:35:10.25218Z","iopub.execute_input":"2023-03-20T17:35:10.252805Z","iopub.status.idle":"2023-03-20T17:35:10.282421Z","shell.execute_reply.started":"2023-03-20T17:35:10.252766Z","shell.execute_reply":"2023-03-20T17:35:10.281244Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"***The higher the count of 'image_id', is where we have the yellow and green color which helps in visually identifying the groups with higher counts of 'image_id'.***\n\n","metadata":{}},{"cell_type":"code","source":"data = train_data.groupby('gleason_score').count()['image_id'].reset_index().sort_values(by='image_id',ascending=False)\ndata.style.background_gradient(cmap='Greens')","metadata":{"execution":{"iopub.status.busy":"2023-03-20T17:35:16.009688Z","iopub.execute_input":"2023-03-20T17:35:16.01001Z","iopub.status.idle":"2023-03-20T17:35:16.042373Z","shell.execute_reply.started":"2023-03-20T17:35:16.009975Z","shell.execute_reply":"2023-03-20T17:35:16.041711Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Gleason ratings of **5 or less** are generally not used. A diagnosis with a Gleason score of **6** is considered to be of low grade. Medium-grade cancer has a Gleason score of **7**, while high-grade cancer has a score of **8, 9, or 10. *A lower-grade cancer develops more gradually and has a reduced chance of cancerous growth than a high-grade cancer.***","metadata":{}},{"cell_type":"code","source":"import plotly.express as px\nimport plotly.graph_objects as go\n\n# Create the funnel chart\nfig = go.Figure(go.Funnel(\n    y = train_data['isup_grade'],\n    x = train_data['image_id'],\n    textposition = \"inside\",\n    marker = {\"color\": data['image_id'],\n              \"color\": ['Red','Green','Orange','cadetblue','Crimson', 'Purple']},\n    textfont = {\"color\": \"white\"}\n))\n\n# Set the layout\nfig.update_layout(\n    title = \"ISUP_grade Distribution Funnel Chart\",\n    yaxis_title = \"ISUP_grade\",\n    xaxis_title = \"Number of Images\",\n    font_family = \"Arial\",\n)\n\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-03-20T17:37:06.291296Z","iopub.execute_input":"2023-03-20T17:37:06.291699Z","iopub.status.idle":"2023-03-20T17:37:06.551987Z","shell.execute_reply.started":"2023-03-20T17:37:06.291659Z","shell.execute_reply":"2023-03-20T17:37:06.551076Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The plotly library is used to create the funnel chart. The y and x arguments are set to the ***'isup_grade'*** and ***'image_id'*** columns of the data, respectively. The marker argument is set to the 'image_id' column of the data, which sets the color of each segment of the funnel chart based on the number of images for each ISUP_grade. The textposition and textfont arguments are used to display the labels inside the funnel chart.\n\n","metadata":{}},{"cell_type":"code","source":"\nfig = go.Figure(data=[go.Bar(\n    x=train_data['isup_grade'], y=train_data['image_id'],\n    hoverinfo='y',\n    marker=dict(color=['Red','Green','Orange','Purple','Blue','Yellow'])\n)])\n\n# Set the layout\nfig.update_layout(\n    title='ISUP_grade and image_id Bar Chart',\n    xaxis_title='ISUP_grade',\n    yaxis_title='Number of Images'\n)\n\n# Show the chart\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-03-20T17:39:25.133961Z","iopub.execute_input":"2023-03-20T17:39:25.134376Z","iopub.status.idle":"2023-03-20T17:39:25.404214Z","shell.execute_reply.started":"2023-03-20T17:39:25.134334Z","shell.execute_reply":"2023-03-20T17:39:25.402702Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The Gleason score-based prior grading system was replaced by the more recent, more straightforward **ISUP grading system****. It uses a five-tiered scale to classify prostate cancer based on the patterns of cell development visible under a microscope, ***ranging from grade 1 (least aggressive) to grade 5 (most aggressive)***","metadata":{}},{"cell_type":"code","source":"train_data.head()","metadata":{"execution":{"iopub.status.busy":"2023-03-20T16:56:26.210059Z","iopub.execute_input":"2023-03-20T16:56:26.210607Z","iopub.status.idle":"2023-03-20T16:56:26.224255Z","shell.execute_reply.started":"2023-03-20T16:56:26.210555Z","shell.execute_reply":"2023-03-20T16:56:26.223219Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create a histogram plot of the Gleason score\nplt.hist(train_data['gleason_score'], bins=10, alpha=0.5, label='All data')\nplt.hist(train_data[train_data['data_provider'] == 'karolinska']['gleason_score'], bins=10, alpha=0.5, label='karolinska')\nplt.hist(train_data[train_data['data_provider'] == 'radboud']['gleason_score'], bins=10, alpha=0.5, label='radboud')\n# plt.hist(data[data['Data provider'] == 'Provider C']['Gleason score'], bins=10, alpha=0.5, label='Provider C')\n\n# Add a legend and title\nplt.legend(loc='upper right')\nplt.title('Distribution of Gleason score by Data Provider')\n\n# Show the plot\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-03-20T16:56:26.246135Z","iopub.execute_input":"2023-03-20T16:56:26.246469Z","iopub.status.idle":"2023-03-20T16:56:26.476877Z","shell.execute_reply.started":"2023-03-20T16:56:26.246438Z","shell.execute_reply":"2023-03-20T16:56:26.475781Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\"\"\"group the data by isup_grade and data_provider, and count the \nnumber of occurrences using the size() method then pivot \nthe data to create a matrix suitable for plotting, \nwith isup_grade as the index and data_provider as the columns. \nFinally, plot the data as a grouped bar chart using plot(kind='bar', \nstacked=False), and add a title and axis labels \nusing title(), xlabel(), and ylabel()\"\"\"\n\n# Group the data by isup_grade and data_provider\ngrouped_data = train_data.groupby(['isup_grade', 'data_provider']).size().reset_index(name='count')\n\n# Pivot the data to create a matrix for plotting\npivot_data = grouped_data.pivot(index='isup_grade', columns='data_provider', values='count')\n\n# Plot the data as a grouped bar chart\npivot_data.plot(kind='bar', stacked=False)\n\n# Add a title and axis labels\nplt.title('Relative Distribution of ISUP Grade by Data Provider')\nplt.xlabel('ISUP Grade')\nplt.ylabel('Count')\n\n# Show the plot\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-03-20T16:56:26.478845Z","iopub.execute_input":"2023-03-20T16:56:26.479145Z","iopub.status.idle":"2023-03-20T16:56:26.676838Z","shell.execute_reply.started":"2023-03-20T16:56:26.479113Z","shell.execute_reply":"2023-03-20T16:56:26.675874Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The majority of the data for isup grade categories **0 and 1** is provided by **Karolinska**. \nThe majority of the data for isup grade categories **3, 4, and 5** is provided by **radbound**.","metadata":{}},{"cell_type":"code","source":"# Group the data by gleason_score and data_provider and count the number of image_id\ngrouped_data = train_data.groupby(['gleason_score', 'data_provider'])['image_id'].count().reset_index()\n\n# Pivot the data to create a matrix of gleason_score vs data_provider counts\npivoted_data = grouped_data.pivot(index='gleason_score', columns='data_provider', values='image_id')\n\n# Create a stacked bar chart of the data\npivoted_data.plot(kind='bar', stacked=False)\n\n# Add axis labels and a title\nplt.xlabel('Gleason Score')\nplt.ylabel('Count')\nplt.title('Relative Distribution of Gleason Score and Data Provider')\n\n# Show the plot\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-03-20T16:56:26.678371Z","iopub.execute_input":"2023-03-20T16:56:26.67887Z","iopub.status.idle":"2023-03-20T16:56:26.917877Z","shell.execute_reply.started":"2023-03-20T16:56:26.678828Z","shell.execute_reply":"2023-03-20T16:56:26.916984Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Every piece of information for the gleason score **category (0+0)** comes from **Karolinska**. \n\nAll the information for the gleason score category **(negative)** is given by **radbound**. \n\n**Karolinska** is a significant data provider for the gleason score category **(3+3)**. \n\nHowever, **radbound** is a significant source of data for **(4+4), (4+3), (4+5), (5+4), (5+5), (5+3), and (3+5).**","metadata":{}},{"cell_type":"code","source":"\"\"\"the crosstab() function is used to create a cross-tabulation of \nisup_grade vs gleason_score, with the normalize parameter set to \n'index' to normalize the counts by row. The resulting cross-tabulation \nis then plotted using the heatmap() function from the seaborn library.\n\nThe resulting heatmap will show the relative distribution of \nisup_grade and gleason_score, with each cell representing the proportion \nof images in the corresponding isup_grade and gleason_score group. \nThe color of each cell indicates the proportion, with darker colors \nindicating higher proportions. The annotations in each cell show the \nproportions as percentages. This plot can be useful for identifying any \npatterns or trends in the distribution of the data.\"\"\"\n\n\n# Create a cross-tabulation of isup_grade vs gleason_score\ncross_tab = pd.crosstab(train_data['isup_grade'], train_data['gleason_score'], normalize='index')\n\n# Create a larger figure with a 10-inch width and 8-inch height\nplt.figure(figsize=(10, 8))\n\n\n# Create a heatmap of the cross-tabulation\nsns.heatmap(cross_tab, cmap='coolwarm', annot=True, fmt='.2f')\n\n\n# Add axis labels and a title\nplt.xlabel('Gleason Score')\nplt.ylabel('ISUP Grade')\nplt.title('Relative Distribution of ISUP Grade and Gleason Score')\n\n# Show the plot\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-03-20T16:56:26.919225Z","iopub.execute_input":"2023-03-20T16:56:26.919468Z","iopub.status.idle":"2023-03-20T16:56:27.400811Z","shell.execute_reply.started":"2023-03-20T16:56:26.919442Z","shell.execute_reply":"2023-03-20T16:56:27.399923Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Image Exploratory Data Analysis (EDA)\n\nI am still learning about image EDA.\n\n## What is .tff format and Why it is used?\n\nTagged Image File Format (TIFF) is a variable-resolution bitmapped image format developed by Aldus (now part of Adobe) in 1986. TIFF is very common for transporting color or gray-scale images into page layout applications, but is less suited to delivering web content.\n\nhttps://www.adobe.com/creativecloud/file-types/image/raster/tiff-file.html\n\n## What is Down-sampling and Up-sampling in Image processing?\nDigital image processing methods such as downsampling and upsampling are used to alter the resolution of an image. \n\nhttps://www.geeksforgeeks.org/spatial-resolution-down-sampling-and-up-sampling-in-image-processing/\n\n\n<br> openslide to display images:\nhttps://www.kaggle.com/wouterbulten/getting-started-with-the-panda-dataset\n<br>The benefit of OpenSlide is that we can load arbitrary regions of the slide, without loading the whole image in memory. \n\nOpenSlide python  documentation: https://openslide.org/api/python/","metadata":{}},{"cell_type":"code","source":"#locating the images \n\nfolder_path = '/kaggle/input/prostate-cancer-grade-assessment'\n\n# image and mask directories\n\ndata_dir = f'{folder_path}/train_images'\nmask_dir = f'{folder_path}/train_label_masks'","metadata":{"execution":{"iopub.status.busy":"2023-03-20T16:56:27.403393Z","iopub.execute_input":"2023-03-20T16:56:27.403656Z","iopub.status.idle":"2023-03-20T16:56:27.40844Z","shell.execute_reply.started":"2023-03-20T16:56:27.403627Z","shell.execute_reply":"2023-03-20T16:56:27.407245Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Understanding Masks\n\n### What are masks?\n\nApart from the slide-level label (present in the csv file), almost all slides in the training set have an associated mask with additional label information. These masks directly indicate which parts of the tissue are healthy and which are cancerous.hese masks are provided to assist with the development of strategies for selecting the most useful subsamples of the images. The mask values depend on the data provider:\n\n* Radboud: Prostate glands are individually labelled, Valid values are:\n           0: background (non tissue) or unknown\n           1: stroma (connective tissue, non-epithelium tissue)\n           2: healthy (benign) epithelium\n           3: cancerous epithelium (Gleason 3)\n           4: cancerous epithelium (Gleason 4)\n           5: cancerous epithelium (Gleason 5)\n\n* Karolinska: Regions are labelled, Valid values are:\n              1: background (non tissue) or unknown\n              2: benign tissue (stroma and epithelium combined)\n              3: cancerous tissue (stroma and epithelium combined)","metadata":{}}]}