{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"![](https://external-content.duckduckgo.com/iu/?u=https%3A%2F%2Fd1nakyqvxb9v71.cloudfront.net%2Fwp-content%2Fuploads%2F2021%2F05%2Fcovid-19-blood-clots-hero.jpg&f=1&nofb=1)","metadata":{"id":"MyVaUQz0su3Z"}},{"cell_type":"markdown","source":"# Mayo Clinic - STRIP AI - EDA🔎\n\n## Image Classification of Stroke Blood Clot Origin\n<hr>","metadata":{"id":"iFegpvTT5cN-"}},{"cell_type":"markdown","source":"## 1. Introduction\n### 1.1. What is it?\n\n### Actuality of the problem / Background\nYour work will enable healthcare providers to better identify the origins of blood clots in deadly strokes, making it easier for physicians to prescribe the best post-stroke therapeutic management and reducing the likelihood of a second stroke.\n\n### Technical formulation of the problem\nGiven image find probability of it being cardiac(CE) or  large artery atherosclerosis(LAA).\n\n### Evaluation\nSubmissions are evaluated using a weighted multi-class logarithmic loss. The overall effect is such that each class is roughly equally important for the final score.\n\nEach image has been labeled with an etiology class, either CE or LAA. For each image, you must submit a probability for each class. The formula is then:\n$ logloss = - \\frac{\\sum_{i=1}^M w_i \\sum_{j=1}^{N_i} y_{ij}/N_i * ln p_{ij} }{\\sum_{i=1}^M w_i }$,\nwhere  \n\n\n\n*   $N$ is the number of images in the class set,\n*   $M$ is the number of classes,\n*   $y_{j}$ is 1 if observation $i$ belongs to the class $j$ and 0 otherwise,  \n*   $p_{ij}$ is the predicted probability that image $i$ belongs to class $j$\n\n\n\n\nMore about why this type of loss is used in classification in [the medium article](https://medium.com/gumgum-tech/handling-class-imbalance-by-introducing-sample-weighting-in-the-loss-function-3bdebd8203b4).\n### The plan:\n* ✅ Load files\n* ✅ Study the features' datatypes\n* ✅ Study features distribution\n* ✅ Find distribution patterns in data\n* ✅ Formulate results \n* 🛠️ Any additional ideas are welcome","metadata":{"id":"3xPkoU616FT5"}},{"cell_type":"markdown","source":"## 2. Libraries and modules 📗\n\nYou can find more about listing lybraries you use in a more organized way in [this article](https://towardsdev.com/from-kagglers-best-project-setup-for-ds-and-ml-ffb253485f98). The arthor is a Kaggle Master.","metadata":{"id":"r59_Ni1bnMQK"}},{"cell_type":"code","source":"# blocks output in Colab\n# %%capture\n\nimport os\nimport pandas as pd\n\n# plots\nimport plotly.express as px\nimport plotly.io as pio\nimport plotly.graph_objects as go\nimport seaborn as sns\n\n# dataset\nimport torch\nfrom torch.utils.data import Dataset, DataLoader\nfrom torchvision import transforms, utils\nfrom skimage import io, transform\nimport numpy as  np","metadata":{"id":"f8DfydSqoADC","execution":{"iopub.status.busy":"2022-10-24T17:24:23.932652Z","iopub.execute_input":"2022-10-24T17:24:23.933042Z","iopub.status.idle":"2022-10-24T17:24:28.338008Z","shell.execute_reply.started":"2022-10-24T17:24:23.933012Z","shell.execute_reply":"2022-10-24T17:24:28.3369Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# setting plotly reused constants:\n# colorshemes\ncontinuous_color_scheme = px.colors.sequential.Sunset\ndiscrete_color_scheme = px.colors.qualitative.Pastel1\n# opacity\nplots_opacity = 0.8","metadata":{"id":"RYz6OHp3jwyb","execution":{"iopub.status.busy":"2022-10-24T17:24:31.562509Z","iopub.execute_input":"2022-10-24T17:24:31.563667Z","iopub.status.idle":"2022-10-24T17:24:31.568833Z","shell.execute_reply.started":"2022-10-24T17:24:31.563623Z","shell.execute_reply":"2022-10-24T17:24:31.567547Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 3. Loading, Exploration And Preprocessing","metadata":{"id":"ArD3kXqinvt6"}},{"cell_type":"code","source":"# List files, they were taken from kaggle and contain dataset description\nALL_DATA_LOCATION = '../input/mayo-clinic-strip-ai'\nprint('Data to be studied:')\nprint(\", \".join(os.listdir(ALL_DATA_LOCATION)))","metadata":{"id":"JBKxSSYTnxic","outputId":"d14e5b23-54db-4dc6-acc9-59bb1d1aa350","execution":{"iopub.status.busy":"2022-10-24T17:26:40.0147Z","iopub.execute_input":"2022-10-24T17:26:40.015155Z","iopub.status.idle":"2022-10-24T17:26:40.023601Z","shell.execute_reply.started":"2022-10-24T17:26:40.015119Z","shell.execute_reply":"2022-10-24T17:26:40.022174Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# reading from csv to dataframe\n# these files contain pathes and information about train data\n# most of the data is in train and other\n# on them the training will be based\n# a question appears: it there any substantial difference between data in other and train?\ndf_train = pd.read_csv(ALL_DATA_LOCATION + '/train.csv')\ndf_test = pd.read_csv(ALL_DATA_LOCATION + '/test.csv')\ndf_other = pd.read_csv(ALL_DATA_LOCATION + '/other.csv')\ndf_sample_submission = pd.read_csv(ALL_DATA_LOCATION + '/sample_submission.csv')","metadata":{"id":"uROy67lmoer5","execution":{"iopub.status.busy":"2022-10-24T17:26:27.011389Z","iopub.execute_input":"2022-10-24T17:26:27.011817Z","iopub.status.idle":"2022-10-24T17:26:27.031034Z","shell.execute_reply.started":"2022-10-24T17:26:27.01178Z","shell.execute_reply":"2022-10-24T17:26:27.029886Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"dataframe's shape: \" + str(df_train.shape))\ndf_train.sample(5)","metadata":{"id":"qWxit0pNrkRC","outputId":"083c1fb4-6b0d-482a-fd89-b09aef67da03","execution":{"iopub.status.busy":"2022-10-24T17:26:44.122124Z","iopub.execute_input":"2022-10-24T17:26:44.122656Z","iopub.status.idle":"2022-10-24T17:26:44.150845Z","shell.execute_reply.started":"2022-10-24T17:26:44.122614Z","shell.execute_reply":"2022-10-24T17:26:44.149602Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"📌**Note:**\nExplanation of the columns\n* image_id - A unique identifier for this instance having the form \n{patient_id}_{image_num}. Corresponds to the image {image_id}.tif.\n\n* center_id - Identifies the medical center where the slide was obtained.\n\n* patient_id - Identifies the patient from whom the slide was obtained.\n\n* image_num - Enumerates images of clots obtained from the same patient.\n\n* label - The etiology of the clot, either CE or LAA. This field is the classification target. \n\n📌**Note:**\nIf you use colab the dataframe can be converted into an interactive table.\nIt is worth exploring the data!","metadata":{"id":"YvhA352VnJZz"}},{"cell_type":"markdown","source":"Now its time to find columns datatypes, possible values and ranges.","metadata":{"id":"N_teJksQvuLj"}},{"cell_type":"code","source":"df_train.info()","metadata":{"id":"tCJWLFVlvZBR","outputId":"ff748fbf-b328-4619-97de-85c4389f5039","execution":{"iopub.status.busy":"2022-10-24T17:26:51.030116Z","iopub.execute_input":"2022-10-24T17:26:51.031493Z","iopub.status.idle":"2022-10-24T17:26:51.060148Z","shell.execute_reply.started":"2022-10-24T17:26:51.031448Z","shell.execute_reply":"2022-10-24T17:26:51.059319Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# what are the possible values for the cols?\nfor col_name in ['center_id' , 'image_num', 'label']:\n  print('{} types:\\n'.format(col_name))\n  print(df_train[[col_name, 'image_id']].groupby(by=col_name, as_index=False).count().rename(columns={'image_id' : 'Count'}).sort_values(by='Count', ascending=False).to_string(index=False))\n  print('\\n')","metadata":{"id":"aFyxfGnIaRRs","outputId":"c5cfa2a1-fa55-4045-be63-e9150d7588da","execution":{"iopub.status.busy":"2022-10-24T17:26:55.178141Z","iopub.execute_input":"2022-10-24T17:26:55.179404Z","iopub.status.idle":"2022-10-24T17:26:55.206592Z","shell.execute_reply.started":"2022-10-24T17:26:55.179349Z","shell.execute_reply":"2022-10-24T17:26:55.205357Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":" **📌Note:** *Although Center_id has type int64 it is a categorical variable and not numerical, because it the center where the image was made.* ","metadata":{"id":"_EhfaCfR52lz"}},{"cell_type":"code","source":"# investigating missing values\ndf_train.isnull().sum()","metadata":{"id":"HOl5_nwKxl9x","outputId":"79ed4d32-f739-4a9d-d1fc-e80ee41148d0","execution":{"iopub.status.busy":"2022-10-24T17:27:02.14909Z","iopub.execute_input":"2022-10-24T17:27:02.149508Z","iopub.status.idle":"2022-10-24T17:27:02.158699Z","shell.execute_reply.started":"2022-10-24T17:27:02.149476Z","shell.execute_reply":"2022-10-24T17:27:02.15771Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# add new col label_id. if label == CE than it is 1, if LLA than 0\ndf_train['label_id'] = (df_train['label'] == 'CE').astype('int64', copy=False)\ndf_train[['label_id','label']].sample(5)","metadata":{"id":"oZAiaN4e-8sn","outputId":"89ba277a-54e6-407b-d9bb-b41ff0406723","execution":{"iopub.status.busy":"2022-10-24T17:27:11.49175Z","iopub.execute_input":"2022-10-24T17:27:11.49213Z","iopub.status.idle":"2022-10-24T17:27:11.507666Z","shell.execute_reply.started":"2022-10-24T17:27:11.492099Z","shell.execute_reply":"2022-10-24T17:27:11.506344Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"All the cols are not nulls in train.csv!🎆🎇\n\nNow lets explore other.csv","metadata":{"id":"ikgScDZUx0lA"}},{"cell_type":"code","source":"# the second dataframe containing data worth exploring\nprint(\"dataframe's shape: \" + str(df_other.shape))\ndf_other.head(4)","metadata":{"id":"sLqOwB2xoeYj","outputId":"430dc963-8af3-4a56-d72a-0aa7d603e7fa","execution":{"iopub.status.busy":"2022-10-24T17:27:17.662503Z","iopub.execute_input":"2022-10-24T17:27:17.663465Z","iopub.status.idle":"2022-10-24T17:27:17.67755Z","shell.execute_reply.started":"2022-10-24T17:27:17.66342Z","shell.execute_reply":"2022-10-24T17:27:17.675889Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"other.csv contains missing data, also many values are Nan or Unknown which do not contain logical information.","metadata":{"id":"L-tz67RKVDp3"}},{"cell_type":"code","source":"# col types and number of non-nulls\nprint(df_other.shape)\ndf_other.info()","metadata":{"id":"BUemQW1_QPFU","outputId":"8203e573-c440-4568-dcc8-6a2733e40b5b","execution":{"iopub.status.busy":"2022-10-24T17:27:27.056863Z","iopub.execute_input":"2022-10-24T17:27:27.057262Z","iopub.status.idle":"2022-10-24T17:27:27.071751Z","shell.execute_reply.started":"2022-10-24T17:27:27.057231Z","shell.execute_reply":"2022-10-24T17:27:27.070331Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# df_other has 334 nulls in col other_specified\ndf_other.isnull().sum()","metadata":{"id":"gdLMJmJkYkEz","outputId":"a39ecde5-6523-40aa-bc01-36c42a513295","execution":{"iopub.status.busy":"2022-10-24T17:27:33.705629Z","iopub.execute_input":"2022-10-24T17:27:33.706029Z","iopub.status.idle":"2022-10-24T17:27:33.715411Z","shell.execute_reply.started":"2022-10-24T17:27:33.705997Z","shell.execute_reply":"2022-10-24T17:27:33.714331Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# what are the possible values for the cols?\nfor col_name in ['image_num' , 'label', 'other_specified']:\n    print('{} types:\\n'.format(col_name))\n    print(df_other[[col_name, 'image_id']].groupby(by=col_name, as_index=False).count().rename(columns={'image_id' : 'Count'}).sort_values(by='Count', ascending=False).to_string(index=False))\n    print('\\n')","metadata":{"id":"vkuP0s7eVj9z","outputId":"57c6890b-7bcd-483f-fa45-d0a19cf9e67d","execution":{"iopub.status.busy":"2022-10-24T17:27:54.159035Z","iopub.execute_input":"2022-10-24T17:27:54.159466Z","iopub.status.idle":"2022-10-24T17:27:54.183611Z","shell.execute_reply.started":"2022-10-24T17:27:54.159431Z","shell.execute_reply":"2022-10-24T17:27:54.182085Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"other_specified col has 334 missing values, Label can be Unknown or Other, other_specified has 9 types. If more data is gathered further analisys can be done.","metadata":{"id":"ibvwHpLhZdEL"}},{"cell_type":"markdown","source":"## 4. Exploration\n * What are the diiferences between other and test?\n\n Train is labels data, other is not. Train contains additional cols about center where image was taken, this can have impact on the images taken.\n\n\n * What is the distribution among target classes?\n * How is the distribution of images made across centers?\n * How many imagescans the same pacient has?\n * What class has consequative retries from the same patient? Is the result always the same class? Are they taken in the same center?\n\n","metadata":{"id":"tBsdwnDB422W"}},{"cell_type":"code","source":"# the classes are inbalanced\n\n# bar\ndistribution_over_classes = df_train['label'].value_counts().reset_index().rename(columns={'index':'Label', 'label':'Count'})\nfig = px.bar(distribution_over_classes, x=\"Label\", y=\"Count\", title='Number of examples in test dataset by label type'\n             , text='Count', color = 'Label', color_discrete_sequence=discrete_color_scheme, opacity=plots_opacity)  \n\nfig.update_traces(texttemplate='%{text:.0s}', textposition='outside') # prnts values above bars\nfig.show()\n\n# pie chart\nfig = px.pie(distribution_over_classes, values='Count', names='Label', title='Number examples in test dataset by label type', color_discrete_sequence=discrete_color_scheme, opacity=plots_opacity)\nfig.show()","metadata":{"id":"XXH6Uyn6Onwo","outputId":"8f3e7766-4c77-416b-ca8c-68f6c474845a","execution":{"iopub.status.busy":"2022-10-24T17:28:02.421779Z","iopub.execute_input":"2022-10-24T17:28:02.422203Z","iopub.status.idle":"2022-10-24T17:28:03.777143Z","shell.execute_reply.started":"2022-10-24T17:28:02.422167Z","shell.execute_reply":"2022-10-24T17:28:03.775956Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Q: Which center is the most polular one?\n# A: 11\n# some centers have only 20-30 images taken\n# all centers have data of both classes\n\n# distribution over centers\ndistribution_over_classes = df_train['center_id'].value_counts().reset_index().rename(columns={'index':'Center', 'center_id':'Count'})\nfig = px.bar(distribution_over_classes, x=\"Center\", y=\"Count\", title='Comparison of number of examples in test dataset by Center'\n             , text='Count', color = 'Center', \n             color_continuous_scale= continuous_color_scheme\n             )  \n             #color_discrete_sequence = generate_color_sequence(2) )#cmocean.cm.thermal)\nfig.update_traces(texttemplate='%{text:.0s}', textposition='outside') # prints values above bars\nfig.show()\n\n# distribution over centers and classes \nfig = px.bar(df_train.groupby(by=['center_id', 'label'], as_index=False).count().rename(columns={'image_id':'Count'}), x=\"center_id\", y=\"Count\", \n             color=\"label\", \n             title=\"Data distribution over centers and labels\",  \n             color_discrete_sequence=discrete_color_scheme,\n             opacity=plots_opacity\n             )\nfig.show()\n\n# the same data can be shown as sunburst plot\nfig = px.sunburst(df_train[['center_id', 'label', 'image_id']].groupby(by=['center_id', 'label'], as_index=False).count().rename(columns={'image_id':'Count'}),\n                  path=['center_id', 'label'], values='Count', color_discrete_sequence=discrete_color_scheme)\nfig.show()","metadata":{"id":"2MTuAg4Ofy_w","outputId":"19c54538-3cc0-422d-a6f6-8d4e3b4f8b01","execution":{"iopub.status.busy":"2022-10-24T17:28:21.03152Z","iopub.execute_input":"2022-10-24T17:28:21.031953Z","iopub.status.idle":"2022-10-24T17:28:21.298934Z","shell.execute_reply.started":"2022-10-24T17:28:21.031919Z","shell.execute_reply":"2022-10-24T17:28:21.297747Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The plot shows that most of images(34%) were made at center 11. Is there is differnce between data generated by different center this may lead to lower accuracy on data generated by center 8 and 9 that make only 4% of the dataset.\n\n","metadata":{"id":"nLuyxWXFj-OS"}},{"cell_type":"code","source":"# Q: how many images does a patient usually have?\n# A: 1\ndf_patient_to_num_of_images = df_train[['patient_id', 'image_num']].groupby(by=['patient_id'], as_index=False).count().rename(columns={'image_num' : 'Number of images taken for the same patient'})\nprint(df_patient_to_num_of_images.groupby(by='Number of images taken for the same patient', as_index=False).count().to_string(index=False))\nfig = px.scatter(df_patient_to_num_of_images, x=\"patient_id\", y=\"Number of images taken for the same patient\", \n                 marginal_y=\"violin\", # possible marginal distribution types: histogram violin box rug\n                 title='Number of images of the same patient',\n                 color_discrete_sequence=discrete_color_scheme,\n                 opacity=plots_opacity) \nfig.show()","metadata":{"id":"K0nGHQOCpigm","outputId":"64dab100-d12e-4b50-8919-fd813cf8b376","execution":{"iopub.status.busy":"2022-10-24T17:28:28.803468Z","iopub.execute_input":"2022-10-24T17:28:28.803901Z","iopub.status.idle":"2022-10-24T17:28:28.962387Z","shell.execute_reply.started":"2022-10-24T17:28:28.803864Z","shell.execute_reply":"2022-10-24T17:28:28.961165Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Q: do patients have the same label result after reexaminations?\n# A: Yes\n# patients with more that 1 image\npatients_with_several_images_grouped = df_train[['patient_id', 'image_num']].groupby(by=['patient_id'], as_index=False).count().query('image_num > 1')\n# patients with more that 1 image grouped by patient_id, label\npatients_with_several_images = df_train[df_train['patient_id'].isin(patients_with_several_images_grouped['patient_id'])][['patient_id', 'label', 'image_id']].groupby(by=['patient_id', 'label'], as_index=False).count()\n# sql-like join\npatients_with_several_images = pd.merge(left=patients_with_several_images, right=patients_with_several_images_grouped, how=\"inner\", on = 'patient_id')\n# if image_id <> image number that the same patient has images with different labels\npatients_with_several_images.query('image_id != image_num')# there are no such cases in dataset","metadata":{"id":"-WOo0dG5sO03","outputId":"64f34b41-5d15-4345-d627-2ab6c17d0d69","execution":{"iopub.status.busy":"2022-10-24T17:28:34.600539Z","iopub.execute_input":"2022-10-24T17:28:34.601012Z","iopub.status.idle":"2022-10-24T17:28:34.639082Z","shell.execute_reply.started":"2022-10-24T17:28:34.600973Z","shell.execute_reply":"2022-10-24T17:28:34.637828Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Q: Are the reexaminations taken in the same center?\n# A: Yes\n# patients with more that 1 image\npatients_with_several_images_grouped = df_train[['patient_id', 'image_num']].groupby(by=['patient_id'], as_index=False).count().query('image_num > 1') \n# patients with more that 1 image grouped by patient_id, center_id\npatients_with_several_images = df_train[df_train['patient_id'].isin(patients_with_several_images_grouped['patient_id'])][['patient_id', 'center_id', 'image_id']].groupby(by=['patient_id', 'center_id'], as_index=False).count()\n# sql-like join\npatients_with_several_images = pd.merge(left=patients_with_several_images, right=patients_with_several_images_grouped, how=\"inner\", on = 'patient_id')\n# if image_id <> image number that the same patient has images with different labels\npatients_with_several_images.query('image_id != image_num')# there are no such cases in dataset","metadata":{"id":"qj10pqYN-t_s","outputId":"a912ebb2-1421-4d82-bd02-33f72ff191f1","execution":{"iopub.status.busy":"2022-10-24T17:28:38.349506Z","iopub.execute_input":"2022-10-24T17:28:38.35051Z","iopub.status.idle":"2022-10-24T17:28:38.384952Z","shell.execute_reply.started":"2022-10-24T17:28:38.350461Z","shell.execute_reply":"2022-10-24T17:28:38.383603Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# looking at the test dataframe\nprint(\"dataframe's shape: \" + str(df_test.shape))\ndf_test.head(4)","metadata":{"id":"a9E7jdHVoekZ","outputId":"15cce47e-33b3-4406-a55b-fdfd73414fc9","execution":{"iopub.status.busy":"2022-10-24T17:28:41.225573Z","iopub.execute_input":"2022-10-24T17:28:41.225996Z","iopub.status.idle":"2022-10-24T17:28:41.241899Z","shell.execute_reply.started":"2022-10-24T17:28:41.22596Z","shell.execute_reply":"2022-10-24T17:28:41.240493Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"For the submission predicting the class of 4 images is needed. All these images were take in the most popular center #11 and it is the first image for these patients.","metadata":{"id":"qUt3MaUa95nE"}},{"cell_type":"markdown","source":"## 5. Results\n* There is a class disbalance in the dataset with CE class having more examples \n* Center 11 is the most popular. The model may overfit on data generated by it\n* Usually(543) patient go though only one check, however some(89) need aditional images taken in the same center. It may result in high simularity of some examples.\n\nThank you for reading)","metadata":{"id":"ZRBueTow_Tk4"}}]}