{"cells":[{"metadata":{},"cell_type":"markdown","source":"# Pandas data tricks and baseline\n\nSometimes it is nice to turn the data you recieve into formats that are more easily fed to your models of interest. Let's not worry about the actual imaging data and just make it easier to work with the labels.\n\nTL;DR: Just use these functions below do convert your sample_submission or train_csv dataframes:"},{"metadata":{"trusted":true},"cell_type":"code","source":"import pandas as pd","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def rsna_to_pivot(df, sub_type_name='HemType'):\n    \"\"\"Convert RSNA data frame to pivoted table with\n    each subtype as a binary encoded column.\"\"\"\n    df2 = df.copy()\n    ids, sub_types = zip(*df['ID'].str.rsplit('_', n=1).values)\n    df2.loc[:, 'ID'] = ids\n    df2.loc[:, sub_type_name] = sub_types\n    return df2.pivot(index='ID', columns=sub_type_name, values='Label')\n\ndef pivot_to_rsna(df, sub_type_name='HemType'):\n    \"\"\"Converted pivoted table back to RSNA spec for submission.\"\"\"\n    df2 = df.copy()\n    df2 = df2.reset_index()\n    unpivot_vars = df2.columns[1:]\n    df2 = pd.melt(df2, id_vars='ID', value_vars=unpivot_vars, var_name=sub_type_name, value_name='Label')\n    df2['ID'] = df2['ID'].str.cat(df2[sub_type_name], sep='_')\n    df2.drop(columns=sub_type_name, inplace=True)\n    return df2","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sample_sub = pd.read_csv('/kaggle/input/rsna-intracranial-hemorrhage-detection/stage_1_sample_submission.csv')\ntrain = pd.read_csv('/kaggle/input/rsna-intracranial-hemorrhage-detection/stage_1_train.csv')\n#First step is to get rid of duplicated entries. Luckily they are consistent.","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Let's check that that is true just to be sure.\n# If there were any groups that were not consistent,\n# the set of labels should be more than 0 and 1, therefore we are safe.\nset(train.groupby('ID').mean()['Label'].values)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Actually remove duplicates\ntrain = train.groupby('ID').first().reset_index()\ntrain.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"If we look at the data in the train CSV, we find that each image has an ID with a corresponding sub-type and a binary label. It isn't fun to have to work with classes buried in the ID. It would be much better to separate these out into lines where each line simply contained the ID and then having separate columns for each class contained at the end of ID, delineating whether or not that type of hemmorhage was present. Luckily the pandas package has some magical functionality to do this."},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"print(train.shape) # Notice all the rows\ntrain.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"First we will separate out the class name from the ID by using an rsplit:"},{"metadata":{"trusted":true},"cell_type":"code","source":"split_series = train['ID'].str.rsplit('_', n=1)\nsplit_series.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Now we have the split we need. We just now need to package that back into the original dataframe. This is easy, we just need to grab these and asign them as new columns:"},{"metadata":{"trusted":true},"cell_type":"code","source":"ids, sub_types = zip(*train['ID'].str.rsplit('_', n=1).values)\ntrain.loc[:, 'ID'] = ids\ntrain.loc[:, 'HemType'] = sub_types # We are using HemType as our column name for our sub_types\ntrain.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The next line is the real magic. We can use a pivot to package everything up as we described."},{"metadata":{"trusted":true},"cell_type":"code","source":"train = train.pivot(index='ID', columns='HemType', values='Label')\nprint(train.shape)\ntrain.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Yay! That's what I like to see. Let's grab some stats real quick.\n# We can save these for later.\ntrain.mean()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Of course though, if we can transform in this direction. We want to transform in the other. This backwards operation is called a melt, and we can do that just as easily. First we will clean up the index:"},{"metadata":{"trusted":true},"cell_type":"code","source":"train = train.reset_index()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Now we will do our melt, which will be roughly the inverse operation of what we just performed."},{"metadata":{"trusted":true},"cell_type":"code","source":"unpivot_vars = train.columns[1:] # Here we need the names of categories so we can push them back in the ID\ntrain = pd.melt(train, id_vars='ID', value_vars=unpivot_vars, var_name='HemType', value_name='Label')\ntrain.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Almoost there, now we just need to convert ID and HemType into one column:"},{"metadata":{"trusted":true},"cell_type":"code","source":"train['ID'] = train['ID'].str.cat(train['HemType'], sep='_')\ntrain.drop(columns='HemType', inplace=True)\ntrain.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Wow look at that! We went from one representation and back again really easily. If you want to learn more about these manipulations of data, I found Hadley Wickham's Tidy Data paper to be extremely helpful. (https://vita.had.co.nz/papers/tidy-data.pdf) \n\nWith all said, our code is really short and fits in these two functions:"},{"metadata":{"trusted":true},"cell_type":"code","source":"# The only additions I\"m adding is copying dataframes so we don't accidentally change data we want to keep.\n\ndef rsna_to_pivot(df, sub_type_name='HemType'):\n    \"\"\"Convert RSNA data frame to pivoted table with\n    each subtype as a binary encoded column.\"\"\"\n    df2 = df.copy()\n    ids, sub_types = zip(*df['ID'].str.rsplit('_', n=1).values)\n    df2.loc[:, 'ID'] = ids\n    df2.loc[:, sub_type_name] = sub_types\n    return df2.pivot(index='ID', columns=sub_type_name, values='Label')\n\ndef pivot_to_rsna(df, sub_type_name='HemType'):\n    \"\"\"Converted pivoted table back to RSNA spec for submission.\"\"\"\n    df2 = df.copy()\n    df2 = df2.reset_index()\n    unpivot_vars = df2.columns[1:]\n    df2 = pd.melt(df2, id_vars='ID', value_vars=unpivot_vars, var_name=sub_type_name, value_name='Label')\n    df2['ID'] = df2['ID'].str.cat(df2[sub_type_name], sep='_')\n    df2.drop(columns=sub_type_name, inplace=True)\n    return df2","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Averaged baseline\n\nNow let's use these functions to create a dead simple baseline that we can use without waiting for all the CT data to unzip.\n\nWe will look at the training data, and just assign the probability of a sub_type of hemmorhage to be the frequency of each type of hemmorhage in the dataset."},{"metadata":{"trusted":true},"cell_type":"code","source":"# Just prep work we did before\nsample_sub = pd.read_csv('/kaggle/input/rsna-intracranial-hemorrhage-detection/stage_1_sample_submission.csv')\ntrain = pd.read_csv('/kaggle/input/rsna-intracranial-hemorrhage-detection/stage_1_train.csv')\ntrain = train.groupby('ID').first().reset_index()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Easy. Abstraction makes life great.\ntrain_pivot = rsna_to_pivot(train)\nsample_sub_pivot = rsna_to_pivot(sample_sub)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_pivot.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sample_sub_pivot.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# We did this before.\naverages = train_pivot.mean()\naverages","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Go through the averages and deliver them to the columns of the submission.\nfor label, value in averages.items():\n    sample_sub_pivot.loc[:, label] = value","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sample_sub_pivot.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# We pivoted and now let's melt this back on in.\nsubmission = pivot_to_rsna(sample_sub_pivot)\nsubmission.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# What's easier than this?\nsubmission.to_csv('submission.csv', index=False)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"This is a really nice way to work with your data, not only for feeding data to a neural network, but also just to make it more expressive for your exploratory data analysis. Hope this helps!"}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":1}