{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Weighted Mean Baseline\n\n**Goal:** Create a simple baseline using a constant prediction for all test sample. This will come from the mean value of each target variable scaled by proposed sample weights: \n\n- 1 for all healthy labels.\n- 2 for low grade solid organ injuries (liver, spleen, kidney).\n- 4 for high grade solid organ injuries.\n- 2 for bowel injuries.\n- 6 for extravasation.\n- 6 for the auto-generated any_injury label.\n\nIn addition, we provide a method to evaluate the score for the training data and investigate alternative scale factors. This will be used to highlight the challenges of unbalanced data and weighted scoring metrics. ","metadata":{}},{"cell_type":"markdown","source":"# Imports","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport pandas.api.types\nimport sklearn.metrics","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-07-30T08:00:44.388668Z","iopub.execute_input":"2023-07-30T08:00:44.389403Z","iopub.status.idle":"2023-07-30T08:00:45.668106Z","shell.execute_reply.started":"2023-07-30T08:00:44.389358Z","shell.execute_reply":"2023-07-30T08:00:45.66688Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load Data","metadata":{}},{"cell_type":"code","source":"# Only requires the training target data. \ny_train = pd.read_csv('/kaggle/input/rsna-2023-abdominal-trauma-detection/train.csv')\n\ny_train.head()","metadata":{"execution":{"iopub.status.busy":"2023-07-30T08:00:45.670681Z","iopub.execute_input":"2023-07-30T08:00:45.671172Z","iopub.status.idle":"2023-07-30T08:00:45.721049Z","shell.execute_reply.started":"2023-07-30T08:00:45.67113Z","shell.execute_reply":"2023-07-30T08:00:45.719765Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# List of Targets\nInjuries = ['bowel_healthy', 'bowel_injury', \n            'extravasation_healthy', 'extravasation_injury', \n            'kidney_healthy', 'kidney_low', 'kidney_high', \n            'liver_healthy', 'liver_low', 'liver_high', \n            'spleen_healthy', 'spleen_low', 'spleen_high', \n            'any_injury']","metadata":{"execution":{"iopub.status.busy":"2023-07-30T08:00:45.722341Z","iopub.execute_input":"2023-07-30T08:00:45.722697Z","iopub.status.idle":"2023-07-30T08:00:45.728958Z","shell.execute_reply.started":"2023-07-30T08:00:45.722658Z","shell.execute_reply":"2023-07-30T08:00:45.727721Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Target EDA","metadata":{}},{"cell_type":"code","source":"y_train[Injuries].describe()","metadata":{"execution":{"iopub.status.busy":"2023-07-30T08:00:45.731871Z","iopub.execute_input":"2023-07-30T08:00:45.733008Z","iopub.status.idle":"2023-07-30T08:00:45.801407Z","shell.execute_reply.started":"2023-07-30T08:00:45.732975Z","shell.execute_reply":"2023-07-30T08:00:45.800285Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# [Score](https://www.kaggle.com/code/metric/rsna-trauma-metric/notebook)","metadata":{}},{"cell_type":"code","source":"# I'm not sure if this is needed!\n# class ParticipantVisibleError(Exception):\n#     pass\n\ndef normalize_probabilities_to_one(df: pd.DataFrame, group_columns: list) -> pd.DataFrame:\n    # Normalize the sum of each row's probabilities to 100%.\n    # 0.75, 0.75 => 0.5, 0.5\n    # 0.1, 0.1 => 0.5, 0.5\n    row_totals = df[group_columns].sum(axis=1)\n    if row_totals.min() == 0:\n        raise ParticipantVisibleError('All rows must contain at least one non-zero prediction')\n    for col in group_columns:\n        df[col] /= row_totals\n    return df\n\n\ndef score(solution: pd.DataFrame, submission: pd.DataFrame, row_id_column_name: str) -> float:\n    '''\n    Pseudocode:\n    1. For every label group (liver, bowel, etc):\n        - Normalize the sum of each row's probabilities to 100%.\n        - Calculate the sample weighted log loss.\n    2. Derive a new any_injury label by taking the max of 1 - p(healthy) for each label group\n    3. Calculate the sample weighted log loss for the new label group\n    4. Return the average of all of the label group log losses as the final score.\n    '''\n    del solution[row_id_column_name]\n    del submission[row_id_column_name]\n\n    # Run basic QC checks on the inputs\n    if not pandas.api.types.is_numeric_dtype(submission.values):\n        raise ParticipantVisibleError('All submission values must be numeric')\n\n    if not np.isfinite(submission.values).all():\n        raise ParticipantVisibleError('All submission values must be finite')\n\n    if solution.min().min() < 0:\n        raise ParticipantVisibleError('All labels must be at least zero')\n    if submission.min().min() < 0:\n        raise ParticipantVisibleError('All predictions must be at least zero')\n\n    # Calculate the label group log losses\n    binary_targets = ['bowel', 'extravasation']\n    triple_level_targets = ['kidney', 'liver', 'spleen']\n    all_target_categories = binary_targets + triple_level_targets\n\n    label_group_losses = []\n    for category in all_target_categories:\n        if category in binary_targets:\n            col_group = [f'{category}_healthy', f'{category}_injury']\n        else:\n            col_group = [f'{category}_healthy', f'{category}_low', f'{category}_high']\n\n        solution = normalize_probabilities_to_one(solution, col_group)\n\n        for col in col_group:\n            if col not in submission.columns:\n                raise ParticipantVisibleError(f'Missing submission column {col}')\n        submission = normalize_probabilities_to_one(submission, col_group)\n        label_group_losses.append(\n            sklearn.metrics.log_loss(\n                y_true=solution[col_group].values,\n                y_pred=submission[col_group].values,\n                sample_weight=solution[f'{category}_weight'].values\n            )\n        )\n\n    # Derive a new any_injury label by taking the max of 1 - p(healthy) for each label group\n    healthy_cols = [x + '_healthy' for x in all_target_categories]\n    any_injury_labels = (1 - solution[healthy_cols]).max(axis=1)\n    any_injury_predictions = (1 - submission[healthy_cols]).max(axis=1)\n    any_injury_loss = sklearn.metrics.log_loss(\n        y_true=any_injury_labels.values,\n        y_pred=any_injury_predictions.values,\n        sample_weight=solution['any_injury_weight'].values\n    )\n\n    label_group_losses.append(any_injury_loss)\n    return np.mean(label_group_losses)","metadata":{"execution":{"iopub.status.busy":"2023-07-30T08:00:45.802962Z","iopub.execute_input":"2023-07-30T08:00:45.803314Z","iopub.status.idle":"2023-07-30T08:00:45.822408Z","shell.execute_reply.started":"2023-07-30T08:00:45.803284Z","shell.execute_reply":"2023-07-30T08:00:45.821182Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In order to evaluate the score, the appropriate weights need to be assigned. The score function defined above expects sample weights for each category of injury (bowel, extravasation, kidney, liver, spleen, any). The sample weight is assigned based on the true target for a given category. For example, if a sample has a low grade kidney injury, then \n\n    [kidney_healthy, kidney_low, kidney_high] = [0,1,0]\n\nand we would set kidney_weight = 2 for for that sample.","metadata":{}},{"cell_type":"code","source":"# Assign the appropriate weights to each category\ndef create_training_solution(y_train):\n    sol_train = y_train.copy()\n    \n    # bowel healthy|injury sample weight = 1|2\n    sol_train['bowel_weight'] = np.where(sol_train['bowel_injury'] == 1, 2, 1)\n    \n    # extravasation healthy/injury sample weight = 1|6\n    sol_train['extravasation_weight'] = np.where(sol_train['extravasation_injury'] == 1, 6, 1)\n    \n    # kidney healthy|low|high sample weight = 1|2|4\n    sol_train['kidney_weight'] = np.where(sol_train['kidney_low'] == 1, 2, np.where(sol_train['kidney_high'] == 1, 4, 1))\n    \n    # liver healthy|low|high sample weight = 1|2|4\n    sol_train['liver_weight'] = np.where(sol_train['liver_low'] == 1, 2, np.where(sol_train['liver_high'] == 1, 4, 1))\n    \n    # spleen healthy|low|high sample weight = 1|2|4\n    sol_train['spleen_weight'] = np.where(sol_train['spleen_low'] == 1, 2, np.where(sol_train['spleen_high'] == 1, 4, 1))\n    \n    # any healthy|injury sample weight = 1|6\n    sol_train['any_injury_weight'] = np.where(sol_train['any_injury'] == 1, 6, 1)\n    return sol_train","metadata":{"execution":{"iopub.status.busy":"2023-07-30T08:00:45.823884Z","iopub.execute_input":"2023-07-30T08:00:45.824228Z","iopub.status.idle":"2023-07-30T08:00:45.838328Z","shell.execute_reply.started":"2023-07-30T08:00:45.824198Z","shell.execute_reply":"2023-07-30T08:00:45.83704Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"solution_train = create_training_solution(y_train)\n\n# predict a constant using the mean of the training data\ny_pred = y_train.copy()\ny_pred[Injuries] = y_train[Injuries].mean().tolist()\n\nno_scale_score = score(solution_train,y_pred,'patient_id')\nprint(f'Training score without scaling: {no_scale_score}')","metadata":{"execution":{"iopub.status.busy":"2023-07-30T08:05:58.102976Z","iopub.execute_input":"2023-07-30T08:05:58.103585Z","iopub.status.idle":"2023-07-30T08:05:58.17581Z","shell.execute_reply.started":"2023-07-30T08:05:58.10355Z","shell.execute_reply":"2023-07-30T08:05:58.174656Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Because the score weighs each target differently, we expect that the score can be improved by scaling by each target by their respective weight. ","metadata":{}},{"cell_type":"code","source":"# Group by different sample weights\nscale_by_2 = ['bowel_injury','kidney_low','liver_low','spleen_low']\nscale_by_4 = ['kidney_high','liver_high','spleen_high']\nscale_by_6 = ['extravasation_injury','any_injury']\n\n# Scale factors based on described metric \nsf_2 = 3\nsf_4 = 7\nsf_6 = 17\n\n# The score function deletes the ID column so we remake it\nsolution_train = create_training_solution(y_train)\n\n# Reset the prediction\ny_pred = y_train.copy()\ny_pred[Injuries] = y_train[Injuries].mean().tolist()\n\n# Scale each target \ny_pred[scale_by_2] *=sf_2\ny_pred[scale_by_4] *=sf_4\ny_pred[scale_by_6] *=sf_6\n\nweight_scale_score = score(solution_train,y_pred,'patient_id')\nprint(f'Training score with weight scaling: {weight_scale_score}')","metadata":{"execution":{"iopub.status.busy":"2023-07-30T08:07:05.286669Z","iopub.execute_input":"2023-07-30T08:07:05.287917Z","iopub.status.idle":"2023-07-30T08:07:05.369863Z","shell.execute_reply.started":"2023-07-30T08:07:05.28788Z","shell.execute_reply":"2023-07-30T08:07:05.368668Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can do even better by increasing the scale factors. We will investigate one alternative option. The goal is to highlight that the most accurate prediction doesn't necessarily give the highest score!","metadata":{}},{"cell_type":"code","source":"# Update scale factors to improve score\n# sf_2 = 2\n# sf_4 = 4\n# sf_6 = 14\n\n# The score function deletes the ID column so we remake it\nsolution_train = create_training_solution(y_train)\n\n# Reset the prediction, again\ny_pred = y_train.copy()\ny_pred[Injuries] = y_train[Injuries].mean().tolist()\n\n# Scale each target \ny_pred[scale_by_2] *=sf_2\ny_pred[scale_by_4] *=sf_4\ny_pred[scale_by_6] *=sf_6\n\nimproved_scale_score = score(solution_train,y_pred,'patient_id')\nprint(f'Training score with better scaling: {improved_scale_score}')","metadata":{"execution":{"iopub.status.busy":"2023-07-30T08:07:12.472637Z","iopub.execute_input":"2023-07-30T08:07:12.473126Z","iopub.status.idle":"2023-07-30T08:07:12.561491Z","shell.execute_reply.started":"2023-07-30T08:07:12.473088Z","shell.execute_reply":"2023-07-30T08:07:12.560404Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We will stick with this choice of scale factors as it appears to be better. Note, this appears to be near the optimal choice of scale factors for the mean, it is not necessarily the optimal choice of constant solution even for the training data. The optimal choice may given equal scaling to each injury which should have the same weight!\n\nWhen producing proper predictions, a constant scale factor may not be the optimal choice!","metadata":{}},{"cell_type":"markdown","source":"# Submission","metadata":{}},{"cell_type":"code","source":"# Load submission template \nsubmission = pd.read_csv('/kaggle/input/rsna-2023-abdominal-trauma-detection/sample_submission.csv')\n\n# Set output to mean of training data\nsubmission[Injuries] = y_train[Injuries].mean().tolist()\n\n# Scale each category by desired scale factor\nsubmission[scale_by_2] *=sf_2\nsubmission[scale_by_4] *=sf_4\nsubmission[scale_by_6] *=sf_6\n\n# Save Submission!\nsubmission.to_csv('submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2023-07-30T08:07:15.770685Z","iopub.execute_input":"2023-07-30T08:07:15.771096Z","iopub.status.idle":"2023-07-30T08:07:15.793358Z","shell.execute_reply.started":"2023-07-30T08:07:15.771063Z","shell.execute_reply":"2023-07-30T08:07:15.792491Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}