{"metadata":{"kernelspec":{"name":"python3","display_name":"Python 3","language":"python"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":84896,"databundleVersionId":10305135,"sourceType":"competition"},{"sourceId":9178166,"sourceType":"datasetVersion","datasetId":5547076}],"dockerImageVersionId":30804,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Regression with an Insurance Dataset 👨‍👩‍👧‍👦\nWelcome! This notebook is going to be a detailed one, so I hope you enjoy it! 📚\n\n> **v7 Update:** The target variable is log transformed for the training process, and the predictions are transformed back by exponent method \n\n## Table of Contents \n\n1. [Importing Libraries and Data](#1)\n2. [Train & Original Similarity Report 📝](#2)\n    - 2.1 [Insights from Similarity Report 🧪](#3)\n    - 2.2 [Suspection 🕵️‍♂ ](#4)\n3. [What should we do next? 🤔](#5)\n    - 3.1 [Choosing the Right Grouping 🤔](#right-grouping)\n        - [My Criteria for Choosing the Grouping Column 📊](#my-criteria)\n        - [Mean and Standard Deviation of Target Column 🎯](#mean-std)\n        - [Plotting Value Counts 📊](#plotting)\n5. [My Suggestions 📝](#6)\n    - 4.1 [Feature Engineering 🛠](#7)\n    - 4.2 [EDA Conclusions](#8)\n6. [Grouping the Data 📊](#9)\n7. [Model Building 🏗](#10)\n    - 6.1 [Hyperparameter Tuning 🎯](#11)\n    - 6.2 [Making Predictions 🔮](#12)\n    - 6.3 [Consolidating Predictions 👩‍💻](#13)","metadata":{}},{"cell_type":"markdown","source":"## Importing Libraries and Data <a id=\"1\"></a>","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\n\nimport matplotlib.pyplot as plt\nimport seaborn as sns","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-08T12:37:32.919855Z","iopub.execute_input":"2024-12-08T12:37:32.920272Z","iopub.status.idle":"2024-12-08T12:37:32.925891Z","shell.execute_reply.started":"2024-12-08T12:37:32.920235Z","shell.execute_reply":"2024-12-08T12:37:32.924453Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train = pd.read_csv('/kaggle/input/playground-series-s4e12/train.csv')\ntest = pd.read_csv('/kaggle/input/playground-series-s4e12/test.csv')\n\norig = pd.read_csv('/kaggle/input/insurance-premium-prediction/Insurance Premium Prediction Dataset.csv')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-08T12:37:35.616209Z","iopub.execute_input":"2024-12-08T12:37:35.616659Z","iopub.status.idle":"2024-12-08T12:37:44.781452Z","shell.execute_reply.started":"2024-12-08T12:37:35.616623Z","shell.execute_reply":"2024-12-08T12:37:44.780398Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train.shape, test.shape","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-08T12:30:48.68415Z","iopub.execute_input":"2024-12-08T12:30:48.684559Z","iopub.status.idle":"2024-12-08T12:30:48.692458Z","shell.execute_reply.started":"2024-12-08T12:30:48.684524Z","shell.execute_reply":"2024-12-08T12:30:48.691186Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Train & Original Similarity Report 📝 <a id=\"2\"></a>\n\nMy custom report to compare the original and train datasets based on the columns and their values.\n\n- [\"A significant difference between train and original dataset\"](https://www.kaggle.com/competitions/playground-series-s4e12/discussion/549332) by [aldparis](https://www.kaggle.com/adaubas) states that the target in the original dataset and the target in the training dataset are having a different distribution\n\n\n![scores](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F275525%2F4f55f7d2e516d36fa94ea7a5bd4be562%2FScreenshot_20241201_220802.png?generation=1733087399896644&alt=media)\n\nAbove is the comparison demonstrated by [aldparis](https://www.kaggle.com/adaubas) in his discussion. `Premium Amount` in the original dataset is having a different distribution than the `Premium Amount` in the training dataset. Persumably this is the cause of high RMSLE scores if we mix train and original datasets.\n\n--- \n\nBut let's also remember the discussion from the last episode [\"Beware of some oblong values across text columns\"](https://www.kaggle.com/competitions/playground-series-s4e11/discussion/543722) by [Ravi Ramakrishnan](https://www.kaggle.com/https://www.kaggle.com/ravi20076) which brought light to the fact that there was a ton of noise in the mental health competition's dataset.\n\nHere, let's see if we can find any noise in the dataset 🧐","metadata":{}},{"cell_type":"code","source":"def similarity_report(df_a, df_b, names=['df_a', 'df_b']):\n\n    report = {\n        f'{names[0]}_dtype': [],\n        f'{names[1]}_dtype': [],\n\n        f'{names[0]}_unique': [],\n        f'{names[1]}_unique': [],\n\n        f'unique_diff': [],\n        f'mean_diff': [],\n\n        f'{names[0]}_min': [],\n        f'{names[1]}_min': [],\n\n        f'{names[0]}_max': [],\n        f'{names[1]}_max': [],\n    }\n\n    a_columns = set(df_a.columns)\n    b_columns = set(df_b.columns)\n\n    # Obtain matching and non-matching columns\n    column_match = list(a_columns.intersection(b_columns))\n    cols_excluded = list(a_columns.symmetric_difference(b_columns))\n\n    print(f'[!]: Columns excluded due to mismatch: {cols_excluded}')\n\n\n    for col in column_match:\n        a_unique = df_a[col].nunique()\n        b_unique = df_b[col].nunique()\n\n        report[f'{names[0]}_unique'].append(a_unique)\n        report[f'{names[1]}_unique'].append(b_unique)\n\n        a_unique_vals = df_a[col].dropna().unique()\n        b_unique_vals = df_b[col].dropna().unique()\n\n        # find any differences in unique values\n        unique_diff = set(a_unique_vals).symmetric_difference(set(b_unique_vals))\n        report['unique_diff'].append(len(unique_diff))\n\n        # Obtain data types of each df's columns\n        report[f'{names[0]}_dtype'].append(df_a[col].dtype)\n        report[f'{names[1]}_dtype'].append(df_b[col].dtype)\n        \n\n        if df_a.dtypes[col] != df_b.dtypes[col]:\n            print(f\"[?]: Column {col} ({df_a[col].dtype}/{df_b[col].dtype}) has different data types in the two dataframes\")\n\n            if ['object', 'category', 'date'] in [df_a[col].dtype, df_b[col].dtype]:\n                print(f\"[?]: Column {col} is of type object or category\")\n                continue\n\n            report[f'{names[0]}_min'].append(df_a[col].min())\n            report[f'{names[1]}_min'].append(df_b[col].min())\n\n            report[f'{names[0]}_max'].append(df_a[col].max())\n            report[f'{names[1]}_max'].append(df_b[col].max())\n\n            report['mean_diff'].append(df_a[col].mean() - df_b[col].mean())\n\n            continue\n\n        if df_a[col].dtype == 'object':\n            report[f'{names[0]}_min'].append('')\n            report[f'{names[1]}_min'].append('')\n            report[f'{names[0]}_max'].append('')\n            report[f'{names[1]}_max'].append('')\n            report['mean_diff'].append('')\n            continue\n\n        a_min = df_a[col].min()\n        b_min = df_b[col].min()\n\n        report[f'{names[0]}_min'].append(a_min)\n        report[f'{names[1]}_min'].append(b_min)\n\n        a_max = df_a[col].max()\n        b_max = df_b[col].max()\n\n        report[f'{names[0]}_max'].append(a_max)\n        report[f'{names[1]}_max'].append(b_max)\n\n        report['mean_diff'].append(df_a[col].mean() - df_b[col].mean())\n\n    report = pd.DataFrame(report, index=column_match)\n    return report\n\n\nreport = similarity_report(train, orig, names=['train', 'original'])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-08T08:50:12.840575Z","iopub.execute_input":"2024-12-08T08:50:12.840998Z","iopub.status.idle":"2024-12-08T08:50:17.704279Z","shell.execute_reply.started":"2024-12-08T08:50:12.840963Z","shell.execute_reply":"2024-12-08T08:50:17.703112Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"report.head(20)","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Log Transformation of Target Variable \n\nLet's consider the following discussion:\n\n- [\"How to make your life easier by transforming the target variable & other ideas\"](https://www.kaggle.com/competitions/playground-series-s4e12/discussion/549336) by [Björn](https://www.kaggle.com/bjoernholzhauer)\n\nThere are more related to this topic, but let's see this portion from the discussion for now:\n\n> If you just create a transformed variable train['y_trans'] = np.log1p(df['Premium Amount'].values), then linear regression or anything else with in-buildt (r)mse loss will directly optimize the competition metric. Sure, sklearn has root_mean_squared_log_error, but that won't work as a loss function for most models. Just remember to transform back before submitting (numpy even provides the function you need: np.expm1()).\n\nTaking this advice/suggestion into account, let's log transform the target variable. But we also need to transform the predictions back to normal, which is done in `tune_lgb` function when oof generation is set to `True`","metadata":{}},{"cell_type":"code","source":"train['Premium Amount'] = np.log1p(train['Premium Amount'])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-08T12:37:52.890618Z","iopub.execute_input":"2024-12-08T12:37:52.891013Z","iopub.status.idle":"2024-12-08T12:37:52.922408Z","shell.execute_reply.started":"2024-12-08T12:37:52.890979Z","shell.execute_reply":"2024-12-08T12:37:52.921176Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Insights from Similarity Report 🧪 <a id=\"3\"></a>\n\n#### 1. One column missing in original ❓\n\nThe column `id` is not present in the original dataset. This is a unique identifier for each row in the training dataset.\n\n\n#### 2. Different Distribution of Target Column 🎯\n\nThe target column `Premium` is having a different distribution in the original dataset compared to the training dataset. This could be a potential source of noise in the dataset.\n\n\n#### 3. Same Columns, Different Data Types 📊\n\n- `Insaurance Duration` carries `float64` data type in the train dataset, while it is `int64` in the original dataset.\n\n- `Vehicle Age` carries `float64` data type in the train dataset, while it is `int64` in the original dataset.\n\n\n#### 4. Noise in the Dataset 📢\n\nIt's not that there are \"oblong values\", but some columns have a lot of difference in the values. For example,\n\n1. For `Annual Income`, the train dataset contains `88593` unique values while original dataset contains `101693` unique values. \n    - As per intuition, the difference should be of `13100` BUT it is `13274` \n    \n    - It is quite opposite of what happened in the mental health competition where the original dataset had less unique values than the train dataset.\n\n\n2. `Health Score` has `532657` unique values in train dataset and `268263` unique values in original dataset. \n    - Again, counter intuitively the difference is `563966` rather than `264394`.\n    \n    - This is a significant difference and could be a potential source of noise in the dataset.\n\n\nThere is more to explore from the report. Gain more insights by checking out the notebook! 📚\n\n### 5. Suspection 🕵️‍♂ <a id=\"4\"></a>\n\nWondering why is the difference in the unique values of the columns so counter-intuitive that it is more than the actual difference in the unique values? See, let's consider a simple example:\n\n- `Policy Type` has 3 unique values both in train and original datasets. `['Premium', 'Comprehensive', 'Basic']` both in the datasets. That makes the difference in unique values `0`.\n\nBut, if **hypothetically** this was the case:\n\n- Unique values in train's col: `['Premium', 'Comprehensive', 'Basic']`\n- Unique values in original's col: `['Premium', 'Popular', 'Basic']`\n\nIn this case, the difference in unique values would be `1`, even though both the datasets have 3 unique values. This is a simple example to demonstrate the point through a hypothetical scenario.","metadata":{}},{"cell_type":"markdown","source":"## What should we do next? 🤔 <a id=\"5\"></a>\n\nSee, columns like `Health Score`, `Policy Start Date`, `Annual Income`, etc. are highly cardinal, maybe not for such a large dataset but still they have a lot of unique values.\n\nThankfully we have the `Policy Start Date` column which is a date column. We can extract features like `year`, `month`, `day`, `dayofweek`, etc. because each value looks like `YYYY-MM-DD HH:MM:SS.SSSSSS`. (woah, that's a lot of precision!)\n","metadata":{}},{"cell_type":"markdown","source":"## My Suggestions 📝 <a id=\"6\"></a>\n\nMaybe fitting a single model on the entire dataset or its folds is not a good idea.\n\n> 💡 Some people suggested an approach in the last episode about creating different models for students and working professionals. I compiled related articles to that [here](https://www.kaggle.com/competitions/playground-series-s4e11/discussion/546744), also mentioning the method I built which ranked in top 5% on the private LB. **Let's try a similar one here too!**\n\n- **Create different models for different groups of data:**. For example, we can first group data according to `Occupation`, or `Policy Type` whichever seems more promising. Then we can train a model on each group separately. This way, we can avoid the noise in the dataset and also make the model more robust.\n\n- **Feature Engineering:** As mentioned above, we can extract features from the `Policy Start Date` column. We can also try to create new features from the existing ones if its a good idea for synthetic data.\n\n- **Model Selection & Hyperparameter Tuning:** We can try different models and tune their hyperparameters to get the best results. We can also try stacking models to get better results.\n\n- **Cross-Validation:** We can use cross-validation to get a better estimate of the model's performance. We can also use different cross-validation strategies to get a better estimate of the model's performance.\n\n\n## Choosing the Right Grouping 🤔 <a id=\"right-grouping\"></a>\n\nThe main question is, which column should we use to group the data? Because as per me, we also need to watch out for columns having significant amount of null values.\n\n\n### My Criteria for Choosing the Grouping Column 📊 <a id=\"my-criteria\"></a>\n\n1. **The column that is not too imbalanced:** The unique values in the column should not be too imbalanced. For example, if a column has 90% of the values as `0` and 10% as `1`, then it is too imbalanced.\n\n2. **The column that is not too noisy:** The column should not have too much noise. For example, if a column has a lot of unique values, then it is too noisy.\n\n3. **Unique Values:** The column should have a good number of unique values. If a column has only 2 unique values, then it is not a good choice for grouping.\n\n4. **Null Values:** The column should not have too many null values. If a column has too many null values, then it is not a good choice for grouping.\n\n**Distinct mean values help identify meaningful groups, while similar standard deviations help ensure that the groups are not too noisy.**","metadata":{}},{"cell_type":"markdown","source":"### Mean and Standard Deviation of Target Column 🎯 <a id=\"mean-std\"></a>\n\nI will also calculate the mean and standard deviation of the target column for each unique value in the column we choose for grouping the data. This will help us in knowing if the mean and standard deviation of the target column are significantly different across the groups; acting as the 5th criteria.","metadata":{}},{"cell_type":"markdown","source":"### Plotting Value Counts 📊 <a id=\"plotting\"></a>\n\nAbove is the code to plot the value counts for columns compared to those of original dataset, which can help us in deciding which column to use for grouping the data.\n\nThe plots are useful for knowing the portion size occupied by each unique value in the column, so let's see which column has the most balanced portion sizes.","metadata":{}},{"cell_type":"code","source":"def plot_mean_std_distr(df, col, title):\n    df = df.copy()\n    target = 'Premium Amount'\n    df[col] = df[col].fillna('Missing')\n    \n    # Group by column and calculate mean and standard deviation\n    stats = df.groupby(col)[target].agg(['mean', 'std']).sort_values('mean')\n    \n    # Percentage coverage of valid values\n    valid_percent = df[col].value_counts().sum() * 100 / df.shape[0]\n    print(f'[!]: {valid_percent:.2f}% of {col} is covered by valid values')\n    \n    # Plot mean and standard deviation\n    fig, ax = plt.subplots(2, 2, figsize=(20, 15))\n    \n    stats['mean'].plot(kind='bar', ax=ax[0, 0], color='firebrick', alpha=0.9)\n    ax[0, 0].set_title(f'Mean {target} by {col}')\n    ax[0, 0].set_ylabel('Mean Premium Amount')\n\n    # label magnitude in each bar\n    for p in ax[0, 0].patches:\n        ax[0, 0].annotate(f'{p.get_height():.2f}', (p.get_x() * 1.005, p.get_height() * 1.005))\n    \n    stats['std'].plot(kind='bar', ax=ax[0, 1], color='crimson', alpha=0.9)\n    ax[0, 1].set_title(f'Standard Deviation of {target} by {col}')\n    ax[0, 1].set_ylabel('Std Dev of Premium Amount')\n\n    # label magnitude in each bar\n    for p in ax[0, 1].patches:\n        ax[0, 1].annotate(f'{p.get_height():.2f}', (p.get_x() * 1.005, p.get_height() * 1.005))\n    \n    # Plot value counts for train and original datasets\n    # df_total_percent = df[col].value_counts().sum() * 100 / df.shape[0]\n    # print(f'[!]: {df_total_percent:.2f}% of {col} is covered by valid values')\n\n    df[col].value_counts(normalize=True).plot(kind='bar', ax=ax[1, 0], color='slateblue')\n    ax[1, 0].set_title(f'Distribution of {title} - Train')\n\n    test[col].value_counts(normalize=True).plot(kind='bar', ax=ax[1, 1], color='darkblue', alpha=0.8)\n    ax[1, 1].set_title(f'Distribution of {title} - Test')\n\n    plt.tight_layout()\n    plt.show()\n\n\ngroup_cols = [ \n    'Policy Type', 'Previous Claims', 'Insurance Duration', \n    'Occupation', 'Location', 'Customer Feedback', \n    'Education Level', 'Exercise Frequency', 'Marital Status', \n    'Number of Dependents', 'Property Type', \n    'Vehicle Age', 'Gender' ]\n\nfor col in group_cols:\n    plot_mean_std_distr(train, col, col)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-08T08:52:02.188901Z","iopub.execute_input":"2024-12-08T08:52:02.189749Z","iopub.status.idle":"2024-12-08T08:52:25.812447Z","shell.execute_reply.started":"2024-12-08T08:52:02.189709Z","shell.execute_reply":"2024-12-08T08:52:25.811208Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The `previous claims` column could be a great fit for the criteria only if it was not such an imbalanced class.\n\nIt has diverse mean, but its standard deviation is also diverse, specially for the 9th unique value. Wait! **The STD for 9th value is dead low because there's only ONE record for that value.** This is a good example of how the standard deviation can be misleading if the portion size is too small\n\n\nBut let's feature engineer the data to see if it helps in choosing the right column for grouping the data.\n\n## Feature Engineering 🛠 <a id=\"7\"></a>\n\nI will be extracting features from the `Policy Start Date` column. I will extract features like `year`, `month`, `day`, `dayofweek`, etc. from the `Policy Start Date` column.","metadata":{}},{"cell_type":"code","source":"def engg_policy_start_date(df):\n    df = df.copy()\n    df['Policy Start Date'] = pd.to_datetime(df['Policy Start Date'])\n    \n    df['Policy Start Year'] = df['Policy Start Date'].dt.year\n    df['Policy Start Month'] = df['Policy Start Date'].dt.month\n    \n    df['Policy Start Day'] = df['Policy Start Date'].dt.day\n    df['Policy Start Day of Week'] = df['Policy Start Date'].dt.dayofweek\n\n    df = df.drop('Policy Start Date', axis=1)\n    return df\n\ntrain = engg_policy_start_date(train)\ntest = engg_policy_start_date(test)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-08T12:38:00.814889Z","iopub.execute_input":"2024-12-08T12:38:00.81528Z","iopub.status.idle":"2024-12-08T12:38:02.624473Z","shell.execute_reply.started":"2024-12-08T12:38:00.815247Z","shell.execute_reply":"2024-12-08T12:38:02.623094Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"new_cols = ['Policy Start Year', 'Policy Start Month', 'Policy Start Day', 'Policy Start Day of Week']\n\nfor col in new_cols:\n    plot_mean_std_distr(train, col, col)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-08T08:53:00.795603Z","iopub.execute_input":"2024-12-08T08:53:00.795993Z","iopub.status.idle":"2024-12-08T08:53:08.962814Z","shell.execute_reply.started":"2024-12-08T08:53:00.79596Z","shell.execute_reply":"2024-12-08T08:53:08.961554Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### EDA Conclusion 📝 <a id=\"8\"></a>\n\nIt seems like the criteria cannot be applied here if we really want to group the data as it's synthetic, even on feature-engineering-produced columns. Let's accept the limitations of such datasets, and still move on to grouping.\n\n\n## Grouping the Data 📊 <a id=\"9\"></a>\n\nI will be grouping the data based on `Policy Start Year` column (produced from feature engineering). Here are the top reasons I have chosen this column:\n\n- The data can be split into 6 groups, which is optimal for myself to handle.\n\n- We can further combine these groups of similar mean and standard deviation to reduce the number of groups and boost variance in the data","metadata":{}},{"cell_type":"code","source":"def split_datasets(df, out_dir='/kaggle/working', label='train'):\n    df = df.copy()\n\n    policy_years = df['Policy Start Year'].unique()\n    datasets = {} \n\n    print(f'== Splitting {label} dataset ==')\n\n    for policy_year in policy_years:\n        datasets[policy_year] = df[df['Policy Start Year'] == policy_year].copy()\n        datasets[policy_year].drop('Policy Start Year', axis=1, inplace=True)\n\n        print(f'[!]: Saving {policy_year} dataset => {datasets[policy_year].shape}')\n        datasets[policy_year].to_csv(f'{out_dir}/{label}_{policy_year}.csv', index=False)\n\n    print('== Splitting complete ==\\n')\n    return datasets\n\ntrain_groups = split_datasets(df=train, label='train') \ntest_groups = split_datasets(df=test, label='test') ","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-08T12:38:09.592766Z","iopub.execute_input":"2024-12-08T12:38:09.593163Z","iopub.status.idle":"2024-12-08T12:38:34.855554Z","shell.execute_reply.started":"2024-12-08T12:38:09.593128Z","shell.execute_reply":"2024-12-08T12:38:34.854419Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Model Building 🏗 <a id=\"10\"></a>\n\nLet's build a model on each group of data and see how it performs. I will be using `LightGBM` for building the model.","metadata":{}},{"cell_type":"code","source":"from sklearn.metrics import mean_squared_error as mse\nfrom sklearn.model_selection import KFold\n\nimport optuna \nfrom lightgbm import LGBMRegressor\n\ndef rmsle(y_true, y_pred):\n    return mse(y_true, y_pred) ** 0.5\n\ndef train_lgb(\n        params={}, \n        folds=10, \n        train_data=None, \n        test_data=None,\n        gen_oof_preds=False):\n    \n    if train_data is None:\n        raise ValueError('train_data is required')\n    \n    if test_data is None and gen_oof_preds is True:\n        raise ValueError('test_data is required when generating out-of-fold predictions')\n    \n    oof_preds_over_test = []\n    oof_scores = []\n\n    feature_importance = pd.DataFrame()\n    feature_importance['feature'] = [col for col in train_data.columns if col != 'Premium Amount']\n    \n    X = train_data.drop('Premium Amount', axis=1)\n    y = train_data['Premium Amount']\n\n    kf = KFold(n_splits=folds, shuffle=True, random_state=42)\n\n    def progress_bar(style='-', current=1, total=10, length=40, label='folds'):\n        progress = current / total * length\n        bar = style * int(progress)\n\n        print(f'[{bar.ljust(length)}] {current}/{total} {label}', end='\\r')\n\n\n    for fold, (train_idx, valid_idx) in enumerate(kf.split(X, y)):\n        progress_bar(current=fold+1, total=folds, label='folds')\n\n        X_train, y_train = X.iloc[train_idx], y.iloc[train_idx]\n        X_valid, y_valid = X.iloc[valid_idx], y.iloc[valid_idx]\n\n        model = LGBMRegressor(**params)\n        model.fit(X_train, y_train, eval_set=[(X_valid, y_valid)])\n\n        oof_preds = model.predict(X_valid)\n        oof_scores.append(rmsle(y_valid, np.abs(oof_preds)))\n\n        if gen_oof_preds:\n            oof_preds_over_test.append(np.expm1(model.predict(test_data)))\n\n        feature_importance[f'fold_{fold+1}'] = model.feature_importances_\n    avg_oof_score = np.mean(oof_scores)\n    \n    if gen_oof_preds:\n        return avg_oof_score, model, oof_preds_over_test, feature_importance\n    return avg_oof_score, model, None, feature_importance","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-08T12:38:44.77184Z","iopub.execute_input":"2024-12-08T12:38:44.772227Z","iopub.status.idle":"2024-12-08T12:38:44.785486Z","shell.execute_reply.started":"2024-12-08T12:38:44.772193Z","shell.execute_reply":"2024-12-08T12:38:44.783977Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Hyperparameter Tuning 🎯 <a id=\"11\"></a>\r\n\r\nFor a quick experiment with the grouping of this data, let's use LightGBM. For hyperparameter selection, I will be using `Optuna`.\r\n\r\nFor Cross-validation, I will be using `KFold` with 10 splits in each group. I will also be using `Early Stopping` to prevent overfitting.","metadata":{}},{"cell_type":"code","source":"import warnings \nwarnings.filterwarnings('ignore')\n\ndef tune_lgb(train_groups, test_groups):\n\n    group_names = train_groups.keys()\n    train_dfs, test_dfs = train_groups.values(), test_groups.values()\n\n\n    # == Objective function for Optuna ==\n    def obj_lgb(trial):\n        params = {\n            # 'n_estimators': 1000, -> let's not tune this.\n            'learning_rate': trial.suggest_float('learning_rate', 1e-2, 0.5),\n            'max_depth': trial.suggest_int('max_depth', 3, 10),\n            'num_leaves': trial.suggest_int('num_leaves', 10, 100),\n            'subsample': trial.suggest_uniform('subsample', 0.1, 1.0),\n            'colsample_bytree': trial.suggest_uniform('colsample_bytree', 0.1, 1.0),\n            'reg_alpha': trial.suggest_loguniform('reg_alpha', 1e-3, 10.0),\n            'reg_lambda': trial.suggest_loguniform('reg_lambda', 1e-3, 10.0),\n            'random_state': 84,\n            'n_jobs': -1,\n            'early_stopping_round': 60,\n            'metric': 'rmse',\n            'verbosity': -1\n        }\n\n        avg_oof_score, _, _, _ = train_lgb(params, folds=10, train_data=train_spl, test_data=test_spl)\n        return avg_oof_score\n    \n\n    # == Data Preperation Step ==\n    def prepare_data(train_split, test_split):\n\n        for col in train_split.select_dtypes('object').columns:\n            train_split[col] = train_split[col].astype('category').cat.codes\n            test_split[col] = test_split[col].astype('category').cat.codes\n\n        return train_split, test_split\n\n\n    # == Finally, let's hyperparameter tune the model for each group == \n    for grp, train_df, test_df in zip(group_names, train_dfs, test_dfs):\n\n        print(f'[!]: {grp} | Training on {train_df.shape} and testing on {test_df.shape}')\n        train_spl, test_spl = prepare_data(train_df, test_df)\n\n        study = optuna.create_study(direction='minimize')\n        study.optimize(obj_lgb, n_trials=10)\n\n        print(f'[!]: Best trial for {grp} => {study.best_trial.value}')\n        print(f'[!]: Best params for {grp} => {study.best_trial.params}')\n        print('==' * 20)\n\n# tune_lgb(train_groups, test_groups) #-> Uncomment","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-08T12:38:51.282244Z","iopub.execute_input":"2024-12-08T12:38:51.28268Z","iopub.status.idle":"2024-12-08T13:00:41.128391Z","shell.execute_reply.started":"2024-12-08T12:38:51.282642Z","shell.execute_reply":"2024-12-08T13:00:41.127333Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Making Predictions 🔮 <a id=\"12\"></a>\n\n\nNow, let's predict over the groups in our dataset. But wait! let me tell you about a huge blunder I did while hyperparameter tuning... I forgot to exclude the ID column before inputing it into the hyperparameter selection process. Let's avoid that while making predictions?\n","metadata":{}},{"cell_type":"code","source":"lgbm_params = {\n    '2023': {'verbosity': -1},\n    '2024': {'verbosity': -1},\n    '2021': {'verbosity': -1},\n    '2022': {'verbosity': -1},\n    '2020': {'verbosity': -1},\n    '2019': {'verbosity': -1}\n}\n\n\ndef prepare_data(train_split, test_split):\n    for col in train_split.select_dtypes('object').columns:\n        train_split[col] = train_split[col].astype('category').cat.codes\n        test_split[col] = test_split[col].astype('category').cat.codes\n\n    return train_split, test_split\n\navg_oof_preds = {}\ngroups_oof_preds = {}\n\nfor grp, params in lgbm_params.items():\n    print(f'\\n[!]: Training on {grp} dataset')\n\n    train_df = pd.read_csv(f'/kaggle/working/train_{grp}.csv')\n    test_df = pd.read_csv(f'/kaggle/working/test_{grp}.csv')\n\n    train_df, test_df = prepare_data(train_df, test_df)\n\n    test_ids = test_df['id']\n    test = test_df.drop('id', axis=1)\n\n    avg_oof_score, model, oof_preds, feature_importance = train_lgb(params, folds=10, train_data=train_df, test_data=test_df, gen_oof_preds=True)\n    print(f'\\n[!]: Average OOF score for {grp} => {avg_oof_score:.6f}')\n\n    grp_oof = np.mean(oof_preds, axis=0)\n    avg_oof_preds[grp] = (test_ids, np.mean(oof_preds, axis=0))\n\n    groups_oof_preds[grp] = oof_preds\n\n    print('\\n\\n')\n    print('==' * 20)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-08T13:12:57.140381Z","iopub.execute_input":"2024-12-08T13:12:57.140798Z","iopub.status.idle":"2024-12-08T13:16:26.587006Z","shell.execute_reply.started":"2024-12-08T13:12:57.140763Z","shell.execute_reply":"2024-12-08T13:16:26.585848Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Consolidating Predictions 👩‍💻 <a id=\"13\"></a>\n\nNow let's combine these predictions into one submission file. But wait again! Won't you see what I did in the above code block?\n\nThe dataset contains records from years 2019 to 2024, so we have 6 groups. For each group, a model gathers 10 predictions (because it predicts on the test dataset over each fold). We don't need 60 predictions to confuse ourselves, so for each group we append *average* of the 10 predictions along with corresponding IDs to each predicted value from the test set. \n\nNow we have 6 sets of predictions which we have to combine. I will be sorting them as well, so let's get started with this too!","metadata":{}},{"cell_type":"code","source":"def consolidate_submission(avg_preds):\n    avg_preds = avg_preds.copy() # so that I don't mess up the original dictionary due to my silly mistakes\n\n    groups = ['2019', '2020', '2021', '2022', '2023', '2024']\n    sub = pd.DataFrame()\n\n    for grp in groups:\n        sub = pd.concat([sub, pd.DataFrame({'id': avg_preds[grp][0], 'Premium Amount': avg_preds[grp][1]})])\n\n    sub = sub.sort_values('id')\n    return sub\n\nsubmission = consolidate_submission(avg_oof_preds)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-08T13:16:35.753409Z","iopub.execute_input":"2024-12-08T13:16:35.753822Z","iopub.status.idle":"2024-12-08T13:16:35.84308Z","shell.execute_reply.started":"2024-12-08T13:16:35.753788Z","shell.execute_reply":"2024-12-08T13:16:35.842118Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(submission.shape[0])\nsubmission.to_csv('/kaggle/working/submission.csv', index=False)\n\nsubmission","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-08T13:16:38.785081Z","iopub.execute_input":"2024-12-08T13:16:38.785502Z","iopub.status.idle":"2024-12-08T13:16:40.466453Z","shell.execute_reply.started":"2024-12-08T13:16:38.785465Z","shell.execute_reply":"2024-12-08T13:16:40.465312Z"}},"outputs":[],"execution_count":null}]}