{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":84896,"databundleVersionId":10305135,"sourceType":"competition"},{"sourceId":10191187,"sourceType":"datasetVersion","datasetId":6296073}],"dockerImageVersionId":30786,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## 🔎 Overview dataset\n\nThe origiginal [dataset](https://www.kaggle.com/datasets/schran/insurance-premium-prediction) was created to facilitate the development and testing of regression models for predicting insurance premiums based on various customer characteristics and policy details. The dataset in this competition was generated by a deep learning model trained on the original dataset. Therefore the dataset contains all the columns of the original dataset but the distribution of the features might differ. The features in the dataset are:\n\n* **Age**: age of the insured individual.\n* **Gender**: gender of the insured individual.\n* **Annual Income**: annual income of the insured individual\n* **Marital Status**: Marital status of the insured individual,\n* **Number of Dependents**: Number of dependents.\n* **Education Level**: Highest education level attained\n* **Occupation**: Occupation of the insured individual.\n* **Health Score**: A score representing the health status.\n* **Location**: Type of location.\n* **Policy Type**: Type of insurance policy.\n* **Previous Claims**: Number of previous claims made.\n* **Vehicle Age**: Age of the vehicle insured.\n* **Credit Score**: Credit score of the insured individual.\n* **Insurance Duration**: Duration of the insurance policy.\n* **Premium Amount**: Target variable representing the insurance premium amount.\n* **Policy Start Date**: Start date of the insurance policy.\n* **Customer Feedback**: Short feedback comments from customers.\n* **Smoking Status**: Smoking status of the insured individual.\n* **Exercise Frequency**: Frequency of exercise.\n* **Property Type**: Type of property owned.","metadata":{}},{"cell_type":"markdown","source":"| date | version| description |\r\n| --- | --- | --- |\r\n20124210 | 14 | only fill missing for categoricals. no encoding for categoricals. |n only fill missing for categoricals. no \n20241213 | 19 | add non-log as feature and adjusted training code to create the non-log dataset | |","metadata":{}},{"cell_type":"markdown","source":"## 📊 Datatypes and missings","metadata":{}},{"cell_type":"code","source":"import joblib\nimport pandas as pd\nimport numpy as np\nimport pathlib\nimport warnings\nimport torch\n\nif torch.cuda.is_available():\n    device = 'GPU'\nelse:\n    device = 'CPU'\n\nwarnings.filterwarnings(\"ignore\")\n\ninput_path = pathlib.Path('/kaggle/input/playground-series-s4e12')\n\ntrain_df = pd.read_csv(input_path / 'train.csv', index_col='id')\ntrain_df.columns = train_df.columns.str.lower().str.split(' ').str.join('_')\n\ntest_df = pd.read_csv(input_path / 'test.csv', index_col='id')\ntest_df.columns = test_df.columns.str.lower().str.split(' ').str.join('_')\n\nsample_submission_df = pd.read_csv(input_path / 'sample_submission.csv')\n\nnon_log_oof, non_log_test = joblib.load(\n    '/kaggle/input/playground-s4e12-non-log/cat_non_loged.pkl'\n)\n\ntrain_df.shape, test_df.shape, non_log_oof.shape, non_log_test.shape","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T14:24:16.58197Z","iopub.execute_input":"2024-12-30T14:24:16.58276Z","iopub.status.idle":"2024-12-30T14:24:22.230439Z","shell.execute_reply.started":"2024-12-30T14:24:16.582722Z","shell.execute_reply":"2024-12-30T14:24:22.229523Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(f'Using {device} device.')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T14:24:22.231714Z","iopub.execute_input":"2024-12-30T14:24:22.232026Z","iopub.status.idle":"2024-12-30T14:24:22.236981Z","shell.execute_reply.started":"2024-12-30T14:24:22.231986Z","shell.execute_reply":"2024-12-30T14:24:22.236146Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_df['non_log_premium_amount'] = non_log_oof\ntest_df['non_log_premium_amount'] = non_log_test","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T14:24:22.238175Z","iopub.execute_input":"2024-12-30T14:24:22.238791Z","iopub.status.idle":"2024-12-30T14:24:22.251554Z","shell.execute_reply.started":"2024-12-30T14:24:22.238752Z","shell.execute_reply":"2024-12-30T14:24:22.250792Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_df.head(5)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T14:24:22.253721Z","iopub.execute_input":"2024-12-30T14:24:22.253961Z","iopub.status.idle":"2024-12-30T14:24:22.274094Z","shell.execute_reply.started":"2024-12-30T14:24:22.253936Z","shell.execute_reply":"2024-12-30T14:24:22.273215Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_df_missing = train_df.isna().sum().to_frame().rename(columns={0: 'n_missing'})\ntrain_df_missing['pct_missing'] = np.round(train_df_missing['n_missing'] / train_df.shape[0], 5)\ntrain_df_missing","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T14:24:22.275097Z","iopub.execute_input":"2024-12-30T14:24:22.27537Z","iopub.status.idle":"2024-12-30T14:24:22.818094Z","shell.execute_reply.started":"2024-12-30T14:24:22.275344Z","shell.execute_reply":"2024-12-30T14:24:22.817257Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"test_df.head(5)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T14:24:22.81906Z","iopub.execute_input":"2024-12-30T14:24:22.819367Z","iopub.status.idle":"2024-12-30T14:24:22.838482Z","shell.execute_reply.started":"2024-12-30T14:24:22.81934Z","shell.execute_reply":"2024-12-30T14:24:22.837549Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"test_df.info()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T14:24:22.839619Z","iopub.execute_input":"2024-12-30T14:24:22.839869Z","iopub.status.idle":"2024-12-30T14:24:23.202804Z","shell.execute_reply.started":"2024-12-30T14:24:22.839844Z","shell.execute_reply":"2024-12-30T14:24:23.20196Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"test_df_missing = test_df.isna().sum().to_frame().rename(columns={0: 'n_missing'})\ntest_df_missing['pct_missing'] = np.round(test_df_missing['n_missing'] / test_df.shape[0], 5)\ntest_df_missing","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T14:24:23.203884Z","iopub.execute_input":"2024-12-30T14:24:23.204144Z","iopub.status.idle":"2024-12-30T14:24:23.563466Z","shell.execute_reply.started":"2024-12-30T14:24:23.204117Z","shell.execute_reply":"2024-12-30T14:24:23.562692Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport seaborn as sns\n\ncolor_palette=sns.color_palette(['#f8766d','#00bfc4', '#00b8e7','#00c19a', '#8494ff','#ed68ed'])\n\nfig, ax = plt.subplots(1, 2, figsize=(15, 5))\n\ntrain_df_missing['pct_missing'].plot(kind='bar', ax=ax[0], color='#f8766d')\nax[0].set_title('% missing train data')\n\ntest_df_missing['pct_missing'].plot(kind='bar', ax=ax[1], color='#00bfc4')\nax[1].set_title('% missing test data')\n\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T14:24:23.564341Z","iopub.execute_input":"2024-12-30T14:24:23.564581Z","iopub.status.idle":"2024-12-30T14:24:24.078677Z","shell.execute_reply.started":"2024-12-30T14:24:23.564557Z","shell.execute_reply":"2024-12-30T14:24:24.07777Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🔠 Categorical features\n\nLets start this analysis with inspecting the categorical features in the dataset. The goal is the check if there is a statistical significant relationship between the categories and the target 'premium_amount'. A way we can do this is by using a one-way ANOVA test or the Kruskal-Wallis H-test. The one-way ANOVA tests the null hypothesis that two or more groups have the same population mean. The Kruskal-Wallis H-test tests the null hypothesis that the population median of all of the groups are equal. It is a non-parametric version of ANOVA. Because the ANOVA has some underlying assumptions we use both to check for a statistical significant relaton between the categorical features and the target. See the Scipy documentation for more information on both tests [here](https://docs.scipy.org/doc/scipy/reference/generated/scipy.stats.kruskal.html) and [here](https://docs.scipy.org/doc/scipy/reference/generated/scipy.stats.f_oneway.html).","metadata":{}},{"cell_type":"code","source":"from scipy.stats import f_oneway\nfrom scipy.stats import kruskal\n\ntest_cat_list = []\ntest_cat_cols = []\n\ncat_cols = train_df.select_dtypes(include=['object', 'category']).columns.drop('policy_start_date')\n\nfor col in cat_cols:\n    test_group = train_df.groupby(col)['premium_amount'].apply(list)\n    \n    f_oneway_result = f_oneway(*test_group)\n    kruskal_result = kruskal(*test_group)\n\n    test_cat_list.append(\n        [f_oneway_result.statistic, f_oneway_result.pvalue, kruskal_result.statistic, kruskal_result.pvalue]\n    )\n\n    test_cat_cols.append(\n        col\n    )\n\npd.DataFrame(test_cat_list, index=test_cat_cols, columns=['anova_statistic','anova_pvalue','kruskal_statistic','kruskal_pvalue'])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T14:24:24.081566Z","iopub.execute_input":"2024-12-30T14:24:24.08189Z","iopub.status.idle":"2024-12-30T14:24:28.098954Z","shell.execute_reply.started":"2024-12-30T14:24:24.081858Z","shell.execute_reply":"2024-12-30T14:24:28.098121Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Based on the the output of the tests there is only one categorical feature (occupation) for which the H0 hypothesis can be rejected. This means that there is a statistical significant difference between premium amount and someones occupation. Maybe we could put this to our benefit in our feature engineering or modelling process.","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport seaborn as sns\n\nn_rows = len(cat_cols)\nn_cols = 3\n\nfor idx, cat_col in enumerate(cat_cols):\n    \n    fix, ax = plt.subplots(1, n_cols, figsize=(15,5))\n\n    mean_premium_amount = train_df.groupby(cat_col)['premium_amount'].mean().reset_index()\n\n    sns.kdeplot(\n        data=train_df, x='premium_amount', hue=cat_col, ax=ax[0], palette=color_palette\n    )\n    \n    sns.countplot(\n        data=train_df, x=cat_col, ax=ax[1], palette=color_palette\n    )\n    \n    sns.barplot(\n        data=mean_premium_amount, x=cat_col, y='premium_amount', ax=ax[2], palette=color_palette\n    )\n    ax[0].set_title('kernel density')\n    ax[1].set_title('count of categories')\n    ax[2].set_title('mean premium amount by category')\n\nplt.tight_layout()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T14:24:28.10026Z","iopub.execute_input":"2024-12-30T14:24:28.100661Z","iopub.status.idle":"2024-12-30T14:25:30.488574Z","shell.execute_reply.started":"2024-12-30T14:24:28.100612Z","shell.execute_reply":"2024-12-30T14:25:30.487715Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"As the output from the tests already suggested; there isn't really a big difference between the different categorical features. Visual inspection endorsis this.","metadata":{}},{"cell_type":"markdown","source":"## 🔢 Numerical features\n\nWe continu by inspecting the numerical features in the dataset. We again start by checking if there is a statistical significant relation between the target (premium amount) and each numerical feature in the dataset. This time we can use the Spearman rank-order correlation coefficient. The Spearman rank-order correlation coefficient is a nonparametric measure of the monotonicity of the relationship between two datasets. Like other correlation coefficients, this one varies between -1 and +1 with 0 implying no correlation. Correlations of -1 or +1 imply an exact monotonic relationship. Positive correlations imply that as x increases, so does y. Negative correlations imply that as x increases, y decreases. See the Scipy [documenation](https://docs.scipy.org/doc/scipy/reference/generated/scipy.stats.spearmanr.html) for more information.","metadata":{}},{"cell_type":"code","source":"from scipy.stats import spearmanr\n\ntest_num_list = []\ntest_num_cols = []\n\nnum_cols = train_df.select_dtypes(include=['int64','float64']).columns.drop('premium_amount')\n\nfor num_col in num_cols:\n\n    # Drop NaN manually because param omit_nan is really slow\n    correlation, p_value = spearmanr(\n        train_df.loc[\n            train_df[num_col].notna(), num_col], train_df.loc[train_df[num_col].notna(), 'premium_amount'\n        ]\n    )\n\n    test_num_list.append([correlation, np.round(p_value, 5)])\n    test_num_cols.append(num_col)\n\npd.DataFrame(test_num_list, index=test_num_cols, columns=['correlation', 'p_value'])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T14:25:30.490112Z","iopub.execute_input":"2024-12-30T14:25:30.490935Z","iopub.status.idle":"2024-12-30T14:25:32.184084Z","shell.execute_reply.started":"2024-12-30T14:25:30.490904Z","shell.execute_reply":"2024-12-30T14:25:32.183181Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The result of test is that there are a couple of numerical features which have a p-value smaller than 0.05. This means that for these features (age, annual income, health score, previous claims and credit score) there is a statistical significant relation between them and the target premium amount. However, the correlation coefficient for these features is so small that it is really hard to find out where the differences are. Therefore we continu or analysis with a visual inspection of each numerical feature.\n\n","metadata":{}},{"cell_type":"code","source":"n_rows = len(num_cols)\nn_cols = 2\n\nfor idx, num_col in enumerate(num_cols):\n\n    plot_num_df = train_df.copy()\n    plot_num_df[f'bin_{num_col}'] = pd.qcut(plot_num_df[num_col], q=50, duplicates='drop')\n    mean_premium_amount = plot_num_df.groupby(f'bin_{num_col}')['premium_amount'].mean().reset_index()\n    \n    fig, ax = plt.subplots(1, n_cols, figsize=(20, 5))\n\n    sns.histplot(\n        data=train_df, x=num_col, ax=ax[0], bins=50, palette=color_palette, color='#f8766d'\n    )\n\n    sns.barplot(\n        data=mean_premium_amount, x=f'bin_{num_col}', y='premium_amount', ax=ax[1], color='#00bfc4'\n    )\n\n    ax[1].tick_params(axis='x', labelrotation=90)\n\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T14:25:32.185056Z","iopub.execute_input":"2024-12-30T14:25:32.185414Z","iopub.status.idle":"2024-12-30T14:25:43.360822Z","shell.execute_reply.started":"2024-12-30T14:25:32.185385Z","shell.execute_reply":"2024-12-30T14:25:43.359985Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Visual inspection of the numerical features shows that there are some (again small) differences between each feature and the target premium amount. The most interesting one probably is the (forced by cutting) linear relation between previous claims and premium amount.  ","metadata":{}},{"cell_type":"markdown","source":"## 🎯 Target\n\nWe continu our analysis by inspecting the distribution of the target premium amount. We can see that premium amount is right skewed. Therefore we can try to transform our target to see if we transform it to a (more) normal distribution. To do this we apply a log transformation. As noted in this discussion [thread](https://www.kaggle.com/competitions/playground-series-s4e12/discussion/549336) it could also be helpful in our modeling process.","metadata":{}},{"cell_type":"code","source":"fig, ax = plt.subplots(1, 2, figsize=(15, 5))\n\ntrain_df['premium_amount'].plot(kind='hist', bins=200, ax=ax[0], color='#f8766d')\n\nnp.log1p(train_df['premium_amount']).plot(kind='hist', bins=200, ax=ax[1], color='#00bfc4')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T14:25:43.362042Z","iopub.execute_input":"2024-12-30T14:25:43.362682Z","iopub.status.idle":"2024-12-30T14:25:44.248852Z","shell.execute_reply.started":"2024-12-30T14:25:43.362639Z","shell.execute_reply":"2024-12-30T14:25:44.247946Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Classifier for premium bins","metadata":{}},{"cell_type":"code","source":"train_df['year'] = pd.to_datetime(train_df['policy_start_date']).dt.year\ntest_df['year'] = pd.to_datetime(test_df['policy_start_date']).dt.year","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T14:25:44.250039Z","iopub.execute_input":"2024-12-30T14:25:44.250424Z","iopub.status.idle":"2024-12-30T14:25:44.890099Z","shell.execute_reply.started":"2024-12-30T14:25:44.250372Z","shell.execute_reply":"2024-12-30T14:25:44.889184Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"q1 = np.percentile(train_df['premium_amount'], q=25)\nq3 = np.percentile(train_df['premium_amount'], q=75)\niqr = q3 - q1\n\nlower_bound = q1 - 1.5 * iqr\nupper_bound = q3 + 1.5 * iqr","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T14:25:44.891234Z","iopub.execute_input":"2024-12-30T14:25:44.891583Z","iopub.status.idle":"2024-12-30T14:25:44.925357Z","shell.execute_reply.started":"2024-12-30T14:25:44.891553Z","shell.execute_reply":"2024-12-30T14:25:44.92472Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_df['premium_bin'] = np.where(\n    train_df.premium_amount <= q1, 1.0,\n    np.where(\n        (train_df.premium_amount > q1) & (train_df.premium_amount <= q3), 2.0, 3.0\n    )\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T14:25:44.926407Z","iopub.execute_input":"2024-12-30T14:25:44.926672Z","iopub.status.idle":"2024-12-30T14:25:44.946064Z","shell.execute_reply.started":"2024-12-30T14:25:44.926646Z","shell.execute_reply":"2024-12-30T14:25:44.945168Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"X_clf = train_df.drop(columns=['premium_bin','policy_start_date', 'premium_amount'])\ny_clf = train_df['premium_bin'].values.reshape(-1,1)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T14:25:44.947069Z","iopub.execute_input":"2024-12-30T14:25:44.94732Z","iopub.status.idle":"2024-12-30T14:25:45.090293Z","shell.execute_reply.started":"2024-12-30T14:25:44.947267Z","shell.execute_reply":"2024-12-30T14:25:45.089566Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"test_df_clf = test_df.drop(columns=['policy_start_date'])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T14:25:45.091351Z","iopub.execute_input":"2024-12-30T14:25:45.091612Z","iopub.status.idle":"2024-12-30T14:25:45.149299Z","shell.execute_reply.started":"2024-12-30T14:25:45.091587Z","shell.execute_reply":"2024-12-30T14:25:45.148327Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"for col in cat_cols:\n  X_clf[col] = X_clf[col].fillna('Unknown')\n  X_clf[col] = X_clf[col].astype('category')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T14:25:45.150428Z","iopub.execute_input":"2024-12-30T14:25:45.150707Z","iopub.status.idle":"2024-12-30T14:25:46.411333Z","shell.execute_reply.started":"2024-12-30T14:25:45.15068Z","shell.execute_reply":"2024-12-30T14:25:46.410345Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from sklearn.model_selection import StratifiedKFold\nfrom sklearn.preprocessing import StandardScaler\nfrom catboost import CatBoostClassifier\nfrom sklearn.metrics import roc_auc_score\n\nclf_params = {\n    'iterations': 698, \n    'depth': 10, \n    'l2_leaf_reg': 7.317083451091762, \n    'bagging_temperature': 0.7928793050151599, \n    'random_strength': 3.165441634122922, \n    'min_data_in_leaf': 9, \n    'border_count': 226,\n    'loss_function': 'MultiClass',\n    'task_type': device.upper(),\n    'cat_features': [c for c in cat_cols],\n    'verbose': False,\n    'use_best_model': True\n}\n\ncb_clf = CatBoostClassifier(**clf_params)\n\ndef train_catboost_classifier(model, X, y):\n\n    skf = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)\n\n    auc_scores = []\n    oof_scores = np.zeros((len(X), 3))\n    clf_models = []\n\n    for i, (train_idx, val_idx) in enumerate(skf.split(X, y)):\n\n        X_train, X_val = X.iloc[train_idx], X.iloc[val_idx]\n        y_train, y_val = y[train_idx], y[val_idx]\n\n        # Train the model on the current fold\n        model.fit(X_train, y_train, eval_set=(X_val, y_val), verbose=200)\n    \n        # Get predicted probabilities\n        y_pred_proba = cb_clf.predict_proba(X_val)\n        oof_scores[val_idx] = y_pred_proba\n    \n        # Calculate AUC for this fold\n        auc = roc_auc_score(y_val, y_pred_proba, multi_class='ovr')\n        auc_scores.append(auc)\n        clf_models.append(model)\n    \n        print(f'=== Fold {i + 1} AUC Score {auc} ===')\n    \n    average_auc = np.mean(auc_scores)\n    print(f\"Average AUC score across folds: {average_auc:.4f}\")\n\n    return clf_models, oof_scores","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T14:25:46.412507Z","iopub.execute_input":"2024-12-30T14:25:46.412779Z","iopub.status.idle":"2024-12-30T14:25:46.42132Z","shell.execute_reply.started":"2024-12-30T14:25:46.412753Z","shell.execute_reply":"2024-12-30T14:25:46.420512Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"clf_models, oof_scores = train_catboost_classifier(cb_clf, X_clf, y_clf)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T14:25:46.422243Z","iopub.execute_input":"2024-12-30T14:25:46.422509Z","iopub.status.idle":"2024-12-30T14:28:49.843678Z","shell.execute_reply.started":"2024-12-30T14:25:46.422484Z","shell.execute_reply":"2024-12-30T14:28:49.842645Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"for col in cat_cols:\n  test_df_clf[col] = test_df_clf[col].fillna('Unknown')\n  test_df_clf[col] = test_df_clf[col].astype('category')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T14:28:49.845052Z","iopub.execute_input":"2024-12-30T14:28:49.845413Z","iopub.status.idle":"2024-12-30T14:28:50.656722Z","shell.execute_reply.started":"2024-12-30T14:28:49.845381Z","shell.execute_reply":"2024-12-30T14:28:50.655825Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"clf_preds_test = []\nclf_preds_train = []\n\nfor clf_model in clf_models:\n    # Predict bin for train\n    clf_pred_train = clf_model.predict_proba(X_clf)\n    clf_preds_train.append(clf_pred_train)\n\n    # Predict bin for test\n    clf_pred_test = clf_model.predict_proba(test_df_clf)\n    clf_preds_test.append(clf_pred_test)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T14:28:50.657793Z","iopub.execute_input":"2024-12-30T14:28:50.658074Z","iopub.status.idle":"2024-12-30T14:29:13.770442Z","shell.execute_reply.started":"2024-12-30T14:28:50.658046Z","shell.execute_reply":"2024-12-30T14:29:13.769671Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_df['premium_bin'] = np.argmax(\n    oof_scores, axis=1\n) + 1\n\ntest_df['premium_bin'] = np.argmax(\n    np.sum(clf_preds_test, axis=0) / len(clf_preds_test), axis=1\n) + 1","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T14:29:13.771626Z","iopub.execute_input":"2024-12-30T14:29:13.771996Z","iopub.status.idle":"2024-12-30T14:29:13.850991Z","shell.execute_reply.started":"2024-12-30T14:29:13.771957Z","shell.execute_reply":"2024-12-30T14:29:13.850007Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## ⏳ Preprocessing\n\nWe start by handling the columns with missing values. The first step will be checking if the columns with missing values are equal in both the training and testing data.","metadata":{}},{"cell_type":"code","source":"train_missing_cols = set(train_df_missing.loc[train_df_missing.n_missing > 0].index)\ntest_missing_cols = set(test_df_missing.loc[test_df_missing.n_missing > 0].index)\n\nprint(f'Number of columns with missings train minus test: {len(train_missing_cols - test_missing_cols)}')\nprint(f'Number of columns with missings test minus train: {len(test_missing_cols - train_missing_cols)}')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T14:29:13.852121Z","iopub.execute_input":"2024-12-30T14:29:13.852425Z","iopub.status.idle":"2024-12-30T14:29:13.858608Z","shell.execute_reply.started":"2024-12-30T14:29:13.852398Z","shell.execute_reply":"2024-12-30T14:29:13.857773Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def preprocess_categorical_features(df):\n    \"\"\"\n    Preprocess the categorical features by filling in missings.\n\n    Args:\n        df: pandas dataframe\n\n    Returns: pandas dataframe with filled missings\n    \n    \"\"\"\n\n    df = df.copy()\n\n    df['marital_status'] = df['marital_status'].fillna('Unknown')\n    df['occupation'] = df['occupation'].fillna('Unkown')\n    df['customer_feedback'] = df['customer_feedback'].fillna('Unknown')\n\n    return df\n\n\ndef preprocess_numerical_features(df):\n    \"\"\"\n    Preprocess the numerical features by filling in missings.\n\n    Args:\n        df: pandas dataframe\n\n    Returns: pandas dataframe with filled missings\n    \n    \"\"\"\n    df = df.copy()\n    \n    df['number_of_dependents'].fillna(df['number_of_dependents'].median(), inplace=True)\n    df['age'].fillna(df['age'].median(), inplace=True)\n    df['vehicle_age'].fillna(df['vehicle_age'].median(), inplace=True)\n    df['previous_claims'].fillna(df['previous_claims'].median(), inplace=True)\n    df['annual_income'].fillna(df['annual_income'].median(), inplace=True)\n    df['credit_score'].fillna(df['credit_score'].median(), inplace=True)\n    df['insurance_duration'].fillna(df['insurance_duration'].median(), inplace=True)\n    df['health_score'].fillna(df['health_score'].median(), inplace=True)\n\n    return df\n    ","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T14:29:13.859685Z","iopub.execute_input":"2024-12-30T14:29:13.859925Z","iopub.status.idle":"2024-12-30T14:29:13.869087Z","shell.execute_reply.started":"2024-12-30T14:29:13.8599Z","shell.execute_reply":"2024-12-30T14:29:13.868328Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Omit this step for now\n\n# train_df = preprocess_categorical_features(train_df)\n# train_df = preprocess_numerical_features(train_df)\n\n# test_df = preprocess_categorical_features(test_df)\n# test_df = preprocess_numerical_features(test_df)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T14:29:13.874896Z","iopub.execute_input":"2024-12-30T14:29:13.875413Z","iopub.status.idle":"2024-12-30T14:29:13.880483Z","shell.execute_reply.started":"2024-12-30T14:29:13.875387Z","shell.execute_reply":"2024-12-30T14:29:13.879728Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(f'Number of missing in training data: {train_df.isnull().sum().sum()}')\nprint(f'Number of missing in testing data: {test_df.isnull().sum().sum()}')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T14:29:13.881374Z","iopub.execute_input":"2024-12-30T14:29:13.881682Z","iopub.status.idle":"2024-12-30T14:29:14.766879Z","shell.execute_reply.started":"2024-12-30T14:29:13.881636Z","shell.execute_reply":"2024-12-30T14:29:14.766086Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Good! No more missing data in both dataframes 😊","metadata":{}},{"cell_type":"code","source":"def get_policy_start_date_features(df):\n\n    feature_cols = [\n        'policy_start_date',\n        'previous_claims',\n        'credit_score',\n        'health_score',\n        'year' # quick fix for duplicate\n    ]\n\n    df = df.copy()[feature_cols]\n\n    df['policy_start_date'] = pd.to_datetime(df['policy_start_date'])\n    today = pd.Timestamp.now()\n\n    df['year'] = df['policy_start_date'].dt.year\n    df['month'] = df['policy_start_date'].dt.month\n    df['week'] = df['policy_start_date'].dt.isocalendar().week.astype('int')\n    df['day'] = df['policy_start_date'].dt.day\n    df['dayofweek'] = df['policy_start_date'].dt.dayofweek\n    df['quarter'] = df['policy_start_date'].dt.quarter\n    df['is_weekend'] = df['dayofweek'].isin([5, 6])\n    df['is_sunday'] = df['dayofweek'].eq(6)\n    df['day_sin'] = np.sin(2 * np.pi * df['day']/31)\n    df['day_cos'] = np.cos(2 * np.pi * df['day']/31)\n    df['dayofweek_sin'] = np.sin(2 * np.pi * df['dayofweek']/6)\n    df['dayofweek_cos'] = np.cos(2 * np.pi * df['dayofweek']/6)\n    df['week_sin'] = np.sin(2 * np.pi * df['week']/52).astype('float64')\n    df['week_cos'] = np.cos(2 * np.pi * df['week']/52).astype('float64')\n    df['month_sin'] = np.sin(2 * np.pi * df['month']/12)\n    df['month_cos'] = np.cos(2 * np.pi * df['month']/12)\n\n    df['policy_age'] = (today - df['policy_start_date']).dt.days\n    df['policy_age_month'] = df['policy_age'] // 30\n    df['policy_age_years'] = df['policy_age'] // 365\n\n    df['claim_rate'] = df['previous_claims'] / df['policy_age']\n    df['combined_risk'] = df['health_score'] + df['credit_score'].fillna(0.0) + df['claim_rate']\n\n    df['n_nan'] = df.isna().sum(axis=1)\n\n    df['2019_policy'] = (df['year'] == 2019).astype('int')\n    \n    df = df.drop(columns=feature_cols)\n\n    return df","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T14:29:14.767876Z","iopub.execute_input":"2024-12-30T14:29:14.76811Z","iopub.status.idle":"2024-12-30T14:29:14.778263Z","shell.execute_reply.started":"2024-12-30T14:29:14.768086Z","shell.execute_reply":"2024-12-30T14:29:14.777437Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_date_features_df = get_policy_start_date_features(train_df)\ntest_date_features_df = get_policy_start_date_features(test_df)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T14:29:14.779329Z","iopub.execute_input":"2024-12-30T14:29:14.779586Z","iopub.status.idle":"2024-12-30T14:29:17.183822Z","shell.execute_reply.started":"2024-12-30T14:29:14.77956Z","shell.execute_reply":"2024-12-30T14:29:17.182807Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# train_df['premium_amount'] = train_df['premium_amount'] / train_df.groupby('year')['premium_amount'].transform('mean')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T14:29:17.185094Z","iopub.execute_input":"2024-12-30T14:29:17.185789Z","iopub.status.idle":"2024-12-30T14:29:17.18986Z","shell.execute_reply.started":"2024-12-30T14:29:17.185747Z","shell.execute_reply":"2024-12-30T14:29:17.188965Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"date_cols = list(train_date_features_df.columns)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T14:29:17.190958Z","iopub.execute_input":"2024-12-30T14:29:17.191333Z","iopub.status.idle":"2024-12-30T14:29:17.200344Z","shell.execute_reply.started":"2024-12-30T14:29:17.191268Z","shell.execute_reply":"2024-12-30T14:29:17.199604Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_df = pd.concat([train_df, train_date_features_df], axis=1).drop(columns=['policy_start_date'])\ntest_df = pd.concat([test_df, test_date_features_df], axis=1).drop(columns=['policy_start_date'])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T14:29:17.201427Z","iopub.execute_input":"2024-12-30T14:29:17.201668Z","iopub.status.idle":"2024-12-30T14:29:18.146266Z","shell.execute_reply.started":"2024-12-30T14:29:17.201644Z","shell.execute_reply":"2024-12-30T14:29:18.145562Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🌐 Modeling\n\nFirst transform the target by applying a log transformation. We also explore the option to add the prediction on the non-transformed target as a feature. As demonstrated by Backpacker, this approach could give as significant improvent on our leaderboard score while maintaining a steady CV-score accross folds. See his work [here](https://www.kaggle.com/code/backpaker/rid-catboost-nonlog-as-feature). Don't forget to upvote if you also like his work. 👍","metadata":{}},{"cell_type":"code","source":"combined_df = pd.concat([train_df, test_df], axis=0, ignore_index=True)\n\nfor cat_col in cat_cols:\n    freq_encoding = combined_df[cat_col].value_counts().to_dict()\n\n    train_df[f\"{cat_col}__freq\"] = train_df[cat_col].map(freq_encoding).astype('float')\n    test_df[f\"{cat_col}__freq\"] = test_df[cat_col].map(freq_encoding).astype('float')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T14:29:18.147368Z","iopub.execute_input":"2024-12-30T14:29:18.147671Z","iopub.status.idle":"2024-12-30T14:29:20.910079Z","shell.execute_reply.started":"2024-12-30T14:29:18.147641Z","shell.execute_reply":"2024-12-30T14:29:20.909075Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"X = train_df.drop(columns=['premium_amount'])\ny = train_df['premium_amount']\ny_log = np.log1p(train_df['premium_amount'])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T14:29:20.911347Z","iopub.execute_input":"2024-12-30T14:29:20.912087Z","iopub.status.idle":"2024-12-30T14:29:21.131905Z","shell.execute_reply.started":"2024-12-30T14:29:20.912043Z","shell.execute_reply":"2024-12-30T14:29:21.131171Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"cat_cols = [c for c in cat_cols] + ['premium_bin']\nnum_cols = [n for n in num_cols] + [c for c in train_df.columns if '__freq' in c]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T14:29:21.132961Z","iopub.execute_input":"2024-12-30T14:29:21.133238Z","iopub.status.idle":"2024-12-30T14:29:21.137919Z","shell.execute_reply.started":"2024-12-30T14:29:21.133209Z","shell.execute_reply":"2024-12-30T14:29:21.136974Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"X = X[cat_cols + num_cols + date_cols]\ntest_df = test_df[cat_cols + num_cols + date_cols]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T14:29:21.13892Z","iopub.execute_input":"2024-12-30T14:29:21.139166Z","iopub.status.idle":"2024-12-30T14:29:21.539185Z","shell.execute_reply.started":"2024-12-30T14:29:21.13914Z","shell.execute_reply":"2024-12-30T14:29:21.538219Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from sklearn.metrics import mean_squared_log_error\n\ndef rmsle(y_true, y_pred):\n    return np.sqrt(mean_squared_log_error(y_true, y_pred))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T14:29:21.540335Z","iopub.execute_input":"2024-12-30T14:29:21.5406Z","iopub.status.idle":"2024-12-30T14:29:21.545098Z","shell.execute_reply.started":"2024-12-30T14:29:21.540575Z","shell.execute_reply":"2024-12-30T14:29:21.544155Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from category_encoders import CatBoostEncoder, TargetEncoder\nfrom sklearn import set_config\nfrom sklearn.compose import ColumnTransformer\nfrom sklearn.impute import SimpleImputer\nfrom sklearn.preprocessing import StandardScaler, OneHotEncoder\nfrom sklearn.pipeline import Pipeline\n\nset_config(transform_output='pandas')\n\nnum_transformer = Pipeline(\n    steps=[\n        ('scaler', StandardScaler())\n    ]\n)\n\ncat_transformer = Pipeline(\n    steps=[\n        ('impute_cat', SimpleImputer(strategy='constant', fill_value='Unknown'))\n    ]\n)\n\npreprocessor = ColumnTransformer(\n    transformers=[\n        ('num', num_transformer, num_cols + date_cols),\n        ('cat', cat_transformer, cat_cols)\n    ], verbose_feature_names_out=False, remainder='passthrough'\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T14:29:21.546263Z","iopub.execute_input":"2024-12-30T14:29:21.546623Z","iopub.status.idle":"2024-12-30T14:29:21.553518Z","shell.execute_reply.started":"2024-12-30T14:29:21.546586Z","shell.execute_reply":"2024-12-30T14:29:21.552716Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"X = preprocessor.fit_transform(X, y_log)\ntest_df = preprocessor.transform(test_df)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T14:29:21.554374Z","iopub.execute_input":"2024-12-30T14:29:21.554652Z","iopub.status.idle":"2024-12-30T14:29:30.643544Z","shell.execute_reply.started":"2024-12-30T14:29:21.554625Z","shell.execute_reply":"2024-12-30T14:29:30.642807Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"for cat_col in cat_cols:\n    X[cat_col] = X[cat_col].astype('category')\n    test_df[cat_col] = test_df[cat_col].astype('category')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T14:29:30.644641Z","iopub.execute_input":"2024-12-30T14:29:30.644988Z","iopub.status.idle":"2024-12-30T14:29:31.819765Z","shell.execute_reply.started":"2024-12-30T14:29:30.644944Z","shell.execute_reply":"2024-12-30T14:29:31.819028Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Because most of the estimators in the Python modeling ecosystem follow the same implementation pattern we are going to define one training function. This way we can reduce the overhead of duplicating the cross-validation code and reduce the risk of potential errors by only implementing model specific code in the training function.","metadata":{}},{"cell_type":"code","source":"from xgboost import XGBRegressor\nfrom catboost import CatBoostRegressor, Pool\nfrom lightgbm import LGBMRegressor, early_stopping, log_evaluation, Dataset\nfrom sklearn.model_selection import KFold\nfrom sklearn.metrics import mean_squared_log_error\nfrom sklearn.metrics import mean_squared_log_error\n\ndef rmsle(y_true, y_pred):\n    return np.sqrt(mean_squared_log_error(y_true, y_pred))\n\n\ndef train_model(model, X, y, non_log=False):\n    \"\"\"\n    Train regression models in a cross-validation loop.\n\n    Args:\n        model: regression estimator (one of cb or lgbm)\n        X: training features\n        y: training target\n        non-log: boolean indication whether the target is \n            log-transformed or not.\n\n    Returns:\n        models: list of models obtained from a KF training run\n        scores: list of scores obtained from a KF training run\n    \"\"\"\n\n    kf = KFold(n_splits=5, shuffle=True, random_state=42)\n\n    models = []\n    scores = np.zeros(len(X))\n\n    for i, (train_idx, val_idx) in enumerate(kf.split(X)):\n        print(f'Fold {i + 1}')\n\n        X_train, X_val = X.iloc[train_idx], X.iloc[val_idx]\n        y_train, y_val = y.iloc[train_idx], y.iloc[val_idx]\n\n        # CatBoostRegressor\n        if isinstance(model, CatBoostRegressor):\n            model.fit(\n                X_train,\n                y_train,\n                eval_set=(X_val, y_val)\n            )\n\n        # LGBMRegressor\n        if isinstance(model, LGBMRegressor):\n            model.fit(\n                X_train,\n                y_train,\n                eval_set=(X_val, y_val)\n            )\n\n        if isinstance(model, XGBRegressor):\n            model.fit(\n                X_train,\n                y_train,\n                eval_set=[(X_val, y_val)],\n                verbose=0\n            )\n\n        models.append(model)\n        \n        y_pred = model.predict(X_val)\n        scores[val_idx] = np.maximum(0, y_pred)\n\n        if non_log:\n            score = rmsle(y_val, scores[val_idx])\n        else:\n            score = rmsle(np.expm1(y_val), np.expm1(scores[val_idx]))\n\n        print(f'=== Fold {i + 1} RSMLE Score {score} ===')\n\n    return models, scores","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T14:29:31.821093Z","iopub.execute_input":"2024-12-30T14:29:31.821401Z","iopub.status.idle":"2024-12-30T14:29:31.831694Z","shell.execute_reply.started":"2024-12-30T14:29:31.82137Z","shell.execute_reply":"2024-12-30T14:29:31.830866Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### CatBoost","metadata":{}},{"cell_type":"code","source":"# Parameters obtained after hyperparameter tuning\ncatboost_params = {\n    'iterations': 754,\n    'learning_rate': 0.02305535125965052,\n    'depth': 10,\n    'l2_leaf_reg': 1.171945061015989,\n    'bagging_temperature': 0.06625187054378748,\n    'random_strength': 0.29495752525094404,\n    'min_data_in_leaf': 52,\n    'border_count': 250,\n    'verbose': 200,\n    'task_type': device,\n    'random_seed': 42,\n    'cat_features': cat_cols,\n    'eval_metric': 'RMSE'\n}\n\ncb_model = CatBoostRegressor(**catboost_params)\n\ncb_models, cb_scores = train_model(\n    model=cb_model,\n    X=X,\n    y=y_log,\n    non_log=False\n)\n\nprint(f'Mean score accross all folds: {rmsle(y, np.expm1(cb_scores))}')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T14:29:31.832663Z","iopub.execute_input":"2024-12-30T14:29:31.8329Z","iopub.status.idle":"2024-12-30T14:37:21.610041Z","shell.execute_reply.started":"2024-12-30T14:29:31.832875Z","shell.execute_reply":"2024-12-30T14:37:21.609041Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"fig, ax = plt.subplots(figsize=(10, 10))\n\npd.Series(\n    np.mean([m.feature_importances_ for m in cb_models], axis=0), X.columns\n).sort_values().plot(kind='barh')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T14:37:21.61123Z","iopub.execute_input":"2024-12-30T14:37:21.611572Z","iopub.status.idle":"2024-12-30T14:37:22.232716Z","shell.execute_reply.started":"2024-12-30T14:37:21.611542Z","shell.execute_reply":"2024-12-30T14:37:22.231704Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import optuna\nfrom optuna.samplers import TPESampler\n\ndef catboost_objective(trial):\n    \n    kf = KFold(n_splits=5, shuffle=True, random_state=42) \n    scores = np.zeros(len(X)) \n    \n    params = {\n        \"iterations\": trial.suggest_int(\"iterations\", 500, 1000), \n        \"learning_rate\": trial.suggest_loguniform(\"learning_rate\", 0.01, 0.3), \n        \"depth\": trial.suggest_int(\"depth\", 4, 10), \n        \"l2_leaf_reg\": trial.suggest_loguniform(\"l2_leaf_reg\", 1, 10), \n        \"bagging_temperature\": trial.suggest_uniform(\"bagging_temperature\", 0.0, 1.0), \n        \"random_strength\": trial.suggest_loguniform(\"random_strength\", 0.1, 10.0), \n        \"subsample\": trial.suggest_uniform(\"subsample\", 0.5, 1.0), \n        \"min_data_in_leaf\": trial.suggest_int(\"min_data_in_leaf\", 1, 100), \n        \"border_count\": trial.suggest_int(\"border_count\", 32, 255), \n        \"eval_metric\": \"RMSE\", \n        \"cat_features\": cat_cols, \n        \"random_seed\": 42, \n        \"task_type\": device, \n        \"verbose\": 0\n    } \n\n    models = []\n    overall_score = 0 \n    \n    for i, (train_idx, val_idx) in enumerate(kf.split(X)):\n        \n        X_train, X_val = X.iloc[train_idx], X.iloc[val_idx]\n        y_train, y_val = y_log.iloc[train_idx], y_log.iloc[val_idx] \n        \n        train_pool = Pool(X_train, y_train, cat_features=cat_cols) \n        val_pool = Pool(X_val, y_val, cat_features=cat_cols) \n        \n        model = CatBoostRegressor(**params) \n        model.fit(train_pool, eval_set=val_pool) \n        \n        models.append(model) \n        \n        y_pred = model.predict(X_val) \n        scores[val_idx] = np.maximum(y_pred) \n        \n        score = rmsle(np.expm1(y_val), np.expm1(scores[val_idx])) \n        \n        overall_score += score \n        \n        print(f'=== Fold {i + 1} RMSLE Score: {score:} ===') \n        \n    avg_score = overall_score / kf.n_splits \n    print(f\"Overall RMSLE: {avg_score}\") \n        \n    return avg_score ","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T14:37:22.23392Z","iopub.execute_input":"2024-12-30T14:37:22.234249Z","iopub.status.idle":"2024-12-30T14:37:22.244457Z","shell.execute_reply.started":"2024-12-30T14:37:22.234212Z","shell.execute_reply":"2024-12-30T14:37:22.243426Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# study = optuna.create_study(direction='minimize')\n# study.optimize(catboost_objective, n_trials=100)\n# study.best_params","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T14:37:22.245507Z","iopub.execute_input":"2024-12-30T14:37:22.245812Z","iopub.status.idle":"2024-12-30T14:37:22.254737Z","shell.execute_reply.started":"2024-12-30T14:37:22.245777Z","shell.execute_reply":"2024-12-30T14:37:22.253995Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### LGBM","metadata":{}},{"cell_type":"code","source":"# Parameters not tuned yet, but basic \nlgbm_params = {\n    'n_estimators': 1000,\n    'learning_rate': 0.01,\n    'max_depth': 10,\n    'num_leaves': 256,\n    'reg_lambda': 1.17,\n    'reg_alpha': 0.1,\n    'categorical_feature': cat_cols,\n    'metric': 'rmse',\n    'objective': 'regression',\n    'random_state': 42,\n    'verbosity': 1,\n    'device': device.lower(),\n    'gpu_platform_id': 0,\n    'gpu_device_id': 0\n}\n\nlgbm_model = LGBMRegressor(**lgbm_params)\n\nlgbm_models, lgbm_scores = train_model(\n    model=lgbm_model,\n    X=X,\n    y=y_log,\n    non_log=False\n)\n\nprint(f'Mean score accross all folds: {rmsle(y, np.expm1(lgbm_scores))}')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T14:37:22.255775Z","iopub.execute_input":"2024-12-30T14:37:22.256143Z","iopub.status.idle":"2024-12-30T14:44:52.521613Z","shell.execute_reply.started":"2024-12-30T14:37:22.256109Z","shell.execute_reply":"2024-12-30T14:44:52.520691Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"fig, ax = plt.subplots(figsize=(10, 10))\n\npd.Series(\n    np.mean([m.feature_importances_ for m in lgbm_models], axis=0), X.columns\n).sort_values().plot(kind='barh')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T14:44:52.522708Z","iopub.execute_input":"2024-12-30T14:44:52.522979Z","iopub.status.idle":"2024-12-30T14:44:53.064467Z","shell.execute_reply.started":"2024-12-30T14:44:52.522951Z","shell.execute_reply":"2024-12-30T14:44:53.063554Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import optuna\n\ndef lgbm_objective(trial):\n    \n    kf = KFold(n_splits=5, shuffle=True, random_state=42) \n    scores = np.zeros(len(X)) \n    \n    params = {\n        \"n_estimators\": trial.suggest_int(\"n_estimators\", 500, 1000),\n        \"learning_rate\": trial.suggest_loguniform(\"learning_rate\", 0.01, 0.3),\n        \"max_depth\": trial.suggest_int(\"max_depth\", 4, 10),\n        \"reg_lambda\": trial.suggest_loguniform(\"reg_lambda\", 1, 10),\n        \"subsample\": trial.suggest_uniform(\"subsample\", 0.5, 1.0),\n        \"subsample_freq\": trial.suggest_int(\"subsample_freq\", 1, 5),\n        \"min_child_samples\": trial.suggest_int(\"min_child_samples\", 1, 100),\n        \"categorical_feature\": cat_cols,\n        \"seed\": 42,\n        \"eval_metric\": \"rmse\",\n        \"verbose\": -1\n    }\n\n    models = []\n    overall_score = 0 \n    \n    for i, (train_idx, val_idx) in enumerate(kf.split(X)):\n        \n        X_train, X_val = X.iloc[train_idx], X.iloc[val_idx]\n        y_train, y_val = y_log.iloc[train_idx], y_log.iloc[val_idx] \n        \n        model = LGBMRegressor(**params) \n        model.fit(\n            X_train,\n            y_train,\n            eval_set=(X_val, y_val)\n        ) \n        \n        models.append(model) \n        \n        y_pred = model.predict(X_val) \n        scores[val_idx] = np.maximum(0, y_pred) \n        \n        score = rmsle(np.expm1(y_val), np.expm1(scores[val_idx])) \n        \n        overall_score += score \n        \n        print(f'=== Fold {i + 1} RMSLE Score: {score:} ===') \n        \n    avg_score = overall_score / kf.n_splits \n    print(f\"Overall RMSLE: {avg_score}\") \n        \n    return avg_score ","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T14:44:53.065654Z","iopub.execute_input":"2024-12-30T14:44:53.065932Z","iopub.status.idle":"2024-12-30T14:44:53.073993Z","shell.execute_reply.started":"2024-12-30T14:44:53.065904Z","shell.execute_reply":"2024-12-30T14:44:53.073113Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# study = optuna.create_study(direction='minimize', sampler=TPESampler())\n# study.optimize(lgbm_objective, n_trials=100)\n# study.best_params","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T14:44:53.075112Z","iopub.execute_input":"2024-12-30T14:44:53.075474Z","iopub.status.idle":"2024-12-30T14:44:53.085958Z","shell.execute_reply.started":"2024-12-30T14:44:53.075435Z","shell.execute_reply":"2024-12-30T14:44:53.085091Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## XGB","metadata":{}},{"cell_type":"code","source":"# Parameters not tuned yet, but basic \nxgb_params = {\n    'n_estimators': 1000,\n    'learning_rate': 0.01,\n    'max_depth': 10,\n    'reg_lambda': 1.17,\n    'reg_alpha': 0.1,\n    'random_state': 42,\n    'num_leaves': None,\n    'min_child_weight': 1,\n    'objective': 'reg:squarederror',\n    'eval_metric': 'rmse',\n    'device': device.lower(),\n    'tree_method': 'gpu_hist' if device.lower() == 'gpu' else 'auto',\n    'verbosity': 0,\n    'enable_categorical': True,\n    'categorical_feature': cat_cols\n}\n\nxgb_model = XGBRegressor(**xgb_params)\n\nxgb_models, xgb_scores = train_model(\n    model=xgb_model,\n    X=X,\n    y=y_log,\n    non_log=False\n)\n\nprint(f'Mean score accross all folds: {rmsle(y, np.expm1(xgb_scores))}')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T14:44:53.086976Z","iopub.execute_input":"2024-12-30T14:44:53.087209Z","iopub.status.idle":"2024-12-30T14:49:11.972025Z","shell.execute_reply.started":"2024-12-30T14:44:53.087185Z","shell.execute_reply":"2024-12-30T14:49:11.971114Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"fig, ax = plt.subplots(figsize=(10, 10))\n\npd.Series(\n    np.mean([m.feature_importances_ for m in xgb_models], axis=0), X.columns\n).sort_values().plot(kind='barh')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T14:49:11.973296Z","iopub.execute_input":"2024-12-30T14:49:11.973669Z","iopub.status.idle":"2024-12-30T14:49:12.611639Z","shell.execute_reply.started":"2024-12-30T14:49:11.973629Z","shell.execute_reply":"2024-12-30T14:49:12.610749Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Keras / Ridge","metadata":{}},{"cell_type":"code","source":"meta_features = np.column_stack((cb_scores, lgbm_scores, xgb_scores))\nmeta_features.shape","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T14:49:12.612843Z","iopub.execute_input":"2024-12-30T14:49:12.613179Z","iopub.status.idle":"2024-12-30T14:49:12.629318Z","shell.execute_reply.started":"2024-12-30T14:49:12.613139Z","shell.execute_reply":"2024-12-30T14:49:12.628479Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# import tensorflow as tf\n# from keras.models import Sequential\n# from keras.layers import Dense, Dropout, BatchNormalization\n# from tensorflow.keras.callbacks import EarlyStopping\n\n# def tf_rmse(y_true, y_pred):\n#     squared_diff = tf.square(y_true - y_pred)\n#     mean_squared_error = tf.reduce_mean(squared_diff)\n#     return tf.sqrt(mean_squared_error)\n\n# meta_model = Sequential([\n#     Dense(128, activation='relu', input_dim=2),\n#     BatchNormalization(),\n#     Dropout(0.1),\n    \n#     Dense(64, activation='relu'),\n#     BatchNormalization(),\n#     Dropout(0.1),\n    \n#     Dense(32, activation='relu'),\n#     Dense(1)\n# ])\n\n# meta_model.compile(\n#     optimizer='adam', loss='mse', metrics=[tf_rmse]\n# )\n\n# early_stopping = EarlyStopping(\n#     monitor='val_loss',\n#     patience=10,\n#     restore_best_weights=True\n# )","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T14:49:12.630447Z","iopub.execute_input":"2024-12-30T14:49:12.630738Z","iopub.status.idle":"2024-12-30T14:49:12.634909Z","shell.execute_reply.started":"2024-12-30T14:49:12.630703Z","shell.execute_reply":"2024-12-30T14:49:12.634035Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import tensorflow as tf\nfrom tensorflow import keras\nfrom keras.models import Sequential\nfrom keras.layers import Dense, Dropout, BatchNormalization\nfrom tensorflow.keras.callbacks import EarlyStopping\nfrom sklearn.linear_model import Ridge\n\nmeta_model = Ridge(alpha=1.0)\n\nkf = KFold(n_splits=5, shuffle=True, random_state=42)\n\nmeta_models = []\nmeta_scores = np.zeros(len(X))\n\nfor i, (train_idx, val_idx) in enumerate(kf.split(meta_features)):\n    print(f'Fold {i + 1}')\n\n    X_meta_train, X_meta_val = meta_features[train_idx], meta_features[val_idx]\n    y_meta_train, y_meta_val = y_log.iloc[train_idx], y_log.iloc[val_idx]\n\n    meta_model.fit(\n        X_meta_train, y_meta_train\n    )\n\n    # with tf.device('/GPU:0'):\n    #     meta_model.fit(\n    #         X_meta_train,\n    #         y_meta_train,\n    #         validation_data=(X_meta_val, y_meta_val),\n    #         epochs=1, \n    #         batch_size=32,\n    #         callbacks=[early_stopping]\n    #     )\n\n    meta_models.append(meta_model)\n    \n    y_meta_pred = meta_model.predict(X_meta_val).flatten()\n    meta_scores[val_idx] = np.maximum(0, y_meta_pred)\n\n    score = rmsle(np.expm1(y_meta_val), np.expm1(meta_scores[val_idx]))\n\n    print(f'=== Fold {i + 1} RSMLE Score {score} ===')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T14:49:12.636122Z","iopub.execute_input":"2024-12-30T14:49:12.63659Z","iopub.status.idle":"2024-12-30T14:49:24.380925Z","shell.execute_reply.started":"2024-12-30T14:49:12.636552Z","shell.execute_reply":"2024-12-30T14:49:24.379451Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(f'Mean score accross all folds: {rmsle(y, np.expm1(meta_scores))}')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T14:49:24.38209Z","iopub.execute_input":"2024-12-30T14:49:24.383184Z","iopub.status.idle":"2024-12-30T14:49:24.414886Z","shell.execute_reply.started":"2024-12-30T14:49:24.383114Z","shell.execute_reply":"2024-12-30T14:49:24.413479Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🔥 Predict","metadata":{}},{"cell_type":"code","source":"cb_preds = np.zeros(len(test_df))\nfor cb_model in cb_models:\n    cb_preds += np.maximum(0, cb_model.predict(test_df)) / len(cb_models)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T14:49:24.41796Z","iopub.execute_input":"2024-12-30T14:49:24.418493Z","iopub.status.idle":"2024-12-30T14:49:31.664193Z","shell.execute_reply.started":"2024-12-30T14:49:24.418431Z","shell.execute_reply":"2024-12-30T14:49:31.663431Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"lgbm_preds = np.zeros(len(test_df))\nfor lgbm_model in lgbm_models:\n    lgbm_preds += np.maximum(0, lgbm_model.predict(test_df)) / len(lgbm_models)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T14:49:31.665414Z","iopub.execute_input":"2024-12-30T14:49:31.665704Z","iopub.status.idle":"2024-12-30T14:55:49.126156Z","shell.execute_reply.started":"2024-12-30T14:49:31.665677Z","shell.execute_reply":"2024-12-30T14:55:49.125436Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"xgb_preds = np.zeros(len(test_df))\nfor xgb_model in xgb_models:\n    xgb_preds += np.maximum(0, xgb_model.predict(test_df)) / len(xgb_models)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T14:55:49.127234Z","iopub.execute_input":"2024-12-30T14:55:49.12761Z","iopub.status.idle":"2024-12-30T14:55:59.295826Z","shell.execute_reply.started":"2024-12-30T14:55:49.127574Z","shell.execute_reply":"2024-12-30T14:55:59.295108Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"test_pred = np.zeros(len(test_df))\ntest_pred.shape","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T14:55:59.296843Z","iopub.execute_input":"2024-12-30T14:55:59.297124Z","iopub.status.idle":"2024-12-30T14:55:59.30318Z","shell.execute_reply.started":"2024-12-30T14:55:59.297097Z","shell.execute_reply":"2024-12-30T14:55:59.302453Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"test_meta_features = np.column_stack((cb_preds, lgbm_preds, xgb_preds))\ntest_meta_features.shape","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T14:55:59.304467Z","iopub.execute_input":"2024-12-30T14:55:59.305073Z","iopub.status.idle":"2024-12-30T14:55:59.322608Z","shell.execute_reply.started":"2024-12-30T14:55:59.305033Z","shell.execute_reply":"2024-12-30T14:55:59.321789Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"for meta_model in meta_models:\n    test_pred += np.maximum(0, np.expm1(meta_model.predict(test_meta_features))) / len(meta_models)\n\ntest_pred","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T14:55:59.323684Z","iopub.execute_input":"2024-12-30T14:55:59.323905Z","iopub.status.idle":"2024-12-30T14:55:59.386015Z","shell.execute_reply.started":"2024-12-30T14:55:59.323882Z","shell.execute_reply":"2024-12-30T14:55:59.384844Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"test_pred.shape","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T14:55:59.387369Z","iopub.execute_input":"2024-12-30T14:55:59.387869Z","iopub.status.idle":"2024-12-30T14:55:59.402952Z","shell.execute_reply.started":"2024-12-30T14:55:59.387812Z","shell.execute_reply":"2024-12-30T14:55:59.400113Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🚀 Sumbission","metadata":{}},{"cell_type":"code","source":"submission_df = pd.DataFrame(test_pred).rename(columns={0:'Premium Amount'})\nsubmission_df.index = test_df.index\nsubmission_df = submission_df.reset_index()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T14:55:59.405359Z","iopub.execute_input":"2024-12-30T14:55:59.405803Z","iopub.status.idle":"2024-12-30T14:55:59.42554Z","shell.execute_reply.started":"2024-12-30T14:55:59.405751Z","shell.execute_reply":"2024-12-30T14:55:59.424488Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"submission_df.to_csv('submission.csv', index = False)\nsubmission_df.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T14:55:59.430953Z","iopub.execute_input":"2024-12-30T14:55:59.431391Z","iopub.status.idle":"2024-12-30T14:56:00.848541Z","shell.execute_reply.started":"2024-12-30T14:55:59.431344Z","shell.execute_reply":"2024-12-30T14:56:00.847611Z"}},"outputs":[],"execution_count":null}]}