{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":84896,"databundleVersionId":10305135,"sourceType":"competition"},{"sourceId":9178166,"sourceType":"datasetVersion","datasetId":5547076}],"dockerImageVersionId":30786,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<center style= \"font-family: Calibri; font-weight:bold; letter-spacing: 0px; color:#2e4053 ; border-radius:5px; font-size:450%; text-align:center;padding:3.0px; background: #C57804; border-bottom: 3px solid #2c3e50   ; border-top: 8px solid k; border-top: 8px solid k\" > Loan Regression with an Insurance Dataset </center>\n    \n<center style= \"font-family: Calibri; font-weight:bold; letter-spacing: 0px; color:#2e4053  ; border-radius:5px; font-size:180%; text-align:center;padding:3.0px; background: silver; border-bottom: 6px solid #2c3e50   ; border-top: 8px solid k\" > Playground Series - Season 4, Episode 12 </center>","metadata":{}},{"cell_type":"markdown","source":"<center>\n<img src=\"https://www.kaggle.com/competitions/84896/images/header\" width=\"800\"/>\n</center>","metadata":{}},{"cell_type":"markdown","source":"# <p style= \"font-family: Calibri; font-weight:bold; letter-spacing: 0px; color:#2e4053; border-radius:5px; font-size:120%; text-align:left;padding:3.0px; background: #C57804; border-bottom: 4px solid #2c3e50; border-top: 4px solid k\" > 1. Introduction </p> ","metadata":{}},{"cell_type":"markdown","source":"Can a software decode on our insurance? Yes it can and has since been doing that. How good can we build a model that predict insurance premium amount base on some information fed to it.\nThat will be the purpose of this project.","metadata":{}},{"cell_type":"markdown","source":"### Let's run the numbers\n\n<center>\n<img src=\"https://encrypted-tbn0.gstatic.com/images?q=tbn:ANd9GcSh3vUTO34GvpyOihsVD7XQkExwJrZ_ZI-cFA&s\" width=\"800\"/>\n</center>","metadata":{}},{"cell_type":"markdown","source":"# <p style= \"font-family: Calibri; font-weight:bold; letter-spacing: 0px; color:#2e4053; border-radius:5px; font-size:120%; text-align:left;padding:3.0px; background: #C57804; border-bottom: 4px solid #2c3e50; border-top: 4px solid k\" > 2. Load the Tools and Datasets </p> ","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nfrom scipy.stats import iqr\nfrom datetime import datetime\nimport warnings\nwarnings.filterwarnings('ignore')\nplt.style.use('ggplot')\n# change default colormap\nplt.rcParams['image.cmap'] = 'Dark2'\n\n# Import the various sklear tools\nfrom sklearn.base import BaseEstimator, TransformerMixin\nfrom sklearn.pipeline import make_pipeline, Pipeline\nfrom sklearn.compose import make_column_transformer\nfrom sklearn.decomposition import PCA\nfrom sklearn.cluster import KMeans\nfrom sklearn.metrics import mean_squared_log_error\nfrom sklearn.metrics import mean_squared_log_error, accuracy_score, roc_auc_score, roc_curve\n\n# from mlxtend.feature_selection import SequentialFeatureSelector as SFS\n# from sklearn.feature_selection import SequentialFeatureSelector as sk_sfs\nfrom sklearn.model_selection import (train_test_split, GridSearchCV, KFold, RepeatedKFold,\n                                     RepeatedStratifiedKFold, RandomizedSearchCV, cross_val_score,\n                                     StratifiedKFold)\nfrom sklearn.ensemble import (RandomForestRegressor, HistGradientBoostingRegressor,\n                              GradientBoostingRegressor, ExtraTreesRegressor, \n                              StackingRegressor, BaggingRegressor,VotingRegressor)\nimport xgboost as xgb\nfrom xgboost import XGBRegressor, XGBClassifier, plot_importance, cv\n\nimport tensorflow as tf\nfrom tensorflow import keras\nfrom keras import Sequential\nfrom keras import layers\n\nfrom sklearn.svm import LinearSVC\nfrom sklearn.naive_bayes import GaussianNB\nfrom lightgbm import LGBMRegressor\nfrom catboost import CatBoostRegressor, Pool\nfrom sklearn.linear_model import Ridge\nfrom sklearn.neighbors import KNeighborsRegressor\nfrom sklearn.preprocessing import (MaxAbsScaler, MinMaxScaler, Normalizer,\n                                   PowerTransformer, QuantileTransformer, LabelEncoder,\n                                   RobustScaler, StandardScaler, minmax_scale,\n                                   OneHotEncoder, FunctionTransformer)\n\nimport yellowbrick\nfrom yellowbrick.classifier import ClassificationReport, DiscriminationThreshold, confusion_matrix\nfrom yellowbrick.regressor import PredictionError\nfrom imblearn.over_sampling import RandomOverSampler\nfrom imblearn.under_sampling import RandomUnderSampler\nfrom yellowbrick.regressor import ResidualsPlot\nfrom yellowbrick.cluster import KElbowVisualizer, intercluster_distance\n\nimport optuna\nfrom optuna.samplers import TPESampler\nimport plotly.express as px\n\n# Set the color scheme \nmy_scheem = 'flare_r'\nsns.set_palette(my_scheem)\n# sns.color_palette('\"blend:#7AB,#EDA\", as_cmap=True')\n\npd.set_option('display.max_columns', 100)\n# verify the versions\nprint(f'pandas version: {pd.__version__}')\nprint(f'numpy version: {np.__version__}')\nprint(f'seaborn version: {sns.__version__}')\nprint(f'optuna version : {optuna.__version__}')\nprint(f'yellowbrick version: {yellowbrick.__version__}')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-28T02:47:28.536049Z","iopub.execute_input":"2024-12-28T02:47:28.536647Z","iopub.status.idle":"2024-12-28T02:47:28.559552Z","shell.execute_reply.started":"2024-12-28T02:47:28.536596Z","shell.execute_reply":"2024-12-28T02:47:28.558416Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Define the root_mean_squared_log_error score\ndef rmsle_scorer(y_true, y_hat):\n    rmsle = np.sqrt(mean_squared_log_error(y_true, y_hat))\n    return rmsle","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-28T02:47:28.563015Z","iopub.execute_input":"2024-12-28T02:47:28.563463Z","iopub.status.idle":"2024-12-28T02:47:28.573712Z","shell.execute_reply.started":"2024-12-28T02:47:28.563417Z","shell.execute_reply":"2024-12-28T02:47:28.572587Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Load the datasets\n\norig_00 = pd.read_csv('/kaggle/input/insurance-premium-prediction/Insurance Premium Prediction Dataset.csv', parse_dates=['Policy Start Date'])\ntrain_00 = pd.read_csv('/kaggle/input/playground-series-s4e12/train.csv', index_col='id', parse_dates=['Policy Start Date'])\ntest_00 = pd.read_csv('/kaggle/input/playground-series-s4e12/test.csv', index_col='id', parse_dates=['Policy Start Date'])\nsubmission_00 = pd.read_csv('/kaggle/input/playground-series-s4e12/sample_submission.csv')\n\n\nprint(f'Shapes of the datasets:\\nTrain: {train_00.shape}\\nTest: {test_00.shape}\\nOriginal: {orig_00.shape}')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-28T02:47:28.575065Z","iopub.execute_input":"2024-12-28T02:47:28.575501Z","iopub.status.idle":"2024-12-28T02:47:36.251664Z","shell.execute_reply.started":"2024-12-28T02:47:28.575457Z","shell.execute_reply":"2024-12-28T02:47:36.250633Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"orig_00 = orig_00[-orig_00[target].isna()]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-28T02:47:36.253318Z","iopub.execute_input":"2024-12-28T02:47:36.253763Z","iopub.status.idle":"2024-12-28T02:47:36.288075Z","shell.execute_reply.started":"2024-12-28T02:47:36.253716Z","shell.execute_reply":"2024-12-28T02:47:36.287095Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Drop the rows in the original dataset that are missing Premium Amount.","metadata":{}},{"cell_type":"code","source":"train_comb = pd.concat([train_00, orig_00], ignore_index=True)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-28T02:47:36.29156Z","iopub.execute_input":"2024-12-28T02:47:36.292149Z","iopub.status.idle":"2024-12-28T02:47:36.366271Z","shell.execute_reply.started":"2024-12-28T02:47:36.292112Z","shell.execute_reply":"2024-12-28T02:47:36.365423Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## How close are the train, test and original datasets?","metadata":{}},{"cell_type":"code","source":"# Define a function to perform the adversarial validation of two datasets\ndef adversarial_validation(df_1, df_2, name_1, name_2):\n    adv_df_1 = df_1[num_features].copy()\n    adv_df_2 = df_2[num_features].copy()\n\n\n    # label the test and train data with 0 and 1 (it doesn't really matter which is which)\n    adv_df_1 = adv_df_1.assign(adv=1)\n    adv_df_2 = adv_df_2.assign(adv=0)\n\n\n    # combine the training and test data into one big dataset\n    combined = pd.concat([adv_df_1, adv_df_2], axis=0)\n\n    # Shuffle\n    combined = combined.sample(frac=1, random_state=64)\n\n    # perform the binary classification, for example using XGboost\n    X_combined = combined.drop('adv', axis=1)\n    y_combined = combined.adv\n\n    # Define cv spliter\n    cv = StratifiedKFold(n_splits = 5,\n                        shuffle = True,\n                        random_state = 64)\n    \n    # Define the classifier\n    xgb_model = XGBClassifier(max_depth=3,\n                              learning_rate = 0.1,\n                              n_estimators = 100,\n                              objective = 'binary:logistic',\n                              random_state = 64)\n\n    # Get the cross validation scores\n    adv_scores = []\n    for i, _ in enumerate(cv.split(X_combined, y_combined)):\n        X_train, X_valid, y_train, y_valid = train_test_split(X_combined, \n                                                              y_combined, \n                                                              test_size=0.3)\n        xgb_model.fit(X_train, y_train)\n        y_pred = xgb_model.predict_proba(X_valid)[:,1]\n        score = roc_auc_score(y_valid, y_pred)\n        adv_scores.append(score)\n\n#         print(f\"Fold {i+1} AUC Score: {score:.5f}\")\n\n    #Plot the roc_curve\n    mean_auc = np.mean(adv_scores)\n    fpr, tpr, _ = roc_curve(y_valid, y_pred)\n    plt.plot(fpr, tpr, label = 'roc_curve (AUC = %0.4f)' % mean_auc)\n    plt.plot([0,1], [0,1], linestyle = '--', color = 'gray', label = 'Random Guess')\n    plt.xlabel('False Positive Rate')\n    plt.ylabel('True Positive Rate')\n    plt.title(f'roc_curve {name_1} vs {name_2}', weight='bold')\n    plt.legend()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-28T02:47:36.367466Z","iopub.execute_input":"2024-12-28T02:47:36.367754Z","iopub.status.idle":"2024-12-28T02:47:36.378409Z","shell.execute_reply.started":"2024-12-28T02:47:36.367725Z","shell.execute_reply":"2024-12-28T02:47:36.377297Z"},"_kg_hide-output":true,"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"num_features = list(test_00.select_dtypes('number'))\nplt.figure(figsize=(18,5))\nplt.subplot(1,3,1)\nadversarial_validation(train_00, test_00, 'train', 'test')\nplt.subplot(1,3,2)\nadversarial_validation(test_00, orig_00, 'train', 'original')\nplt.subplot(1,3,3)\nadversarial_validation(train_comb, test_00, 'train_comb', 'test')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-28T02:47:36.380041Z","iopub.execute_input":"2024-12-28T02:47:36.38049Z","iopub.status.idle":"2024-12-28T02:50:25.253137Z","shell.execute_reply.started":"2024-12-28T02:47:36.380443Z","shell.execute_reply":"2024-12-28T02:50:25.252162Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"As said the feature distributions in train and test datasets are close to, but not exactly the same, as the original.","metadata":{}},{"cell_type":"markdown","source":"# <p style= \"font-family: Calibri; font-weight:bold; letter-spacing: 0px; color:#2e4053; border-radius:5px; font-size:120%; text-align:left;padding:3.0px; background: #C57804; border-bottom: 4px solid #2c3e50; border-top: 4px solid k\" > 3. Preview and Examine the Datasets </p> ","metadata":{}},{"cell_type":"markdown","source":"### What does the data say?\n\n<center>\n<img src=\"https://beyondtheory.co.uk/storage/images/other/2016/08/Beyond-Theory-Data-Analysis-Landing-Page-graphic.png\" width=\"700\"/>\n</center>","metadata":{}},{"cell_type":"markdown","source":"## Dataset Preview","metadata":{}},{"cell_type":"code","source":"# Preview the train dataset\ntrain_00.head(3)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-28T02:50:25.254533Z","iopub.execute_input":"2024-12-28T02:50:25.254863Z","iopub.status.idle":"2024-12-28T02:50:25.276374Z","shell.execute_reply.started":"2024-12-28T02:50:25.254827Z","shell.execute_reply":"2024-12-28T02:50:25.275313Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Dataset Info","metadata":{}},{"cell_type":"code","source":"train_00.info()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-28T02:50:25.278133Z","iopub.execute_input":"2024-12-28T02:50:25.278619Z","iopub.status.idle":"2024-12-28T02:50:25.838539Z","shell.execute_reply.started":"2024-12-28T02:50:25.278575Z","shell.execute_reply":"2024-12-28T02:50:25.837424Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Define the target\ntarget = 'Premium Amount'","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-28T02:50:25.839943Z","iopub.execute_input":"2024-12-28T02:50:25.840397Z","iopub.status.idle":"2024-12-28T02:50:25.845417Z","shell.execute_reply.started":"2024-12-28T02:50:25.840331Z","shell.execute_reply":"2024-12-28T02:50:25.844223Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Check if there are duplicates in the datasets\nfor df_name, df in [('Train', train_00), ('Test', test_00)]:\n    nunb_of_duplicates = df.duplicated().sum()\n    if nunb_of_duplicates != 0:\n        print(f'{df_name} dataset has {nunb_of_duplicates} duplicates.')\n    else:\n        print(f'{df_name} dataset has no duplicates')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-28T02:50:25.846875Z","iopub.execute_input":"2024-12-28T02:50:25.847539Z","iopub.status.idle":"2024-12-28T02:50:27.940228Z","shell.execute_reply.started":"2024-12-28T02:50:25.847492Z","shell.execute_reply":"2024-12-28T02:50:27.939147Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Count and Fill Missing Values","metadata":{}},{"cell_type":"code","source":"# Count the missing values in the datasets\nnull_count = pd.DataFrame({'NaN in train': train_00.isna().sum(), \n                           'NaN in test': test_00.isna().sum(), \n                           # 'NaN in orig': orig_00.isna().sum(),\n                           '% NaN in train': train_00.isna().mean()*100, \n                           '% NaN in test': test_00.isna().mean()*100, \n                           # '% NaN in orig': orig_00.isna().mean()*100\n                          }\n                         ).drop(index=[target]).astype('int')\n# pickup only the features with missing values\nnull_count.sort_values(by='NaN in train', ascending=False).head(11).style.background_gradient(cmap='Reds')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-28T02:50:27.941589Z","iopub.execute_input":"2024-12-28T02:50:27.94188Z","iopub.status.idle":"2024-12-28T02:50:29.755298Z","shell.execute_reply.started":"2024-12-28T02:50:27.941852Z","shell.execute_reply":"2024-12-28T02:50:29.754307Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"There are lot of missing values in the train and test datasets. Some of the features, ­­`­Previous Claims` and `Occupation` are missing close to 30% of their data in both sets.","metadata":{}},{"cell_type":"code","source":"# Get the year, month and age of the account\ndef get_the_month_year(df):\n    # df['month'] = pd.to_datetime(df['Policy Start Date']).dt.month\n    # df['year'] = pd.to_datetime(df['Policy Start Date']).dt.year.astype('string')\n    # current_date = datetime.now()\n    # df['account_age'] = df['Policy Start Date'].apply(lambda x: current_date.year - x.year - ((current_date.month, current_date.day) < (x.month, x.day)))\n    df = df.drop(columns = ['Policy Start Date'])\n    return df\n\n\ntrain_01 = get_the_month_year(train_00)\ntrain_comb_01 = get_the_month_year(train_comb)\ntest_01 = get_the_month_year(test_00)\norig_01 = get_the_month_year(orig_00)\ntrain_01.head(3)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-28T02:50:29.756834Z","iopub.execute_input":"2024-12-28T02:50:29.757254Z","iopub.status.idle":"2024-12-28T02:50:30.264464Z","shell.execute_reply.started":"2024-12-28T02:50:29.757209Z","shell.execute_reply":"2024-12-28T02:50:30.263305Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# View numeric features\ntest_01.select_dtypes('number').head(3)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-28T02:50:30.270277Z","iopub.execute_input":"2024-12-28T02:50:30.2707Z","iopub.status.idle":"2024-12-28T02:50:30.29641Z","shell.execute_reply.started":"2024-12-28T02:50:30.270667Z","shell.execute_reply":"2024-12-28T02:50:30.295387Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# What are the numerical features?\nnum_cols = test_01.select_dtypes('number').columns.tolist()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-28T02:50:30.299595Z","iopub.execute_input":"2024-12-28T02:50:30.299918Z","iopub.status.idle":"2024-12-28T02:50:30.32646Z","shell.execute_reply.started":"2024-12-28T02:50:30.299887Z","shell.execute_reply":"2024-12-28T02:50:30.325088Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# View categorical features\ntest_01.select_dtypes(exclude='number').head(3)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-28T02:50:30.327718Z","iopub.execute_input":"2024-12-28T02:50:30.32801Z","iopub.status.idle":"2024-12-28T02:50:30.424087Z","shell.execute_reply.started":"2024-12-28T02:50:30.327982Z","shell.execute_reply":"2024-12-28T02:50:30.423096Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# What are the categorical features?\ncat_cols = test_01.select_dtypes(exclude='number').columns.tolist()\ncat_cols ","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-28T02:50:30.425207Z","iopub.execute_input":"2024-12-28T02:50:30.425535Z","iopub.status.idle":"2024-12-28T02:50:30.582348Z","shell.execute_reply.started":"2024-12-28T02:50:30.425506Z","shell.execute_reply":"2024-12-28T02:50:30.581421Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_01.isna().sum()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-28T02:50:30.583906Z","iopub.execute_input":"2024-12-28T02:50:30.584781Z","iopub.status.idle":"2024-12-28T02:50:31.132976Z","shell.execute_reply.started":"2024-12-28T02:50:30.584733Z","shell.execute_reply":"2024-12-28T02:50:31.132005Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"for feat in cat_cols:\n    try:\n        print('\\n{} has {} unique values and has {} missing values which accounts for {:.2f} % of the data.'.format(feat, train_01[feat].nunique(), train_01[feat].isna().sum(), train_01[feat].isna().mean()*100))\n    except:\n        pass","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-28T02:50:31.134045Z","iopub.execute_input":"2024-12-28T02:50:31.134338Z","iopub.status.idle":"2024-12-28T02:50:32.756041Z","shell.execute_reply.started":"2024-12-28T02:50:31.13431Z","shell.execute_reply":"2024-12-28T02:50:32.75497Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# def missing_handler(df):\n#     for feat in cat_cols:\n#         try:\n#             df[feat] = df[feat].fillna(df[feat].mode())\n#         except:\n#             pass\n#     for feat in num_cols:\n#         try:\n#             df[feat] = df[feat].fillna(df[feat].median())\n#         except:\n#             pass\n#     return df","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-28T02:50:32.757715Z","iopub.execute_input":"2024-12-28T02:50:32.758016Z","iopub.status.idle":"2024-12-28T02:50:32.761846Z","shell.execute_reply.started":"2024-12-28T02:50:32.757986Z","shell.execute_reply":"2024-12-28T02:50:32.760917Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def missing_handler(df):\n    for feat in cat_cols:\n        df[feat] = df[feat].fillna('unknown')\n    for feat in num_cols:\n        # df[feat] = df[feat].fillna(df[feat].median())\n        df[feat] = df[feat].fillna(df[feat].min())\n    return df","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-28T02:50:32.762971Z","iopub.execute_input":"2024-12-28T02:50:32.763337Z","iopub.status.idle":"2024-12-28T02:50:32.776272Z","shell.execute_reply.started":"2024-12-28T02:50:32.763292Z","shell.execute_reply":"2024-12-28T02:50:32.77529Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_02 = missing_handler(train_01)\ntrain_comb_02 = missing_handler(train_comb)\ntest_02 = missing_handler(test_01)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-28T02:50:32.77812Z","iopub.execute_input":"2024-12-28T02:50:32.778573Z","iopub.status.idle":"2024-12-28T02:50:35.215714Z","shell.execute_reply.started":"2024-12-28T02:50:32.778534Z","shell.execute_reply":"2024-12-28T02:50:35.214572Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_02.isna().sum()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-28T02:50:35.217067Z","iopub.execute_input":"2024-12-28T02:50:35.217423Z","iopub.status.idle":"2024-12-28T02:50:35.771489Z","shell.execute_reply.started":"2024-12-28T02:50:35.217384Z","shell.execute_reply":"2024-12-28T02:50:35.770543Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <p style= \"font-family: Calibri; font-weight:bold; letter-spacing: 0px; color:#2e4053; border-radius:5px; font-size:120%; text-align:left;padding:3.0px; background: #C57804; border-bottom: 4px solid #2c3e50; border-top: 4px solid k\" > 3. Examine the Datasets </p> ","metadata":{}},{"cell_type":"code","source":"train_02[target].plot.hist(bins=20)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-28T02:50:35.772755Z","iopub.execute_input":"2024-12-28T02:50:35.773162Z","iopub.status.idle":"2024-12-28T02:50:36.129917Z","shell.execute_reply.started":"2024-12-28T02:50:35.773117Z","shell.execute_reply":"2024-12-28T02:50:36.128798Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plt.figure(figsize=(10,4))\nfor f, feat in enumerate(num_cols, start=1):\n    plt.subplot(2,5,f)\n    sns.boxenplot(train_02, x=feat)\nplt.tight_layout()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-28T02:50:36.131054Z","iopub.execute_input":"2024-12-28T02:50:36.131344Z","iopub.status.idle":"2024-12-28T02:50:39.951652Z","shell.execute_reply.started":"2024-12-28T02:50:36.131316Z","shell.execute_reply":"2024-12-28T02:50:39.95065Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plt.figure(figsize=(10,4))\nfor f, feat in enumerate(num_cols, start=1):\n    plt.subplot(2,5,f)\n    sns.violinplot(train_02, x=feat)\nplt.tight_layout()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-28T02:50:39.953175Z","iopub.execute_input":"2024-12-28T02:50:39.953619Z","iopub.status.idle":"2024-12-28T02:50:58.611303Z","shell.execute_reply.started":"2024-12-28T02:50:39.953573Z","shell.execute_reply":"2024-12-28T02:50:58.610299Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# for feat in ['Number of Dependents', 'Previous Claims', 'Vehicle Age', 'Insurance Duration']:\n#     display(train_02[feat].value_counts().to_frame().style.background_gradient())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-28T02:50:58.612744Z","iopub.execute_input":"2024-12-28T02:50:58.613204Z","iopub.status.idle":"2024-12-28T02:50:58.618003Z","shell.execute_reply.started":"2024-12-28T02:50:58.613121Z","shell.execute_reply":"2024-12-28T02:50:58.617044Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plt.figure(figsize=(16, 12))\nfor f, feat in enumerate(['Number of Dependents', 'Previous Claims', 'Vehicle Age', 'Insurance Duration'], start=1):\n    plt.subplot(1, 4, f)\n    pd.Series({' ': 1}).plot.pie(colors=['grey'], radius=0.2, shadow=False)\n    train_02[feat].value_counts().plot.pie(autopct='%.1f%%', radius=1.2, pctdistance=0.44, shadow=False, \n                                         textprops={'color':'black', 'rotation':True, 'weight':'bold', 'size': 8},\n                                         startangle=90 , rotatelabels=True,\n                                         labeldistance=0.7, wedgeprops={'width':0.9}, frame=True)\n    plt.ylabel('')\n    plt.title(feat, color='black', fontsize=10, weight='bold')\n    plt.tight_layout()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-28T02:50:58.619343Z","iopub.execute_input":"2024-12-28T02:50:58.61975Z","iopub.status.idle":"2024-12-28T02:50:59.908386Z","shell.execute_reply.started":"2024-12-28T02:50:58.619706Z","shell.execute_reply":"2024-12-28T02:50:59.907303Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# for f, feat in enumerate(cat_cols, start=1):\n#     if train_00[feat].nunique() < 20:\n#         print(f'\\n==== {feat} ====')\n#         feat_count = train_00[feat].value_counts().to_frame().style.background_gradient()\n#         display(feat_count)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-28T02:50:59.909949Z","iopub.execute_input":"2024-12-28T02:50:59.910387Z","iopub.status.idle":"2024-12-28T02:50:59.914822Z","shell.execute_reply.started":"2024-12-28T02:50:59.910322Z","shell.execute_reply":"2024-12-28T02:50:59.913904Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Generate a mask for the upper triangle \nmask = np.triu(np.ones_like(train_02[num_cols].corr(), dtype=bool))\n# Generate a custom diverging colormap \ncmap = my_scheem[:-2]\n\nsns.heatmap(train_02[num_cols].corr(), annot=True, cmap=cmap, mask=mask, fmt='.3f')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-28T02:50:59.916095Z","iopub.execute_input":"2024-12-28T02:50:59.916516Z","iopub.status.idle":"2024-12-28T02:51:00.808278Z","shell.execute_reply.started":"2024-12-28T02:50:59.916472Z","shell.execute_reply":"2024-12-28T02:51:00.807472Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"We cannot identify any pair of features with high correlation.","metadata":{}},{"cell_type":"markdown","source":"# <p style= \"font-family: Calibri; font-weight:bold; letter-spacing: 0px; color:#2e4053; border-radius:5px; font-size:120%; text-align:left;padding:3.0px; background: #C57804; border-bottom: 4px solid #2c3e50; border-top: 4px solid k\" > 4. Preprocessor </p> ","metadata":{}},{"cell_type":"code","source":"# Define X and y\nuse_comb=False\n\nif use_comb: # include orig in the train set\n    X = train_comb_02.copy()\nelse: # do not include orig in the train set\n    X = train_02.copy()\n# get the target \ny = X.pop(target)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-28T02:51:00.809439Z","iopub.execute_input":"2024-12-28T02:51:00.809733Z","iopub.status.idle":"2024-12-28T02:51:01.290075Z","shell.execute_reply.started":"2024-12-28T02:51:00.809704Z","shell.execute_reply":"2024-12-28T02:51:01.289176Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# What are the categorical features?\ncat_cols = X.select_dtypes(exclude='number').columns.tolist()\ncat_cols ","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-28T02:51:01.291133Z","iopub.execute_input":"2024-12-28T02:51:01.291468Z","iopub.status.idle":"2024-12-28T02:51:01.741422Z","shell.execute_reply.started":"2024-12-28T02:51:01.291436Z","shell.execute_reply":"2024-12-28T02:51:01.740421Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"X.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-28T02:51:01.742736Z","iopub.execute_input":"2024-12-28T02:51:01.743006Z","iopub.status.idle":"2024-12-28T02:51:01.764293Z","shell.execute_reply.started":"2024-12-28T02:51:01.742978Z","shell.execute_reply":"2024-12-28T02:51:01.763281Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"X.shape","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-28T02:51:01.765661Z","iopub.execute_input":"2024-12-28T02:51:01.765958Z","iopub.status.idle":"2024-12-28T02:51:01.778557Z","shell.execute_reply.started":"2024-12-28T02:51:01.765927Z","shell.execute_reply":"2024-12-28T02:51:01.77759Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# # feat_to_scale = X_ts.select_dtypes(include='number').columns.tolist()\n# features_trans = make_column_transformer(\n#     (MinMaxScaler(), num_cols),\n#     (OneHotEncoder(), cat_cols),\n#     remainder='passthrough', \n#     sparse_threshold=0)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-28T02:51:01.779829Z","iopub.execute_input":"2024-12-28T02:51:01.780133Z","iopub.status.idle":"2024-12-28T02:51:01.788799Z","shell.execute_reply.started":"2024-12-28T02:51:01.780102Z","shell.execute_reply":"2024-12-28T02:51:01.787778Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# scaler = MinMaxScaler()\nimport category_encoders as ce\n\nscaler = PowerTransformer(method='yeo-johnson', standardize=True, copy=True)\n\nfeatures_trans = make_column_transformer(\n    (ce.CatBoostEncoder(), cat_cols),\n    # (LabelEncoder(), cat_cols),\n    (scaler, num_cols),\n    remainder='passthrough'\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-28T02:51:01.790063Z","iopub.execute_input":"2024-12-28T02:51:01.790501Z","iopub.status.idle":"2024-12-28T02:51:01.801176Z","shell.execute_reply.started":"2024-12-28T02:51:01.79047Z","shell.execute_reply":"2024-12-28T02:51:01.800253Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"X_prep = X.copy()\n\nX_prep = features_trans.fit_transform(X_prep, y)\n\nX_prep","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-28T02:51:01.802575Z","iopub.execute_input":"2024-12-28T02:51:01.802874Z","iopub.status.idle":"2024-12-28T02:51:19.133058Z","shell.execute_reply.started":"2024-12-28T02:51:01.802844Z","shell.execute_reply":"2024-12-28T02:51:19.131965Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# X_p = X.copy()\n\n# X_p = features_trans.fit_transform(X_p)\n\n# pd.DataFrame(X_p).head(3)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-28T02:51:19.134311Z","iopub.execute_input":"2024-12-28T02:51:19.134669Z","iopub.status.idle":"2024-12-28T02:51:19.140478Z","shell.execute_reply.started":"2024-12-28T02:51:19.134638Z","shell.execute_reply":"2024-12-28T02:51:19.139566Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"model = KMeans(n_init='auto')\n\nX_num = X[num_cols]\n\nviz = KElbowVisualizer(model, k=(4,16))\nviz.fit(X_num)\nviz.show()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-28T02:51:19.141585Z","iopub.execute_input":"2024-12-28T02:51:19.141864Z","iopub.status.idle":"2024-12-28T02:51:45.767677Z","shell.execute_reply.started":"2024-12-28T02:51:19.141836Z","shell.execute_reply":"2024-12-28T02:51:45.766636Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"intercluster_distance(KMeans(6, random_state=42), X_num)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-28T02:51:45.76881Z","iopub.execute_input":"2024-12-28T02:51:45.769106Z","iopub.status.idle":"2024-12-28T02:51:56.636057Z","shell.execute_reply.started":"2024-12-28T02:51:45.769075Z","shell.execute_reply":"2024-12-28T02:51:56.635096Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"X_tr, X_ts, y_tr, y_ts = train_test_split(X, np.log1p(y), test_size=0.25, random_state=42)\n\n[d.shape for d in [X_tr, X_ts, y_tr, y_ts]]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-28T02:51:56.643703Z","iopub.execute_input":"2024-12-28T02:51:56.644306Z","iopub.status.idle":"2024-12-28T02:51:57.332034Z","shell.execute_reply.started":"2024-12-28T02:51:56.64427Z","shell.execute_reply":"2024-12-28T02:51:57.331051Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"pd.DataFrame(X_tr).head(5)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-28T02:51:57.333141Z","iopub.execute_input":"2024-12-28T02:51:57.333469Z","iopub.status.idle":"2024-12-28T02:51:57.356331Z","shell.execute_reply.started":"2024-12-28T02:51:57.333437Z","shell.execute_reply":"2024-12-28T02:51:57.355259Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <p style= \"font-family: Calibri; font-weight:bold; letter-spacing: 0px; color:#2e4053; border-radius:5px; font-size:120%; text-align:left;padding:3.0px; background: #C57804; border-bottom: 4px solid #2c3e50; border-top: 4px solid k\" > 5. Modeling </p> ","metadata":{}},{"cell_type":"markdown","source":"<center>\n<img src=\"https://media.geeksforgeeks.org/wp-content/uploads/20240215112547/Data-Modeling-in-Analysis.webp\" width=\"800\"/>\n</center>","metadata":{}},{"cell_type":"code","source":"import lightgbm as lgb\n\ndef objective(trial): \n    # Define LightGBM parameters\n    params = {\n        \"objective\": 'regression',\n        # \"metric\": \"l2\",\n        \"lambda_l1\": trial.suggest_float(\"lambda_l1\", 1e-8, 10.0, log=True),\n        \"lambda_l2\": trial.suggest_float(\"lambda_l2\", 1e-8, 10.0, log=True),\n        \"learning_rate\": trial.suggest_float(\"learning_rate\", 1e-3, 0.5, log=True),\n        \"num_leaves\": trial.suggest_int(\"num_leaves\", 2, 256),\n        \"max_depth\": trial.suggest_int(\"max_depth\", 2, 10),\n        \"feature_fraction\": trial.suggest_float(\"feature_fraction\", 0.4, 1.0),\n        \"bagging_fraction\": trial.suggest_float(\"bagging_fraction\", 0.4, 1.0),\n        \"bagging_freq\": trial.suggest_int(\"bagging_freq\", 1, 7),\n        \"min_child_samples\": trial.suggest_int(\"min_child_samples\", 5, 100),\n    }\n    \n    # Create LightGBM dataset\n    model_pipe = make_pipeline(features_trans,\n                               LGBMRegressor(**params, verbose=-1))\n    \n    # Train the model\n    model_pipe.fit(X_tr, y_tr)\n    \n    # Make predictions\n    preds = model_pipe.predict(X_ts)\n    preds = np.expm1(preds)\n\n    # Calculate the mean squared error\n    rmsle = rmsle_scorer(np.expm1(y_ts), preds)\n    return rmsle\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-28T02:51:57.357543Z","iopub.execute_input":"2024-12-28T02:51:57.357889Z","iopub.status.idle":"2024-12-28T02:51:57.368859Z","shell.execute_reply.started":"2024-12-28T02:51:57.357857Z","shell.execute_reply":"2024-12-28T02:51:57.36783Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"study = optuna.create_study(direction=\"minimize\")\nstudy.optimize(objective, n_trials=60)\n\nprint(\"Number of finished trials: {}\".format(len(study.trials)))\nprint(\"Best trial:\")\ntrial = study.best_trial\nprint(\"  Value: {}\".format(trial.value))\nprint(\"  Params: \")\nfor key, value in trial.params.items():\n    print(\"    {}: {}\".format(key, value))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-28T02:51:57.370289Z","iopub.execute_input":"2024-12-28T02:51:57.370627Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# best_lgbm_params = {'lambda_l1': 0.007787038174438944, \n#                     'lambda_l2': 0.007475379286754892, \n#                     'learning_rate': 0.06469248141029064, \n#                     'num_leaves': 101, \n#                     'max_depth': 9, \n#                     'feature_fraction': 0.9838734620932038, \n#                     'bagging_fraction': 0.871730858029871, \n#                     'bagging_freq': 3, \n#                     'min_child_samples': 50}","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"model = LGBMRegressor(**best_lgbm_params, verbose=-1)\n# model = HistGradientBoostingRegressor()\n# model = CatBoostRegressor(n_estimators=400, eval_fraction=0.2)\n\n\nmodel_pipe = make_pipeline(features_trans, model)\n\nmodel_pipe.fit(X_tr, y_tr)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"rmsle_scorer(np.expm1(y_ts), np.expm1(model_pipe.predict(X_ts)))","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The residual is the difference between the predicted and actual values. Ideal there shouldn't be any residual i.e residual = 0. Hence on the graph of residuals we want our points to be closer to the `zero line`.","metadata":{}},{"cell_type":"code","source":"# Instantiate the linear model and visualizer\nmodel = make_pipeline(features_trans, model)\n\nvisualizer = ResidualsPlot(model)\n\nvisualizer.fit(X_tr, y_tr)  # Fit the training data to the visualizer\nvisualizer.score(X_ts, y_ts)  # Evaluate the model on the test data\nvisualizer.poof()                 # Finalize and render the figure\nprint('not the best')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"visualizer = ResidualsPlot(model, hist=False, qqplot=True)\n\nvisualizer.fit(X_tr, y_tr)  # Fit the training data to the visualizer\nvisualizer.score(X_ts, y_ts)  # Evaluate the model on the test data\nvisualizer.poof()                 # Finalize and render the figure\nprint('not the best')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"visualizer = ResidualsPlot(model, hist=False, qqplot=True)\n\nvisualizer.fit(X_tr, y_tr)  # Fit the training data to the visualizer\nvisualizer.score(X_ts, y_ts)  # Evaluate the model on the test data\nvisualizer.poof()                 # Finalize and render the figure\nprint('not the best')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The Residual chart show that the model is not performing well:\n * The residuals range from about `-5000 to 2000`. which is a very wide range for values that are in the range `0, 5000`.\n * The R2 score are too low and suggest a random model.","metadata":{}},{"cell_type":"code","source":"pred_ts = np.expm1(model_pipe.predict(X_ts))","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"ax = pd.DataFrame(pred_ts).plot.hist(bins=20)\nnp.expm1(y_ts).plot.hist(bins=20, ax=ax, alpha=0.5)\nplt.title('Compare the distribution of true vs predicted target in validation set')\nplt.show()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"As opposed to our expectation, the predicted values fall in a norrower range compared to that of the train data.","metadata":{}},{"cell_type":"code","source":"viz = PredictionError(model, line_colors='green')\nviz.fit(X_tr, y_tr)\nviz.score(X_ts, y_ts)\nviz.show()\nplt.show()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"It looks more like a random prediction.","metadata":{}},{"cell_type":"markdown","source":"# <p style= \"font-family: Calibri; font-weight:bold; letter-spacing: 0px; color:#2e4053; border-radius:5px; font-size:120%; text-align:left;padding:3.0px; background: #C57804; border-bottom: 4px solid #2c3e50; border-top: 4px solid k\" > 6. Prediction and submission of test </p> ","metadata":{}},{"cell_type":"code","source":"# Prediction on test set\npreds = model_pipe.predict(test_02)\n\nsubmission_00[target] = np.expm1(preds)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"submission_00","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"submission_00.to_csv('submission.csv', index=False)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### <p style= \"font-family: Calibri; font-weight:bold; letter-spacing: 0px; color:#2e4053; border-radius:5px; font-size:120%; text-align:left;padding:3.0px; background: #C57804; border-bottom: 4px solid #2c3e50; border-top: 4px solid k\" > We don't expect this atempt to give a good result. More is still to come... </p> ","metadata":{}}]}