{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"nvidiaTeslaT4","dataSources":[{"sourceId":84896,"databundleVersionId":10305135,"sourceType":"competition"}],"dockerImageVersionId":30822,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# 🌟 **S4E12 | Algorithm Spot Check** 🌟\n\nWelcome to the **Algorithm Spot Check** stage of our journey! 🚀 \n\nWe’ve been progressing step-by-step to better understand and analyze our dataset. Here’s a quick recap of where we’ve been so far:  \n\n### 📚 **Journey So Far**\n1. **[Starter Pack](https://www.kaggle.com/code/isaranja/s4e12-starter-pack-by-ama)**  \n   *We kicked off with an overview and setup of the dataset.*  \n\n2. **[EDA and Transformation](https://www.kaggle.com/code/isaranja/s4e12-eda-transformation-by-ama)**  \n   *Explored the data, visualized patterns, and applied transformations to prepare it for modeling.*  \n\n3. **[Feature Engineering](https://www.kaggle.com/code/isaranja/s4e12-feature-engineering-by-ama)**  \n   *Created and transformed features to capture hidden patterns in the data.*  \n\n---\n\n### 🧠 **What's Next?**\nNow, it’s time to dive into the exciting phase of **Algorithm Spot Check**! 🔍  \nWe’ll analyze which algorithm performs best with the given dataset and uncover the strongest contenders for our problem.\n\n---\n\n### 🤖 **Algorithm Spot Check**  \nThere are many AutoML packages designed to test multiple algorithms efficiently. Among them, **PyCaret** stands out as a user-friendly and versatile tool for such tasks.\n\nIn this notebook, we’ll leverage **PyCaret** to:\n- Test multiple algorithms with minimal effort.\n- Compare performance metrics to identify the best model.\n- Gain insights into the suitability of different algorithms for our dataset.\n\nLet’s dive in and discover which algorithms work best! 🌟","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"markdown","source":"---\n### 🌟 **Preparing the Environment**  \n\nBefore we dive into algorithm testing, let’s make sure our environment is ready! 🛠️  \n\nIn this step, we’ll:  \n- **Load Libraries**: Import all the necessary libraries for modeling and analysis.  \n- **Set Configurations**: Ensure the environment is configured for seamless execution.  \n- **Apply Previous Transformations**: Bring forward all the transformations and configurations applied during earlier steps.  \n- **Feature Engineering**: Load the feature engineering outputs to leverage the improved dataset.\n\nThis ensures we’re building on a solid foundation and have everything in place to proceed smoothly! 🚀  ","metadata":{}},{"cell_type":"code","source":"# initiation\n#*********** Installing mission librarries******************************************************************\n!pip install pycaret -q\n\n#***********************************************************************************************************\n# Importing libraries, Helper functions, Loading files and Transformations identified during EDA\n\nfrom IPython.core.interactiveshell import InteractiveShell #\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nfrom scipy.stats import f_oneway, pointbiserialr, pearsonr, spearmanr # statistical analysis\n\nfrom sklearn.preprocessing import LabelEncoder,OneHotEncoder, MinMaxScaler # transformation\n\nfrom tabulate import tabulate # tabulate printing\n\nimport seaborn as sns # plots\nimport matplotlib.pyplot as plt # plots\n\nimport warnings #\n\nfrom pycaret.regression import * # AutoML\n\n#********** Settings **************************************************************************************\n# settings for jupyter envioronment\n\n# This ensures that plots are rendered inline\n%matplotlib inline\n\n# This ensures that all output, including text and plots, is shown automatically\nInteractiveShell.ast_node_interactivity = \"all\"\n\n# Switching off the future warrnings\nwarnings.simplefilter(action='ignore', category=FutureWarning)\n\n#********** Loading Dataset ********************************************************************************\n# Loading dataset\ntrain_df = pd.read_csv(\"/kaggle/input/playground-series-s4e12/train.csv\",parse_dates=['Policy Start Date'])\ntest_df = pd.read_csv(\"/kaggle/input/playground-series-s4e12/test.csv\",parse_dates=['Policy Start Date'])\n\n# Merging two dataframes after adding trn and tst tag\ntrain_df['src']='trn'\ntest_df['src']='tst'\n\ndf = pd.concat([train_df, test_df], ignore_index=True)\n\n#********** Transformation *********************************************************************************\n# Replace 'inf' and '-inf' with NaN\ndf.replace([np.inf, -np.inf], np.nan, inplace=True)\n\n# Missing value imputation\nfill_values = {'Age': df['Age'].median(), \n               'Annual Income': df['Annual Income'].min(),\n               'Marital Status': 'No-Data',\n               'Number of Dependents':df['Number of Dependents'].median(),\n               'Occupation':'No-Data',\n               'Health Score':df['Health Score'].median(),\n               'Previous Claims':df['Previous Claims'].median(),\n               'Vehicle Age':df['Vehicle Age'].median(),\n               'Credit Score':860,\n               'Insurance Duration':df['Insurance Duration'].median(),\n               'Customer Feedback':'No-Data',\n               'Premium Amount':0.0\n              }\n\ndf.fillna(value=fill_values, inplace=True)\n\n# Log transformation\ndf['annual_income_log'] = np.log1p(df['Annual Income'])\ndf['premium_amount_log'] = np.log1p(df['Premium Amount'])\n\n#********** Feature Engineering *****************************************************************************\n\n# deriving year from the policy start date\ndf['policy_start_year'] = df['Policy Start Date'].dt.year\n\n# deriving the month from the policy start date \ndf['policy_start_month'] = df['Policy Start Date'].dt.month\n\n# deriving the day of policy start date\ndf['policy_start_day'] = df['Policy Start Date'].dt.day\n\n# deriving policy age\nmax_date = df['Policy Start Date'].max()\ndf['policy_age'] = (max_date - df['Policy Start Date']).dt.days\ndf['policy_age_bins'] = pd.cut(df['policy_age'], bins=[df['policy_age'].min()-1,1700,df['policy_age'].max()], labels=['new','old'])\n\n# Annual Income\ndf['annual_income_bins'] = pd.cut(df['annual_income_log'], bins=[df['annual_income_log'].min()-1,1,10.9,df['annual_income_log'].max()], labels=['no-data','mid','high'])\n\n# Credit Score\ndf['credit_score_bins'] = pd.cut(df['Credit Score'], bins=[df['Credit Score'].min()-1,380,550,805,850,df['Credit Score'].max()], labels=['S','M','L','XL','no-data'])\n\n# Annual Income per dependents\ndf['annual_income_to_dependent'] = np.log1p(df['Annual Income']/(df['Number of Dependents']+1))\n\n# Health score and Credit Score combined\ndf['health_and_credit_score'] = df['Health Score']+(df['Credit Score'])\ndf['health_and_credit_score_bins'] = pd.cut(df['health_and_credit_score'], bins=[df['health_and_credit_score'].min()-1,440,575,df['health_and_credit_score'].max()], labels=['S','M','L'])\n\n#Annual Income per Insurance Duration\ndf['annual_income_to_insurance_duration'] = np.log1p((df['Annual Income']/(df['Insurance Duration']))+1)\ndf['annual_income_to_insurance_duration_bins'] = pd.cut(df['annual_income_to_insurance_duration'], bins=[df['annual_income_to_insurance_duration'].min()-1,8.5,9.75,10.9,df['annual_income_to_insurance_duration'].max()], labels=['S','M','L','XL'])\n\n# Annual Income per Age\ndf['annual_income_to_age'] = np.log1p((df['Annual Income']/(df['Age']))+1)\ndf['annual_income_to_age_bins'] = pd.cut(df['annual_income_to_age'], bins=[df['annual_income_to_age'].min()-1,6.6,8.2,df['annual_income_to_age'].max()], labels=['S','M','L'])\n\n# Annual Income to Health Score ratio\ndf['annual_income_to_health_score'] = np.log1p((df['Annual Income']/(df['Health Score']))+1)\n\n# Credit Score to Dependents ratio\ndf['credit_score_to_dependent'] = np.log1p(df['Credit Score']/(df['Number of Dependents']+1))\ndf['credit_score_to_dependent_bins'] = pd.cut(df['credit_score_to_dependent'], bins=[df['credit_score_to_dependent'].min()-1,5,6.25,df['credit_score_to_dependent'].max()], labels=['S','M','L'])\n\n# Annual Income to Policy Age\ndf['annual_income_to_policy_age'] = np.log1p((df['Annual Income']/(df['policy_age']+1))+1)\n\n# Credit Score to Previous Claims\ndf['credit_score_to_previous_claims'] = df['Credit Score']/(df['Previous Claims']+1)\ndf['credit_score_to_previous_claims_bins'] = pd.cut(df['credit_score_to_previous_claims'], bins=[df['credit_score_to_previous_claims'].min()-1,300,440,565,df['credit_score_to_previous_claims'].max()], labels=['S','M','L','XL'])\n\n# Annual Income to Previous Claims\ndf['annual_income_to_previous_claims'] = np.log1p(df['Annual Income']/(df['Previous Claims']+1))\n\n# Dependents and previous claims\ndf['dependents_and_previous_claims'] = (df['Number of Dependents'])*(df['Previous Claims'])\n\n#********** Feature list *****************************************************************************\nfeature_list = [\n#    'id',\n    'Age',\n    'Gender',\n#    'Annual Income',\n    'Marital Status',\n    'Number of Dependents',\n    'Education Level',\n    'Occupation',\n    'Health Score',\n    'Location',\n    'Policy Type',\n    'Previous Claims',\n    'Vehicle Age',\n    'Credit Score',\n    'Insurance Duration',\n#    'Policy Start Date',\n    'Customer Feedback',\n    'Smoking Status',\n    'Exercise Frequency',\n    'Property Type',\n#    'Premium Amount',\n#    'src',\n    'annual_income_log',\n    'premium_amount_log',\n    'policy_start_year',\n    'policy_start_month',\n    'policy_start_day',\n    'policy_age',\n    'policy_age_bins',\n    'annual_income_bins',\n    'credit_score_bins',\n    'annual_income_to_dependent',\n    'health_and_credit_score',\n    'health_and_credit_score_bins',\n    'annual_income_to_insurance_duration',\n    'annual_income_to_insurance_duration_bins',\n    'annual_income_to_age',\n    'annual_income_to_age_bins',\n    'annual_income_to_health_score',\n    'credit_score_to_dependent',\n    'credit_score_to_dependent_bins',\n    'annual_income_to_policy_age',\n    'credit_score_to_previous_claims',\n    'credit_score_to_previous_claims_bins',\n    'annual_income_to_previous_claims',\n    'dependents_and_previous_claims'\n]\ncat_cols = ['Gender','Marital Status','Education Level','Occupation','Location','Policy Type','Customer Feedback','Smoking Status','Exercise Frequency','Property Type',\n            'policy_age_bins','annual_income_bins','credit_score_bins','health_and_credit_score_bins','annual_income_to_age_bins','credit_score_to_dependent_bins','credit_score_to_previous_claims_bins',\n            'annual_income_to_insurance_duration_bins']\n\n#********** Final Dataset *****************************************************************************\n# Dataset looks like                                             \nwith pd.option_context('display.max_columns', None): # setting the max rows\n    display(df.head())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-27T12:30:08.862666Z","iopub.execute_input":"2024-12-27T12:30:08.862982Z","iopub.status.idle":"2024-12-27T12:30:30.796266Z","shell.execute_reply.started":"2024-12-27T12:30:08.862959Z","shell.execute_reply":"2024-12-27T12:30:30.795503Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"---\n## **🚀 Algorithm Spot Check**\n\nIn this step, we will leverage the power of **PyCaret**, an automated machine learning library, to evaluate multiple machine learning models simultaneously. PyCaret simplifies the process of testing different algorithms, saving us time and effort.\n\n### **Why PyCaret?**\n- It automates model training and evaluation.\n- Handles preprocessing steps, including transformations for categorical columns, missing values, and scaling.\n- Provides easy-to-use tools for comparing multiple models based on performance metrics.\n\n### **Optimization for Faster Evaluation**\nTo reduce the time required for model evaluation:\n1. **Downsampling the Dataset**: We'll use a smaller subset of the dataset for faster training and evaluation while maintaining the integrity of the dataset's structure.\n2. **3-Fold Cross-Validation**: Instead of the default 10-fold CV, we'll reduce the number of folds to 3 to speed up the process.\n3. **Turbo Mode**: Enabling PyCaret's **turbo mode** ensures faster execution by focusing on a smaller subset of models during the initial comparison.\n\n### **Categorical Column Handling**\nSince PyCaret provides built-in handling for categorical columns, there is no need for us to preprocess them manually. PyCaret will automatically:\n- Encode categorical variables using appropriate encoding methods (e.g., one-hot or ordinal encoding).\n- But we have handled missing values as we did in previos steps\n\nThis allows us to focus on evaluating models without worrying about the complexities of data transformation.\n\n### **Goal**\nThe aim is to discover the best-performing algorithm for our dataset quickly and efficiently while letting PyCaret manage the preprocessing tasks for us.","metadata":{}},{"cell_type":"code","source":"# Searching best performing models\n\n# Define the target column\ntarget_column = 'premium_amount_log'\n\n# Initialize the PyCaret regression setup\nregression_setup = setup(\n    data=df.loc[df['src']=='trn',feature_list].sample(frac=0.1, random_state=42), #down sampling to reduce computational time\n    target=target_column,     # Target column\n    normalize=True,           # Normalize data for better performance\n    session_id=123,           # Set random seed for reproducibility\n    verbose=False,             # Suppress unnecessary output\n    fold=3,                   # 3-fold cross-validation\n    n_jobs=-1,\n    #use_gpu=True,\n    log_experiment=False,\n    ordinal_features={\n        'Education Level': ['High School', 'Bachelor\\'s', 'Master\\'s', 'PhD'],\n        'Policy Type': ['Basic','Comprehensive','Premium'],\n        'Customer Feedback': ['No-Data','Poor','Average','Good'],\n        'Exercise Frequency':['Rarely','Monthly','Weekly','Daily']\n    },\n    categorical_features=['Gender', 'Marital Status','Occupation','Location','Smoking Status','Property Type',\n                          'policy_age_bins','annual_income_bins','credit_score_bins','health_and_credit_score_bins','annual_income_to_age_bins','credit_score_to_dependent_bins','credit_score_to_previous_claims_bins',\n                          'annual_income_to_insurance_duration_bins']\n)\n\n# Compare all models and rank them by RMSLE\nmodels_comparison = compare_models(sort='RMSLE',exclude=['knn','rf','lar','et','dt'],turbo=True) # lightgbm was removed to avoid warrining massages.\n\n# Finalize the best model for further testing\nbest_model = finalize_model(models_comparison)","metadata":{"_kg_hide-input":true,"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### **Best performing models training with full dataset**\nLet's conclude the algorithm spotcheck with identifyng which model works better for the given dataset in highlevel","metadata":{}},{"cell_type":"code","source":"# Single model\nimport lightgbm as lgb\n# ML Training with full dataset\ntarget_column = 'premium_amount_log'\n\nregression_setup = setup(\n    data=df.loc[df['src']=='trn',feature_list],\n    target=target_column,              # Target column\n    normalize=True,                    # Normalize data for better performance\n    transformation=True,               # Apply transformation (Box-Cox, Yeo-Johnson)\n    session_id=123,                    # Set random seed for reproducibility\n    verbose=False,                     # Suppress unnecessary output\n    fold=3,                            # 3-fold cross-validation\n    n_jobs=-1,\n    log_experiment=False,\n    use_gpu=True,\n    numeric_features=['Age', 'Number of Dependents','Health Score','Previous Claims','Vehicle Age','Credit Score','Insurance Duration','annual_income_log','policy_start_year',\n                      'policy_start_month','policy_start_day','policy_age','annual_income_to_dependent','health_and_credit_score','annual_income_to_insurance_duration','annual_income_to_age',\n                      'annual_income_to_health_score','credit_score_to_dependent','annual_income_to_policy_age','credit_score_to_previous_claims','annual_income_to_previous_claims','dependents_and_previous_claims'],\n    ordinal_features={\n        'Education Level': ['High School', 'Bachelor\\'s', 'Master\\'s', 'PhD'],\n        'Policy Type': ['Basic','Comprehensive','Premium'],\n        'Customer Feedback': ['No-Data','Poor','Average','Good'],\n        'Exercise Frequency':['Rarely','Monthly','Weekly','Daily']\n    },\n    categorical_features=['Gender', 'Marital Status','Occupation','Location','Smoking Status','Property Type',\n                          'policy_age_bins','annual_income_bins','credit_score_bins','health_and_credit_score_bins','annual_income_to_age_bins','credit_score_to_dependent_bins','credit_score_to_previous_claims_bins',\n                          'annual_income_to_insurance_duration_bins']\n)\n\n# Create a specific models\n#gbr = create_model('gbr')\ncatb = create_model('catboost')\nxgbt = create_model('xgboost')\nlgbm = create_model('lightgbm',verbosity=-1)\nlgbm.set_params(verbosity=-1) # supressing warning massages\n\n# Optionally, tune the hyperparameters of the model\n#tuned_lgb = tune_model(lgbm)\n#tuned_xgb = tune_model(xgbt)\n#tuned_cat = tune_model(catb)\n\n# Blend the selected models\n#ensemble_model = blend_models([lgbm,catb,xgbt])\n\n# Stack the top 3 models\nstacked_model = stack_models([lgbm,catb,xgbt])\n\n# Finalize the model (train it on the full training dataset)\nfinal_model = finalize_model(stacked_model)","metadata":{"execution":{"iopub.status.busy":"2024-12-27T12:31:13.453253Z","iopub.execute_input":"2024-12-27T12:31:13.453603Z"},"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## **Submission file generation**\nLet's generate a submission file in this step too. The next step is to finetune a model further and error analysis.","metadata":{}},{"cell_type":"code","source":"# Make predictions on the test dataset\n#predictions = predict_model(final_model, data=df.loc[df['src']=='tst',feature_list].drop(columns=['Premium Amount']))\npredictions = predict_model(final_model, data=df[df['src']=='tst'].drop(columns=['Premium Amount','premium_amount_log']))\n\n# Submission file generation\n\nsubmission_df = predictions[['id','prediction_label']]\nsubmission_df.columns=['id','Premium Amount']\nsubmission_df['Premium Amount'] = np.expm1(submission_df['Premium Amount'])\nsubmission_df['Premium Amount'] = submission_df['Premium Amount'].round(decimals=3)\n\nsubmission_df.to_csv('submission.csv', index=False)\n\nprint(\"submission.csv file generation completed\")","metadata":{"trusted":true,"_kg_hide-input":true},"outputs":[],"execution_count":null}]}