{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":84896,"databundleVersionId":10305135,"sourceType":"competition"}],"dockerImageVersionId":30804,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# introduction","metadata":{}},{"cell_type":"markdown","source":"## Declaration and Focus of This Notebook 📘\n\n**Declaration:**  \nI would like to state upfront that this notebook will not delve into Exploratory Data Analysis (EDA). If you are interested in understanding how to perform EDA, I recommend referring to a more detailed guide available at [this notebook](https://www.kaggle.com/code/zyh1104/insurance-analyze), which provides a comprehensive framework. Please feel free to explore it for a deeper understanding.\n\n---\n\n**Objective:**  \nThe primary focus of our efforts here will be on refining the intricacies of our modeling process. This includes aspects such as:\n\n1. **Data Preprocessing:** Preparing the data by cleaning, normalizing, and transforming it to improve the performance of our models.\n2. **Feature Engineering:** Creating new features or modifying existing ones to better represent the underlying problem to the model, which can lead to more accurate predictions.\n3. **Hyperparameter Optimization:** Fine-tuning the parameters of our models to achieve the best performance.\n\nWe aim to discuss and brainstorm on how we can further enhance the model's score within a general framework.\n\n---\n\n**Collaboration Invitation:**  \nIf you have any excellent suggestions or innovative ideas on these topics, please don’t hesitate to share them in the comments section below. Your input could be invaluable in our collective pursuit of improving model performance.\n\n---","metadata":{}},{"cell_type":"markdown","source":"# Import Libraries 🎇 ","metadata":{}},{"cell_type":"code","source":"import numpy as np \nimport pandas as pd \nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport plotly.graph_objs as go\nfrom sklearn.preprocessing import OneHotEncoder\nfrom sklearn.preprocessing import StandardScaler\nfrom sklearn.model_selection import train_test_split\nimport optuna\nfrom sklearn.model_selection import KFold\nimport lightgbm as lgb\nfrom sklearn.metrics import mean_squared_log_error\nfrom optuna.visualization import plot_optimization_history, plot_param_importances\nfrom sklearn.metrics import r2_score, mean_absolute_error, mean_squared_error\nimport sys\nimport os\nimport logging\nimport gc\n\nimport warnings\nwarnings.filterwarnings('ignore')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-17T05:41:37.972366Z","iopub.execute_input":"2024-12-17T05:41:37.972745Z","iopub.status.idle":"2024-12-17T05:41:38.52073Z","shell.execute_reply.started":"2024-12-17T05:41:37.972713Z","shell.execute_reply":"2024-12-17T05:41:38.520058Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Data Preprocessing 🎈","metadata":{}},{"cell_type":"markdown","source":"## load data","metadata":{}},{"cell_type":"code","source":"train = pd.read_csv(\"/kaggle/input/playground-series-s4e12/train.csv\")\ntest = pd.read_csv(\"/kaggle/input/playground-series-s4e12/test.csv\")\n\ngc.collect()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-17T05:41:38.52244Z","iopub.execute_input":"2024-12-17T05:41:38.522954Z","iopub.status.idle":"2024-12-17T05:41:49.44845Z","shell.execute_reply.started":"2024-12-17T05:41:38.522926Z","shell.execute_reply":"2024-12-17T05:41:49.447486Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-17T05:41:49.44936Z","iopub.execute_input":"2024-12-17T05:41:49.449602Z","iopub.status.idle":"2024-12-17T05:41:49.481959Z","shell.execute_reply.started":"2024-12-17T05:41:49.449578Z","shell.execute_reply":"2024-12-17T05:41:49.481031Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train.shape","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-17T05:41:49.483217Z","iopub.execute_input":"2024-12-17T05:41:49.4836Z","iopub.status.idle":"2024-12-17T05:41:49.491097Z","shell.execute_reply.started":"2024-12-17T05:41:49.483563Z","shell.execute_reply":"2024-12-17T05:41:49.490339Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"test.shape","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-17T05:41:49.493111Z","iopub.execute_input":"2024-12-17T05:41:49.493441Z","iopub.status.idle":"2024-12-17T05:41:49.503658Z","shell.execute_reply.started":"2024-12-17T05:41:49.493415Z","shell.execute_reply":"2024-12-17T05:41:49.502834Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Missing value processing ","metadata":{}},{"cell_type":"markdown","source":"## Handling Missing Values Strategy \n\n### Deletion of Highly Missing Features\nIn addressing missing value processing within our dataset, we've adopted a strategic approach. Specifically, any feature with missing values exceeding 50% will be removed from the dataset. This decision is made to maintain data integrity and to prevent overfitting due to potential bias in imputed values.\n\n### Median Imputation for Important Features\nFor features deemed important, we will employ median imputation to fill in the missing values. The median is a robust measure of central tendency that is less affected by outliers compared to the mean, making it an ideal choice for preserving the underlying distribution of these critical variables.\n\n### Mean Imputation for General Features\nConversely, for general features that do not carry as much weight in the model, we will use mean imputation. This method is straightforward and helps maintain the overall balance of the dataset by replacing missing values with the average value of the feature.\n\nBy applying these tailored strategies, we aim to minimize the impact of missing data on our model's performance while ensuring that our dataset remains as informative and representative as possible.","metadata":{}},{"cell_type":"code","source":"y_train  = train['Premium Amount']\n\ncombined_data = pd.concat([train.drop('Premium Amount', axis=1), test], ignore_index=True)\n\nmissing_data = combined_data.isnull().sum()\nmissing_percentage = missing_data / len(combined_data) * 100\nprint(missing_percentage)\n\n# Set a threshold for deleting missing values (for example, missing more than 50%)\nthreshold = 50\ncolumns_to_drop = missing_percentage[missing_percentage > threshold].index\ncombined_data = combined_data.drop(columns=columns_to_drop)\nprint(f\"Deleted column: {columns_to_drop.tolist()}\")\n\n\nfor col in combined_data.columns:\n    if combined_data[col].dtype in ['float64', 'int64']:  \n        if col in ['Age', 'Annual Income', 'Health Score', 'Credit Score']:  \n            # Fill important columns with medians\n            combined_data[col] = combined_data[col].fillna(combined_data[col].median())\n        else:\n            # The general feature column is filled with the mean\n            combined_data[col] = combined_data[col].fillna(combined_data[col].mean())\n    elif combined_data[col].dtype == 'object':  \n        # Fill missing values with the mode\n        combined_data[col] = combined_data[col].fillna(combined_data[col].mode()[0])\n\ntrain = combined_data.iloc[:len(train)]\ntest = combined_data.iloc[len(train):]\n\nprint(\"Missing value after processing:\")\nprint(train.isnull().sum())\nprint(test.isnull().sum())\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-17T05:41:49.504757Z","iopub.execute_input":"2024-12-17T05:41:49.505083Z","iopub.status.idle":"2024-12-17T05:41:55.435964Z","shell.execute_reply.started":"2024-12-17T05:41:49.505047Z","shell.execute_reply":"2024-12-17T05:41:55.434987Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Data classification ","metadata":{}},{"cell_type":"markdown","source":"### Segmentation of Dataset\nIn this analysis, we approach the data by dividing it into three distinct categories to facilitate targeted preprocessing and feature engineering strategies:\n\n1. **Numerical Features**  \n   These are the continuous or discrete numerical values that can be used directly or may require scaling/normalization to be effectively utilized by machine learning models.\n\n2. **Categorical Features**  \n   Categorical data represents variables that are names or labels, such as country names, brand names, or yes/no answers. We will employ techniques like one-hot encoding or label encoding to convert these into a format suitable for model training.\n\n3. **Datetime Features**  \n   Datetime features require special handling to extract meaningful information such as year, month, day, or even specific times of the day which can be crucial for time-series analysis or models that depend on temporal dynamics.\n\nBy meticulously categorizing the data, we ensure that each type of feature is processed in a manner that best preserves its inherent information and maximizes its contribution to the model's predictive power.","metadata":{}},{"cell_type":"code","source":"num_cols = list(combined_data.select_dtypes(include=['number']).columns)\ncat_cols = list(combined_data.select_dtypes(exclude=['number', 'datetime64[ns]']).columns)\n\ndatetime_cols = ['Policy Start Date']\n\nif 'id' in num_cols:\n    num_cols.remove('id')\n\nif 'Policy Start Date' in cat_cols:\n    cat_cols.remove('Policy Start Date')\n\nprint(\"Numerical columns:\")\nprint(num_cols)\nprint(\"\\nCategorical columns excluding datetime columns:\")\nprint(cat_cols)\nprint(\"\\nDatetime column:\")\nprint(datetime_cols)\n\ntrain = combined_data.iloc[:len(train)]\ntest = combined_data.iloc[len(train):]\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-17T05:41:55.437003Z","iopub.execute_input":"2024-12-17T05:41:55.437276Z","iopub.status.idle":"2024-12-17T05:41:56.511298Z","shell.execute_reply.started":"2024-12-17T05:41:55.43725Z","shell.execute_reply":"2024-12-17T05:41:56.510335Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Feature engineering 🔍","metadata":{}},{"cell_type":"markdown","source":"## Feature Treatment Methods 🌐\n\nBy examining the datasets, we can categorize the data into three main types: **Numerical features**, **Classification features**, and **Datetime features**. Below are the general treatment methods for each type:\n\n### Numerical Characteristics\n\n| Treatment Method | Description |\n| ---------------- | ----------- |\n| **Standardization** | Converts data into a distribution with a mean of 0 and a standard deviation of 1. |\n| **Dealing with Missing Values** | Fill in missing values (e.g., with the mean, median, or mode) or remove samples containing missing values. |\n| **Feature Engineering** | Create new features, such as polynomial features, logarithmic transformations, etc. |\n\n### Classification Features\n\n| Treatment Method | Description |\n| ---------------- | ----------- |\n| **One-Hot Encoding** | Converts each category to a new binary column. |\n| **Label Encoding** | Converts category labels to integers. |\n| **Frequency Coding** | Uses the frequency of occurrence of a category as its code value. |\n| **Target Encoding** | Encodes class features using the mean of the target variable; be cautious of overfitting. |\n\n### Datetime Features\n\n| Treatment Method | Description |\n| ---------------- | ----------- |\n| **Parse Datetime** | Converts datetime from string format to Python datetime objects. |\n| **Extract Features** | Extracts components such as year, month, day, hour, etc., from datetime. |\n| **Time Period Coding** | Converts datetime into periodic features like day of the week, weekend, etc. |\n| **Time Difference** | Calculates the difference between two timestamps (e.g., how long since insurance began). |\n| **Seasonality and Trend Analysis** | Applies time series analysis to capture seasonality and trends in the data. |\n\n### General Advice\n\nFor the above three types of features, it's important to avoid overprocessing the features. Overfitting our model due to excessive feature manipulation can be detrimental to its performance.","metadata":{}},{"cell_type":"markdown","source":"## Numerical feature processing ","metadata":{}},{"cell_type":"markdown","source":"### Data Standardization\n\nPrincipal Component Analysis (PCA) is a dimensionality reduction technique that is sensitive to the scale of the data. Therefore, it is crucial to standardize the numerical features before applying PCA.\n\n#### Why Standardize Data?\n\n- **PCA Sensitivity**: PCA is affected by the variances of the initial variables. Features with larger variances will dominate the principal components, which can lead to less meaningful outcomes.\n- **Equal Importance**: Standardization ensures that all features are treated equally during the transformation process, regardless of their original scale.\n\n#### How to Standardize Data?\n\nTo standardize the data, you can use the following formula for each feature:\n\n$$ z = \\frac{x - \\mu}{\\sigma} $$\n\nWhere:\n- \\( x \\) is the original value.\n- \\( \\mu \\) is the mean of the feature.\n- \\( \\sigma \\) is the standard deviation of the feature.\n- \\( z \\) is the standardized value.\n\n#### Why do we split data？\n\r\nThis is because before any pre-processing (such as coding, standardization), the data set should be split into training sets and test sets to avoid data leakagee\n","metadata":{}},{"cell_type":"code","source":"# Applying Normalization\nscaler = StandardScaler()\ncombined_data[num_cols] = scaler.fit_transform(combined_data[num_cols])\n\ntrain = combined_data.iloc[:len(train)]\ntest = combined_data.iloc[len(train):]\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-17T05:41:56.512304Z","iopub.execute_input":"2024-12-17T05:41:56.51258Z","iopub.status.idle":"2024-12-17T05:41:56.933145Z","shell.execute_reply.started":"2024-12-17T05:41:56.512554Z","shell.execute_reply":"2024-12-17T05:41:56.932469Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-17T05:41:56.934365Z","iopub.execute_input":"2024-12-17T05:41:56.93461Z","iopub.status.idle":"2024-12-17T05:41:56.950239Z","shell.execute_reply.started":"2024-12-17T05:41:56.934587Z","shell.execute_reply":"2024-12-17T05:41:56.949416Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"test.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-17T05:41:56.951387Z","iopub.execute_input":"2024-12-17T05:41:56.951726Z","iopub.status.idle":"2024-12-17T05:41:56.969874Z","shell.execute_reply.started":"2024-12-17T05:41:56.95169Z","shell.execute_reply":"2024-12-17T05:41:56.969043Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Categorical feature processing ","metadata":{}},{"cell_type":"markdown","source":"#### Choice of Encoding Technique\nFor handling categorical features in classification tasks, I have opted for **One-Hot Encoding** as the method of processing. This technique is particularly effective for converting categorical variables into a form that can be provided to machine learning algorithms, allowing them to work with the data more efficiently.\n\nOne-Hot Encoding creates a new binary column for each category of a feature, which helps prevent the model from making assumptions about the numerical relationship between categories. This approach is especially crucial in scenarios where the order of categories is arbitrary and does not represent a sequence or rank.\n\nBy employing One-Hot Encoding, we ensure that our model can accurately interpret the categorical data, thereby enhancing the predictive performance of our classification models.","metadata":{}},{"cell_type":"code","source":"# One-hot encoding\nencoder = OneHotEncoder(sparse=False, drop='first')\ncombined_cat = combined_data[cat_cols]\n\nencoder.fit(combined_cat)\ncombined_encoded = encoder.transform(combined_cat)\nencoded_feature_names = encoder.get_feature_names_out(cat_cols)\ncombined_encoded_df = pd.DataFrame(combined_encoded, columns=encoded_feature_names, index=combined_data.index)\n\ncombined_data = combined_data.drop(columns=cat_cols)\ncombined_data = pd.concat([combined_data, combined_encoded_df], axis=1)\n\ntrain = combined_data.iloc[:len(train)]\ntest = combined_data.iloc[len(train):]\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-17T05:41:56.97107Z","iopub.execute_input":"2024-12-17T05:41:56.971416Z","iopub.status.idle":"2024-12-17T05:42:03.548504Z","shell.execute_reply.started":"2024-12-17T05:41:56.97138Z","shell.execute_reply":"2024-12-17T05:42:03.54778Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-17T05:42:03.549586Z","iopub.execute_input":"2024-12-17T05:42:03.549913Z","iopub.status.idle":"2024-12-17T05:42:03.572207Z","shell.execute_reply.started":"2024-12-17T05:42:03.549879Z","shell.execute_reply":"2024-12-17T05:42:03.571568Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"test.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-17T05:42:03.573537Z","iopub.execute_input":"2024-12-17T05:42:03.573859Z","iopub.status.idle":"2024-12-17T05:42:03.598664Z","shell.execute_reply.started":"2024-12-17T05:42:03.573823Z","shell.execute_reply":"2024-12-17T05:42:03.59785Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Datetime feature processing ","metadata":{}},{"cell_type":"markdown","source":"### Feature Extraction Approach\nIn the realm of datetime feature processing, I have chosen to **extract features** rather than using the raw datetime values directly. This strategic choice is motivated by the fact that extracting meaningful components from datetime objects can significantly enhance the model's ability to capture temporal patterns and trends.\n\nBy decomposing datetime features into more interpretable parts such as year, month, day, hour, and more, we provide the model with a richer set of inputs that can lead to better performance in predictive tasks. This method allows us to uncover seasonal variations, cyclic trends, and other time-related effects that are often critical in time series analysis and other datetime-dependent models.\n\nThrough this meticulous feature engineering, we aim to tap into the full potential of our datetime data, thereby bolstering the accuracy and reliability of our models.","metadata":{}},{"cell_type":"code","source":"for col in datetime_cols:\n    combined_data[col] = pd.to_datetime(combined_data[col])\n    combined_data[f'{col}_year'] = combined_data[col].dt.year\n    combined_data[f'{col}_month'] = combined_data[col].dt.month\n    combined_data[f'{col}_day'] = combined_data[col].dt.day\n    combined_data[f'{col}_hour'] = combined_data[col].dt.hour\n    combined_data[f'{col}_minute'] = combined_data[col].dt.minute\n\n    combined_data = combined_data.drop(columns=[col])\n\ntrain = combined_data.iloc[:len(train)]\ntest = combined_data.iloc[len(train):]\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-17T05:42:03.601918Z","iopub.execute_input":"2024-12-17T05:42:03.602495Z","iopub.status.idle":"2024-12-17T05:42:04.658802Z","shell.execute_reply.started":"2024-12-17T05:42:03.602467Z","shell.execute_reply":"2024-12-17T05:42:04.6581Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-17T06:24:57.849635Z","iopub.execute_input":"2024-12-17T06:24:57.850005Z","iopub.status.idle":"2024-12-17T06:24:57.874836Z","shell.execute_reply.started":"2024-12-17T06:24:57.849975Z","shell.execute_reply":"2024-12-17T06:24:57.873715Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"test.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-17T06:25:16.349039Z","iopub.execute_input":"2024-12-17T06:25:16.349441Z","iopub.status.idle":"2024-12-17T06:25:16.372765Z","shell.execute_reply.started":"2024-12-17T06:25:16.349408Z","shell.execute_reply":"2024-12-17T06:25:16.371618Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Hyperparameter optimization 🔧","metadata":{}},{"cell_type":"markdown","source":"### Further Enhancing Our Model\nBuilding on our previous efforts, the next phase in our modeling journey involves **hyperparameter optimization**. This crucial step is aimed at refining our model's performance by systematically searching for the most optimal set of parameters.\n\nThrough this process, we will be able to fine-tune our model, adjusting everything from learning rates to the complexity of the model itself. The goal is to strike the right balance that maximizes predictive accuracy while avoiding overfitting.\n\nStay tuned as we delve into the world of hyperparameters, where small adjustments can lead to significant improvements in our model's predictive power.","metadata":{}},{"cell_type":"code","source":"\nlogging.getLogger('lightgbm').setLevel(logging.ERROR)\noptuna.logging.set_verbosity(optuna.logging.WARNING)\n\ntrain.columns = train.columns.str.replace(\" \", \"_\", regex=True)\ntest.columns = test.columns.str.replace(\" \", \"_\", regex=True)\n\n\ntrain['Premium Amount'] = y_train\nX_train = train.drop('Premium Amount', axis=1)\n\n\ny_train_log = np.log1p(y_train)\n\n\ntest = test[X_train.columns]\n\ndef objective(trial):\n    param = {\n        \"objective\": \"regression\",\n        \"metric\": \"rmse\",\n        \"boosting_type\": trial.suggest_categorical(\"boosting_type\", [\"gbdt\", \"dart\"]),\n        \"num_leaves\": trial.suggest_int(\"num_leaves\", 200, 512),\n        \"learning_rate\": trial.suggest_loguniform(\"learning_rate\", 1e-4, 1e-1),\n        \"feature_fraction\": trial.suggest_uniform(\"feature_fraction\", 0.6, 1.0),\n        \"bagging_fraction\": trial.suggest_uniform(\"bagging_fraction\", 0.6, 1.0),\n        \"bagging_freq\": trial.suggest_int(\"bagging_freq\", 5, 12),\n        \"min_data_in_leaf\": trial.suggest_int(\"min_data_in_leaf\", 20, 100),\n        \"max_depth\": trial.suggest_int(\"max_depth\", -1, 16), \n        \"lambda_l1\": trial.suggest_loguniform(\"lambda_l1\", 1e-4, 10.0),\n        \"lambda_l2\": trial.suggest_loguniform(\"lambda_l2\", 1e-4, 10.0),\n        \"device_type\": \"gpu\", \n        \"seed\": 42,\n        \"verbose\": -1 \n    }\n\n    n_splits = 5\n    kf = KFold(n_splits=n_splits, shuffle=True, random_state=42)\n    cv_scores = []\n    \n    for train_index, val_index in kf.split(X_train):\n        X_train_fold, X_val_fold = X_train.iloc[train_index], X_train.iloc[val_index]\n        y_train_fold, y_val_fold = y_train_log.iloc[train_index], y_train_log.iloc[val_index]\n\n        dtrain = lgb.Dataset(X_train_fold, label=y_train_fold)\n        dval = lgb.Dataset(X_val_fold, label=y_val_fold, reference=dtrain)\n\n        early_stopping = lgb.early_stopping(stopping_rounds=50, verbose=False)\n\n        model = lgb.train(\n            param,\n            dtrain,\n            valid_sets=[dval],\n            callbacks=[early_stopping],  \n        )\n\n        y_val_pred = model.predict(X_val_fold)\n        \n        y_val_pred = np.expm1(y_val_pred)\n        rmsle = np.sqrt(mean_squared_log_error(np.expm1(y_val_fold), np.maximum(y_val_pred, 0)))\n        cv_scores.append(rmsle)\n    return np.mean(cv_scores)\n\nstudy = optuna.create_study(direction='minimize')\nstudy.optimize(objective, n_trials=10)\n\nprint(\"Best parameters:\", study.best_params)\nprint(f\"Best RMSLE: {study.best_value:.4f}\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-17T05:42:40.494356Z","iopub.execute_input":"2024-12-17T05:42:40.49503Z","iopub.status.idle":"2024-12-17T05:56:58.252601Z","shell.execute_reply.started":"2024-12-17T05:42:40.494998Z","shell.execute_reply":"2024-12-17T05:56:58.251248Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Model Train 📉","metadata":{}},{"cell_type":"markdown","source":"### Outputting RMSLE\nDuring the model training process, I will be focusing on outputting the validation of RMSLE (Root Mean Squared Logarithmic Error). This crucial step allows us to quantitatively assess the model's predictive accuracy by measuring the difference between predicted values and actual outcomes on a logarithmic scale.\n\nBy monitoring the RMSLE, we can ensure that our model not only performs well on paper but also generalizes effectively to unseen data, providing us with a robust measure of its real-world predictive power.","metadata":{}},{"cell_type":"code","source":"logging.getLogger('lightgbm').setLevel(logging.ERROR)  \nwarnings.filterwarnings(\"ignore\", category=UserWarning, message=\".*Found whitespace in feature_names.*\")\n\nX_train.columns = X_train.columns.str.replace(\" \", \"_\", regex=True)\ntest.columns = test.columns.str.replace(\" \", \"_\", regex=True)\n\n\nbest_params = study.best_params\n\nbest_params.update({\n    \"objective\": \"regression\",\n    \"metric\": \"rmse\",\n    \"device_type\": \"gpu\",  \n    \"seed\": 42,\n    \"verbose\": -1  \n})\n\n\ny_train_log = np.log1p(y_train)\n\nn_splits = 5\nkf = KFold(n_splits=n_splits, shuffle=True, random_state=42)\ncv_scores = []\n\n# Lists to store out-of-fold predictions\noof_predictions = np.zeros(len(X_train))\nfinal_test_predictions = np.zeros(len(test))  \n\nbest_model = None\n\nfor train_index, val_index in kf.split(X_train):\n    X_train_fold, X_val_fold = X_train.iloc[train_index], X_train.iloc[val_index]\n    y_train_fold, y_val_fold = y_train_log.iloc[train_index], y_train_log.iloc[val_index]\n\n    dtrain = lgb.Dataset(X_train_fold, label=y_train_fold)\n    dval = lgb.Dataset(X_val_fold, label=y_val_fold, reference=dtrain)\n\n    early_stopping = lgb.early_stopping(stopping_rounds=50, verbose=False) \n\n    model = lgb.train(\n        best_params,\n        dtrain,\n        valid_sets=[dval],\n        callbacks=[early_stopping],  \n    )\n\n    best_model = model\n\n    oof_predictions[val_index] = np.expm1(model.predict(X_val_fold))\n    final_test_predictions += np.expm1(model.predict(test)) / n_splits  \n\n    y_val_pred = model.predict(X_val_fold)\n    \n    y_val_pred = np.expm1(y_val_pred)\n    fold_rmsle = np.sqrt(mean_squared_log_error(np.expm1(y_val_fold), np.maximum(y_val_pred, 0)))\n    cv_scores.append(fold_rmsle)\n\nprint(f\"Mean RMSLE across {n_splits} folds: {np.mean(cv_scores):.4f}\")\n\nfinal_rmsle = np.sqrt(mean_squared_log_error(y_train, np.maximum(oof_predictions, 0)))\nprint(f\"Final RMSLE on out-of-fold predictions: {final_rmsle:.4f}\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-17T06:08:32.129091Z","iopub.execute_input":"2024-12-17T06:08:32.129979Z","iopub.status.idle":"2024-12-17T06:11:45.083851Z","shell.execute_reply.started":"2024-12-17T06:08:32.129946Z","shell.execute_reply":"2024-12-17T06:11:45.082381Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Parameter visualization 🎨","metadata":{}},{"cell_type":"markdown","source":"## Optimization History and Parameter Importance Analysis\n\nDuring the model optimization process, we utilized two key visualization tools to gain insights into the optimization process and the impact of parameters.\n\n### Optimization History\nWe presented the optimization history through a dedicated chart that meticulously records the performance of each trial. This allows us to track the progress of parameter optimization. By observing changes in learning rates, loss functions, and other key metrics, we can visually assess how model performance varies with different parameter combinations. This not only helps us understand which parameters significantly affect model performance but also reveals trends and patterns in the optimization process.\n\n### Parameter Importance\nAdditionally, we identified the most influential parameters on model performance using a parameter importance plot. This chart ranks parameters based on their contribution to model improvement, highlighting the most significant ones. In this way, we can prioritize adjustments to parameters that have a substantial positive impact on model performance, making more efficient use of our resources and time.\n\nBy combining the analysis of optimization history and parameter importance, we gain a deeper understanding of model behavior and make more informed decisions to enhance model performance.","metadata":{}},{"cell_type":"code","source":"# Optimization history\nplot_optimization_history(study).show()\n\n# Parameter importance\nplot_param_importances(study).show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-17T05:42:04.714192Z","iopub.status.idle":"2024-12-17T05:42:04.714464Z","shell.execute_reply.started":"2024-12-17T05:42:04.71433Z","shell.execute_reply":"2024-12-17T05:42:04.714344Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Feature importance 🏁","metadata":{}},{"cell_type":"markdown","source":"## Understanding the Importance of Visual Features\n\nAt this critical juncture, it's essential to recognize the pivotal role that visual features play in our data analysis process. Visual features have the unparalleled ability to render complex data interactions comprehensible at a glance.\n\n### Why Visual Features Matter\n\n- **Intuitive Insights**: Visual features allow us to grasp the significance of various elements within our dataset intuitively, making it easier to digest and interpret information.\n- **Pattern Recognition**: They facilitate the identification of patterns and trends that might otherwise remain hidden amidst raw numbers and text.\n- **Decision Making**: By providing a clear and concise representation of data, visual features enhance our decision-making capabilities, ensuring that choices are informed and strategic.\n\n### Harnessing Visual Features for Better Understanding\n\n- **Data Visualization**: Utilize charts, graphs, and heatmaps to visualize data distributions, correlations, and outliers.\n- **Feature Importance**: Employ techniques such as bar charts or forest plots to illustrate the impact and relevance of different features.\n- **Interactive Exploration**: Leverage interactive visualization tools to delve deeper into the data, allowing for a more dynamic exploration of features.\n\nBy integrating visual features into our analysis toolkit, we empower ourselves with a powerful means to understand and communicate the intricacies of our data.\n\n","metadata":{}},{"cell_type":"code","source":"feature_importance = pd.DataFrame({\n    'feature': X_train.columns,\n    'importance': best_model.feature_importance()\n}).sort_values('importance', ascending=False)\n\n\nplt.figure(figsize=(10, 6))\nsns.barplot(x='importance', y='feature', data=feature_importance.head(10))\nplt.title('Top 10 Important Features')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-17T05:42:04.715569Z","iopub.status.idle":"2024-12-17T05:42:04.715859Z","shell.execute_reply.started":"2024-12-17T05:42:04.715718Z","shell.execute_reply":"2024-12-17T05:42:04.715733Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Generate Predictions & Submission 🏆","metadata":{}},{"cell_type":"markdown","source":"## Create a Prediction Test \n\n### Objective\nIn this phase of our project, we need to ensure the integrity and correctness of our submission format. Specifically, we will conduct a thorough check to validate that the CSV file we have submitted adheres to the required specifications.\n\n### Importance of Format Verification\n- **Accuracy in Submission**: Ensuring that our CSV format is correct is crucial for maintaining the consistency and reliability of our data submissions.\n- **Avoidance of Errors**: A proper format check helps prevent errors that could lead to misinterpretation or rejection of our results.\n- **Alignment with Standards**: Adhering to the specified CSV format guarantees that our submissions are in line with the standards set by the platform or competition requirements.\n\n### Steps for CSV Format Verification\n1. **Schema Check**: Verify that each column in the CSV file corresponds to the expected data type and format.\n2. **Data Integrity**: Ensure that there are no missing values or inconsistencies that could affect the performance of our models.\n3. **Header Validation**: Confirm that the header row accurately reflects the content of the data columns.\n4. **Conformance to Specifications**: Cross-reference the CSV structure against the submission guidelines provided to ensure full compliance.\n\nBy meticulously carrying out these checks, we can be confident that our submissions are not only correct but also poised to deliver the best possible results.\n\n","metadata":{}},{"cell_type":"code","source":"test_submission = pd.read_csv('/kaggle/input/playground-series-s4e12/sample_submission.csv') \n\nsubmission = pd.DataFrame({'id': test_submission['id'], 'Premium Amount': final_test_predictions})\nsubmission.to_csv(\"submission.csv\", index=False)\n\nprint(\"Submission file created:\")\nprint(submission.head())\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-17T05:42:04.717203Z","iopub.status.idle":"2024-12-17T05:42:04.717509Z","shell.execute_reply.started":"2024-12-17T05:42:04.717367Z","shell.execute_reply":"2024-12-17T05:42:04.717382Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Conclusion","metadata":{}},{"cell_type":"markdown","source":"## 🌟 **Thank You for Reading This Notebook in Its Entirety** 🌟\n\nDear Kaggler,\n\nI sincerely appreciate your dedication to reading through this notebook from beginning to end. It's readers like you who make our community vibrant and insightful.\n\n## 🚀 **I Hope This Has Been Helpful** 🚀\n\nMy goal was to provide you with valuable knowledge and tools to aid in your data science journey. If there's anything specific you've gained from this notebook, or if you feel there's an area we could expand upon, please don't hesitate to reach out.\n\n## 💡 **Your Input is Valued** 💡\n\nI am eager to learn from your experiences and expertise as well. Should you have any suggestions, feedback, or insights, please feel free to share them in the comments section below. Your contributions help make this a collaborative and enriching environment for all.\n\n## 🌈 **Wishing You the Best** 🌈\n\nAs you continue to explore, learn, and innovate, I wish you all the best in your endeavors. May your models be free of overfitting and your GPUs never run out of memory!\n\n---\n\nFeel free to engage with the community and happy data mining!","metadata":{}}]}