{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"codemirror_mode":{"name":"ipython","version":3},"file_extension":".py","mimetype":"text/x-python","name":"python","nbconvert_exporter":"python","pygments_lexer":"ipython3","version":"3.9.11"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":84896,"databundleVersionId":10305135,"sourceType":"competition"}],"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# 🚀 ClaimSense: Gradient Boosting for Insurance Premium Prediction with Robust EDA 🔍\n\n## Notebook Outline\n\n1. [Overview]\n2. [Exploratory Data Analysis (EDA)]\n3. [Data Preprocessing & Feature Engineering]\n4. [Model Training]\n5. [Model Evaluation & Selection]\n6. [Final Submission]\n\n------\n\n## 1. Overview\n\n### 1.1 Objective\n\nThe goal is to predict the insurance premium amount for each policyholder in the test set.\n\n### 1.2 Data\n\nWe have two CSV files:\n\n- **train.csv**: Contains features + the target variable (**Premium Amount**).\n- **test.csv**: Contains the same features (minus the target) for which we need to predict Premium Amount.\n\n### 1.3 Evaluation Metric\n\nCommon metrics for regression tasks include RMSE, MAE, or R². (If the competition defines something else, follow that.)\n\n------\n\n## 2. Exploratory Data Analysis (EDA)\n\nIn this section, we’ll:\n\n- Load the datasets.\n- Take a quick look at shape, columns, and basic stats.\n- Check for missing values.\n- Visualize the target distribution.\n\n### 2.1 Imports","metadata":{"vscode":{"languageId":"plaintext"}}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n# For model training later\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.preprocessing import LabelEncoder\nfrom sklearn.linear_model import LinearRegression\nfrom sklearn.ensemble import RandomForestRegressor\nfrom sklearn.metrics import mean_squared_error, r2_score\n\nimport xgboost as xgb\nimport lightgbm as lgb\n\n# Let's ensure plots show in the notebook\n%matplotlib inline","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 2.2 Load the data","metadata":{}},{"cell_type":"code","source":"train_data = pd.read_csv('/kaggle/input/playground-series-s4e12/train.csv')\ntest_data = pd.read_csv('/kaggle/input/playground-series-s4e12/test.csv')\n\nprint(\"Train shape:\", train_data.shape)\nprint(\"Test shape:\", test_data.shape)","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 2.3 Inspect columns","metadata":{}},{"cell_type":"code","source":"print(\"Train columns:\", train_data.columns.tolist())\nprint(\"Test columns:\", test_data.columns.tolist())","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 2.4 Peek at data","metadata":{}},{"cell_type":"code","source":"display(train_data.head())\ndisplay(test_data.head())","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 2.5 Check missing values","metadata":{}},{"cell_type":"code","source":"print(\"Missing (Train):\\n\", train_data.isnull().sum())\nprint(\"-\" * 30)\nprint(\"Missing (Test):\\n\", test_data.isnull().sum())","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 2.6 Basic descriptive stats (for train)","metadata":{}},{"cell_type":"code","source":"display(train_data.describe())","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 2.7 Distribution of target","metadata":{}},{"cell_type":"code","source":"sns.histplot(train_data['Premium Amount'], kde=True)\nplt.title('Distribution of Premium Amount')\nplt.show()","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 3. Data Preprocessing & Feature Engineering\n\nWe will:\n\n1. Convert `Policy Start Date` to datetime and extract year, month.\n2. Handle missing values (numeric → median, categorical → mode).\n3. Label encode categorical columns (fit on combined train+test to avoid unseen labels).\n\n\n\n### 3.1 Convert to datetime and extract date features","metadata":{}},{"cell_type":"code","source":"train_data['Policy Start Date'] = pd.to_datetime(train_data['Policy Start Date'], errors='coerce')\ntest_data['Policy Start Date'] = pd.to_datetime(test_data['Policy Start Date'], errors='coerce')\n\ntrain_data['Policy_Year'] = train_data['Policy Start Date'].dt.year\ntrain_data['Policy_Month'] = train_data['Policy Start Date'].dt.month\n\ntest_data['Policy_Year'] = test_data['Policy Start Date'].dt.year\ntest_data['Policy_Month'] = test_data['Policy Start Date'].dt.month","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 3.2 Drop the original date column if we won't use it","metadata":{}},{"cell_type":"code","source":"train_data.drop('Policy Start Date', axis=1, inplace=True)\ntest_data.drop('Policy Start Date', axis=1, inplace=True)","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 3.3 Handle missing values","metadata":{}},{"cell_type":"code","source":"num_features = train_data.select_dtypes(include=[np.number]).columns.tolist()\ncat_features = train_data.select_dtypes(exclude=[np.number]).columns.tolist()\n\n# Remove target from numeric features if present\nif 'Premium Amount' in num_features:\n    num_features.remove('Premium Amount')\n\n# Numeric: fill with median\nfor col in num_features:\n    median_val = train_data[col].median()\n    train_data[col] = train_data[col].fillna(median_val)\n    if col in test_data.columns:\n        test_data[col] = test_data[col].fillna(median_val)\n\n# Categorical: fill with mode\nfor col in cat_features:\n    if not train_data[col].mode().empty:\n        mode_val = train_data[col].mode()[0]\n    else:\n        mode_val = 'Unknown'\n    train_data[col] = train_data[col].fillna(mode_val)\n    if col in test_data.columns:\n        test_data[col] = test_data[col].fillna(mode_val)","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 3.4 Label Encoding","metadata":{}},{"cell_type":"code","source":"for col in cat_features:\n    # Combine train + test for consistent encoding\n    combined_data = pd.concat([\n        train_data[col].astype(str),\n        test_data[col].astype(str)\n    ], axis=0)\n    le = LabelEncoder()\n    le.fit(combined_data)\n\n    train_data[col] = le.transform(train_data[col].astype(str))\n    test_data[col] = le.transform(test_data[col].astype(str))","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 3.5 Finalize features","metadata":{}},{"cell_type":"code","source":"X = train_data.drop(['id', 'Premium Amount'], axis=1)\ny = train_data['Premium Amount']\n\nX_test = test_data.drop(['id'], axis=1)\n\nprint(\"Final training feature shape:\", X.shape)\nprint(\"Final test feature shape:\", X_test.shape)","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 4. Model Training\n\nWe’ll try:\n\n1. A simple Linear Regression (baseline).\n2. XGBoost.\n3. LightGBM.\n4. (Optional) Random Forest.\n\nThen, we’ll pick the best model.\n\n\n\n### 4.1 Train/Validation Split","metadata":{}},{"cell_type":"code","source":"X_train, X_val, y_train, y_val = train_test_split(\n  X, y, test_size=0.2, random_state=42\n)\n\nprint(\"Train Split:\", X_train.shape, y_train.shape)\nprint(\"Validation Split:\", X_val.shape, y_val.shape)","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 4.2 Linear Regression","metadata":{}},{"cell_type":"code","source":"lr = LinearRegression()\nlr.fit(X_train, y_train)\n\ny_pred_lr = lr.predict(X_val)\nmse_lr = mean_squared_error(y_val, y_pred_lr)\nr2_lr = r2_score(y_val, y_pred_lr)\nrmse_lr = np.sqrt(mse_lr)\n\nprint(f\"[Linear Regression] MSE: {mse_lr:.4f}, R2: {r2_lr:.4f}\")\nprint(f\"[Linear Regression] RMSE: {rmse_lr:.4f}\")","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 4.3 XGBoost","metadata":{}},{"cell_type":"code","source":"xgb_model = xgb.XGBRegressor(\n  n_estimators=500,\n  learning_rate=0.05,\n  max_depth=6,\n  subsample=0.8,\n  colsample_bytree=0.8,\n  random_state=42\n)\n\nxgb_model.fit(\n  X_train, y_train,\n  eval_set=[(X_val, y_val)],\n  verbose=False\n)\n\ny_pred_xgb = xgb_model.predict(X_val)\nmse_xgb = mean_squared_error(y_val, y_pred_xgb)\nr2_xgb = r2_score(y_val, y_pred_xgb)\nrmse_xgb = np.sqrt(mse_xgb)\n\nprint(f\"[XGBoost] MSE: {mse_xgb:.4f}, R2: {r2_xgb:.4f}\")\nprint(f\"[XGBoost] RMSE: {rmse_xgb:.4f}\")","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 4.4 LightGBM","metadata":{}},{"cell_type":"code","source":"lgb_model = lgb.LGBMRegressor(\n  n_estimators=500,\n  learning_rate=0.05,\n  max_depth=6,\n  subsample=0.8,\n  colsample_bytree=0.8,\n  random_state=42\n)\n\nlgb_model.fit(\n  X_train, y_train,\n  eval_set=[(X_val, y_val)]\n)\n\ny_pred_lgb = lgb_model.predict(X_val)\nmse_lgb = mean_squared_error(y_val, y_pred_lgb)\nr2_lgb = r2_score(y_val, y_pred_lgb)\nrmse_lgb = np.sqrt(mse_lgb)\n\nprint(f\"[LightGBM] MSE: {mse_lgb:.4f}, R2: {r2_lgb:.4f}\")\nprint(f\"[LightGBM] RMSE: {rmse_lgb:.4f}\")","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 4.5 Optional: Random Forest","metadata":{}},{"cell_type":"code","source":"rf = RandomForestRegressor(\n  n_estimators=200,\n  max_depth=10,\n  random_state=42\n)\nrf.fit(X_train, y_train)\n\ny_pred_rf = rf.predict(X_val)\nmse_rf = mean_squared_error(y_val, y_pred_rf)\nr2_rf = r2_score(y_val, y_pred_rf)\nrmse_rf = np.sqrt(mse_rf)\n\nprint(f\"[Random Forest] MSE: {mse_rf:.4f}, R2: {r2_rf:.4f}\")\nprint(f\"RMSE: {rmse_rf:.4f}\")","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 5. Model Evaluation & Selection\n\nAt this point, we compare metrics (MSE, R²) for each model. For example:\n\n- **Linear Regression**: MSE = ..., R² = ...\n- **XGBoost**: MSE = ..., R² = ...\n- **LightGBM**: MSE = ..., R² = ...\n- **Random Forest**: MSE = ..., R² = ...\n\nWhichever yields the **lowest M\n\nSE** (and/or highest R²) is considered the best model. Let’s assume **LightGBM** performed the best here.\n\n------\n\n## 6. Final Submission\n\nWe retrain the best model on the **entire** training set (X, y) and then predict on **X_test**.","metadata":{}},{"cell_type":"code","source":"best_model = lgb_model  # Suppose LightGBM performed best\n\n# Retrain on the full data\nbest_model.fit(X, y)\n\n# Predict on test\ntest_preds = best_model.predict(X_test)\n\n# Create submission\nsubmission = pd.DataFrame({\n  'id': test_data['id'],\n  'Premium Amount': test_preds\n})\n\n# Round predictions if necessary\nsubmission['Premium Amount'] = submission['Premium Amount'].round(3)\n\n# Save CSV\nsubmission.to_csv('submission.csv', index=False)\nprint(\"Submission file created! 🚀\")","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🎉 Notebook Recap\n\nBy splitting the code into **multiple cells** and adding **markdown** sections, we’ve created a **reader-friendly** approach:\n\n1. We performed **EDA** to get a handle on the data.\n2. We **cleaned** the data, handled **missing values**, and **engineered** date features.\n3. We tested multiple **models**: Linear Regression, XGBoost, LightGBM, and Random Forest.\n4. We chose the **best** model (LightGBM) and produced our **submission** file.\n\nWith **further tuning** (hyperparameters, cross-validation, ensembling), we can boost performance even more. Good luck, and happy modeling!\n\n**Big thanks** for following along – if you have any questions, feel free to leave a comment below! 🚀","metadata":{}}]}