{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":84896,"databundleVersionId":10305135,"sourceType":"competition"},{"sourceId":10297953,"sourceType":"datasetVersion","datasetId":6373952},{"sourceId":10298115,"sourceType":"datasetVersion","datasetId":6374059}],"dockerImageVersionId":30822,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"![](/kaggle/input/image-data/1e29ddb72c077c53b87ee743701f0bb9.jpg)","metadata":{}},{"cell_type":"markdown","source":"# **🌟 Introduction:**\n\n**Welcome to Part 1 of this series! This notebook focuses on data preprocessing and model training for the insurance premium prediction competition. Exploratory Data Analysis (EDA) and data visualization will be presented in Part 2.**\n\n**🎯 Goal**\n\n\n**1.Predict insurance premium amounts based on training data features.**\n\n**2.The evaluation metric is RMSLE, which penalizes under-predictions more than over-predictions. Predictions must be non-negative since premiums cannot be negative.**\n\n**3.The premium is typically calculated based on factors that:Reflect the risk of insuring a customer (e.g., age, health score, vehicle age, etc.).Affect the policy cost (e.g., credit score, insurance duration, policy type).**\n\n\n\n\n\n\n\n\n\n","metadata":{}},{"cell_type":"markdown","source":"# **🧠 Data Understanding:**","metadata":{}},{"cell_type":"markdown","source":"**📚 Import Necessary Libraries**","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nfrom sklearn.model_selection import train_test_split,GridSearchCV\nfrom sklearn.linear_model import LinearRegression\nfrom sklearn.ensemble import RandomForestRegressor\nfrom sklearn.tree import DecisionTreeRegressor\nimport xgboost\nfrom xgboost import XGBRegressor,plot_importance\nfrom sklearn.metrics import mean_squared_log_error\nfrom sklearn.preprocessing import OneHotEncoder,StandardScaler,LabelEncoder\nfrom sklearn.pipeline import Pipeline\nfrom sklearn.compose import ColumnTransformer\nfrom sklearn.impute import SimpleImputer\nimport warnings\nwarnings.filterwarnings('ignore')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-03T19:42:07.432031Z","iopub.execute_input":"2025-01-03T19:42:07.432266Z","iopub.status.idle":"2025-01-03T19:42:09.688398Z","shell.execute_reply.started":"2025-01-03T19:42:07.432242Z","shell.execute_reply":"2025-01-03T19:42:09.687348Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_data=pd.read_csv(\"/kaggle/input/playground-series-s4e12/train.csv\")\ntest_data=pd.read_csv(\"/kaggle/input/playground-series-s4e12/test.csv\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-03T19:42:09.689846Z","iopub.execute_input":"2025-01-03T19:42:09.690261Z","iopub.status.idle":"2025-01-03T19:42:20.894141Z","shell.execute_reply.started":"2025-01-03T19:42:09.690235Z","shell.execute_reply":"2025-01-03T19:42:20.893301Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_data","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-03T19:42:20.896412Z","iopub.execute_input":"2025-01-03T19:42:20.896777Z","iopub.status.idle":"2025-01-03T19:42:21.706606Z","shell.execute_reply.started":"2025-01-03T19:42:20.896752Z","shell.execute_reply":"2025-01-03T19:42:21.705532Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"test_data","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-03T19:42:21.708544Z","iopub.execute_input":"2025-01-03T19:42:21.70901Z","iopub.status.idle":"2025-01-03T19:42:21.736905Z","shell.execute_reply.started":"2025-01-03T19:42:21.708966Z","shell.execute_reply":"2025-01-03T19:42:21.735766Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print('train data : ', train_data.shape)\nprint('test data : ', test_data.shape)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-03T19:42:21.738102Z","iopub.execute_input":"2025-01-03T19:42:21.738409Z","iopub.status.idle":"2025-01-03T19:42:21.751295Z","shell.execute_reply.started":"2025-01-03T19:42:21.738383Z","shell.execute_reply":"2025-01-03T19:42:21.750253Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(\"train data:\", train_data.columns)\nprint(\"test data:\", test_data.columns)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-03T19:42:21.752363Z","iopub.execute_input":"2025-01-03T19:42:21.752749Z","iopub.status.idle":"2025-01-03T19:42:21.77189Z","shell.execute_reply.started":"2025-01-03T19:42:21.752709Z","shell.execute_reply":"2025-01-03T19:42:21.770724Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_data.describe().round(2).T","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-03T19:42:21.773005Z","iopub.execute_input":"2025-01-03T19:42:21.773374Z","iopub.status.idle":"2025-01-03T19:42:22.470693Z","shell.execute_reply.started":"2025-01-03T19:42:21.773344Z","shell.execute_reply":"2025-01-03T19:42:22.469707Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**🔍 Explore Columns and Data Types**","metadata":{}},{"cell_type":"markdown","source":"**📝 Column Description **\n\n- id: Unique identifier for each customer.\n\n- Age: Age of the customer in years.\n\n- Gender: Gender of the customer (e.g., Male, Female).\n\n- Annual Income: Annual income of the customer in currency units.\n\n- Marital Status: Marital status of the customer (e.g., Single, Married).\n\n- Number of Dependents: Number of dependents the customer has.(Number of Dependents refers to the number of individuals financially dependent on the customer. These could include children, elderly parents, or any other family members who rely on the customer for financial support.)\n\n- Education Level: Customer's highest level of education.\n\n- Occupation: Customer's occupation or job role.\n\n- Health Score: A numerical representation of the customer's health condition.\n\n- Location: Geographic location of the customer.\n\n- Policy Type: Type of insurance policy the customer holds.\n\n- Previous Claims: Number of previous insurance claims made by the customer.\n\n- Vehicle Age: Age of the vehicle in years.\n\n- Credit Score: Customer's creditworthiness score.\n\n- Insurance Duration: Duration of the current insurance policy.\n\n- Policy Start Date: Start date of the insurance policy.\n\n- Customer Feedback: Feedback or ratings provided by the customer.\n\n- Smoking Status: Whether the customer is a smoker or not.\n\n- Exercise Frequency: Frequency of physical exercise by the customer.\n\n- Property Type: Type of property the customer owns (e.g., Apartment, House).\n\n- Premium Amount: The amount paid as the insurance premium.","metadata":{}},{"cell_type":"code","source":"train_data.info()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-03T19:42:22.471601Z","iopub.execute_input":"2025-01-03T19:42:22.471884Z","iopub.status.idle":"2025-01-03T19:42:23.108979Z","shell.execute_reply.started":"2025-01-03T19:42:22.47186Z","shell.execute_reply":"2025-01-03T19:42:23.107899Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_data.drop(columns=['id'],inplace=True)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-03T19:42:23.112276Z","iopub.execute_input":"2025-01-03T19:42:23.112567Z","iopub.status.idle":"2025-01-03T19:42:23.32096Z","shell.execute_reply.started":"2025-01-03T19:42:23.112542Z","shell.execute_reply":"2025-01-03T19:42:23.319926Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_data[\"Policy Start Date\"] = pd.to_datetime(train_data[\"Policy Start Date\"], errors='coerce', format='%Y-%m-%d')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-03T19:42:23.322881Z","iopub.execute_input":"2025-01-03T19:42:23.323216Z","iopub.status.idle":"2025-01-03T19:42:25.273649Z","shell.execute_reply.started":"2025-01-03T19:42:23.323174Z","shell.execute_reply":"2025-01-03T19:42:25.272615Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_data","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-03T19:42:25.274535Z","iopub.execute_input":"2025-01-03T19:42:25.274781Z","iopub.status.idle":"2025-01-03T19:42:25.303575Z","shell.execute_reply.started":"2025-01-03T19:42:25.274759Z","shell.execute_reply":"2025-01-03T19:42:25.302621Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**🚨 Missing Values and Duplicate Values**","metadata":{}},{"cell_type":"code","source":"train_data.isnull().sum()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-03T19:42:25.304576Z","iopub.execute_input":"2025-01-03T19:42:25.304845Z","iopub.status.idle":"2025-01-03T19:42:25.871725Z","shell.execute_reply.started":"2025-01-03T19:42:25.304822Z","shell.execute_reply":"2025-01-03T19:42:25.870621Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_data.duplicated().sum()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-03T19:42:25.872757Z","iopub.execute_input":"2025-01-03T19:42:25.873067Z","iopub.status.idle":"2025-01-03T19:42:27.276961Z","shell.execute_reply.started":"2025-01-03T19:42:25.873034Z","shell.execute_reply":"2025-01-03T19:42:27.276023Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**🔢 Numerical and Categorical Columns**","metadata":{}},{"cell_type":"code","source":"tar_col ='Premium Amount';\nnum_col = train_data.select_dtypes(include = ['number']).columns\ncat_col = train_data.select_dtypes(include = ['object']).columns\nprint(\"Target Column :\" ,tar_col)\nprint( \"\\nNumerical Columns :\" , num_col.tolist())\nprint( \"\\nCategorical Columns :\" , cat_col.tolist())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-03T19:42:27.277888Z","iopub.execute_input":"2025-01-03T19:42:27.278234Z","iopub.status.idle":"2025-01-03T19:42:27.83706Z","shell.execute_reply.started":"2025-01-03T19:42:27.278198Z","shell.execute_reply":"2025-01-03T19:42:27.836012Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"num_data=train_data.select_dtypes(include=['number'])\ncat_data=train_data.select_dtypes(include=['object'])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-03T19:42:27.837948Z","iopub.execute_input":"2025-01-03T19:42:27.838208Z","iopub.status.idle":"2025-01-03T19:42:28.357499Z","shell.execute_reply.started":"2025-01-03T19:42:27.838186Z","shell.execute_reply":"2025-01-03T19:42:28.356349Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print('Numerical Data Distribution!')\nnum_data.describe().round(2).T","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-03T19:42:28.358553Z","iopub.execute_input":"2025-01-03T19:42:28.358922Z","iopub.status.idle":"2025-01-03T19:42:29.009056Z","shell.execute_reply.started":"2025-01-03T19:42:28.358884Z","shell.execute_reply":"2025-01-03T19:42:29.008053Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(\"Categorical Data Dsicription!\")\ncat_data.describe().T","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-03T19:42:29.01004Z","iopub.execute_input":"2025-01-03T19:42:29.01031Z","iopub.status.idle":"2025-01-03T19:42:30.402357Z","shell.execute_reply.started":"2025-01-03T19:42:29.010286Z","shell.execute_reply":"2025-01-03T19:42:30.401316Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"for c in cat_data:\n    col_count=train_data[c].nunique()\n    print(f'{c} has {col_count} unqiue_values: ')\n    print(\"--\"*20)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-03T19:42:30.403443Z","iopub.execute_input":"2025-01-03T19:42:30.403798Z","iopub.status.idle":"2025-01-03T19:42:30.996436Z","shell.execute_reply.started":"2025-01-03T19:42:30.403764Z","shell.execute_reply":"2025-01-03T19:42:30.995536Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"for i in cat_col:\n    cat_value=train_data[i].value_counts()\n    print(f\"Value Count for {i}\")\n    print(cat_value)\n    print(\"*\"*40)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-03T19:42:30.997349Z","iopub.execute_input":"2025-01-03T19:42:30.997642Z","iopub.status.idle":"2025-01-03T19:42:31.866057Z","shell.execute_reply.started":"2025-01-03T19:42:30.997618Z","shell.execute_reply":"2025-01-03T19:42:31.864927Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# **🛠️ Data Preprocessing**","metadata":{}},{"cell_type":"markdown","source":"**📊 Check the Percentage of Rows Affected**","metadata":{}},{"cell_type":"code","source":"# Percentage of rows with nulls in 'Occupation' and 'Previous Claims'\naffected_rows = train_data[train_data['Occupation'].isnull() | train_data['Previous Claims'].isnull()]\nprint(len(affected_rows) / len(train_data) * 100)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-03T19:42:31.866912Z","iopub.execute_input":"2025-01-03T19:42:31.867167Z","iopub.status.idle":"2025-01-03T19:42:32.014323Z","shell.execute_reply.started":"2025-01-03T19:42:31.867145Z","shell.execute_reply":"2025-01-03T19:42:32.013365Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"- If a large percentage of rows >20% are affected, dropping these rows might harm the dataset's quality.","metadata":{}},{"cell_type":"markdown","source":"**⚙️ Feature Engineering:**","metadata":{}},{"cell_type":"code","source":"def date_trans(df):\n    df['Policy Start Date']= pd.to_datetime(df['Policy Start Date'])\n    df['Year'] = df['Policy Start Date'].dt.year\n    df['Day'] = df['Policy Start Date'].dt.day\n    df['Month'] = df['Policy Start Date'].dt.month\n    df.drop('Policy Start Date' , axis =1, inplace = True)\n    return df\n\n\ntrain_data = date_trans(train_data)\ntest_data = date_trans(test_data)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-03T19:42:32.015297Z","iopub.execute_input":"2025-01-03T19:42:32.015676Z","iopub.status.idle":"2025-01-03T19:42:32.88101Z","shell.execute_reply.started":"2025-01-03T19:42:32.015638Z","shell.execute_reply":"2025-01-03T19:42:32.88012Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"X=train_data.drop(columns=['Premium Amount' ])\ny=train_data['Premium Amount']","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-03T19:42:32.881991Z","iopub.execute_input":"2025-01-03T19:42:32.882327Z","iopub.status.idle":"2025-01-03T19:42:33.045155Z","shell.execute_reply.started":"2025-01-03T19:42:32.882292Z","shell.execute_reply":"2025-01-03T19:42:33.044233Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"num_col=num_col.drop(['Premium Amount'])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-03T19:42:33.046124Z","iopub.execute_input":"2025-01-03T19:42:33.046492Z","iopub.status.idle":"2025-01-03T19:42:33.051205Z","shell.execute_reply.started":"2025-01-03T19:42:33.046455Z","shell.execute_reply":"2025-01-03T19:42:33.050252Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**⚖️ Feature Scaling & Encoding**","metadata":{}},{"cell_type":"code","source":"num_pipeline = Pipeline(steps=[\n    ('imputer', SimpleImputer(strategy='median')),\n    # ('scaler', StandardScaler())                       # Scale numerical features\n])\n\n# Preprocessing pipeline for categorical features\ncat_pipeline = Pipeline(steps=[\n    ('imputer', SimpleImputer(strategy='constant', fill_value='Unknown')),  # Handle missing values\n    ('onehot', OneHotEncoder(handle_unknown='ignore'))                      # Encode categorical features\n])\npreprocessor = ColumnTransformer(\n    transformers=[\n        ('num', num_pipeline, num_col),\n        ('cat', cat_pipeline, cat_col)\n    ]\n)\nX_processed = preprocessor.fit_transform(X)\ntest_transformed = preprocessor.transform(test_data.drop(columns=['id']))\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-03T19:42:33.052198Z","iopub.execute_input":"2025-01-03T19:42:33.052505Z","iopub.status.idle":"2025-01-03T19:42:43.230491Z","shell.execute_reply.started":"2025-01-03T19:42:33.052466Z","shell.execute_reply":"2025-01-03T19:42:43.22966Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# **📐 Train-Test Split:**","metadata":{}},{"cell_type":"code","source":"X_train,X_test,y_train,y_test=train_test_split(X_processed,y,test_size=0.2,random_state=42)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-03T19:42:43.231462Z","iopub.execute_input":"2025-01-03T19:42:43.23205Z","iopub.status.idle":"2025-01-03T19:42:43.478807Z","shell.execute_reply.started":"2025-01-03T19:42:43.232013Z","shell.execute_reply":"2025-01-03T19:42:43.47788Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# **🤖 Model Training with Multiple Models**","metadata":{}},{"cell_type":"code","source":"models={'linear Regression': LinearRegression(),\n       'Decisoin Tree ': DecisionTreeRegressor(),\n       'XGBoost':XGBRegressor()\n       }\n\n#train and evaluate each model\nfor model_name,model in models.items():\n    model.fit(X_train,y_train)\n    y_pred=model.predict(X_test)\n    rmsle= np.sqrt(mean_squared_log_error(y_test,y_pred))\n    \n    print(f'{model_name} Accuracy : {rmsle}')\n    ","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-03T19:42:43.479766Z","iopub.execute_input":"2025-01-03T19:42:43.480132Z","iopub.status.idle":"2025-01-03T19:43:21.234089Z","shell.execute_reply.started":"2025-01-03T19:42:43.480095Z","shell.execute_reply":"2025-01-03T19:43:21.232183Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**XGBoost is the best model among the ones tested, as it achieves the lowest RMSLE, meaning it makes the most accurate predictions (with minimal penalty for under-predictions)**","metadata":{}},{"cell_type":"markdown","source":"# **🎛️ Hyperparameter Tuning:**","metadata":{}},{"cell_type":"code","source":"pip install optuna","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-03T19:43:21.241582Z","iopub.execute_input":"2025-01-03T19:43:21.242527Z","iopub.status.idle":"2025-01-03T19:43:26.776217Z","shell.execute_reply.started":"2025-01-03T19:43:21.24248Z","shell.execute_reply":"2025-01-03T19:43:26.775135Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from sklearn.metrics import mean_squared_log_error\nfrom sklearn.model_selection import train_test_split\n\ny_shifted = y - y.min() + 1   # This ensures all values are positive\n\ndef objective(trial):\n    # Split the data\n    X_train, X_test, y_train, y_test = train_test_split(X_processed, y_shifted, test_size=0.2, random_state=42)\n\n    # Suggest hyperparameters\n    n_estimators = trial.suggest_int('n_estimators', 50, 500)\n    max_depth = trial.suggest_int('max_depth', 3, 20)\n    learning_rate = trial.suggest_float('learning_rate', 0.01, 0.3)\n    subsample = trial.suggest_float('subsample', 0.6, 1.0)\n\n    # Train the model (e.g., XGBoost)\n    from xgboost import XGBRegressor\n    model = XGBRegressor(\n        n_estimators=n_estimators,\n        max_depth=max_depth,\n        learning_rate=learning_rate,\n        subsample=subsample,\n        random_state=42\n    )\n\n    model.fit(X_train, y_train)\n\n    # Predict and calculate RMSLE\n    y_pred = model.predict(X_test)\n    rmsle = mean_squared_log_error(y_test, y_pred, squared=False)\n    return rmsle  # Optuna minimizes this metric\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-03T19:43:26.77783Z","iopub.execute_input":"2025-01-03T19:43:26.778202Z","iopub.status.idle":"2025-01-03T19:43:26.794091Z","shell.execute_reply.started":"2025-01-03T19:43:26.778172Z","shell.execute_reply":"2025-01-03T19:43:26.793166Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import optuna\n\nstudy = optuna.create_study(direction='minimize', study_name='RMSLE_Optimization')\n\n# Set n_trials to 50 and timeout to 1800 seconds (30 minutes)\nstudy.optimize(objective, n_trials=10, timeout=600)  # 50 trials, 30 minutes max time\n# Best parameters and score\n#print(\"Best Parameters:\", study.best_params)\n#print(\"Best RMSLE:\", study.best_value)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-03T19:46:07.181887Z","iopub.execute_input":"2025-01-03T19:46:07.182261Z","iopub.status.idle":"2025-01-03T19:49:16.518104Z","shell.execute_reply.started":"2025-01-03T19:46:07.18223Z","shell.execute_reply":"2025-01-03T19:49:16.516401Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"best_parameter={'n_estimators': 480, 'max_depth': 4, 'learning_rate': 0.14750972264306628, 'subsample': 0.6858914931937488}","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-03T19:49:25.892858Z","iopub.execute_input":"2025-01-03T19:49:25.893209Z","iopub.status.idle":"2025-01-03T19:49:25.897706Z","shell.execute_reply.started":"2025-01-03T19:49:25.893182Z","shell.execute_reply":"2025-01-03T19:49:25.896497Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"best_params = study.best_params\nfinal_model = XGBRegressor(**best_params, random_state=42)\nfinal_model.fit(X_processed, y)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-03T19:49:29.276885Z","iopub.execute_input":"2025-01-03T19:49:29.277239Z","iopub.status.idle":"2025-01-03T19:49:49.566528Z","shell.execute_reply.started":"2025-01-03T19:49:29.277201Z","shell.execute_reply":"2025-01-03T19:49:49.564557Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# **📊 Model Evaluation:**","metadata":{}},{"cell_type":"code","source":"y_test = np.nan_to_num(y_test)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-03T19:49:54.206658Z","iopub.execute_input":"2025-01-03T19:49:54.207043Z","iopub.status.idle":"2025-01-03T19:49:54.213241Z","shell.execute_reply.started":"2025-01-03T19:49:54.207003Z","shell.execute_reply":"2025-01-03T19:49:54.212084Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"y_pred = final_model.predict(X_test)\nrmsle= np.sqrt(mean_squared_log_error(y_test,y_pred))\nprint(f\"RMSLE : {rmsle} \" )\n\n\nplot_importance(final_model, importance_type='gain', title='XGB Feature Importance', max_num_features=10, color='purple')\nsns.set_palette('viridis')\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-03T19:49:57.692127Z","iopub.execute_input":"2025-01-03T19:49:57.69253Z","iopub.status.idle":"2025-01-03T19:49:58.836278Z","shell.execute_reply.started":"2025-01-03T19:49:57.692479Z","shell.execute_reply":"2025-01-03T19:49:58.835217Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# **📤 Submission:**","metadata":{}},{"cell_type":"code","source":"output= pd.DataFrame(test_data['id'])\nxgb_output = final_model.predict(test_transformed)\noutput['Premium Amount']= xgb_output\noutput.to_csv(\"/kaggle/working/submission.csv\", index = None)\noutput.head(10)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-03T19:50:14.608165Z","iopub.execute_input":"2025-01-03T19:50:14.608537Z","iopub.status.idle":"2025-01-03T19:50:18.438914Z","shell.execute_reply.started":"2025-01-03T19:50:14.60851Z","shell.execute_reply":"2025-01-03T19:50:18.437917Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# **🔚 Conclusion:**","metadata":{}},{"cell_type":"markdown","source":"**After performing data preprocessing and hyperparameter tuning using Optuna, we improved the XGBoost model's RMSLE from 1.14 to 1.079.**\n\n**Initial Performance:** The model before tuning had an RMSLE of 1.14, indicating some under-predictions.\n\n**After Optuna Tuning:** Hyperparameter optimization reduced the RMSLE to 1.079, reflecting better model accuracy and more precise predictions of insurance premiums.\n\n**Dataset Insights:** The model is now better at capturing the relationships between features like age, health score, and credit score, improving premium prediction accuracy.\n\n**Overall, Optuna tuning significantly improved model performance, making the model more reliable for predicting insurance premiums.**","metadata":{}},{"cell_type":"markdown","source":"# **🚀 If you found this notebook helpful, insightful, or inspiring, I would truly appreciate your upvote! Your support can help bring this notebook to the spotlight and make it shine in the competition🏆✨\n\n# **Thank you for your time and encouragement! 😊**\n\n","metadata":{}}]}