{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":84896,"databundleVersionId":10305135,"sourceType":"competition"}],"dockerImageVersionId":30787,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<div style=\"border: 1px solid #a4b7d4; padding: 10px; background-color: #f5f7fa; color: #333a56; text-align: center; font-size: 1.5em; font-weight: bold;\">\n  WORK IN PROGRESS\n</div>","metadata":{}},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"<div style=\"border: 1px solid #a4b7d4; padding: 10px; background-color: #f5f7fa; color: #333a56; text-align: center; font-size: 1.5em; font-weight: bold;\">\n  Regression with an Insurance Dataset<br>\n  Playground Series - Season 4, Episode 12\n</div>","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2024-12-23T19:30:04.001774Z","iopub.execute_input":"2024-12-23T19:30:04.002133Z","iopub.status.idle":"2024-12-23T19:30:04.313852Z","shell.execute_reply.started":"2024-12-23T19:30:04.002101Z","shell.execute_reply":"2024-12-23T19:30:04.313005Z"},"_kg_hide-input":true,"_kg_hide-output":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import sklearn\nsklearn.__version__","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-23T19:30:55.598704Z","iopub.execute_input":"2024-12-23T19:30:55.599518Z","iopub.status.idle":"2024-12-23T19:30:56.051248Z","shell.execute_reply.started":"2024-12-23T19:30:55.599481Z","shell.execute_reply":"2024-12-23T19:30:56.050409Z"},"_kg_hide-input":true,"_kg_hide-output":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport seaborn as sns\nimport matplotlib.gridspec as gridspec\n\nimport warnings\nwarnings.filterwarnings(\"ignore\", category=UserWarning, module=\"seaborn\")\nwarnings.filterwarnings(\"ignore\", category=FutureWarning, module=\"seaborn\")\n\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.preprocessing import StandardScaler, OneHotEncoder\nfrom sklearn.compose import ColumnTransformer\nfrom sklearn.impute import SimpleImputer\nfrom sklearn.metrics import mean_squared_error, mean_absolute_error, r2_score\n\nimport torch\nfrom sklearn.pipeline import Pipeline","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-23T19:30:59.315139Z","iopub.execute_input":"2024-12-23T19:30:59.315612Z","iopub.status.idle":"2024-12-23T19:31:02.912573Z","shell.execute_reply.started":"2024-12-23T19:30:59.315576Z","shell.execute_reply":"2024-12-23T19:31:02.911637Z"},"_kg_hide-input":true,"_kg_hide-output":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Exploratory Data Analysis (EDA)\n\nExploratory Data Analysis (EDA) is a critical step in any data project. It allows us to understand, summarize, and visualize the dataset effectively, paving the way for further analysis or modeling.\n\n---\n\n### 🎯 Objectives\n\n- 🕵️‍♀️ Gain insights into the data.\n- 📊 Visualize distributions, relationships, and patterns.\n- 🧹 Identify missing values, outliers, and data inconsistencies.etection.\n\n---\n\nLet’s dive into the EDA! 🚀\n","metadata":{}},{"cell_type":"code","source":"df_eda_train = pd.read_csv('/kaggle/input/playground-series-s4e12/train.csv')\ndf_eda_test = pd.read_csv('/kaggle/input/playground-series-s4e12/test.csv')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-23T19:31:38.437502Z","iopub.execute_input":"2024-12-23T19:31:38.438623Z","iopub.status.idle":"2024-12-23T19:31:46.406367Z","shell.execute_reply.started":"2024-12-23T19:31:38.438587Z","shell.execute_reply":"2024-12-23T19:31:46.405414Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Training set ","metadata":{}},{"cell_type":"code","source":"# Check dataset shape and first rows\nprint(f\"Dataset contains {df_eda_train.shape[0]} rows and {df_eda_train.shape[1]} columns.\")\ndf_eda_train.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-23T19:32:58.802221Z","iopub.execute_input":"2024-12-23T19:32:58.802557Z","iopub.status.idle":"2024-12-23T19:32:58.825667Z","shell.execute_reply.started":"2024-12-23T19:32:58.802528Z","shell.execute_reply":"2024-12-23T19:32:58.82461Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df_eda_train.describe().round(3)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-23T19:33:03.704176Z","iopub.execute_input":"2024-12-23T19:33:03.704795Z","iopub.status.idle":"2024-12-23T19:33:04.276357Z","shell.execute_reply.started":"2024-12-23T19:33:03.70476Z","shell.execute_reply":"2024-12-23T19:33:04.275438Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df_eda_train.info()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-23T19:33:04.277836Z","iopub.execute_input":"2024-12-23T19:33:04.278163Z","iopub.status.idle":"2024-12-23T19:33:04.835781Z","shell.execute_reply.started":"2024-12-23T19:33:04.278132Z","shell.execute_reply":"2024-12-23T19:33:04.834869Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"numerical_columns = df_eda_train.select_dtypes(exclude=['object']).columns","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-23T19:33:07.797544Z","iopub.execute_input":"2024-12-23T19:33:07.798232Z","iopub.status.idle":"2024-12-23T19:33:07.836175Z","shell.execute_reply.started":"2024-12-23T19:33:07.798193Z","shell.execute_reply":"2024-12-23T19:33:07.835266Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"missing_values = df_eda_train.isnull().sum()\nmissing_values","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-23T19:33:08.107086Z","iopub.execute_input":"2024-12-23T19:33:08.107381Z","iopub.status.idle":"2024-12-23T19:33:08.649556Z","shell.execute_reply.started":"2024-12-23T19:33:08.107355Z","shell.execute_reply":"2024-12-23T19:33:08.64862Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"total_rows = len(df_eda_train)\nmissing_percentage = (missing_values / total_rows) * 100\nmissing_percentage.round(2)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-23T19:33:10.636341Z","iopub.execute_input":"2024-12-23T19:33:10.636675Z","iopub.status.idle":"2024-12-23T19:33:10.64482Z","shell.execute_reply.started":"2024-12-23T19:33:10.636645Z","shell.execute_reply":"2024-12-23T19:33:10.643772Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"for column in df_eda_train.columns:\n    if df_eda_train[column].dtype == 'O': \n        size = df_eda_train[column].nunique()\n        print(f\"Column '{column}' has {size} unique elements.\")\n        if size <= 5:\n            print(f\"\\t Unique elements: \\n\\t{df_eda_train[column].unique()}\")\n            value_counts = df_eda_train[column].value_counts(normalize=True)  # Mit normalize=True in Prozent\n            print(\"\\t Value distribution (counts and percentages):\")\n            for value, proportion in value_counts.items():\n                count = df_eda_train[column].value_counts()[value]\n                print(f\"\\t\\t{value}: {count} ({proportion:.2%})\")\n        print()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-23T19:46:29.562282Z","iopub.execute_input":"2024-12-23T19:46:29.562848Z","iopub.status.idle":"2024-12-23T19:46:33.795266Z","shell.execute_reply.started":"2024-12-23T19:46:29.562814Z","shell.execute_reply":"2024-12-23T19:46:33.794073Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plt.figure(figsize=(15,9))\nplt.title(\"Visualizing Missing Values\")\nsns.heatmap(df_eda_train.isnull(), cbar=False, cmap=sns.color_palette('magma'), yticklabels=False);\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-23T18:18:37.274456Z","iopub.execute_input":"2024-12-23T18:18:37.274685Z","iopub.status.idle":"2024-12-23T18:18:57.616466Z","shell.execute_reply.started":"2024-12-23T18:18:37.274662Z","shell.execute_reply":"2024-12-23T18:18:57.615578Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Calculate the correlation matrix\ncorrelation_matrix = df_eda_train[numerical_columns].corr()\n\n# Plot the heatmap\nplt.figure(figsize=(12, 8))\nsns.heatmap(correlation_matrix, annot=True, fmt=\".2f\", cmap=\"crest\", cbar=True, linewidths=0.5)\nplt.title(\"Correlation Heatmap of Numerical Variables\", fontsize=16)\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-23T18:18:57.617562Z","iopub.execute_input":"2024-12-23T18:18:57.61786Z","iopub.status.idle":"2024-12-23T18:18:58.499917Z","shell.execute_reply.started":"2024-12-23T18:18:57.617813Z","shell.execute_reply":"2024-12-23T18:18:58.499118Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Test set","metadata":{}},{"cell_type":"code","source":"# Check dataset shape and first rows\nprint(f\"Dataset contains {df_eda_test.shape[0]} rows and {df_eda_test.shape[1]} columns.\")\ndf_eda_test.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-23T18:18:58.504495Z","iopub.execute_input":"2024-12-23T18:18:58.504902Z","iopub.status.idle":"2024-12-23T18:18:58.523458Z","shell.execute_reply.started":"2024-12-23T18:18:58.504858Z","shell.execute_reply":"2024-12-23T18:18:58.522663Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df_eda_test.describe().round(3)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-23T18:18:58.524824Z","iopub.execute_input":"2024-12-23T18:18:58.525179Z","iopub.status.idle":"2024-12-23T18:18:58.864606Z","shell.execute_reply.started":"2024-12-23T18:18:58.525142Z","shell.execute_reply":"2024-12-23T18:18:58.86374Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df_eda_test.info()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-23T18:18:58.865727Z","iopub.execute_input":"2024-12-23T18:18:58.866011Z","iopub.status.idle":"2024-12-23T18:18:59.234758Z","shell.execute_reply.started":"2024-12-23T18:18:58.865985Z","shell.execute_reply":"2024-12-23T18:18:59.233862Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"missing_values = df_eda_test.isnull().sum()\nmissing_values","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-23T18:18:59.235804Z","iopub.execute_input":"2024-12-23T18:18:59.236115Z","iopub.status.idle":"2024-12-23T18:18:59.594634Z","shell.execute_reply.started":"2024-12-23T18:18:59.236088Z","shell.execute_reply":"2024-12-23T18:18:59.593791Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"total_rows = len(df_eda_test)\nmissing_percentage = (missing_values / total_rows) * 100\nmissing_percentage.round(2)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-23T18:18:59.595635Z","iopub.execute_input":"2024-12-23T18:18:59.595921Z","iopub.status.idle":"2024-12-23T18:18:59.603301Z","shell.execute_reply.started":"2024-12-23T18:18:59.595895Z","shell.execute_reply":"2024-12-23T18:18:59.602208Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"for column in df_eda_test.columns:\n    if df_eda_test[column].dtype == 'O': \n        size = df_eda_test[column].nunique()\n        print(f\"Column '{column}' has {size} unique elements.\")\n        if size <= 5:\n            print(f\"\\t Unique elements: \\n\\t{df_eda_test[column].unique()}\")\n        print()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-23T18:18:59.604434Z","iopub.execute_input":"2024-12-23T18:18:59.604691Z","iopub.status.idle":"2024-12-23T18:19:00.436711Z","shell.execute_reply.started":"2024-12-23T18:18:59.604667Z","shell.execute_reply":"2024-12-23T18:19:00.435816Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Data Preprocessing\n\nData preprocessing is a crucial step in preparing the dataset for analysis and modeling. It ensures the data is clean, consistent, and ready for machine learning algorithms.\n\n---\n\n### 🔍 Objectives\n\n- Handle missing values.\n- Encode categorical features.\n- Standardize or normalize numerical features.\n- Create new features or transform existing ones if necessary.","metadata":{}},{"cell_type":"code","source":"train_df = pd.read_csv('/kaggle/input/playground-series-s4e12/train.csv')\ntest_df = pd.read_csv('/kaggle/input/playground-series-s4e12/test.csv')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-23T18:19:00.43792Z","iopub.execute_input":"2024-12-23T18:19:00.438557Z","iopub.status.idle":"2024-12-23T18:19:06.303238Z","shell.execute_reply.started":"2024-12-23T18:19:00.438518Z","shell.execute_reply":"2024-12-23T18:19:06.302458Z"},"_kg_hide-input":true,"_kg_hide-output":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"- **Binary coding:** For binary columns (`gender`, `smoking status`).\n- **Single-factor coding:** For nominal categories without order (`location`, `marital status`).\n- **Ordinal coding:** For ordered categories (`education level`, `customer feedback`).\n- **Date:** Extracts important time components from date fields.\n- **Missing values:** Treatment with `fillna`.","metadata":{}},{"cell_type":"markdown","source":"**Binary Endcoding**\n- Gender: `Female` -> 0, `Male` -> 1\n- Smoking Status: `No` -> 0, `Yes` -> 1","metadata":{}},{"cell_type":"code","source":"train_df['Gender'] = train_df['Gender'].map({'Female': 0, 'Male': 1})\ntrain_df['Smoking Status'] = train_df['Smoking Status'].map({'No': 0, 'Yes': 1})","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-23T18:19:06.304282Z","iopub.execute_input":"2024-12-23T18:19:06.304631Z","iopub.status.idle":"2024-12-23T18:19:06.44939Z","shell.execute_reply.started":"2024-12-23T18:19:06.304596Z","shell.execute_reply":"2024-12-23T18:19:06.448673Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**One-Hot Encoding**\n- Marital Status\n- Location","metadata":{}},{"cell_type":"code","source":"train_df = pd.get_dummies(train_df, columns=['Location'], prefix='Location')\ntrain_df = pd.get_dummies(train_df, columns=['Marital Status'], prefix='Marital Status')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-23T18:19:06.450245Z","iopub.execute_input":"2024-12-23T18:19:06.450494Z","iopub.status.idle":"2024-12-23T18:19:07.552527Z","shell.execute_reply.started":"2024-12-23T18:19:06.450469Z","shell.execute_reply":"2024-12-23T18:19:07.551786Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Ordinal Encoding**\n- Education Level:\n    - `High School` -> 0, `Bachelor's` -> 1, `Master's` -> 2, `PhD` -> 3\n- Customer Feedback:\n    - `Poor` -> 0, `Average` -> 1, `Good` -> 2\n","metadata":{}},{"cell_type":"code","source":"education_mapping = {'High School': 0, \"Bachelor's\": 1, \"Master's\": 2, 'PhD': 3}\ntrain_df['Education Level'] = train_df['Education Level'].map(education_mapping)\n\nfeedback_mapping = {'Poor': 0, 'Average': 1, 'Good': 2}\ntrain_df['Customer Feedback'] = train_df['Customer Feedback'].map(feedback_mapping)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-23T18:19:07.555323Z","iopub.execute_input":"2024-12-23T18:19:07.555608Z","iopub.status.idle":"2024-12-23T18:19:07.708632Z","shell.execute_reply.started":"2024-12-23T18:19:07.555581Z","shell.execute_reply":"2024-12-23T18:19:07.707715Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Handling Missing Values**","metadata":{}},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Date Features**","metadata":{}},{"cell_type":"code","source":"def date(df):\n\n    df['Policy Start Date'] = pd.to_datetime(df['Policy Start Date'])\n    df['Year'] = df['Policy Start Date'].dt.year\n    df['Day'] = df['Policy Start Date'].dt.day\n    df['Month'] = df['Policy Start Date'].dt.month\n    df['Month_name'] = df['Policy Start Date'].dt.month_name()\n    df['Day_of_week'] = df['Policy Start Date'].dt.day_name()\n    df['Week'] = df['Policy Start Date'].dt.isocalendar().week\n    df['Year_sin'] = np.sin(2 * np.pi * df['Year'])\n    df['Year_cos'] = np.cos(2 * np.pi * df['Year'])\n    min_year = df['Year'].min()\n    max_year = df['Year'].max()\n    df['Year_sin'] = np.sin(2 * np.pi * (df['Year'] - min_year) / (max_year - min_year))\n    df['Year_cos'] = np.cos(2 * np.pi * (df['Year'] - min_year) / (max_year - min_year))\n    df['Month_sin'] = np.sin(2 * np.pi * df['Month'] / 12) \n    df['Month_cos'] = np.cos(2 * np.pi * df['Month'] / 12)\n    df['Day_sin'] = np.sin(2 * np.pi * df['Day'] / 31)  \n    df['Day_cos'] = np.cos(2 * np.pi * df['Day'] / 31)\n    df['Group']=(df['Year']-2020)*48+df['Month']*4+df['Day']//7\n    \n    df.drop('Policy Start Date', axis=1, inplace=True)\n\n    return df","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-23T18:19:07.709779Z","iopub.execute_input":"2024-12-23T18:19:07.710136Z","iopub.status.idle":"2024-12-23T18:19:07.71874Z","shell.execute_reply.started":"2024-12-23T18:19:07.7101Z","shell.execute_reply":"2024-12-23T18:19:07.717877Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_df = date(train_df)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-23T18:19:07.719755Z","iopub.execute_input":"2024-12-23T18:19:07.720057Z","iopub.status.idle":"2024-12-23T18:19:09.348505Z","shell.execute_reply.started":"2024-12-23T18:19:07.720011Z","shell.execute_reply":"2024-12-23T18:19:09.347493Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Direct Mapping for Small Categories**\n- Policy Type: `Basic` -> 0, `Comprehensive` -> 1, `Premium` -> 2\n- Exercise Frequency: `Rarely` -> 0, `Monthly` -> 1, `Weekly` -> 2, `Daily` -> 3","metadata":{}},{"cell_type":"code","source":"policy_mapping = {'Basic': 0, 'Comprehensive': 1, 'Premium': 2}\nexercise_mapping = {'Rarely': 0, 'Monthly': 1, 'Weekly': 2, 'Daily': 3}\ntrain_df['Policy Type'] = train_df['Policy Type'].map(policy_mapping)\ntrain_df['Exercise Frequency'] = train_df['Exercise Frequency'].map(exercise_mapping)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-23T18:19:09.349745Z","iopub.execute_input":"2024-12-23T18:19:09.350129Z","iopub.status.idle":"2024-12-23T18:19:09.490991Z","shell.execute_reply.started":"2024-12-23T18:19:09.350091Z","shell.execute_reply":"2024-12-23T18:19:09.490334Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_df.info()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-23T18:19:09.491915Z","iopub.execute_input":"2024-12-23T18:19:09.492168Z","iopub.status.idle":"2024-12-23T18:19:09.736411Z","shell.execute_reply.started":"2024-12-23T18:19:09.492144Z","shell.execute_reply":"2024-12-23T18:19:09.735594Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Define features and target\nnumerical_features = [\n    'Age', 'Annual Income', 'Number of Dependents', 'Health Score', \n    'Previous Claims', 'Vehicle Age', 'Credit Score', 'Insurance Duration', \n    'Year_sin', 'Year_cos', 'Month_sin', 'Month_cos', 'Day_sin', 'Day_cos'\n]\ncategorical_features = [\n    'Gender', 'Marital Status', 'Education Level', 'Occupation', 'Location',\n    'Policy Type', 'Customer Feedback', 'Smoking Status', 'Exercise Frequency', \n    'Property Type', 'Month_name', 'Day_of_week'\n]\ntarget_column = 'Premium Amount'","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-23T18:19:09.737686Z","iopub.execute_input":"2024-12-23T18:19:09.738044Z","iopub.status.idle":"2024-12-23T18:19:09.742766Z","shell.execute_reply.started":"2024-12-23T18:19:09.738007Z","shell.execute_reply":"2024-12-23T18:19:09.741981Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Split train data into features and target\nX = train_df.drop(columns=[target_column, 'id', 'Group', 'Year', 'Month', 'Day', 'Week'])\ny = train_df[target_column]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-23T18:19:09.744339Z","iopub.execute_input":"2024-12-23T18:19:09.744657Z","iopub.status.idle":"2024-12-23T18:19:09.872886Z","shell.execute_reply.started":"2024-12-23T18:19:09.744623Z","shell.execute_reply":"2024-12-23T18:19:09.871955Z"},"_kg_hide-input":false},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(f\"{X.shape = }\")\nprint(f\"{y.shape = }\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-23T18:19:09.874095Z","iopub.execute_input":"2024-12-23T18:19:09.874401Z","iopub.status.idle":"2024-12-23T18:19:09.879259Z","shell.execute_reply.started":"2024-12-23T18:19:09.874376Z","shell.execute_reply":"2024-12-23T18:19:09.878353Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Model Training\n\nTraining the model is the core step in any machine learning pipeline. Here, we use the processed features and target variable to fit a predictive model and evaluate its performance on a validation set.\n\n---\n\n### 🔍 Objectives\n- Train the model using the training dataset (`X_train`, `y_train`).\n- Evaluate the model on the validation set (`X_val`, `y_val`).\n- Optimize the model's parameters to improve its performance.\n\n---\n","metadata":{}},{"cell_type":"markdown","source":"before KNN imputer :\n\n___\n\nPerformance Metrics:\n- RMSLE: \n- RMSE: \n- MAE: \n- R²: \n- MAPE: ","metadata":{}},{"cell_type":"markdown","source":"# Generate Predictions & Prepare Submission\n\nFinally, we use the trained model to predict outcomes on the test set and format the results into a submission file for the competition.","metadata":{}}]}