{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":84896,"databundleVersionId":10305135,"sourceType":"competition"}],"dockerImageVersionId":30786,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2024-12-03T14:01:22.796147Z","iopub.execute_input":"2024-12-03T14:01:22.796532Z","iopub.status.idle":"2024-12-03T14:01:23.811325Z","shell.execute_reply.started":"2024-12-03T14:01:22.796497Z","shell.execute_reply":"2024-12-03T14:01:23.810355Z"},"_kg_hide-input":true,"_kg_hide-output":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"!pip install -q scikit-learn==1.5.2","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-03T14:01:23.813076Z","iopub.execute_input":"2024-12-03T14:01:23.813487Z","iopub.status.idle":"2024-12-03T14:01:38.561892Z","shell.execute_reply.started":"2024-12-03T14:01:23.813455Z","shell.execute_reply":"2024-12-03T14:01:38.560894Z"},"_kg_hide-input":true,"_kg_hide-output":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import sklearn\nsklearn.__version__","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-03T14:01:38.563403Z","iopub.execute_input":"2024-12-03T14:01:38.563766Z","iopub.status.idle":"2024-12-03T14:01:39.609679Z","shell.execute_reply.started":"2024-12-03T14:01:38.563732Z","shell.execute_reply":"2024-12-03T14:01:39.608583Z"},"_kg_hide-input":true,"_kg_hide-output":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport seaborn as sns\nimport matplotlib.gridspec as gridspec\n\nimport warnings\nwarnings.filterwarnings(\"ignore\", category=UserWarning, module=\"seaborn\")\nwarnings.filterwarnings(\"ignore\", category=FutureWarning, module=\"seaborn\")\n\nimport numpy as np\nimport pandas as pd\nimport tensorflow as tf\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.preprocessing import StandardScaler, OneHotEncoder\nfrom sklearn.compose import ColumnTransformer\nfrom sklearn.impute import SimpleImputer\nfrom sklearn.metrics import root_mean_squared_log_error,mean_squared_error, mean_absolute_error, r2_score\n\nimport optuna\nimport lightgbm as lgb\n\nimport torch\nfrom sklearn.pipeline import Pipeline","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-03T14:01:39.611709Z","iopub.execute_input":"2024-12-03T14:01:39.612242Z","iopub.status.idle":"2024-12-03T14:01:58.085561Z","shell.execute_reply.started":"2024-12-03T14:01:39.61221Z","shell.execute_reply":"2024-12-03T14:01:58.084758Z"},"_kg_hide-input":true,"_kg_hide-output":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_df=pd.read_csv('/kaggle/input/playground-series-s4e12/train.csv')\ntest_df=pd.read_csv('/kaggle/input/playground-series-s4e12/test.csv')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-03T14:01:58.08675Z","iopub.execute_input":"2024-12-03T14:01:58.088364Z","iopub.status.idle":"2024-12-03T14:02:07.46566Z","shell.execute_reply.started":"2024-12-03T14:01:58.08831Z","shell.execute_reply":"2024-12-03T14:02:07.464625Z"},"_kg_hide-input":true,"_kg_hide-output":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Exploratory Data Analysis (EDA)\n\nExploratory Data Analysis (EDA) is a critical step in any data project. It allows us to understand, summarize, and visualize the dataset effectively, paving the way for further analysis or modeling.\n\n---\n\n### 🎯 Objectives\n\n- 🕵️‍♀️ Gain insights into the data.\n- 📊 Visualize distributions, relationships, and patterns.\n- 🧹 Identify missing values, outliers, and data inconsistencies.etection.\n\n---\n\nLet’s dive into the EDA! 🚀\n","metadata":{}},{"cell_type":"markdown","source":"## 📜 Dataset Overview","metadata":{}},{"cell_type":"code","source":"# Check dataset shape and first rows\nprint(f\"Dataset contains {train_df.shape[0]} rows and {train_df.shape[1]} columns.\")\ntrain_df.head()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-03T14:02:07.467033Z","iopub.execute_input":"2024-12-03T14:02:07.468715Z","iopub.status.idle":"2024-12-03T14:02:07.510491Z","shell.execute_reply.started":"2024-12-03T14:02:07.46866Z","shell.execute_reply":"2024-12-03T14:02:07.509493Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_df.info()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-03T14:02:07.511782Z","iopub.execute_input":"2024-12-03T14:02:07.512045Z","iopub.status.idle":"2024-12-03T14:02:08.128912Z","shell.execute_reply.started":"2024-12-03T14:02:07.512019Z","shell.execute_reply":"2024-12-03T14:02:08.12781Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Save 'id' column for submission\ntest_ids = test_df['id']\n\n# Define the target column\ntarget_column = 'Premium Amount'\n\n# Select categorical and numerical columns (initial)\ncategorical_columns = train_df.select_dtypes(include=['object']).columns\nnumerical_columns = train_df.select_dtypes(exclude=['object']).columns\n\n# Print out column information\nprint(\"Target Column:\", target_column)\nprint(\"\\nCategorical Columns:\", categorical_columns.tolist())\nprint(\"\\nNumerical Columns:\", numerical_columns.tolist())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-03T14:02:08.130037Z","iopub.execute_input":"2024-12-03T14:02:08.130434Z","iopub.status.idle":"2024-12-03T14:02:08.245665Z","shell.execute_reply.started":"2024-12-03T14:02:08.130379Z","shell.execute_reply":"2024-12-03T14:02:08.244508Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 📊 Descriptive Statistics","metadata":{}},{"cell_type":"code","source":"train_df.describe().round(2)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-03T14:02:08.246757Z","iopub.execute_input":"2024-12-03T14:02:08.247045Z","iopub.status.idle":"2024-12-03T14:02:08.893545Z","shell.execute_reply.started":"2024-12-03T14:02:08.247013Z","shell.execute_reply":"2024-12-03T14:02:08.892545Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"for column in categorical_columns:\n    num_unique = train_df[column].nunique()\n    print(f\"'{column}' has {num_unique} unique categories.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-03T14:02:08.896628Z","iopub.execute_input":"2024-12-03T14:02:08.896956Z","iopub.status.idle":"2024-12-03T14:02:09.780769Z","shell.execute_reply.started":"2024-12-03T14:02:08.896926Z","shell.execute_reply":"2024-12-03T14:02:09.779664Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Print top 10 unique value counts for each categorical column\nfor column in categorical_columns:\n    print(f\"\\nTop value counts in '{column}':\\n{train_df[column].value_counts().head(10)}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-03T14:02:09.781969Z","iopub.execute_input":"2024-12-03T14:02:09.782298Z","iopub.status.idle":"2024-12-03T14:02:11.053466Z","shell.execute_reply.started":"2024-12-03T14:02:09.782266Z","shell.execute_reply":"2024-12-03T14:02:11.052396Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(\"The mean of columns:\")\nprint(train_df[numerical_columns].mean())\n\nprint(\"\\nThe std dev of columns:\")\nprint(train_df[numerical_columns].std())\n\nprint(\"\\nThe skewness of columns:\")\nprint(train_df[numerical_columns].skew())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-03T14:02:11.055395Z","iopub.execute_input":"2024-12-03T14:02:11.055882Z","iopub.status.idle":"2024-12-03T14:02:11.529505Z","shell.execute_reply.started":"2024-12-03T14:02:11.055832Z","shell.execute_reply":"2024-12-03T14:02:11.528405Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🧹 Data Cleaning Insights","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(15,9))\nplt.title(\"Visualizing Missing Values\")\nsns.heatmap(train_df.isnull(), cbar=False, cmap=sns.color_palette('magma'), yticklabels=False);\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-03T14:02:11.530865Z","iopub.execute_input":"2024-12-03T14:02:11.531219Z","iopub.status.idle":"2024-12-03T14:02:34.157819Z","shell.execute_reply.started":"2024-12-03T14:02:11.531175Z","shell.execute_reply":"2024-12-03T14:02:34.156787Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🖼️ Visual Exploration","metadata":{}},{"cell_type":"code","source":"# Create a color palette for the columns\npalette = sns.color_palette('tab10', len(numerical_columns))\ncolor_dict = dict(zip(numerical_columns, palette))\n\n# Create a grid of subplots for histograms, boxplots, and scatterplots/violin plots\nfig = plt.figure(figsize=(30, 10 * len(numerical_columns)))\ngs = gridspec.GridSpec(2 * len(numerical_columns), 2, figure=fig)\n\ndf_binned = train_df.copy()\n\nfor i, column in enumerate(numerical_columns):\n\n    if train_df[column].nunique() > 50: discrete = False\n    else : discrete = True\n    \n    # Plot histogram with a unique color\n    ax_hist = fig.add_subplot(gs[2 * i, 0])\n    sns.histplot(\n        data=train_df, x=column, fill=True, common_norm=False, alpha=0.6,\n        linewidth=0.8, color=color_dict[column], ax=ax_hist,  discrete = discrete\n    )\n    \n    # Plot boxplot with the same unique color\n    ax_box = fig.add_subplot(gs[2 * i + 1, 0])\n    sns.boxplot(data=train_df, x=column, ax=ax_box, color=color_dict[column])\n    ax_box.set_title(f'{column} vs Target (Boxplot)', fontsize=14)\n    sns.despine(ax=ax_box)\n\n    # Conditional plot: violin plot or barplot based on unique values, fallback to scatterplot\n    ax_conditional = fig.add_subplot(gs[2 * i:2 * i + 2, 1])  # Merges 2 rows\n    if train_df[column].nunique() <= 10:\n        # If the column has 10 or fewer unique values, use a violin plot\n        sns.violinplot(data=train_df, x=column, y=target_column, ax=ax_conditional, color=color_dict[column], alpha=0.6)\n        ax_conditional.set_title(f'{column} vs {target_column} (Violin Plot)', fontsize=14)\n    else:\n        # Bin the column into 10 intervals, but keep original target column values\n        df_binned['Binned Column'] = pd.cut(train_df[column], bins=10)\n        sns.violinplot(data=df_binned, x='Binned Column', y=target_column, ax=ax_conditional, color=color_dict[column], alpha=0.6)\n        ax_conditional.set_title(f'{column} (Binned) vs {target_column} (Violin Plot)', fontsize=14)\n        ax_conditional.set_xlabel(f'{column} (Binned)', fontsize=12)\n\nplt.tight_layout()  # Adjust subplots to fit into the figure area\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-03T14:02:34.15926Z","iopub.execute_input":"2024-12-03T14:02:34.159751Z","iopub.status.idle":"2024-12-03T14:03:16.592302Z","shell.execute_reply.started":"2024-12-03T14:02:34.159695Z","shell.execute_reply":"2024-12-03T14:03:16.590861Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Filtrer les colonnes catégorielles et exclure 'Policy Start Date'\nfiltered_columns = [col for col in categorical_columns if col != 'Policy Start Date']\n\n# Créer des sous-graphiques pour barplots et boxplots\nfig, axes = plt.subplots(len(filtered_columns), 2, figsize=(15, 5 * len(filtered_columns)))\n\nfor i, column in enumerate(filtered_columns):\n    # Barplot à gauche\n    sns.countplot(data=train_df, x=column, ax=axes[i, 0], palette='tab10')\n    axes[i, 0].set_title(f'Distribution of {column}', fontsize=14)\n    axes[i, 0].set_xlabel(column, fontsize=12)\n    axes[i, 0].set_ylabel('Count', fontsize=12)\n    sns.despine(ax=axes[i, 0])\n\n    # Boxplot à droite\n    sns.boxplot(data=train_df, x=column, y=target_column, ax=axes[i, 1], palette='tab10')\n    axes[i, 1].set_title(f'{column} vs {target_column}', fontsize=14)\n    axes[i, 1].set_xlabel(column, fontsize=12)\n    axes[i, 1].set_ylabel(target_column, fontsize=12)\n    sns.despine(ax=axes[i, 1])\n\nplt.tight_layout()  # Ajustement global des sous-graphiques\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-03T14:03:16.593917Z","iopub.execute_input":"2024-12-03T14:03:16.594232Z","iopub.status.idle":"2024-12-03T14:03:32.038554Z","shell.execute_reply.started":"2024-12-03T14:03:16.594202Z","shell.execute_reply":"2024-12-03T14:03:32.03764Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Calculate the correlation matrix\ncorrelation_matrix = train_df[numerical_columns].corr()\n\n# Plot the heatmap\nplt.figure(figsize=(12, 8))\nsns.heatmap(correlation_matrix, annot=True, fmt=\".2f\", cmap=\"coolwarm\", cbar=True, linewidths=0.5)\nplt.title(\"Correlation Heatmap of Numerical Variables\", fontsize=16)\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-03T14:03:32.039628Z","iopub.execute_input":"2024-12-03T14:03:32.039881Z","iopub.status.idle":"2024-12-03T14:03:33.065446Z","shell.execute_reply.started":"2024-12-03T14:03:32.039855Z","shell.execute_reply":"2024-12-03T14:03:33.064496Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 🏁 Next Steps\n\n- Dive deeper into specific features of interest.\n- Address missing values and.\n- Prepare the dataset for modeling.","metadata":{}},{"cell_type":"markdown","source":"# Data Preprocessing\n\nData preprocessing is a crucial step in preparing the dataset for analysis and modeling. It ensures the data is clean, consistent, and ready for machine learning algorithms.\n\n---\n\n### 🔍 Objectives\n\n- Handle missing values.\n- Encode categorical features.\n- Standardize or normalize numerical features.\n- Create new features or transform existing ones if necessary.","metadata":{}},{"cell_type":"markdown","source":"## 🗓️ Feature Transformation: Date Handling\n\nIn this step, we transform the `Policy Start Date` feature into multiple useful date-related features. This helps to extract temporal patterns and improve the model's ability to learn from the data.","metadata":{}},{"cell_type":"code","source":"def date(df):\n\n    df['Policy Start Date'] = pd.to_datetime(df['Policy Start Date'])\n    df['Year'] = df['Policy Start Date'].dt.year\n    df['Day'] = df['Policy Start Date'].dt.day\n    df['Month'] = df['Policy Start Date'].dt.month\n    df['Month_name'] = df['Policy Start Date'].dt.month_name()\n    df['Day_of_week'] = df['Policy Start Date'].dt.day_name()\n    df['Week'] = df['Policy Start Date'].dt.isocalendar().week\n    df['Year_sin'] = np.sin(2 * np.pi * df['Year'])\n    df['Year_cos'] = np.cos(2 * np.pi * df['Year'])\n    min_year = df['Year'].min()\n    max_year = df['Year'].max()\n    df['Year_sin'] = np.sin(2 * np.pi * (df['Year'] - min_year) / (max_year - min_year))\n    df['Year_cos'] = np.cos(2 * np.pi * (df['Year'] - min_year) / (max_year - min_year))\n    df['Month_sin'] = np.sin(2 * np.pi * df['Month'] / 12) \n    df['Month_cos'] = np.cos(2 * np.pi * df['Month'] / 12)\n    df['Day_sin'] = np.sin(2 * np.pi * df['Day'] / 31)  \n    df['Day_cos'] = np.cos(2 * np.pi * df['Day'] / 31)\n    df['Group']=(df['Year']-2020)*48+df['Month']*4+df['Day']//7\n    \n    df.drop('Policy Start Date', axis=1, inplace=True)\n\n    return df\n\n# Apply the date function to both datasets\ntrain_df = date(train_df)\ntest_df = date(test_df)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-03T14:03:33.066969Z","iopub.execute_input":"2024-12-03T14:03:33.067379Z","iopub.status.idle":"2024-12-03T14:03:36.057829Z","shell.execute_reply.started":"2024-12-03T14:03:33.067334Z","shell.execute_reply":"2024-12-03T14:03:36.056693Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Define features and target\nnumerical_features = [\n    'Age', 'Annual Income', 'Number of Dependents', 'Health Score', \n    'Previous Claims', 'Vehicle Age', 'Credit Score', 'Insurance Duration', \n    'Year_sin', 'Year_cos', 'Month_sin', 'Month_cos', 'Day_sin', 'Day_cos'\n]\ncategorical_features = [\n    'Gender', 'Marital Status', 'Education Level', 'Occupation', 'Location',\n    'Policy Type', 'Customer Feedback', 'Smoking Status', 'Exercise Frequency', \n    'Property Type', 'Month_name', 'Day_of_week'\n]\ntarget_column = 'Premium Amount'","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-03T14:03:36.059223Z","iopub.execute_input":"2024-12-03T14:03:36.05967Z","iopub.status.idle":"2024-12-03T14:03:36.065915Z","shell.execute_reply.started":"2024-12-03T14:03:36.059626Z","shell.execute_reply":"2024-12-03T14:03:36.064655Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## ✂️ Splitting Data: Features and Target\n\nIn this step, we split the training dataset into:\n- **Features (`X`)**: The independent variables used to predict the target.\n- **Target (`y`)**: The dependent variable that the model will learn to predict.","metadata":{}},{"cell_type":"code","source":"# Split train data into features and target\nX = train_df.drop(columns=[target_column, 'id', 'Group', 'Year', 'Month', 'Day', 'Week'])\ny = train_df[target_column]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-03T14:03:36.067129Z","iopub.execute_input":"2024-12-03T14:03:36.067447Z","iopub.status.idle":"2024-12-03T14:03:36.288189Z","shell.execute_reply.started":"2024-12-03T14:03:36.067417Z","shell.execute_reply":"2024-12-03T14:03:36.287Z"},"_kg_hide-input":false},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🛠️ Handle Missing Values & Preprocessing Pipeline\n\nIn this step, we handle missing values and set up a preprocessing pipeline to prepare the data for machine learning.\n\n---\n\n### 🔍 What We Did\n\n1. **Missing Values Imputation**:\n   - **Numerical Features**:\n     - Replaced missing values with the **mean** of the respective columns.\n   - **Categorical Features**:\n     - Replaced missing values with the constant value **\"Unknown\"**.\n\n2. **Feature Scaling & Encoding**:\n   - **Numerical Features**:\n     - Standardized using `StandardScaler` to normalize the values.\n   - **Categorical Features**:\n     - One-Hot Encoded to handle categorical variables as numerical inputs.\n\n3. **Combined Using a ColumnTransformer**:\n   - Applied preprocessing selectively to numerical and categorical features using a single, unified pipeline.","metadata":{}},{"cell_type":"code","source":"# Preprocessing pipeline for numerical features\nnum_pipeline = Pipeline(steps=[\n    ('imputer', SimpleImputer(strategy='median')),\n    #('scaler', StandardScaler())                       # Scale numerical features\n])\n\n# Preprocessing pipeline for categorical features\ncat_pipeline = Pipeline(steps=[\n    ('imputer', SimpleImputer(strategy='constant', fill_value='Unknown')),  # Handle missing values\n    ('onehot', OneHotEncoder(handle_unknown='ignore'))                      # Encode categorical features\n])\n\n# Combine pipelines into a ColumnTransformer\npreprocessor = ColumnTransformer(\n    transformers=[\n        ('num', num_pipeline, numerical_features),\n        ('cat', cat_pipeline, categorical_features)\n    ]\n)\n\n# Preprocess train and test data\nX_processed = preprocessor.fit_transform(X)\ntest_processed = preprocessor.transform(test_df.drop(columns=['id', 'Group', 'Year', 'Month', 'Day', 'Week']))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-03T14:03:36.289559Z","iopub.execute_input":"2024-12-03T14:03:36.2899Z","iopub.status.idle":"2024-12-03T14:03:48.618983Z","shell.execute_reply.started":"2024-12-03T14:03:36.289868Z","shell.execute_reply":"2024-12-03T14:03:48.617845Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## ✂️ Splitting Data: Training and Validation Sets\n\nHere, we split the preprocessed data into training and validation sets to evaluate the model's performance during training.","metadata":{}},{"cell_type":"code","source":"# Split the data\nX_train, X_val, y_train, y_val = train_test_split(X_processed, y, test_size=0.2, random_state=42)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-03T14:03:48.620164Z","iopub.execute_input":"2024-12-03T14:03:48.620431Z","iopub.status.idle":"2024-12-03T14:03:48.931446Z","shell.execute_reply.started":"2024-12-03T14:03:48.620404Z","shell.execute_reply":"2024-12-03T14:03:48.93064Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Model Training\n\nTraining the model is the core step in any machine learning pipeline. Here, we use the processed features and target variable to fit a predictive model and evaluate its performance on a validation set.\n\n---\n\n### 🔍 Objectives\n- Train the model using the training dataset (`X_train`, `y_train`).\n- Evaluate the model on the validation set (`X_val`, `y_val`).\n- Optimize the model's parameters to improve its performance.\n\n---\n","metadata":{}},{"cell_type":"markdown","source":"## 🔧 Hyperparameter Optimization with Optuna\n\nOptuna is a powerful library for hyperparameter optimization. In this step, we use Optuna to fine-tune the hyperparameters of a LightGBM model to achieve optimal performance.","metadata":{}},{"cell_type":"code","source":"# Define Optuna optimization function\ndef objective(trial):\n    # Define parameter search space\n    param = {\n        \"objective\": \"regression\",\n        \"metric\": \"rmse\",\n        \"boosting_type\": trial.suggest_categorical(\"boosting_type\", [\"gbdt\", \"dart\"]),\n        \"num_leaves\": trial.suggest_int(\"num_leaves\", 200, 512),\n        \"learning_rate\": trial.suggest_loguniform(\"learning_rate\", 1e-4, 1e-1),\n        \"feature_fraction\": trial.suggest_uniform(\"feature_fraction\", 0.6, 1.0),\n        \"bagging_fraction\": trial.suggest_uniform(\"bagging_fraction\", 0.6, 1.0),\n        \"bagging_freq\": trial.suggest_int(\"bagging_freq\", 5, 12),\n        \"min_data_in_leaf\": trial.suggest_int(\"min_data_in_leaf\", 20, 100),\n        \"max_depth\": trial.suggest_int(\"max_depth\", -1, 16),  # -1 means no limit\n        \"lambda_l1\": trial.suggest_loguniform(\"lambda_l1\", 1e-4, 10.0),\n        \"lambda_l2\": trial.suggest_loguniform(\"lambda_l2\", 1e-4, 10.0),\n        \"device_type\": \"gpu\",  # Enable GPU support\n        \"seed\" : 42\n\n    }\n\n    # Create a LightGBM dataset\n    dtrain = lgb.Dataset(X_train, label=y_train)\n    dval = lgb.Dataset(X_val, label=y_val, reference=dtrain)\n\n    # Train LightGBM model\n    model = lgb.train(\n        param,\n        dtrain,\n        valid_sets=[dval],\n    )\n\n    # Predict on validation set\n    y_val_pred = model.predict(X_val)\n    \n    # Compute RMSLE using sklearn's root_mean_squared_log_error\n    rmsle = root_mean_squared_log_error(y_val, np.maximum(y_val_pred, 0))\n    return rmsle\n\n# Run Optuna study\nstudy = optuna.create_study(direction=\"minimize\")\nstudy.optimize(objective, n_trials=1)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-03T14:03:48.932623Z","iopub.execute_input":"2024-12-03T14:03:48.932932Z","iopub.status.idle":"2024-12-03T14:03:58.880275Z","shell.execute_reply.started":"2024-12-03T14:03:48.932902Z","shell.execute_reply":"2024-12-03T14:03:58.879159Z"},"_kg_hide-output":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 🔧 Best Hyperparameters\n\nAfter running the hyperparameter optimization process with Optuna, the best combination of parameters was identified. These parameters will be used to train the final LightGBM model.","metadata":{}},{"cell_type":"code","source":"# Initialize or update the best_params dictionary\nbest_params = {\n    'boosting_type': 'dart',\n    'num_leaves': 384,\n    'learning_rate': 0.024680120465142227,\n    'feature_fraction': 0.9883068358315126,\n    'bagging_fraction': 0.7201712704805496,\n    'bagging_freq': 7,\n    'min_data_in_leaf': 50,\n    'max_depth': 15,\n    'lambda_l1': 0.0011290211269753322,\n    'lambda_l2': 3.056310541294088,\n    'seed': 42\n}","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-03T14:03:58.881543Z","iopub.execute_input":"2024-12-03T14:03:58.881903Z","iopub.status.idle":"2024-12-03T14:03:58.888932Z","shell.execute_reply.started":"2024-12-03T14:03:58.881867Z","shell.execute_reply":"2024-12-03T14:03:58.886654Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🚀 Train Final Model with Best Parameters\n\nFinally, we train the final LightGBM model using the optimal hyperparameters identified during the hyperparameter optimization process. This ensures the model is trained with the most effective configuration for achieving the best performance.","metadata":{}},{"cell_type":"code","source":"# Train final model with best parameters\n#best_params = study.best_params\n\nfinal_model = lgb.train(\n    best_params,\n    lgb.Dataset(X_processed, label=y),\n)","metadata":{"trusted":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-12-03T14:03:58.890214Z","iopub.execute_input":"2024-12-03T14:03:58.890548Z","iopub.status.idle":"2024-12-03T14:04:50.178314Z","shell.execute_reply.started":"2024-12-03T14:03:58.890511Z","shell.execute_reply":"2024-12-03T14:04:50.177384Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 📊 Model Evaluation\n\nIn this section, we evaluate the performance of the trained model using multiple metrics and visualizations. This helps us understand how well the model performs and identify areas for improvement.\n\n---\n\n### 🔍 Performance Metrics\nWe compute the following metrics to evaluate the model:\n- **RMSLE (Root Mean Squared Logarithmic Error)**: Measures the ratio-based prediction error, suitable for skewed datasets.\n- **RMSE (Root Mean Squared Error)**: Quantifies the average magnitude of prediction errors.\n- **MAE (Mean Absolute Error)**: Represents the average absolute difference between predicted and actual values.\n- **R² (Coefficient of Determination)**: Indicates the proportion of variance in the target explained by the model.\n- **MAPE (Mean Absolute Percentage Error)**: Expresses prediction errors as a percentage of actual values.","metadata":{}},{"cell_type":"code","source":"# 1. Performance Metrics\ny_pred = final_model.predict(X_processed)\n\n# Calcul des métriques\nrmsle = root_mean_squared_log_error(y, y_pred)\nrmse = np.sqrt(mean_squared_error(y, y_pred))\nmae = mean_absolute_error(y, y_pred)\nr2 = r2_score(y, y_pred)\nmape = np.mean(np.abs((y - y_pred) / y)) * 100\n\n# Display performance metrics\nprint(f\"\\nPerformance Metrics:\\n{'-'*30}\")\nprint(f\"RMSLE: {rmsle:.4f}\")\nprint(f\"RMSE: {rmse:.4f}\")\nprint(f\"MAE: {mae:.4f}\")\nprint(f\"R²: {r2:.4f}\")\nprint(f\"MAPE: {mape:.2f}%\")\n\n# 2. Feature Importance\nimportances = final_model.feature_importance(importance_type='split')  # or 'gain'\nfeatures = preprocessor.get_feature_names_out()\nsorted_indices = importances.argsort()[::-1]\n\n# Create a DataFrame for feature importances\nimportance_df = pd.DataFrame({\n    'Feature': [features[i] for i in sorted_indices],\n    'Importance': importances[sorted_indices]\n})\n\n# Plot top 10 feature importances\nplt.figure(figsize=(12, 6))\nsns.barplot(data=importance_df.head(10), x='Importance', y='Feature', palette=\"coolwarm\")\nplt.title(\"Top 10 Feature Importances\", fontsize=16, fontweight='bold')\nplt.xlabel(\"Feature Importance\", fontsize=12)\nplt.ylabel(\"Feature\", fontsize=12)\nplt.tight_layout()\nplt.show()\n\n# 3. Residual Analysis\nresiduals = y - y_pred\n\n# Residuals vs Predicted Values\nplt.figure(figsize=(12, 6))\nsns.scatterplot(x=y_pred, y=residuals, alpha=0.6, color=\"#007acc\")\nplt.axhline(y=0, color='red', linestyle='--', linewidth=1.5)\nplt.title(\"Residuals vs Predicted Values\", fontsize=16, fontweight='bold')\nplt.xlabel(\"Predicted Values\", fontsize=12)\nplt.ylabel(\"Residuals\", fontsize=12)\nplt.tight_layout()\nplt.show()\n\n# Residual Distribution\nplt.figure(figsize=(10, 6))\nsns.histplot(residuals, bins=30, kde=True, color=\"#55a630\")\nplt.axvline(x=0, color='red', linestyle='--', linewidth=1.5)\nplt.title(\"Distribution of Residuals\", fontsize=16, fontweight='bold')\nplt.xlabel(\"Residuals\", fontsize=12)\nplt.ylabel(\"Frequency\", fontsize=12)\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-03T14:04:50.179898Z","iopub.execute_input":"2024-12-03T14:04:50.18031Z","iopub.status.idle":"2024-12-03T14:05:05.260514Z","shell.execute_reply.started":"2024-12-03T14:04:50.180248Z","shell.execute_reply":"2024-12-03T14:05:05.259508Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"before KNN imputer :\n\n___\n\nPerformance Metrics:\n- RMSLE: 1.0634\n- RMSE: 906.9466\n- MAE: 625.4777\n- R²: -0.0993\n- MAPE: 204.97%","metadata":{}},{"cell_type":"markdown","source":"# Generate Predictions & Prepare Submission\n\nFinally, we use the trained model to predict outcomes on the test set and format the results into a submission file for the competition.","metadata":{}},{"cell_type":"code","source":"# Make predictions on the test set\ntest_predictions = final_model.predict(test_processed, num_iteration=final_model.best_iteration)\n\n# Prepare submission file\nsubmission = pd.DataFrame({'id': test_df['id'], 'Premium Amount': test_predictions})\nsubmission.to_csv(\"submission.csv\", index=False)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-03T14:05:05.261966Z","iopub.execute_input":"2024-12-03T14:05:05.262385Z","iopub.status.idle":"2024-12-03T14:05:10.834989Z","shell.execute_reply.started":"2024-12-03T14:05:05.26234Z","shell.execute_reply":"2024-12-03T14:05:10.834143Z"},"_kg_hide-output":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Conclusion\n\nFinally, we have completed the end-to-end process of building and evaluating a machine learning model. This journey included data exploration, preprocessing, model training, hyperparameter optimization, and preparing the final submission file.\n\n---\n\n### 🔑 Key Highlights\n\n1. **Exploratory Data Analysis (EDA)**:\n   - Gained valuable insights into the dataset through visualization and descriptive statistics.\n   - Identified key patterns, correlations, and potential data quality issues.\n\n2. **Data Preprocessing**:\n   - Handled missing values, encoded categorical features, and standardized numerical features.\n   - Engineered new features, especially time-based features, to improve the model's predictive capabilities.\n\n3. **Model Training and Evaluation**:\n   - Trained a LightGBM model with optimized hyperparameters using Optuna.\n   - Evaluated the model's performance using metrics such as RMSE, RMSLE, and \\( R^2 \\).\n   - Performed residual analysis to validate the model's behavior and assumptions.\n\n4. **Final Submission**:\n   - Used the trained model to predict outcomes on the test dataset.\n   - Prepared and saved the submission file in the required format.\n\n---\n\n### 🚀 Next Steps\n\n- Explore additional feature engineering opportunities to further enhance model performance.\n- Experiment with ensemble methods or other advanced algorithms for potential improvements.\n- Analyze and fine-tune the model using feedback from the competition leaderboard.\n\n---\n\nThank you for exploring this notebook!\n","metadata":{}}]}