{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":84896,"databundleVersionId":10305135,"sourceType":"competition"}],"dockerImageVersionId":30805,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<a id=\"1\"></a>\n# <div style=\"text-align:center; border-radius:15px 50px; padding:7px; color:white; margin:0; font-size:110%; font-family:Pacifico; background-color:#0073e6; overflow:hidden\"><b>Machine learning - Regression insurance</b></div>","metadata":{}},{"cell_type":"markdown","source":"<a id=\"1\"></a>\n<div style=\"text-align:center; border-radius:15px 50px; padding:7px; color:white; margin:0; font-size:110%; font-family:Pacifico; background-color:#0073e6; overflow:hidden\"><b>Part 1 - Business Problem</b></div>\n-\r\n\r\n### Business Problem\r\n\r\n**Objective:**  \r\nAn insurance company wants to better predict the premiums they should charge their customers. The challenge lies in accurately estimating the **Premium Amount** for potential policyholders based on a variety of factors. These factors might include demographic information, policy details, and past claims history.\r\n\r\n**Motivation:**  \r\nAccurate premium predictions are crucial for the company’s profitability and competitiveness. Underpricing policies can lead to financial losses, while overpricing can result in losing customers to competitors. By leveraging machine learning models to predict premiums, the company aims to:  \r\n1. **Optimize Revenue:** Ensure premiums reflect the risk level of individual customers.  \r\n2. **Improve Customer Retention:** Offer fair pricing to attract and retain customers.  \r\n3. **Enhance Efficiency:** Automate premium calculations to reduce manual errors and improve processing speed.\r\n\r\n**Challenge:**  \r\nYou are tasked with developing a regression model that predicts the **Premium Amount** for each customer in the dataset. The prediction accuracy will be evaluated using the **Root Mean Squared Logarithmic Error (RMSLE)** metric.\r\n\r\n**Impact:**  \r\nIf successful, your model could help the insurance company:  \r\n- Gain a competitive edge through better pricing strategies.  \r\n- Enhance customer satisfaction by offering tailored and fair premiums.  \r\n- Improve their decision-making process by understanding the key drivers of real-world business goals.","metadata":{}},{"cell_type":"code","source":"# Import of libraries\n\n# System libraries\nimport re\nimport string\nimport unicodedata\nimport itertools\nfrom collections import Counter\n\n# Library for file manipulation\nimport pandas as pd\nimport numpy as np\nimport pandas\n\n# Data visualization\nimport seaborn as sns\nimport matplotlib.pylab as pl\nimport matplotlib as m\nimport matplotlib as mpl\nimport matplotlib.pyplot as plt\nimport plotly.express as px\nfrom matplotlib import pyplot as plt\n\n# Configuration for graph width and layout\nsns.set_theme(style='whitegrid')\npalette='viridis'\n\n# Warnings remove alerts\nimport warnings\nwarnings.filterwarnings(\"ignore\")\n\n# Python version\nfrom platform import python_version\nprint('Python version in this Jupyter Notebook:', python_version())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-10T16:53:04.335564Z","iopub.execute_input":"2024-12-10T16:53:04.335972Z","iopub.status.idle":"2024-12-10T16:53:04.342494Z","shell.execute_reply.started":"2024-12-10T16:53:04.335945Z","shell.execute_reply":"2024-12-10T16:53:04.341716Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <div style=\"text-align:center; border-radius:15px 50px; padding:7px; color:white; margin:0; font-size:110%; font-family:Pacifico; background-color:#0073e6; overflow:hidden\"><b> Part 2 - Database</b></div>","metadata":{}},{"cell_type":"code","source":"# Database\ntrain_df = pd.read_csv(\"/kaggle/input/playground-series-s4e12/train.csv\")\ntest_df = pd.read_csv(\"/kaggle/input/playground-series-s4e12/test.csv\")\n\n# Viewing dataset\ntrain_df","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-10T16:53:04.348141Z","iopub.execute_input":"2024-12-10T16:53:04.34873Z","iopub.status.idle":"2024-12-10T16:53:13.709392Z","shell.execute_reply.started":"2024-12-10T16:53:04.348691Z","shell.execute_reply":"2024-12-10T16:53:13.708491Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Viewing first 5 data\ntrain_df.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-10T16:53:13.71047Z","iopub.execute_input":"2024-12-10T16:53:13.710764Z","iopub.status.idle":"2024-12-10T16:53:13.730395Z","shell.execute_reply.started":"2024-12-10T16:53:13.710738Z","shell.execute_reply":"2024-12-10T16:53:13.729396Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Viewing 5 latest data\ntrain_df.tail()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-10T16:53:13.731855Z","iopub.execute_input":"2024-12-10T16:53:13.732355Z","iopub.status.idle":"2024-12-10T16:53:13.758051Z","shell.execute_reply.started":"2024-12-10T16:53:13.73228Z","shell.execute_reply":"2024-12-10T16:53:13.757151Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Info data\ntrain_df.info()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-10T16:53:13.759169Z","iopub.execute_input":"2024-12-10T16:53:13.759503Z","iopub.status.idle":"2024-12-10T16:53:14.316585Z","shell.execute_reply.started":"2024-12-10T16:53:13.759466Z","shell.execute_reply":"2024-12-10T16:53:14.315631Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Type data\ntrain_df.dtypes","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-10T16:53:14.317653Z","iopub.execute_input":"2024-12-10T16:53:14.317947Z","iopub.status.idle":"2024-12-10T16:53:14.32508Z","shell.execute_reply.started":"2024-12-10T16:53:14.317919Z","shell.execute_reply":"2024-12-10T16:53:14.324159Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Viewing rows and columns\ntrain_df.shape","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-10T16:53:14.326314Z","iopub.execute_input":"2024-12-10T16:53:14.326877Z","iopub.status.idle":"2024-12-10T16:53:14.336464Z","shell.execute_reply.started":"2024-12-10T16:53:14.326837Z","shell.execute_reply":"2024-12-10T16:53:14.335526Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <div style=\"text-align:center; border-radius:15px 50px; padding:7px; color:white; margin:0; font-size:110%; font-family:Pacifico; background-color:#0073e6; overflow:hidden\"><b> Part 3 - Data cleaning</b></div>","metadata":{}},{"cell_type":"code","source":"print(\"Checking for missing values in each column:\")\nprint(train_df.isnull().sum())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-10T16:53:14.338979Z","iopub.execute_input":"2024-12-10T16:53:14.339265Z","iopub.status.idle":"2024-12-10T16:53:14.881814Z","shell.execute_reply.started":"2024-12-10T16:53:14.339224Z","shell.execute_reply":"2024-12-10T16:53:14.880995Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Drop 'id' column\ntrain_df = train_df.drop(['id'], axis=1)\ntrain_df","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-10T16:53:14.882895Z","iopub.execute_input":"2024-12-10T16:53:14.883239Z","iopub.status.idle":"2024-12-10T16:53:15.06065Z","shell.execute_reply.started":"2024-12-10T16:53:14.88321Z","shell.execute_reply":"2024-12-10T16:53:15.059822Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Checking the number of null values ​​in specific columns\nprint(train_df[['Age', 'Annual Income', 'Marital Status', \n                'Number of Dependents', 'Occupation', 'Health Score',\n                'Previous Claims','Credit Score','Insurance Duration',\n                'Customer Feedback', 'Vehicle Age']].isnull().sum())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-10T16:53:15.061884Z","iopub.execute_input":"2024-12-10T16:53:15.062252Z","iopub.status.idle":"2024-12-10T16:53:15.272014Z","shell.execute_reply.started":"2024-12-10T16:53:15.062223Z","shell.execute_reply":"2024-12-10T16:53:15.271115Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def remove_outliers(df, columns, min_rows=10):\n    \"\"\"\n    Removes outliers from the specified numeric columns using the IQR method,\n    ensuring that at least `min_rows` remain in the DataFrame.\n    \n    Parameters:\n    df (DataFrame): The DataFrame from which outliers will be removed.\n    columns (list): A list of column names where outliers should be removed.\n    min_rows (int): The minimum number of rows that should remain in the DataFrame \n                    after removing outliers. Default is 10.\n    \n    Returns:\n    DataFrame: A DataFrame with outliers removed from the specified columns.\n    \"\"\"\n    \n    for col in columns:\n        # Attempt to convert the column to numeric, ignoring errors\n        df[col] = pd.to_numeric(df[col], errors='coerce')\n        \n        # Fill NaN values with the median of the column (or another appropriate method)\n        df[col].fillna(df[col].median(), inplace=True)\n        \n        # Calculate the first and third quartiles and the interquartile range (IQR)\n        Q1 = df[col].quantile(0.25)\n        Q3 = df[col].quantile(0.75)\n        IQR = Q3 - Q1\n        lower_bound = Q1 - 1.5 * IQR\n        upper_bound = Q3 + 1.5 * IQR\n        \n        # Apply the outlier filter only if it leaves at least `min_rows` in the DataFrame\n        filtered_df = df[(df[col] >= lower_bound) & (df[col] <= upper_bound)]\n        if filtered_df.shape[0] >= min_rows:\n            df = filtered_df\n        else:\n            # Print a message if the filter would remove too many rows\n            print(f\"Column: {col} - Unable to apply outlier filter without removing all rows.\")\n    \n    return df\n\n# Applying the function to the train_df DataFrame\nnumeric_columns = ['Age', 'Annual Income', 'Marital Status', \n                   'Number of Dependents', 'Occupation', 'Health Score',\n                   'Previous Claims','Credit Score','Vehicle Age',\n                   'Insurance Duration',\n                   'Customer Feedback']\n\n# Remove outliers\ntrain_df = remove_outliers(train_df, numeric_columns)\ntrain_df.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-10T16:53:15.273075Z","iopub.execute_input":"2024-12-10T16:53:15.273333Z","iopub.status.idle":"2024-12-10T16:53:19.897824Z","shell.execute_reply.started":"2024-12-10T16:53:15.273307Z","shell.execute_reply":"2024-12-10T16:53:19.897047Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def drop_missing(df, columns):\n    \"\"\"Remove rows that contain missing values in the specified columns.\"\"\"\n    \n    df = df.dropna(subset=columns)\n    \n    return df\n\n# Applying the function to train_df\ncolumns_to_check = ['Annual Income', 'Previous Claims', 'Credit Score']\ntrain_df = drop_missing(train_df, columns_to_check)\n\n# Drop 'id' column\ntrain_df = train_df.drop(['Marital Status', 'Occupation', 'Customer Feedback'], axis=1)\ntrain_df","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-10T16:53:19.899055Z","iopub.execute_input":"2024-12-10T16:53:19.899691Z","iopub.status.idle":"2024-12-10T16:53:20.522867Z","shell.execute_reply.started":"2024-12-10T16:53:19.899632Z","shell.execute_reply":"2024-12-10T16:53:20.522051Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(\"Checking for missing values in each column:\")\nprint(train_df.isnull().sum())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-10T16:53:20.524106Z","iopub.execute_input":"2024-12-10T16:53:20.524467Z","iopub.status.idle":"2024-12-10T16:53:20.880612Z","shell.execute_reply.started":"2024-12-10T16:53:20.524428Z","shell.execute_reply":"2024-12-10T16:53:20.879787Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <div style=\"text-align:center; border-radius:15px 50px; padding:7px; color:white; margin:0; font-size:110%; font-family:Pacifico; background-color:#0073e6; overflow:hidden\"><b> Part 4 - Exploratory data analysis</b></div>","metadata":{}},{"cell_type":"code","source":"# Viewing descriptive statistics\ntrain_df.describe().T","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-10T16:53:20.88166Z","iopub.execute_input":"2024-12-10T16:53:20.881949Z","iopub.status.idle":"2024-12-10T16:53:21.249277Z","shell.execute_reply.started":"2024-12-10T16:53:20.881924Z","shell.execute_reply":"2024-12-10T16:53:21.2484Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Setting up the seaborn style\nsns.set(style=\"whitegrid\")\n\n# Histogram for the 'price' variable\nplt.figure(figsize=(12, 7))\nsns.histplot(train_df['Premium Amount'], kde=True, color='skyblue', bins=25, alpha=0.7, line_kws={'linewidth': 2, 'color': 'blue'})\nplt.title('Distribution Premium Amount', fontsize=16)\nplt.xlabel('Price', fontsize=14)\nplt.ylabel('Frequency', fontsize=14)\n\n# Adding mean and median lines\nmean_price = train_df['Premium Amount'].mean()\nmedian_price = train_df['Premium Amount'].median()\nplt.axvline(mean_price, color='red', linestyle='--', linewidth=2, label=f'Mean: ${mean_price:.2f}')\nplt.axvline(median_price, color='green', linestyle='-', linewidth=2, label=f'Median: ${median_price:.2f}')\n\n# Adding a legend\nplt.legend()\nplt.grid(False)\n\n# Displaying the plot\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-10T16:53:21.250732Z","iopub.execute_input":"2024-12-10T16:53:21.251123Z","iopub.status.idle":"2024-12-10T16:53:25.866286Z","shell.execute_reply.started":"2024-12-10T16:53:21.251079Z","shell.execute_reply":"2024-12-10T16:53:25.865443Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Correlation\ncorrelation_matrix = train_df.corr(numeric_only=True)\n\n# Creating a correlation heatmap for numerical features\nplt.figure(figsize=(10, 6))\nsns.heatmap(correlation_matrix, annot=True, fmt=\".2f\", cmap=\"coolwarm\", cbar=True)\nplt.title(\"Correlation Heatmap of Numerical Features\", fontsize=14)\nplt.xticks(rotation=45)\nplt.yticks(rotation=0)\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-10T16:53:25.86744Z","iopub.execute_input":"2024-12-10T16:53:25.86817Z","iopub.status.idle":"2024-12-10T16:53:26.614595Z","shell.execute_reply.started":"2024-12-10T16:53:25.868126Z","shell.execute_reply":"2024-12-10T16:53:26.613809Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Histogram for Age Distribution\nplt.figure(figsize=(8, 5))\nsns.histplot(data=train_df, x='Age', kde=True, color='blue', bins=20)\nplt.title(\"Age Distribution\", fontsize=14)\nplt.xlabel(\"Age\")\nplt.ylabel(\"Frequency\")\n\n# Disabling the grid lines for a cleaner plot\nplt.grid(False)\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-10T17:48:16.692241Z","iopub.execute_input":"2024-12-10T17:48:16.692604Z","iopub.status.idle":"2024-12-10T17:48:20.32481Z","shell.execute_reply.started":"2024-12-10T17:48:16.692571Z","shell.execute_reply":"2024-12-10T17:48:20.323785Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# histplot\nsns.histplot(train_df[\"Premium Amount\"])\n\n# Adding a legend\nplt.legend()\nplt.grid(False)\n\n# Displaying the plot\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-10T16:53:30.450013Z","iopub.execute_input":"2024-12-10T16:53:30.450733Z","iopub.status.idle":"2024-12-10T16:53:31.843186Z","shell.execute_reply.started":"2024-12-10T16:53:30.450687Z","shell.execute_reply":"2024-12-10T16:53:31.842264Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# barplot Premium\nsns.barplot(x=\"Policy Type\", y=\"Premium Amount\", data=train_df)\n\n# Adding a legend\nplt.legend()\nplt.grid(False)\n\n# Displaying the plot\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-10T16:53:31.844438Z","iopub.execute_input":"2024-12-10T16:53:31.844824Z","iopub.status.idle":"2024-12-10T16:53:39.331871Z","shell.execute_reply.started":"2024-12-10T16:53:31.844784Z","shell.execute_reply":"2024-12-10T16:53:39.330709Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Create a bar plot using Seaborn to show the average 'Premium Amount' across different 'Gender' categories\nsns.barplot(x=\"Gender\", y=\"Premium Amount\", data=train_df)\n\n# Adding a legend to the plot for better clarity\nplt.legend()\n\n# Disabling the grid lines for a cleaner plot\nplt.grid(False)\n\n# Displaying the plot\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-10T16:53:39.333097Z","iopub.execute_input":"2024-12-10T16:53:39.333509Z","iopub.status.idle":"2024-12-10T16:53:47.115921Z","shell.execute_reply.started":"2024-12-10T16:53:39.333458Z","shell.execute_reply":"2024-12-10T16:53:47.115053Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Create a boxplot to visualize the distribution of 'Premium Amount' across different 'Gender' categories\nplt.figure(figsize=(8, 5))\nsns.boxplot(data=train_df, x='Gender', y='Premium Amount', palette='Set3')\nplt.title(\"Premium Amount by Gender\", fontsize=14)\nplt.xlabel(\"Gender\")\nplt.ylabel(\"Premium Amount\")\n\n# Disabling the grid lines for a cleaner plot\nplt.grid(False)\n\n# Displaying the plot\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-10T17:49:22.897495Z","iopub.execute_input":"2024-12-10T17:49:22.897874Z","iopub.status.idle":"2024-12-10T17:49:23.244942Z","shell.execute_reply.started":"2024-12-10T17:49:22.897839Z","shell.execute_reply":"2024-12-10T17:49:23.243949Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Create a histogram to show the distribution of 'Premium Amount' in the dataset\nplt.figure(figsize=(8, 5))\nsns.histplot(data=train_df, x='Premium Amount', kde=True, bins=30, color='green')\nplt.title(\"Premium Amount Distribution\", fontsize=14)\nplt.xlabel(\"Premium Amount\")\nplt.ylabel(\"Frequency\")\n\n# Disabling the grid lines for a cleaner plot\nplt.grid(False)\n\n# Displaying the plot\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-10T17:49:33.993139Z","iopub.execute_input":"2024-12-10T17:49:33.993851Z","iopub.status.idle":"2024-12-10T17:49:37.818565Z","shell.execute_reply.started":"2024-12-10T17:49:33.993815Z","shell.execute_reply":"2024-12-10T17:49:37.817641Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Boxplot of Premiums by Policy Type\nplt.figure(figsize=(10, 6))\nsns.boxplot(data=train_df, x='Policy Type', y='Premium Amount', palette='Set2')\nplt.title(\"Premium Amount by Policy Type\", fontsize=14)\nplt.xlabel(\"Policy Type\")\nplt.ylabel(\"Premium Amount\")\n\n# Disabling the grid lines for a cleaner plot\nplt.grid(False)\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-10T17:49:58.276048Z","iopub.execute_input":"2024-12-10T17:49:58.276387Z","iopub.status.idle":"2024-12-10T17:49:58.855705Z","shell.execute_reply.started":"2024-12-10T17:49:58.276354Z","shell.execute_reply":"2024-12-10T17:49:58.854812Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Distribution of prizes by educational level\nplt.figure(figsize=(10, 6))\nsns.boxplot(data=train_df, x='Education Level', y='Premium Amount', palette='muted')\nplt.title(\"Premium Amount by Education Level\", fontsize=14)\nplt.xlabel(\"Education Level\")\nplt.ylabel(\"Premium Amount\")\n\n# Disabling the grid lines for a cleaner plot\nplt.grid(False)\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-10T17:52:50.895818Z","iopub.execute_input":"2024-12-10T17:52:50.89619Z","iopub.status.idle":"2024-12-10T17:52:51.204224Z","shell.execute_reply.started":"2024-12-10T17:52:50.89615Z","shell.execute_reply":"2024-12-10T17:52:51.203198Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Analysis - Credit Score**\r\n\r\n- This variable impacts insurance premiums and other customer characteristics.","metadata":{}},{"cell_type":"code","source":"# Credit Score Distribution\nplt.figure(figsize=(18, 5))\nsns.histplot(data=train_df, x='Credit Score', kde=True, bins=30, color='blue')\nplt.title(\"Distribution of Credit Score\", fontsize=14)\nplt.xlabel(\"Credit Score\")\nplt.ylabel(\"Frequency\")\n\n# Disabling the grid lines for a cleaner plot\nplt.grid(False)\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-10T17:51:00.934165Z","iopub.execute_input":"2024-12-10T17:51:00.934512Z","iopub.status.idle":"2024-12-10T17:51:04.6475Z","shell.execute_reply.started":"2024-12-10T17:51:00.934481Z","shell.execute_reply":"2024-12-10T17:51:04.64665Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Categorizing Credit Score into bins\nbins = [0, 300, 600, 700, 850]\nlabels = ['Very Low', 'Low', 'Medium', 'High']\ntrain_df['Credit Score Group'] = pd.cut(train_df['Credit Score'], bins=bins, labels=labels)\n\n# Boxplot of Premium Amount by Credit Score Group\nplt.figure(figsize=(10, 6))\nsns.boxplot(data=train_df, x='Credit Score Group', y='Premium Amount', palette='Set3')\nplt.title(\"Premium Amount by Credit Score Group\", fontsize=14)\nplt.xlabel(\"Credit Score Group\")\nplt.ylabel(\"Premium Amount\")\n\n# Disabling the grid lines for a cleaner plot\nplt.grid(False)\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-10T17:51:27.790151Z","iopub.execute_input":"2024-12-10T17:51:27.790501Z","iopub.status.idle":"2024-12-10T17:51:28.169725Z","shell.execute_reply.started":"2024-12-10T17:51:27.790471Z","shell.execute_reply":"2024-12-10T17:51:28.168902Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Boxplot of Previous Claims by Credit Score Group\nplt.figure(figsize=(10, 6))\nsns.boxplot(data=train_df, x='Credit Score Group', y='Previous Claims', palette='Set3')\nplt.title(\"Previous Claims by Credit Score Group\", fontsize=14)\nplt.xlabel(\"Credit Score Group\")\nplt.ylabel(\"Previous Claims\")\n\n# Disabling the grid lines for a cleaner plot\nplt.grid(False)\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-10T17:51:17.214498Z","iopub.execute_input":"2024-12-10T17:51:17.214859Z","iopub.status.idle":"2024-12-10T17:51:17.540374Z","shell.execute_reply.started":"2024-12-10T17:51:17.214828Z","shell.execute_reply":"2024-12-10T17:51:17.53939Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <div style=\"text-align:center; border-radius:15px 50px; padding:7px; color:white; margin:0; font-size:110%; font-family:Pacifico; background-color:#0073e6; overflow:hidden\"><b> Part 5 - Outlier removal</b></div>\n\n- In this chart, we can observe the presence of outliers in the target column. To ensure the quality and accuracy of the analysis, these outliers will be removed. Removing outliers is a crucial step in data preprocessing as they can distort results and negatively impact predictive models.\n\n- This removal will be performed using appropriate statistical methods, such as the interquartile range (IQR) or z-score analysis, ensuring that only the most representative data is considered in the subsequent analysis.","metadata":{}},{"cell_type":"code","source":"# Create a boxplot to show the distribution of 'Premium Amount' across different 'Education Levels'\nplt.figure(figsize=(10, 6))\nsns.boxplot(data=train_df, x='Premium Amount', y='Education Level', palette='Set2')\nplt.title(\"Distribution of Premium Amount by Education Level\", fontsize=16)\nplt.xlabel(\"Premium Amount\")\nplt.ylabel(\"Education Level\")\n\n# Disabling the grid lines for a cleaner plot\nplt.grid(False)\n\n# Displaying the plot\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-10T17:52:06.860457Z","iopub.execute_input":"2024-12-10T17:52:06.860826Z","iopub.status.idle":"2024-12-10T17:52:50.893866Z","shell.execute_reply.started":"2024-12-10T17:52:06.860795Z","shell.execute_reply":"2024-12-10T17:52:50.892991Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Create a boxplot to show the distribution of 'Premium Amount'\nplt.figure(figsize=(8, 5))\nsns.boxplot(train_df[\"Premium Amount\"], palette='Set2')\nplt.title(\"Distribution of Premium Amount\", fontsize=14)\nplt.xlabel(\"Premium Amount\")\n\n# Disabling the grid lines for a cleaner plot\nplt.grid(False)\n\n# Displaying the plot\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-10T17:52:51.205592Z","iopub.execute_input":"2024-12-10T17:52:51.20601Z","iopub.status.idle":"2024-12-10T17:52:51.417846Z","shell.execute_reply.started":"2024-12-10T17:52:51.205965Z","shell.execute_reply":"2024-12-10T17:52:51.416985Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Setting up the figure and axes for subplots with tight layout\nfig, axes = plt.subplots(1, 2, figsize=(18, 6))  # 1 row, 3 columns\n\n# First boxplot - Fuel Type vs Price\nsns.boxplot(x=\"Policy Type\", y=\"Premium Amount\", data=train_df, palette=\"Set2\", ax=axes[0])\naxes[0].set_title('Price Distribution by Fuel Type', fontsize=14)\naxes[0].set_xlabel('Fuel Type', fontsize=12)\naxes[0].set_ylabel('Price (USD)', fontsize=12)\naxes[0].tick_params(axis='x', rotation=45)\n\n# Third boxplot - Model Year vs Price\nsns.boxplot(x=\"Policy Type\", y=\"Premium Amount\", data=train_df, palette=\"Set2\", ax=axes[1])\naxes[1].set_title('Price Distribution by Model Year', fontsize=14)\naxes[1].set_xlabel('Model Year', fontsize=12)\naxes[1].set_ylabel('Price (USD)', fontsize=12)\naxes[1].tick_params(axis='x', rotation=45)\n\n# Automatically adjust subplot parameters for better fit\nplt.tight_layout()\n\n# Display the plots\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-10T16:54:00.704607Z","iopub.execute_input":"2024-12-10T16:54:00.704902Z","iopub.status.idle":"2024-12-10T16:54:02.051481Z","shell.execute_reply.started":"2024-12-10T16:54:00.704876Z","shell.execute_reply":"2024-12-10T16:54:02.050592Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"### Outlier removal\n\n# Assuming 'train_df' is the DataFrame where you want to remove the outliers\nQ1 = train_df['Premium Amount'].quantile(0.25)  # Calculate the first quartile (25th percentile) of the 'price' column\nQ3 = train_df['Premium Amount'].quantile(0.75)  # Calculate the third quartile (75th percentile) of the 'price' column\nIQR = Q3 - Q1  # Calculate the interquartile range (IQR) for 'price'\n\n# Define the boundaries to consider a value as an outlier\nlower_bound = Q1 - 1.5 * IQR  # Calculate the lower bound (below which values will be considered outliers)\nupper_bound = Q3 + 1.5 * IQR  # Calculate the upper bound (above which values will be considered outliers)\n\n# Filter out the outliers from the 'price' column\ntrain_df = train_df[(train_df['Premium Amount'] >= lower_bound) & (train_df['Premium Amount'] <= upper_bound)]  # Keep only the data points within the bounds\ntrain_df.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-10T16:54:02.052636Z","iopub.execute_input":"2024-12-10T16:54:02.05292Z","iopub.status.idle":"2024-12-10T16:54:02.180053Z","shell.execute_reply.started":"2024-12-10T16:54:02.052892Z","shell.execute_reply":"2024-12-10T16:54:02.179114Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"fig, axes = plt.subplots(2, 1, figsize=(12, 10))\n\n# Boxplot\nsns.boxplot(x=train_df['Premium Amount'], ax=axes[0], palette=\"Set2\")\naxes[0].set_title('Boxplot of Price - After Outlier Removal', fontsize=16)\naxes[0].set_xlabel('Price (USD)', fontsize=14)\naxes[0].grid(True, axis='y', linestyle='--', alpha=0.7)  # Adding grid lines for better readability\n\n# Histogram\nsns.histplot(train_df['Premium Amount'], bins=40, kde=True, ax=axes[1], palette=\"Set2\")  # Consistent color\naxes[1].set_title('Histogram of Price - After Outlier Removal', fontsize=16)\naxes[1].set_xlabel('Price (USD)', fontsize=14)\naxes[1].set_ylabel('Frequency', fontsize=14)\naxes[1].grid(True, axis='y', linestyle='--', alpha=0.7)  # Adding grid lines\n\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-10T16:54:02.181107Z","iopub.execute_input":"2024-12-10T16:54:02.181395Z","iopub.status.idle":"2024-12-10T16:54:06.715064Z","shell.execute_reply.started":"2024-12-10T16:54:02.181368Z","shell.execute_reply":"2024-12-10T16:54:06.714131Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Here is a boxplot after the removal of outliers in the target variable. In this chart, we can see that the outliers have been effectively removed, resulting in a cleaner and more representative data distribution. \n\nThe removal of outliers was performed using precise statistical methods such as the interquartile range (IQR) and z-score analysis to ensure that only relevant data is retained. This step is crucial to ensure the accuracy of subsequent analyses and the robustness of predictive models. \n\nWith the outliers removed, we achieve a clearer visualization of data dispersion and central tendencies, allowing for more reliable insights and better-informed decisions.","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"text-align:center; border-radius:15px 50px; padding:7px; color:white; margin:0; font-size:110%; font-family:Pacifico; background-color:#0073e6; overflow:hidden\"><b> Part 6 - Feature engineering</b></div>","metadata":{}},{"cell_type":"code","source":"# Importing library\nfrom sklearn.preprocessing import LabelEncoder\nfrom sklearn.impute import SimpleImputer\n\n# Creating the Label encoder\nLabel_pre = LabelEncoder()\ndata_cols=train_df.select_dtypes(exclude=['int','float']).columns\nlabel_col =list(data_cols)\n\n# Applying encoder\ntrain_df[label_col]=train_df[label_col].apply(lambda col:Label_pre.fit_transform(col))\n\n# Viewing\nLabel_pre","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-10T16:54:06.716171Z","iopub.execute_input":"2024-12-10T16:54:06.716456Z","iopub.status.idle":"2024-12-10T16:54:08.795497Z","shell.execute_reply.started":"2024-12-10T16:54:06.71643Z","shell.execute_reply":"2024-12-10T16:54:08.794827Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Viewing dataset after applying label encoder\ntrain_df.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-10T16:54:18.117954Z","iopub.execute_input":"2024-12-10T16:54:18.118354Z","iopub.status.idle":"2024-12-10T16:54:18.135498Z","shell.execute_reply.started":"2024-12-10T16:54:18.118319Z","shell.execute_reply":"2024-12-10T16:54:18.134596Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Why Apply Feature Engineering:**\n\nFeature engineering is a crucial step in the machine learning pipeline, as it involves transforming raw data into meaningful features that better represent the underlying problem to predictive models. By improving the quality of the features, we can enhance model performance, making it more accurate and reliable.\n\n1. **Label Encoding:**\n   - The categorical variables in `train_df` are encoded using `LabelEncoder`. \n   - **Purpose:** Many machine learning algorithms require numerical input, but real-world datasets often contain categorical data. `LabelEncoder` converts these categorical labels into a numerical format. Each category is assigned a unique integer, which allows the algorithm to process the categorical data.\n\n2. **Handling Missing Data:**\n   - Although not explicitly shown in your snippet, handling missing data is typically a part of feature engineering. This is crucial because many machine learning models cannot handle missing values directly, and leaving them untreated could introduce bias or reduce model accuracy.\n   - **SimpleImputer** is often used to fill missing values with a specific strategy, such as replacing missing values with the mean, median, or mode.\n\n3. **Importance of Feature Engineering:**\n   - **Improves Model Accuracy:** Properly engineered features can significantly enhance the predictive power of models by making patterns in the data more apparent to the algorithm.\n   - **Reduces Overfitting:** By carefully selecting and transforming features, we can reduce the noise in the data, leading to models that generalize better to unseen data.\n   - **Handles Non-Numeric Data:** Feature engineering techniques like label encoding enable models to work with non-numeric data by converting it into a suitable format.\n   - **Optimizes Training Time:** By reducing the complexity of features and ensuring all data is in a consistent format, models can be trained more efficiently.\n\nIn summary, applying feature engineering, including steps like label encoding, allows you to preprocess the dataset into a form that machine learning models can easily interpret and learn from, ultimately leading to better predictions.","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"text-align:center; border-radius:15px 50px; padding:7px; color:white; margin:0; font-size:110%; font-family:Pacifico; background-color:#0073e6; overflow:hidden\"><b> Part 7 - Training and testing division</b></div>","metadata":{}},{"cell_type":"code","source":"# Selecting variables for model\nX = train_df.drop('Premium Amount', axis=1)\ny = train_df['Premium Amount']","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-10T16:54:21.021145Z","iopub.execute_input":"2024-12-10T16:54:21.021462Z","iopub.status.idle":"2024-12-10T16:54:21.09664Z","shell.execute_reply.started":"2024-12-10T16:54:21.021433Z","shell.execute_reply":"2024-12-10T16:54:21.095881Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Visualizing data x\nX.shape","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-10T16:54:21.573576Z","iopub.execute_input":"2024-12-10T16:54:21.573955Z","iopub.status.idle":"2024-12-10T16:54:21.5798Z","shell.execute_reply.started":"2024-12-10T16:54:21.573926Z","shell.execute_reply":"2024-12-10T16:54:21.578828Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Visualizing data y\ny.shape","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-10T16:54:21.955974Z","iopub.execute_input":"2024-12-10T16:54:21.956326Z","iopub.status.idle":"2024-12-10T16:54:21.962073Z","shell.execute_reply.started":"2024-12-10T16:54:21.956296Z","shell.execute_reply":"2024-12-10T16:54:21.961112Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <div style=\"text-align:center; border-radius:15px 50px; padding:7px; color:white; margin:0; font-size:110%; font-family:Pacifico; background-color:#0073e6; overflow:hidden\"><b> Part 8 - Model training</b></div>","metadata":{}},{"cell_type":"code","source":"# Importing libraries\nfrom sklearn.model_selection import train_test_split\n\n# Splitting the data into training and testing sets\nX_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)\n\n# Viewing X_train rows and columns\nprint(\"Viewing X train data:\", X_train.shape)\n\n# Viewing y_train rows and columns\nprint(\"Viewing y train data:\", y_train.shape)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-10T16:54:24.068228Z","iopub.execute_input":"2024-12-10T16:54:24.068571Z","iopub.status.idle":"2024-12-10T16:54:24.322747Z","shell.execute_reply.started":"2024-12-10T16:54:24.068541Z","shell.execute_reply":"2024-12-10T16:54:24.321741Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"At this stage, we proceeded with model training. A crucial step in this process is splitting the data into training and testing sets. This division allows us to evaluate the model's performance on unseen data and avoid potential overfitting issues. Typically, a portion of the data is reserved for training, while another portion is held out for evaluation. Additionally, techniques such as cross-validation can be applied to ensure a more robust assessment of the model.\n\nDuring training, the models are exposed to the training data, adjusting their parameters to optimize performance with respect to the chosen metric, such as accuracy or mean squared error. After training, the models are evaluated using the testing data to assess their generalization ability. This step is crucial to ensure that the model can make accurate predictions on new data. The model training process involves choosing and properly configuring algorithms, selecting hyperparameters, and continuously evaluating the model's performance on different datasets.","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"text-align:center; border-radius:15px 50px; padding:7px; color:white; margin:0; font-size:110%; font-family:Pacifico; background-color:#0073e6; overflow:hidden\"><b> Part 9 - Model machine learning</b></div>","metadata":{}},{"cell_type":"code","source":"import xgboost as xgb\nimport lightgbm as lgb\nfrom lightgbm import log_evaluation\nfrom math import sqrt\nfrom sklearn.metrics import r2_score\nfrom sklearn.metrics import mean_squared_error\n\n# XGBoost model parameters with GPU and other adjustments\nxgb_params = {'tree_method': 'gpu_hist',           # Use GPU-optimized tree method\n              'predictor': 'gpu_predictor',        # Use GPU for prediction\n              'objective': 'reg:squarederror',     # Objective function for regression\n              'n_estimators': 1000,                # Number of trees (estimators)\n              'learning_rate': 0.05,               # Learning rate\n              'max_depth': 6,                      # Maximum depth of trees\n              'subsample': 0.8,                    # Ratio of samples for subsampling\n              'colsample_bytree': 0.8,             # Proportion of columns for subsampling\n              'reg_alpha': 0.1,                    # L1 Regularization\n              'reg_lambda': 0.1,                   # L2 Regularization\n              'verbosity': 1                       # Verbosity level for output (1 for detailed output)\n             }\n\n# LightGBM model parameters with GPU and other adjustments\nlgbm_params = {'boosting_type': 'gbdt',             # Type of boosting (Gradient Boosted Decision Trees)\n               'objective': 'regression',           # Objective function for regression\n               'metric': 'rmse',                    # Evaluation metric: Root Mean Squared Error (RMSE)\n               'device': 'gpu',                     # Use GPU for training\n               'gpu_platform_id': 0,                # GPU platform ID (set as needed)\n               'gpu_device_id': 0,                  # GPU device ID (set as needed)\n               'num_leaves': 31,                    # Number of leaves in the tree\n               'learning_rate': 0.05,               # Learning rate\n               'n_estimators': 1000,                # Number of trees (estimators)\n               'max_depth': -1,                     # Maximum depth of trees (-1 means no limit)\n               'min_child_samples': 20,             # Minimum number of samples in a child node\n               'subsample': 0.8,                    # Ratio of samples for subsampling\n               'colsample_bytree': 0.8,             # Proportion of columns for subsampling\n               'reg_alpha': 0.1,                    # L1 Regularization\n               'reg_lambda': 0.1,                   # L2 Regularization\n               'verbose': -1                        # Verbosity level (-1 for silent)\n              }\n\n# Initializing and training the XGBoost model\nxgb_model = xgb.XGBRegressor(**xgb_params)\nxgb_model.fit(X_train, y_train, \n              eval_set=[(X_test, y_test)], \n              early_stopping_rounds=10, \n              verbose=True)\n\n# Making predictions with the trained XGBoost model\nxgb_model_pred = xgb_model.predict(X_test)\n\n# Initializing and training the LightGBM model\nlgbm_model = lgb.LGBMRegressor(**lgbm_params)\nlgbm_model.fit(X_train, y_train,\n               eval_set=[(X_test, y_test)],  \n               callbacks=[log_evaluation(10)])\n\n# Making predictions with the trained LightGBM model\nlgbm_predictions = lgbm_model.predict(X_test)\n\n# Calculating RMSE for the LightGBM model\nlgbm_rmse = sqrt(mean_squared_error(y_test, lgbm_predictions))\nprint(f\"LightGBM RMSE: {lgbm_rmse:.4f}\")\n\n# Calculating RMSE for the XGBoost model\nxgb_rmse = sqrt(mean_squared_error(y_test, xgb_model_pred))\nprint(f\"XGBoost RMSE: {xgb_rmse:.4f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-10T16:54:27.362869Z","iopub.execute_input":"2024-12-10T16:54:27.363212Z","iopub.status.idle":"2024-12-10T16:54:57.800013Z","shell.execute_reply.started":"2024-12-10T16:54:27.363184Z","shell.execute_reply":"2024-12-10T16:54:57.799103Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Calculating the residuals (difference between actual and predicted values)\nlgbm_residuals = y_test - lgbm_predictions\nxgb_residuals = y_test - xgb_model_pred\n\n# Setting up the figure size\nplt.figure(figsize=(14, 6))\n\n# Residual plot for the LightGBM model\nplt.subplot(1, 2, 1)\nsns.scatterplot(x=lgbm_predictions, y=lgbm_residuals, color='blue', alpha=0.5)\nplt.axhline(y=0, color='red', linestyle='--')\nplt.title('Residual Plot - LightGBM')\nplt.xlabel('Predicted Values')\nplt.ylabel('Residuals')\n\n# Residual plot for the XGBoost model\nplt.subplot(1, 2, 2)\nsns.scatterplot(x=xgb_model_pred, y=xgb_residuals, color='green', alpha=0.5)\nplt.axhline(y=0, color='red', linestyle='--')\nplt.title('Residual Plot - XGBoost')\nplt.xlabel('Predicted Values')\nplt.ylabel('Residuals')\n\n# Adjusting the layout for better visualization\nplt.tight_layout()\n\n# Displaying the plots\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-10T16:56:00.816063Z","iopub.execute_input":"2024-12-10T16:56:00.817105Z","iopub.status.idle":"2024-12-10T16:56:02.466382Z","shell.execute_reply.started":"2024-12-10T16:56:00.817066Z","shell.execute_reply":"2024-12-10T16:56:02.465552Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Plot the linear regression graph LGBM\nplt.figure(figsize=(10, 6))\nplt.scatter(y_test, lgbm_predictions, color='green', alpha=0.5) # Scatter plot of actual values ​​vs predictions\nplt.plot([y_test.min(), y_test.max()], [y_test.min(), y_test.max()], 'k--', lw=2) # Reference line for perfect regression\nplt.legend([\"Current\", \"Forecast\"])\nplt.xlabel('Real Values')\nplt.ylabel('Forecasts')\nplt.title('Linear Regression Graph with LightGBM')\nplt.grid(False)\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-10T16:56:09.73788Z","iopub.execute_input":"2024-12-10T16:56:09.738814Z","iopub.status.idle":"2024-12-10T16:56:11.914128Z","shell.execute_reply.started":"2024-12-10T16:56:09.738765Z","shell.execute_reply":"2024-12-10T16:56:11.91334Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Plot the linear regression graph XGBoost\nplt.figure(figsize=(10, 6))\nplt.scatter(y_test, xgb_model_pred, color='green', alpha=0.5) # Scatter plot of actual values ​​vs predictions\nplt.plot([y_test.min(), y_test.max()], [y_test.min(), y_test.max()], 'k--', lw=2) # Reference line for perfect regression\nplt.legend([\"Current\", \"Forecast\"])\nplt.xlabel('Real Values')\nplt.ylabel('Forecasts')\nplt.title('Linear Regression Graph with XGBoost')\nplt.grid(False)\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-10T16:56:14.305156Z","iopub.execute_input":"2024-12-10T16:56:14.305535Z","iopub.status.idle":"2024-12-10T16:56:16.415225Z","shell.execute_reply.started":"2024-12-10T16:56:14.305485Z","shell.execute_reply":"2024-12-10T16:56:16.414322Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <div style=\"text-align:center; border-radius:15px 50px; padding:7px; color:white; margin:0; font-size:110%; font-family:Pacifico; background-color:#0073e6; overflow:hidden\"><b> Part 10 - Feature Importance</b></div>","metadata":{}},{"cell_type":"code","source":"# Plotting Feature Importance for the XGBoost Model\nplt.figure(figsize=(20, 10))\nxgb.plot_importance(xgb_model, max_num_features=25, importance_type='weight')\nplt.title('Importance of Features - XGBoost')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-10T16:56:43.568302Z","iopub.execute_input":"2024-12-10T16:56:43.569139Z","iopub.status.idle":"2024-12-10T16:56:43.906095Z","shell.execute_reply.started":"2024-12-10T16:56:43.569103Z","shell.execute_reply":"2024-12-10T16:56:43.905257Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Importing the feature importances from the LightGBM model into a DataFrame\nlgbm_feature_importances = pd.DataFrame({'Feature': X_train.columns,\n                                         'Importance': lgbm_model.feature_importances_})\n\n# Sorting the features by their importance in descending order\nlgbm_feature_importances = lgbm_feature_importances.sort_values(by='Importance', ascending=False)\n\n# Plotting the top 20 most important features\nplt.figure(figsize=(10, 8))\nsns.barplot(x='Importance', y='Feature', palette=\"Set2\", data=lgbm_feature_importances.head(20))\nplt.title('Feature Importance - LightGBM')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-10T17:13:52.002135Z","iopub.execute_input":"2024-12-10T17:13:52.00249Z","iopub.status.idle":"2024-12-10T17:13:52.380985Z","shell.execute_reply.started":"2024-12-10T17:13:52.002459Z","shell.execute_reply":"2024-12-10T17:13:52.38014Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <div style=\"text-align:center; border-radius:15px 50px; padding:7px; color:white; margin:0; font-size:110%; font-family:Pacifico; background-color:#0073e6; overflow:hidden\"><b> Part 11 - Final result ML</b></div>","metadata":{}},{"cell_type":"code","source":"from sklearn.metrics import r2_score\nfrom sklearn.metrics import mean_squared_error\nfrom math import sqrt\n\n# Calculating the metrics for the LightGBM model\nlgbm_r2 = r2_score(y_test, lgbm_predictions)\nlgbm_rmse = sqrt(mean_squared_error(y_test, lgbm_predictions))\n\n# Calculating the metrics for the XGBoost model\nxgb_r2 = r2_score(y_test, xgb_model_pred)\nxgb_rmse = sqrt(mean_squared_error(y_test, xgb_model_pred))\n\n# Storing the results in a dictionary\nresults = {\"Model\": [\"LightGBM\", \"XGBoost\"],\n           \"R²\": [lgbm_r2, xgb_r2],\n           \"RMSE\": [lgbm_rmse, xgb_rmse]}\n\n# Converting the dictionary into a DataFrame\nresults_df = pd.DataFrame(results)\n\n# Displaying the DataFrame with the results\nresults_df","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-10T16:57:52.406905Z","iopub.execute_input":"2024-12-10T16:57:52.407846Z","iopub.status.idle":"2024-12-10T16:57:52.425334Z","shell.execute_reply.started":"2024-12-10T16:57:52.407808Z","shell.execute_reply":"2024-12-10T16:57:52.424514Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <div style=\"text-align:center; border-radius:15px 50px; padding:7px; color:white; margin:0; font-size:110%; font-family:Pacifico; background-color:#0073e6; overflow:hidden\"><b> Part 12 - Model 2 - Turing</b></div>","metadata":{}},{"cell_type":"code","source":"# Training and testing division\nX1 = train_df.drop(columns=['Premium Amount'], axis=1)\ny1 = train_df['Premium Amount']\n\n# Importing libraries\nfrom sklearn.model_selection import train_test_split\n\n# Splitting the data into training and testing sets\nX_train, X_test, y_train, y_test = train_test_split(X1, y1, test_size=0.2, random_state=42)\n\n# Viewing X_train rows and columns\nprint(\"Viewing X train data:\", X_train.shape)\n\n# Viewing y_train rows and columns\nprint(\"Viewing y train data:\", y_train.shape)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-10T17:07:26.011068Z","iopub.execute_input":"2024-12-10T17:07:26.01205Z","iopub.status.idle":"2024-12-10T17:07:26.352269Z","shell.execute_reply.started":"2024-12-10T17:07:26.012Z","shell.execute_reply":"2024-12-10T17:07:26.35119Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from xgboost import XGBRegressor\nfrom sklearn.model_selection import KFold\nfrom sklearn.metrics import mean_squared_log_error\n\n# Define the XGBoost model with GPU support\nxgb_model = XGBRegressor(\n    booster='gbtree',\n    objective='reg:squarederror',\n    eval_metric='rmse',  # RMSLE requires log transformation\n    learning_rate=0.05,\n    n_estimators=100,\n    random_state=42,\n    gpu_id=0  # Adding GPU parameter\n)\n\n# Evaluation using K-Fold Cross Validation\nkf = KFold(n_splits=10, shuffle=True, random_state=42)\nfold_metrics = []\n\nfor fold, (train_index, val_index) in enumerate(kf.split(X_train), start=1):\n    # Split the data into training and validation folds\n    X_train_fold, X_val_fold = X_train.iloc[train_index], X_train.iloc[val_index]\n    y_train_fold, y_val_fold = y_train.iloc[train_index], y_train.iloc[val_index]\n    \n    # Train the model on the training fold\n    xgb_model.fit(X_train_fold, y_train_fold)\n    \n    # Make predictions on the validation fold\n    y_val_pred = xgb_model.predict(X_val_fold)\n    \n    # Calculate metrics\n    r2 = r2_score(y_val_fold, y_val_pred)\n    mae = mean_absolute_error(y_val_fold, y_val_pred)\n    rmse = sqrt(mean_squared_error(y_val_fold, y_val_pred))\n    mse = mean_squared_error(y_val_fold, y_val_pred)\n    rmsle_val = np.sqrt(mean_squared_log_error(y_val_fold, y_val_pred))\n    \n    # Append metrics for this fold\n    fold_metrics.append({'Fold': fold,\n                         'R^2': r2,\n                         'MAE': mae,\n                         'RMSE': rmse,\n                         'MSE': mse,\n                         'RMSLE': rmsle_val})\n\n# Evaluate the model's average performance across validation folds\nmean_metrics = pd.DataFrame(fold_metrics).mean().round(4)\nstd_metrics = pd.DataFrame(fold_metrics).std().round(4)\n\n# Create a DataFrame with the metrics\nmetrics_df = pd.DataFrame(fold_metrics)\nmetrics_df.loc['Mean'] = mean_metrics\nmetrics_df.loc['Std Dev'] = std_metrics\n\n# Output the results\nprint(f\"Mean RMSLE (Validation): {mean_rmsle_validation:.4f}\")\nprint(f\"Standard Deviation of RMSLE (Validation): {std_rmsle_validation:.4f}\")\nprint(f\"Mean RMSLE (Training): {mean_rmsle_training:.4f}\")\nprint(f\"Standard Deviation of RMSLE (Training): {std_train_rmsle:.4f}\")\n\nfold_metrics","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-10T17:31:29.581052Z","iopub.execute_input":"2024-12-10T17:31:29.581399Z","iopub.status.idle":"2024-12-10T17:31:44.906338Z","shell.execute_reply.started":"2024-12-10T17:31:29.581369Z","shell.execute_reply":"2024-12-10T17:31:44.905479Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Make predictions on the test set\nxgb_test_predictions = xgb_model.predict(X_test_kaggle)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-10T17:15:55.999365Z","iopub.execute_input":"2024-12-10T17:15:56.000141Z","iopub.status.idle":"2024-12-10T17:15:56.183411Z","shell.execute_reply.started":"2024-12-10T17:15:56.000092Z","shell.execute_reply":"2024-12-10T17:15:56.182725Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Plot the importance of features\nfrom xgboost import XGBRegressor, plot_importance\n\nplt.figure(figsize=(10, 12))\nplot_importance(xgb_model, importance_type='weight', max_num_features=10)\nplt.title('Top 10 Features Based on Weight Importance')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-10T17:16:07.680034Z","iopub.execute_input":"2024-12-10T17:16:07.680357Z","iopub.status.idle":"2024-12-10T17:16:07.951912Z","shell.execute_reply.started":"2024-12-10T17:16:07.68033Z","shell.execute_reply":"2024-12-10T17:16:07.951084Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Residual Scatterplot\nresiduals = y_val_fold - y_val_pred\nplt.figure(figsize=(20, 5))\nplt.subplot(1, 2, 1)\nplt.scatter(y_val_pred, residuals, alpha=0.5)\nplt.axhline(y=0, color='r', linestyle='--')\nplt.xlabel('Predicted Values')\nplt.ylabel('Residuals')\nplt.title('Residuals vs Predicted Values')\n\n# Residual Histogram\nplt.subplot(1, 2, 2)\nplt.hist(residuals, bins=20, color='skyblue', edgecolor='black')\nplt.xlabel('Residuals')\nplt.ylabel('Frequency')\nplt.title('Histogram of Residuals')\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-10T17:17:30.396962Z","iopub.execute_input":"2024-12-10T17:17:30.39731Z","iopub.status.idle":"2024-12-10T17:17:31.195966Z","shell.execute_reply.started":"2024-12-10T17:17:30.397278Z","shell.execute_reply":"2024-12-10T17:17:31.195206Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Regression predictions\nplt.figure(figsize=(20.5, 10))\n\nfor fold, (y_val_pred, y_val_fold) in enumerate(zip(all_val_predictions, all_val_actuals), start=1):\n    plt.subplot(2, 5, fold)\n    plt.scatter(y_val_fold, y_val_pred, alpha=0.5, edgecolors='w', linewidth=1)\n    plt.plot([y_val_fold.min(), y_val_fold.max()], [y_val_fold.min(), y_val_fold.max()], 'k--', lw=2)\n    plt.xlabel('True Values')\n    plt.ylabel('Predicted Values')\n    plt.title(f'Fold {fold}')\n\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-10T17:20:38.20945Z","iopub.execute_input":"2024-12-10T17:20:38.210176Z","iopub.status.idle":"2024-12-10T17:20:42.150721Z","shell.execute_reply.started":"2024-12-10T17:20:38.21014Z","shell.execute_reply":"2024-12-10T17:20:42.149838Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Calculate metrics\n\nfrom sklearn.metrics import mean_squared_error, mean_absolute_error, r2_score, mean_squared_log_error\nrmsle_val = np.sqrt(mean_squared_log_error(y_val_fold, y_val_pred))\nrmsle_val","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-10T17:25:19.371774Z","iopub.execute_input":"2024-12-10T17:25:19.372124Z","iopub.status.idle":"2024-12-10T17:25:19.380869Z","shell.execute_reply.started":"2024-12-10T17:25:19.372094Z","shell.execute_reply":"2024-12-10T17:25:19.380018Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <div style=\"text-align:center; border-radius:15px 50px; padding:7px; color:white; margin:0; font-size:110%; font-family:Pacifico; background-color:#0073e6; overflow:hidden\"><b> Part 13 - Submission</b></div>","metadata":{}},{"cell_type":"code","source":"# Reload the full test dataset\ntest_df = pd.read_csv('/kaggle/input/playground-series-s4e12/test.csv')\n\n# Remove only the 'id' column if necessary\nX_test_kaggle = test_df.drop(columns=['id'])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-10T16:58:44.703732Z","iopub.execute_input":"2024-12-10T16:58:44.704575Z","iopub.status.idle":"2024-12-10T16:58:47.072263Z","shell.execute_reply.started":"2024-12-10T16:58:44.704537Z","shell.execute_reply":"2024-12-10T16:58:47.071313Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from sklearn.preprocessing import LabelEncoder\n\n# Remove additional columns from the test set that are not present in the training set\ncols_to_remove = [col for col in X_test_kaggle.columns if col not in X_train.columns]\nif cols_to_remove:\n    X_test_kaggle = X_test_kaggle.drop(columns=cols_to_remove)\n    \n# Apply Label Encoding to each categorical column that is present in both datasets\nfor col in X_test_kaggle.select_dtypes(include=['object']).columns:\n    le = LabelEncoder()\n    X_test_kaggle[col] = le.fit_transform(X_test_kaggle[col])\n\n# The LabelEncoder instance used for encoding\nle","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-10T16:58:55.088244Z","iopub.execute_input":"2024-12-10T16:58:55.088904Z","iopub.status.idle":"2024-12-10T16:58:56.816817Z","shell.execute_reply.started":"2024-12-10T16:58:55.088871Z","shell.execute_reply":"2024-12-10T16:58:56.815961Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Get the features of the trained model\ntrain_columns = X_train.columns\n\n# Check the columns in the test set\ntest_columns = X_test_kaggle.columns\n\n# Check if there are any missing columns in the test set\nmissing_columns = set(train_columns) - set(test_columns)\n\nif missing_columns:\n    \n    # Add the missing columns back to X_test_kaggle\n    for col in missing_columns:\n\n        # Add missing columns with default values (e.g., zeros)\n        X_test_kaggle[col] = 0  \n\n# Now, make predictions\nlgbm_test_predictions = lgbm_model.predict(X_test_kaggle)\n\n# Create the submission DataFrame with the correct column name\nlgbm_submission_df = pd.DataFrame({'id': test_df['id'],\n                                   'Premium Amount': lgbm_test_predictions})\n\n# Save the submission file\nlgbm_submission_df.to_csv('lgbm_submission.csv', index=False)\n\n# Visualize the submission DataFrame\nlgbm_submission_df","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-10T17:05:47.117013Z","iopub.execute_input":"2024-12-10T17:05:47.117359Z","iopub.status.idle":"2024-12-10T17:06:11.608777Z","shell.execute_reply.started":"2024-12-10T17:05:47.11733Z","shell.execute_reply":"2024-12-10T17:06:11.607803Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <div style=\"text-align:center; border-radius:15px 50px; padding:7px; color:white; margin:0; font-size:110%; font-family:Pacifico; background-color:#0073e6; overflow:hidden\"><b> Part 14 - Submission Turing</b></div>","metadata":{}},{"cell_type":"code","source":"# Final predictions for submission\n\n# Replace with the appropriate method to load the test set data\nX_test_kaggle = pd.read_csv('/kaggle/input/playground-series-s4e12/test.csv')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-10T17:32:06.730231Z","iopub.execute_input":"2024-12-10T17:32:06.730598Z","iopub.status.idle":"2024-12-10T17:32:08.740652Z","shell.execute_reply.started":"2024-12-10T17:32:06.730566Z","shell.execute_reply":"2024-12-10T17:32:08.739655Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from sklearn.preprocessing import LabelEncoder\n\n# Remove columns from the test set that are not present in the training set\ncols_to_remove = [col for col in X_test_kaggle.columns if col not in X_train.columns]\nif cols_to_remove:\n    X_test_kaggle = X_test_kaggle.drop(columns=cols_to_remove)\n\n# Apply Label Encoding to each categorical column that is present in both datasets\nfor col in X_test_kaggle.select_dtypes(include=['object']).columns:\n    le = LabelEncoder()\n    X_test_kaggle[col] = le.fit_transform(X_test_kaggle[col])\n\n# Get the features of the trained model\ntrain_columns = X_train.columns\n\n# Check the columns in the test set\ntest_columns = X_test_kaggle.columns\n\n# Check if there are any missing columns in the test set\nmissing_columns = set(train_columns) - set(test_columns)\n\nif missing_columns:\n    # Add missing columns with default values (e.g., zeros)\n    for col in missing_columns:\n        X_test_kaggle[col] = 0\n\n# The LabelEncoder instance used for encoding\nle","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-10T17:37:51.87992Z","iopub.execute_input":"2024-12-10T17:37:51.880751Z","iopub.status.idle":"2024-12-10T17:37:51.891137Z","shell.execute_reply.started":"2024-12-10T17:37:51.880703Z","shell.execute_reply":"2024-12-10T17:37:51.890293Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Make predictions for the test set\npredictions = xgb_model.predict(X_test_kaggle)\n\n# Check if the 'id' column is in the X_train DataFrame\nif 'id' not in X_train.columns:\n    # Remove the 'id' column from X_test_kaggle if it is not present in X_train\n    X_test_kaggle = X_test_kaggle.drop(columns=['id'])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-10T17:38:55.099663Z","iopub.execute_input":"2024-12-10T17:38:55.100047Z","iopub.status.idle":"2024-12-10T17:38:55.134141Z","shell.execute_reply.started":"2024-12-10T17:38:55.100015Z","shell.execute_reply":"2024-12-10T17:38:55.13345Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Remove the 'id' column from X_test_kaggle if it is present\nif 'id' in X_test_kaggle.columns:\n    X_test_kaggle = X_test_kaggle.drop(columns=['id'])\n\n# Make predictions for the test set\npredictions = xgb_model.predict(X_test_kaggle)\n\n# Create the submission DataFrame\n# Add the original 'id' column from the test set\nsubmission_df = pd.DataFrame({'id': test_df['id'],  \n                              'Premium Amount': predictions})\n\n# Save the DataFrame as a CSV file for submission to Kaggle\nsubmission_df.to_csv('submissionXGBRegressor_2.csv', index=False)\n\n# Display the DataFrame with submission results\nsubmission_df\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-10T17:39:51.524058Z","iopub.execute_input":"2024-12-10T17:39:51.524408Z","iopub.status.idle":"2024-12-10T17:39:52.817924Z","shell.execute_reply.started":"2024-12-10T17:39:51.524376Z","shell.execute_reply":"2024-12-10T17:39:52.817048Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}