{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"codemirror_mode":{"name":"ipython","version":3},"file_extension":".py","mimetype":"text/x-python","name":"python","nbconvert_exporter":"python","pygments_lexer":"ipython3","version":"3.10.12"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":84896,"databundleVersionId":10305135,"sourceType":"competition"}],"dockerImageVersionId":30823,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true},"papermill":{"default_parameters":{},"duration":490.346675,"end_time":"2024-12-30T16:19:04.398231","environment_variables":{},"exception":null,"input_path":"__notebook__.ipynb","output_path":"__notebook__.ipynb","parameters":{},"start_time":"2024-12-30T16:10:54.051556","version":"2.6.0"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<div style=\"border-radius: 15px; border: 2px solid #6A1B9A; padding: 20px; background: linear-gradient(135deg, #9C27B0, #4CAF50); text-align: center; box-shadow: 0px 4px 8px rgba(0, 0, 0, 0.5);\">\n    <h1 style=\"color: #ffffff; text-shadow: 2px 2px 4px rgba(0, 0, 0, 0.7); font-weight: bold; margin-bottom: 10px; font-size: 36px; font-family: 'Roboto', sans-serif;\">\n        2️⃣⚡💡 LightGBM Duo: GBDT + GOSS+OOF 🔍\n    </h1>\n</div>\n\n<!-- Include Google Fonts for a modern font -->\n<link href=\"https://fonts.googleapis.com/css2?family=Roboto:wght@700&display=swap\" rel=\"stylesheet\">\n","metadata":{"papermill":{"duration":0.013407,"end_time":"2024-12-30T16:10:56.293765","exception":false,"start_time":"2024-12-30T16:10:56.280358","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"### 📝 Updated Dataset Overview\n\nThe dataset for this project stems from the **Kaggle Playground Series - Season 4, Episode 12**, focusing on the prediction of insurance premiums. It encompasses a diverse range of features, simulating real-world scenarios in insurance premium determination.\n\n\n#### **Key Highlights**:\n- **Features**: A comprehensive set of **20 features** (excluding the target variable), including numerical, categorical, and temporal data types.\n- **Target Variable**: The **Premium Amount**, a continuous numerical value, serves as the target for prediction.\n\n\n#### **Feature Breakdown**:\n1. **Numerical Features**:\n   - Quantitative attributes such as:\n     - **Age**: Reflecting policyholder age.\n     - **Annual Income**: Indicating financial capability.\n     - **Health Score**: A measure of the individual's health condition.\n     - **Credit Score**: Evaluating financial reliability.\n     - **Vehicle Age**: Representing the insured vehicle's age.\n2. **Categorical Features**:\n   - Qualitative characteristics, including:\n     - **Gender**, **Marital Status**, **Education Level**, **Occupation**, and **Policy Type**, providing demographic and contextual insights.\n3. **Temporal Feature**:\n   - **Policy Start Date**: A date-based feature capturing the inception of the insurance policy.\n\n#### **Dataset Challenges**:\n1. **Missing Values**:\n   - Several features contain missing data points, requiring thoughtful imputation techniques to maintain data integrity and enable robust predictions.\n2. **Skewed Distributions**:\n   - Features like **Annual Income** and **Premium Amount** exhibit significant skewness, necessitating transformations such as logarithmic scaling to normalize the data.\n3. **Diverse Feature Types**:\n   - The dataset includes a mix of numerical, categorical, and temporal data types, requiring comprehensive preprocessing strategies for consistency and model compatibility.\n4. **Multicollinearity**:\n   - Potential correlations among features may affect model performance, necessitating correlation analysis and feature selection.\n\n### 🎯 **Objective**\n\nThe primary objective of this project is to build a **highly accurate and interpretable machine learning model** to predict the **Premium Amount** for insurance policyholders. The model will aim to balance predictive accuracy and usability for actionable insights.\n\n#### **Key Objectives**:\n1. **Data Preprocessing**:\n   - Address missing values, normalize numerical distributions, and encode categorical variables for consistent and robust model training.\n2. **Feature Engineering**:\n   - Enhance predictive capability by:\n     - Deriving new features from existing ones.\n     - Selecting features with the highest predictive importance.\n3. **Model Training and Optimization**:\n   - Leverage cutting-edge models such as:\n     - **LightGBM (GBDT and GOSS)** for gradient boosting.\n   - Employ advanced tuning techniques like **Optuna** for hyperparameter optimization.\n4. **Performance Evaluation**:\n   - Measure model effectiveness using the **Root Mean Squared Logarithmic Error (RMSLE)** metric, ensuring alignment with real-world scenarios.\n5. **Interpretability and Insights**:\n   - Extract feature importance metrics to identify key drivers of premium costs.\n   - Analyze prediction trends to uncover actionable insights for stakeholders.\n\n#### **Project Aspirations**:\nThis project aims to:\n- Achieve a **competitive RMSLE score**, surpassing baseline performance benchmarks.\n- Provide a **scalable and adaptable framework** for insurance premium prediction, applicable across various datasets and contexts.\n- Offer **meaningful insights** into the primary determinants of insurance premiums, aiding strategic decision-making.\n\nBy leveraging advanced **data science techniques**, **rigorous preprocessing**, and **state-of-the-art models**, this project seeks to establish itself as a reliable and innovative solution for predicting insurance premiums in the industry. 🚀","metadata":{}},{"cell_type":"markdown","source":"# <span style=\"color:transparent;\">Import Libraries</span>\n\n<div style=\"border-radius: 15px; border: 2px solid #6A1B9A; padding: 10px; background: linear-gradient(135deg, #9C27B0, #4CAF50); text-align: center; box-shadow: 0px 4px 8px rgba(0, 0, 0, 0.5);\">\n    <h1 style=\"color: #ffffff; text-shadow: 2px 2px 4px rgba(0, 0, 0, 0.7); font-weight: bold; margin-bottom: 5px; font-size: 28px; font-family: 'Roboto', sans-serif;\">\n        Import Libraries\n    </h1>\n</div>\n\n<!-- Include Google Fonts for a modern font -->\n<link href=\"https://fonts.googleapis.com/css2?family=Roboto:wght@700&display=swap\" rel=\"stylesheet\">\n","metadata":{"papermill":{"duration":0.009788,"end_time":"2024-12-30T16:10:56.333711","exception":false,"start_time":"2024-12-30T16:10:56.323923","status":"completed"},"tags":[]}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nfrom matplotlib import cm\nfrom matplotlib.cm import viridis\nimport seaborn as sns\nimport math\nfrom sklearn.preprocessing import LabelEncoder, OrdinalEncoder, OneHotEncoder\nfrom sklearn.compose import ColumnTransformer\nfrom sklearn.pipeline import Pipeline\n\nfrom scipy.signal import find_peaks\nfrom scipy.stats import skew \n\nimport optuna\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.metrics import mean_squared_log_error\nfrom sklearn.preprocessing import StandardScaler\nimport lightgbm as lgb\nimport xgboost as xgb\nfrom catboost import CatBoostRegressor\nfrom lightgbm import early_stopping, log_evaluation\nfrom sklearn.model_selection import KFold\n\n# Ignore general warnings\nimport warnings\nwarnings.filterwarnings(\"ignore\")\n\n# Suppress LightGBM logs\nimport logging\nlogging.getLogger(\"lightgbm\").setLevel(logging.ERROR)\n","metadata":{"execution":{"iopub.status.busy":"2024-12-31T04:48:49.647338Z","iopub.execute_input":"2024-12-31T04:48:49.647656Z","iopub.status.idle":"2024-12-31T04:48:54.381366Z","shell.execute_reply.started":"2024-12-31T04:48:49.647626Z","shell.execute_reply":"2024-12-31T04:48:54.380719Z"},"papermill":{"duration":5.60843,"end_time":"2024-12-30T16:11:01.952182","exception":false,"start_time":"2024-12-30T16:10:56.343752","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <span style=\"color:transparent;\">Data Loading and Initial Exploration</span>\n\n<div style=\"border-radius: 15px; border: 2px solid #6A1B9A; padding: 10px; background: linear-gradient(135deg, #9C27B0, #4CAF50); text-align: center; box-shadow: 0px 4px 8px rgba(0, 0, 0, 0.5);\">\n    <h1 style=\"color: #ffffff; text-shadow: 2px 2px 4px rgba(0, 0, 0, 0.7); font-weight: bold; margin-bottom: 5px; font-size: 28px; font-family: 'Roboto', sans-serif;\">\n        Data Loading and Initial Exploration\n    </h1>\n</div>\n\n<!-- Include Google Fonts for a modern font -->\n<link href=\"https://fonts.googleapis.com/css2?family=Roboto:wght@700&display=swap\" rel=\"stylesheet\">\n","metadata":{"papermill":{"duration":0.012261,"end_time":"2024-12-30T16:11:01.975399","exception":false,"start_time":"2024-12-30T16:11:01.963138","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Load the datasets\ntrain_data = pd.read_csv('/kaggle/input/playground-series-s4e12/train.csv',index_col=[0])\ntest_data = pd.read_csv('/kaggle/input/playground-series-s4e12/test.csv',index_col=[0])\nsample_data = pd.read_csv('/kaggle/input/playground-series-s4e12/sample_submission.csv')\n\n# Verify shapes\nprint(\"Train Data Shape:\", train_data.shape)\nprint(\"Test Data Shape:\", test_data.shape)","metadata":{"execution":{"iopub.status.busy":"2024-12-31T04:48:54.382106Z","iopub.execute_input":"2024-12-31T04:48:54.382671Z","iopub.status.idle":"2024-12-31T04:49:03.108385Z","shell.execute_reply.started":"2024-12-31T04:48:54.382647Z","shell.execute_reply":"2024-12-31T04:49:03.107632Z"},"papermill":{"duration":9.16185,"end_time":"2024-12-30T16:11:11.148519","exception":false,"start_time":"2024-12-30T16:11:01.986669","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Display sample data\nprint(\"Training Dataset: \\n\")\ndisplay(train_data.head())\nprint('\\n')\nprint(\"Test Dataset: \\n\")\ndisplay(test_data.head())","metadata":{"execution":{"iopub.status.busy":"2024-12-31T04:49:03.109428Z","iopub.execute_input":"2024-12-31T04:49:03.109758Z","iopub.status.idle":"2024-12-31T04:49:03.150246Z","shell.execute_reply.started":"2024-12-31T04:49:03.109733Z","shell.execute_reply":"2024-12-31T04:49:03.149426Z"},"papermill":{"duration":0.054812,"end_time":"2024-12-30T16:11:11.214329","exception":false,"start_time":"2024-12-30T16:11:11.159517","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <span style=\"color:transparent;\">Data Inspection and Understanding</span>\n\n<div style=\"border-radius: 15px; border: 2px solid #6A1B9A; padding: 10px; background: linear-gradient(135deg, #9C27B0, #4CAF50); text-align: center; box-shadow: 0px 4px 8px rgba(0, 0, 0, 0.5);\">\n    <h1 style=\"color: #ffffff; text-shadow: 2px 2px 4px rgba(0, 0, 0, 0.7); font-weight: bold; margin-bottom: 5px; font-size: 28px; font-family: 'Roboto', sans-serif;\">\n        Data Inspection and Understanding\n    </h1>\n</div>\n\n<!-- Include Google Fonts for a modern font -->\n<link href=\"https://fonts.googleapis.com/css2?family=Roboto:wght@700&display=swap\" rel=\"stylesheet\">\n","metadata":{"papermill":{"duration":0.010912,"end_time":"2024-12-30T16:11:11.237745","exception":false,"start_time":"2024-12-30T16:11:11.226833","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Display information for the training dataset\nprint(\"Training Dataset Information: \\n\")\ntrain_info = train_data.info()\ndisplay(train_info)\nprint('\\n')\n# Display information for the test dataset\nprint(\"Test Dataset Information: \\n\")\ntest_info = test_data.info()\ndisplay(test_info)","metadata":{"execution":{"iopub.status.busy":"2024-12-31T04:49:03.151085Z","iopub.execute_input":"2024-12-31T04:49:03.151364Z","iopub.status.idle":"2024-12-31T04:49:04.054784Z","shell.execute_reply.started":"2024-12-31T04:49:03.151342Z","shell.execute_reply":"2024-12-31T04:49:04.054025Z"},"papermill":{"duration":0.908064,"end_time":"2024-12-30T16:11:12.156862","exception":false,"start_time":"2024-12-30T16:11:11.248798","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"#### Dataset Shapes:\n1. **Train Dataset**: Contains 1,200,000 rows and 20 columns.\n2. **Test Dataset**: Contains 800,000 rows and 19 columns.\n   - The `Premium Amount` column, which is the target variable, is missing in the test dataset (as expected).\n\n#### Data Types:\n1. **Train Dataset**: \n   - 9 numerical columns (`float64`) and 11 categorical/text columns (`object`).\n2. **Test Dataset**:\n   - 8 numerical columns (`float64`) and 11 categorical/text columns (`object`).\n","metadata":{"papermill":{"duration":0.012249,"end_time":"2024-12-30T16:11:12.180702","exception":false,"start_time":"2024-12-30T16:11:12.168453","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"# <span style=\"color:transparent;\">Handling Missing Data</span>\n\n<div style=\"border-radius: 15px; border: 2px solid #6A1B9A; padding: 10px; background: linear-gradient(135deg, #9C27B0, #4CAF50); text-align: center; box-shadow: 0px 4px 8px rgba(0, 0, 0, 0.5);\">\n    <h1 style=\"color: #ffffff; text-shadow: 2px 2px 4px rgba(0, 0, 0, 0.7); font-weight: bold; margin-bottom: 5px; font-size: 28px; font-family: 'Roboto', sans-serif;\">\n        Handling Missing Data\n    </h1>\n</div>\n\n<!-- Include Google Fonts for a modern font -->\n<link href=\"https://fonts.googleapis.com/css2?family=Roboto:wght@700&display=swap\" rel=\"stylesheet\">\n","metadata":{"papermill":{"duration":0.010921,"end_time":"2024-12-30T16:11:12.202937","exception":false,"start_time":"2024-12-30T16:11:12.192016","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"## Missing Values Overview and Heatmaps","metadata":{"papermill":{"duration":0.011093,"end_time":"2024-12-30T16:11:12.225311","exception":false,"start_time":"2024-12-30T16:11:12.214218","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Check for missing values in the datasets\nmissing_train = train_data.isnull()\nmissing_test = test_data.isnull()\n\n# Create a single figure with subplots for both datasets\nfig, axes = plt.subplots(1, 2, figsize=(18, 6))\n\n# Heatmap for missing values in the training dataset\nsns.heatmap(missing_train, cmap='viridis', cbar=True, yticklabels=False, ax=axes[0])\naxes[0].set_title('Missing Values Heatmap - Training Dataset', fontsize=14)\naxes[0].set_xlabel('Features', fontsize=12)\naxes[0].set_ylabel('Entries', fontsize=12)\n\n# Heatmap for missing values in the test dataset\nsns.heatmap(missing_test, cmap='viridis', cbar=True, yticklabels=False, ax=axes[1])\naxes[1].set_title('Missing Values Heatmap - Test Dataset', fontsize=14)\naxes[1].set_xlabel('Features', fontsize=12)\naxes[1].set_ylabel('Entries', fontsize=12)\n\nplt.tight_layout()\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-12-31T04:49:04.057351Z","iopub.execute_input":"2024-12-31T04:49:04.057626Z","iopub.status.idle":"2024-12-31T04:49:43.653348Z","shell.execute_reply.started":"2024-12-31T04:49:04.057603Z","shell.execute_reply":"2024-12-31T04:49:43.652535Z"},"papermill":{"duration":40.874305,"end_time":"2024-12-30T16:11:53.111286","exception":false,"start_time":"2024-12-30T16:11:12.236981","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"#### Insights from Missing Values Heatmaps:\n\n1. **Training Dataset:**\n   - Several features in the training dataset have missing values, as indicated by the yellow streaks in the heatmap.\n   - Features such as `Number of Dependents`, `Occupation`, `Previous Claims`, and `Credit Score` have a relatively high number of missing values.\n   - Some columns, such as `Policy Start Date` and `Gender`, seem to have no missing values as indicated by their continuous dark bands.\n\n2. **Test Dataset:**\n   - The test dataset also has missing values, with patterns similar to the training dataset.\n   - Features like `Previous Claims`, `Occupation`, and `Number of Dependents` show significant missing values, aligning with the training dataset's pattern.\n   - No additional features have missing values compared to the training dataset, which ensures consistency between the datasets.\n\n3. **Dataset Comparison:**\n   - The distribution of missing values appears consistent across both the training and test datasets, implying similar data collection or preprocessing methods were used.\n   - The percentage of missing values for features like `Credit Score` and `Previous Claims` might influence their impact on predictive models.","metadata":{"papermill":{"duration":0.014819,"end_time":"2024-12-30T16:11:53.139336","exception":false,"start_time":"2024-12-30T16:11:53.124517","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Function to calculate missing values, percentages, and data types\ndef missing_values_table(df):\n    missing_count = df.isnull().sum()\n    missing_percentage = 100 * missing_count / len(df)\n    data_types = df.dtypes\n    return pd.DataFrame({\n        'Missing Values': missing_count,\n        'Percentage (%)': missing_percentage,\n        'Data Type': data_types\n    })\n\n# Create tables for train and test datasets\ntrain_missing_table = missing_values_table(train_data)\ntest_missing_table = missing_values_table(test_data)\n\n# Display the tables\nprint(\"Missing Values Table - Training Dataset:\\n\")\ndisplay(train_missing_table[train_missing_table['Missing Values'] > 0])  # Display only features with missing values\nprint(\"\\n\")\n\nprint(\"Missing Values Table - Test Dataset:\\n\")\ndisplay(test_missing_table[test_missing_table['Missing Values'] > 0])  ","metadata":{"execution":{"iopub.status.busy":"2024-12-31T04:49:43.655394Z","iopub.execute_input":"2024-12-31T04:49:43.655691Z","iopub.status.idle":"2024-12-31T04:49:44.532397Z","shell.execute_reply.started":"2024-12-31T04:49:43.655668Z","shell.execute_reply":"2024-12-31T04:49:44.531516Z"},"papermill":{"duration":0.902978,"end_time":"2024-12-30T16:11:54.056804","exception":false,"start_time":"2024-12-30T16:11:53.153826","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"#### **Observational Insights from Missing Values Table:**\n\n#### Training Dataset:\n1. **Columns with High Missing Values:**\n   - `Occupation` (29.84%) and `Previous Claims` (30.34%) have the highest percentage of missing values. These features might significantly impact the dataset's completeness and require careful handling.\n   - `Credit Score` (11.49%) and `Number of Dependents` (9.14%) also have notable missing percentages.\n\n2. **Columns with Moderate Missing Values:**\n   - Features such as `Health Score` (6.17%) and `Customer Feedback` (6.49%) have moderate levels of missing values.\n\n3. **Columns with Minimal Missing Values:**\n   - Features like `Vehicle Age` (0.0005%) and `Insurance Duration` (0.00008%) have very few missing values, which can be easily imputed without much impact.\n\n4. **Data Types:**\n   - Features with missing values include both numerical (`float64`) and categorical (`object`) data types, indicating the need for distinct imputation strategies.\n\n\n#### Test Dataset:\n1. **Columns with High Missing Values:**\n   - Similar to the training dataset, `Occupation` (29.89%) and `Previous Claims` (30.35%) exhibit the highest percentage of missing values.\n   - `Credit Score` (11.43%) and `Number of Dependents` (9.14%) also show significant missingness.\n\n2. **Columns with Moderate Missing Values:**\n   - Features such as `Health Score` (6.18%) and `Customer Feedback` (6.53%) align closely with the training dataset in terms of missing values.\n\n3. **Columns with Minimal Missing Values:**\n   - Features like `Vehicle Age` (0.000375%) and `Insurance Duration` (0.00025%) have minimal missing data, mirroring the training dataset.\n\n4. **Consistency:**\n   - The percentages of missing values in the test dataset are highly consistent with the training dataset, making it easier to apply uniform imputation strategies.\n","metadata":{"papermill":{"duration":0.013065,"end_time":"2024-12-30T16:11:54.083913","exception":false,"start_time":"2024-12-30T16:11:54.070848","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Filter missing values for train and test datasets\ntrain_missing = train_missing_table[train_missing_table['Missing Values'] > 0].sort_values(by='Percentage (%)', ascending=False)\ntest_missing = test_missing_table[test_missing_table['Missing Values'] > 0].sort_values(by='Percentage (%)', ascending=False)\n\n# Set up the figure and subplots\nfig, axes = plt.subplots(1, 2, figsize=(14, 6), sharey=True)\n\n# Bar plot for train dataset\ntrain_colors = cm.get_cmap('viridis', len(train_missing))(range(len(train_missing)))\naxes[0].barh(train_missing.index, train_missing['Percentage (%)'], color=train_colors)\naxes[0].set_title('Percentage of Missing Values (Train Data)', fontsize=12)\naxes[0].set_xlabel('Percentage (%)', fontsize=10)\naxes[0].set_ylabel('Features', fontsize=10)\naxes[0].grid(axis='x', linestyle='--', alpha=0.6)\naxes[0].invert_yaxis()  \n\n# Bar plot for test dataset\ntest_colors = cm.get_cmap('viridis', len(test_missing))(range(len(test_missing)))\naxes[1].barh(test_missing.index, test_missing['Percentage (%)'], color=test_colors)\naxes[1].set_title('Percentage of Missing Values (Test Data)', fontsize=12)\naxes[1].set_xlabel('Percentage (%)', fontsize=10)\naxes[1].grid(axis='x', linestyle='--', alpha=0.6)\nplt.tight_layout()\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-12-31T04:49:44.533323Z","iopub.execute_input":"2024-12-31T04:49:44.53363Z","iopub.status.idle":"2024-12-31T04:49:45.017298Z","shell.execute_reply.started":"2024-12-31T04:49:44.533595Z","shell.execute_reply":"2024-12-31T04:49:45.016407Z"},"papermill":{"duration":0.514376,"end_time":"2024-12-30T16:11:54.611514","exception":false,"start_time":"2024-12-30T16:11:54.097138","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"#### Key Observations:\n- **Consistency Across Datasets:** The percentage of missing values in features is highly consistent between the training and test datasets, simplifying the imputation strategy.\n- **Critical Features:** High missing percentages in key features like `Previous Claims` and `Occupation` could impact model performance significantly if not addressed carefully.\n- **Potential Strategies:**\n  - Imputation: Use median or mean values for numerical features like `Credit Score` and `Number of Dependents`. For categorical features like `Occupation`, use the mode or `\"Unknown\"`.","metadata":{"papermill":{"duration":0.014479,"end_time":"2024-12-30T16:11:54.642723","exception":false,"start_time":"2024-12-30T16:11:54.628244","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Filter only features with missing values in the training dataset\nfeatures_with_missing = train_missing_table[train_missing_table['Missing Values'] > 0].index.tolist()\n\ndef analyze_nan_with_target_filtered(df, target_column, features):\n    missing_analysis = {}\n    \n    for col in features:\n        # Split the data into missing and non-missing subsets for the column\n        missing_mask = df[col].isnull()\n        non_missing_mask = ~missing_mask\n        \n        # Calculate statistics for Premium Amount\n        stats = {\n            \"Missing Count\": missing_mask.sum(),\n            \"Non-Missing Count\": non_missing_mask.sum(),\n            \"Mean (Missing)\": df.loc[missing_mask, target_column].mean(),\n            \"Mean (Non-Missing)\": df.loc[non_missing_mask, target_column].mean(),\n            \"Median (Missing)\": df.loc[missing_mask, target_column].median(),\n            \"Median (Non-Missing)\": df.loc[non_missing_mask, target_column].median(),\n            \"Std Dev (Missing)\": df.loc[missing_mask, target_column].std(),\n            \"Std Dev (Non-Missing)\": df.loc[non_missing_mask, target_column].std(),\n        }\n        \n        missing_analysis[col] = stats\n    \n    return pd.DataFrame(missing_analysis).T\n\n# Perform the analysis for only features with missing values\nmissing_vs_premium_filtered = analyze_nan_with_target_filtered(train_data, \"Premium Amount\", features_with_missing)\n\n# Display the results\nprint(\"Analysis of Missing Values with Target (Premium Amount):\\n\")\ndisplay(missing_vs_premium_filtered)\n","metadata":{"execution":{"iopub.status.busy":"2024-12-31T04:49:45.018028Z","iopub.execute_input":"2024-12-31T04:49:45.018236Z","iopub.status.idle":"2024-12-31T04:49:45.873131Z","shell.execute_reply.started":"2024-12-31T04:49:45.018217Z","shell.execute_reply":"2024-12-31T04:49:45.872387Z"},"papermill":{"duration":0.959637,"end_time":"2024-12-30T16:11:55.619023","exception":false,"start_time":"2024-12-30T16:11:54.659386","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# List of float-type columns with missing values\nfloat_missing_columns = ['Age', 'Annual Income', 'Number of Dependents', \n                         'Health Score', 'Previous Claims', 'Vehicle Age', \n                         'Credit Score', 'Insurance Duration']\n\nfor column in float_missing_columns:\n    # Drop NaN values for binning, but retain the NaN group separately\n    valid_data = train_data[column].dropna()\n    bins = 10  \n\n    # Bin the non-NaN values\n    binned_data = pd.cut(valid_data, bins)\n    df = pd.DataFrame({\n        column: binned_data,\n        'Premium Amount': train_data.loc[valid_data.index, 'Premium Amount']\n    })\n    \n    # Group by the binned column and calculate mean Premium Amount\n    grouped = df.groupby(column, observed=True, dropna=False).agg(\n        lambda x: np.expm1(np.log1p(x).mean())\n    ).reset_index()\n    \n    # Add the NaN group separately\n    nan_group_mean = np.expm1(np.log1p(train_data.loc[train_data[column].isnull(), 'Premium Amount']).mean())\n    nan_group = pd.DataFrame({column: ['NaN'], 'Premium Amount': [nan_group_mean]})\n    \n    # Concatenate the NaN group with the grouped data\n    grouped = pd.concat([grouped, nan_group], ignore_index=True)\n    \n    def label(x):\n        if isinstance(x, float) or x == 'NaN':\n            return x\n        x = x.mid  # Get the midpoint of the interval\n        s = int(np.floor(np.log10(x)))\n        return int(round(x, -s+1))\n\n    # Select the viridis color\n    viridis_colors = viridis(range(256))\n    line_color = viridis_colors[5]  \n\n    # Plot the non-NaN bins as a line plot\n    plt.plot(grouped[:-1]['Premium Amount'], marker='o', color=line_color, label='non-nan')\n\n    # Plot the NaN bin as a bar\n    plt.bar(len(grouped) - 1, grouped.iloc[-1]['Premium Amount'], color='red', label='nan')\n\n    # Set x-ticks and labels\n    plt.xticks(range(len(grouped)), labels=grouped[column].apply(label), fontsize=6)\n    plt.title(f'Average Premium Amount vs {column}')\n    plt.xlabel(column)\n    plt.ylabel('Premium Amount')\n    plt.legend(loc='upper right')\n    plt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-12-31T04:49:45.873906Z","iopub.execute_input":"2024-12-31T04:49:45.874211Z","iopub.status.idle":"2024-12-31T04:49:48.53153Z","shell.execute_reply.started":"2024-12-31T04:49:45.874177Z","shell.execute_reply":"2024-12-31T04:49:48.530378Z"},"papermill":{"duration":3.398433,"end_time":"2024-12-30T16:11:59.03435","exception":false,"start_time":"2024-12-30T16:11:55.635917","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# List of object-type columns with missing values\nobject_missing_columns = ['Marital Status', 'Occupation', 'Customer Feedback']\n\nfor column in object_missing_columns:\n    # Group by the column and calculate the average log-transformed Premium Amount\n    grouped = train_data.groupby(column)['Premium Amount'].agg(\n        lambda x: np.expm1(np.log1p(x).mean())  # Use log-transform to calculate mean\n    ).reset_index()\n    \n    # Calculate the average log-transformed Premium Amount for missing (NaN) values\n    nan_group_mean = np.expm1(np.log1p(train_data.loc[train_data[column].isnull(), 'Premium Amount']).mean())\n    \n    # Add a row for the NaN group\n    nan_group = pd.DataFrame({column: ['NaN'], 'Premium Amount': [nan_group_mean]})\n    grouped = pd.concat([grouped, nan_group], ignore_index=True)\n    \n    # Use viridis color palette for the bars\n    viridis_colors = viridis(range(256))  \n    bar_colors = [viridis_colors[5]] * (len(grouped) - 1) + ['red']  \n\n    # Plot the bar chart\n    plt.figure(figsize=(10, 6))\n    plt.bar(grouped[column].astype(str), grouped['Premium Amount'], color=bar_colors)\n    \n    # Set labels and title\n    plt.title(f'{column.upper()} Log-Transformed Average Premium Amounts')\n    plt.ylabel('Average Premium Amount')\n    plt.xlabel(column)\n    plt.xticks(rotation=45, ha='right')  \n    plt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-12-31T04:49:48.532559Z","iopub.execute_input":"2024-12-31T04:49:48.532893Z","iopub.status.idle":"2024-12-31T04:49:49.678581Z","shell.execute_reply.started":"2024-12-31T04:49:48.532869Z","shell.execute_reply":"2024-12-31T04:49:49.677488Z"},"papermill":{"duration":1.208018,"end_time":"2024-12-30T16:12:00.263793","exception":false,"start_time":"2024-12-30T16:11:59.055775","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Imputation Strategies for Numeric and Categorical Data","metadata":{"papermill":{"duration":0.026154,"end_time":"2024-12-30T16:12:00.317601","exception":false,"start_time":"2024-12-30T16:12:00.291447","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"### Filling Missing Values in Numeric Columns","metadata":{"papermill":{"duration":0.026055,"end_time":"2024-12-30T16:12:00.368959","exception":false,"start_time":"2024-12-30T16:12:00.342904","status":"completed"},"tags":[]}},{"cell_type":"code","source":"numeric_columns = train_data.select_dtypes(include=['number']).columns\n\nfor col in numeric_columns:\n    if col in test_data.columns:\n        # Impute missing values with -1\n        train_data[col].fillna(-1, inplace=True)\n        test_data[col].fillna(-1, inplace=True)\n","metadata":{"execution":{"iopub.status.busy":"2024-12-31T04:49:49.679541Z","iopub.execute_input":"2024-12-31T04:49:49.679869Z","iopub.status.idle":"2024-12-31T04:49:49.767387Z","shell.execute_reply.started":"2024-12-31T04:49:49.679838Z","shell.execute_reply":"2024-12-31T04:49:49.76649Z"},"papermill":{"duration":0.117554,"end_time":"2024-12-30T16:12:00.513219","exception":false,"start_time":"2024-12-30T16:12:00.395665","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Filling Missing Values in Object Columns","metadata":{"papermill":{"duration":0.024172,"end_time":"2024-12-30T16:12:00.666816","exception":false,"start_time":"2024-12-30T16:12:00.642644","status":"completed"},"tags":[]}},{"cell_type":"code","source":"object_columns = train_data.select_dtypes(include=['object']).columns\nfor col in object_columns:\n    if col in test_data.columns:\n        train_data[col].fillna(\"Unknown\", inplace=True)\n        test_data[col].fillna(\"Unknown\", inplace=True)\n","metadata":{"execution":{"iopub.status.busy":"2024-12-31T04:49:49.773386Z","iopub.execute_input":"2024-12-31T04:49:49.773741Z","iopub.status.idle":"2024-12-31T04:49:50.851302Z","shell.execute_reply.started":"2024-12-31T04:49:49.773711Z","shell.execute_reply":"2024-12-31T04:49:50.850368Z"},"papermill":{"duration":1.093243,"end_time":"2024-12-30T16:12:01.784598","exception":false,"start_time":"2024-12-30T16:12:00.691355","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"- For each categorical column:\n    - Missing values in both the `train_data` and `test_data` for that column are replaced with the string `\"Unknown\"`.","metadata":{"papermill":{"duration":0.024924,"end_time":"2024-12-30T16:12:01.83436","exception":false,"start_time":"2024-12-30T16:12:01.809436","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Verify Missing Values\n\nprint(\"Missing Values After Imputation - Training Dataset:\")\nprint(train_data.isnull().sum())\n\nprint(\"\\nMissing Values After Imputation - Test Dataset:\")\nprint(test_data.isnull().sum())","metadata":{"execution":{"iopub.status.busy":"2024-12-31T04:49:50.852031Z","iopub.execute_input":"2024-12-31T04:49:50.852243Z","iopub.status.idle":"2024-12-31T04:49:51.984601Z","shell.execute_reply.started":"2024-12-31T04:49:50.852223Z","shell.execute_reply":"2024-12-31T04:49:51.983684Z"},"papermill":{"duration":0.933883,"end_time":"2024-12-30T16:12:02.792658","exception":false,"start_time":"2024-12-30T16:12:01.858775","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"#### **Missing Values Successfully Imputed**\n   - After executing the imputation logic for both numeric and categorical columns:\n     - All missing values in the training dataset (`train_data`) have been successfully replaced. \n     - All missing values in the test dataset (`test_data`) have been successfully replaced.\n   - There are **no remaining NaN values** in any column for either dataset.\n\n#### **Key Observations**\n   - **Numeric Columns:** Missing values were filled using the **median** value of the respective column in the training dataset. This approach ensures that the central tendency of the data is preserved while mitigating the impact of outliers.\n   - **Categorical Columns:** Missing values were filled with the placeholder `\"Unknown\"`. This ensures that the absence of data is explicitly marked without distorting the distribution of existing categories.","metadata":{"papermill":{"duration":0.025356,"end_time":"2024-12-30T16:12:02.843124","exception":false,"start_time":"2024-12-30T16:12:02.817768","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Check for duplicate rows in the training dataset\ntrain_duplicates = train_data.duplicated().sum()\nprint(f\"\\nNumber of duplicate rows in the training dataset: {train_duplicates}\")\n\n# Check for duplicate rows in the test dataset\ntest_duplicates = test_data.duplicated().sum()\nprint(f\"Number of duplicate rows in the test dataset: {test_duplicates}\")","metadata":{"execution":{"iopub.status.busy":"2024-12-31T04:49:51.985562Z","iopub.execute_input":"2024-12-31T04:49:51.985927Z","iopub.status.idle":"2024-12-31T04:49:54.273288Z","shell.execute_reply.started":"2024-12-31T04:49:51.985892Z","shell.execute_reply":"2024-12-31T04:49:54.272303Z"},"papermill":{"duration":2.537648,"end_time":"2024-12-30T16:12:05.405634","exception":false,"start_time":"2024-12-30T16:12:02.867986","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"After checking for duplicate rows in both datasets:\n - **Training Dataset (`train_data`)**: There are **0 duplicate rows** detected.\n - **Test Dataset (`test_data`)**: There are **0 duplicate rows** detected.","metadata":{"papermill":{"duration":0.028223,"end_time":"2024-12-30T16:12:05.460018","exception":false,"start_time":"2024-12-30T16:12:05.431795","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"# <span style=\"color:transparent;\">Exploratory Data Analysis (EDA)</span>\n\n<div style=\"border-radius: 15px; border: 2px solid #6A1B9A; padding: 10px; background: linear-gradient(135deg, #9C27B0, #4CAF50); text-align: center; box-shadow: 0px 4px 8px rgba(0, 0, 0, 0.5);\">\n    <h1 style=\"color: #ffffff; text-shadow: 2px 2px 4px rgba(0, 0, 0, 0.7); font-weight: bold; margin-bottom: 5px; font-size: 28px; font-family: 'Roboto', sans-serif;\">\n        Exploratory Data Analysis (EDA)\n    </h1>\n</div>\n\n<!-- Include Google Fonts for a modern font -->\n<link href=\"https://fonts.googleapis.com/css2?family=Roboto:wght@700&display=swap\" rel=\"stylesheet\">\n","metadata":{"papermill":{"duration":0.026508,"end_time":"2024-12-30T16:12:05.513185","exception":false,"start_time":"2024-12-30T16:12:05.486677","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"## Target Column Extraction and Visualizing Distribution","metadata":{"papermill":{"duration":0.024697,"end_time":"2024-12-30T16:12:05.56374","exception":false,"start_time":"2024-12-30T16:12:05.539043","status":"completed"},"tags":[]}},{"cell_type":"code","source":"target_column = (set(train_data.columns) - set(test_data.columns)).pop()\n\nprint(f\"Target column: {target_column}\")\nprint(f\"Data type: {train_data[target_column].dtype}\")\n","metadata":{"execution":{"iopub.status.busy":"2024-12-31T04:49:54.274197Z","iopub.execute_input":"2024-12-31T04:49:54.274475Z","iopub.status.idle":"2024-12-31T04:49:54.279374Z","shell.execute_reply.started":"2024-12-31T04:49:54.274427Z","shell.execute_reply":"2024-12-31T04:49:54.278399Z"},"papermill":{"duration":0.032687,"end_time":"2024-12-30T16:12:05.623301","exception":false,"start_time":"2024-12-30T16:12:05.590614","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Custom colormap using viridis\nviridis_cmap = cm.get_cmap(\"viridis\")\n\ndef visualize_premium_amount_with_peaks(data, feature='Premium Amount'):\n    plt.figure(figsize=(9, 4))\n\n    # Histogram with KDE\n    plt.subplot(1, 2, 1)\n    ax = sns.histplot(data[feature], bins=30, kde=True, color=viridis_cmap(0.5))\n    plt.title(f'Histogram of {feature} with KDE', fontsize=11)\n    plt.xlabel(feature, fontsize=10)\n    plt.ylabel('Frequency', fontsize=10)\n    plt.grid(True, linestyle='--', alpha=0.6) \n\n    # Extract KDE values to find peaks\n    kde = sns.kdeplot(data[feature], ax=ax, color=viridis_cmap(0.7)).lines[0].get_data()\n    kde_x, kde_y = kde[0], kde[1]\n    peaks, _ = find_peaks(kde_y)\n\n    # Highlight peaks\n    for peak_idx in peaks:\n        plt.plot(kde_x[peak_idx], kde_y[peak_idx], \"ro\")  # Red dots on peaks\n\n    # Box Plot\n    plt.subplot(1, 2, 2)\n    sns.boxplot(x=data[feature], color=viridis_cmap(0.5))\n    plt.title(f'Box Plot of {feature}', fontsize=11)\n    plt.xlabel(feature, fontsize=10)\n    plt.grid(True, linestyle='--', alpha=0.6)  \n    plt.tight_layout()\n    plt.show()\n\nvisualize_premium_amount_with_peaks(train_data, feature='Premium Amount')\n","metadata":{"execution":{"iopub.status.busy":"2024-12-31T04:49:54.280387Z","iopub.execute_input":"2024-12-31T04:49:54.280706Z","iopub.status.idle":"2024-12-31T04:50:04.020395Z","shell.execute_reply.started":"2024-12-31T04:49:54.280675Z","shell.execute_reply":"2024-12-31T04:50:04.019549Z"},"papermill":{"duration":9.919777,"end_time":"2024-12-30T16:12:15.568625","exception":false,"start_time":"2024-12-30T16:12:05.648848","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"1. **Distribution Characteristics**\n   - The `Premium Amount` shows a **right-skewed distribution**, meaning most of the values are concentrated towards lower premiums, with fewer instances of higher premium amounts.\n   - The red peaks highlight the **modes** (local maxima), indicating clusters where specific premium amounts are more common. The highest mode appears around 1000, suggesting this is a prevalent premium value in the dataset.\n\n2. **Outliers and Variability**\n   - The box plot demonstrates that a significant number of premium values fall within a lower range, as shown by the interquartile range (IQR).\n   - There are **many outliers** present, representing higher premium amounts that lie far beyond the upper whisker of the box plot. These outliers may require further analysis to understand their nature or consider their treatment in modeling (e.g., log transformation).","metadata":{"papermill":{"duration":0.025826,"end_time":"2024-12-30T16:12:15.621142","exception":false,"start_time":"2024-12-30T16:12:15.595316","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"## Distribution Analysis of Numerical Features","metadata":{"papermill":{"duration":0.025878,"end_time":"2024-12-30T16:12:15.673245","exception":false,"start_time":"2024-12-30T16:12:15.647367","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Define columns to analyze\ncolumns_to_analyze = train_data.select_dtypes(include=['number']).columns.drop('Premium Amount')\n\nviridis_cmap = cm.get_cmap(\"viridis\")\n# Extract three colors from the colormap\nviridis_colors = [viridis_cmap(0.3), viridis_cmap(0.5), viridis_cmap(0.8)]\n\nfig, axes = plt.subplots(len(columns_to_analyze), 3, figsize=(25, len(columns_to_analyze) * 5))\n\nfor i, column in enumerate(columns_to_analyze):\n    # Histogram for train_data\n    sns.histplot(train_data[column], bins=30, kde=True, color=viridis_colors[0], ax=axes[i, 0])\n    axes[i, 0].set_title(f'Distribution of {column} (Train)', fontsize=14)\n    axes[i, 0].set_xlabel(column, fontsize=10)\n    axes[i, 0].set_ylabel('Frequency', fontsize=10)\n    axes[i, 0].grid(visible=True, linestyle='--', alpha=0.6)\n\n    # Boxplot for train_data\n    sns.boxplot(x=train_data[column], color=viridis_colors[1], ax=axes[i, 1])\n    axes[i, 1].set_title(f'Boxplot of {column} (Train)', fontsize=14)\n    axes[i, 1].set_xlabel(column, fontsize=10)\n    axes[i, 1].grid(visible=True, linestyle='--', alpha=0.6)\n\n    # Boxplot for test_data\n    sns.boxplot(x=test_data[column], color=viridis_colors[2], ax=axes[i, 2])\n    axes[i, 2].set_title(f'Boxplot of {column} (Test)', fontsize=12)\n    axes[i, 2].set_xlabel(column, fontsize=10)\n    axes[i, 2].grid(visible=True, linestyle='--', alpha=0.6)\n\nplt.tight_layout()\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-12-31T04:50:04.021254Z","iopub.execute_input":"2024-12-31T04:50:04.021515Z"},"papermill":{"duration":40.534236,"end_time":"2024-12-30T16:12:56.233727","exception":false,"start_time":"2024-12-30T16:12:15.699491","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Select numeric columns \nnumeric_data = train_data.select_dtypes(include=['number'])\n\n# Compute the correlation matrix\ncorrelation_matrix = numeric_data.corr()\nplt.figure(figsize=(8, 6))\n\n# Create the heatmap\nsns.heatmap(\n    correlation_matrix, \n    annot=True, \n    fmt=\".2f\", \n    cmap='viridis',  \n    cbar=True, \n    square=True,\n    mask=np.triu(np.ones_like(correlation_matrix, dtype=bool)),  \n    linewidths=0.5  \n)\n\nplt.title('Correlation Heatmap of Numerical Features (Excluding Target)', fontsize=12)\nplt.xticks(rotation=45, ha='right', fontsize=10)\nplt.yticks(fontsize=10)\nplt.tight_layout()\nplt.show()\n","metadata":{"execution":{"iopub.execute_input":"2024-12-31T04:50:43.407316Z","iopub.status.idle":"2024-12-31T04:50:44.166799Z","shell.execute_reply.started":"2024-12-31T04:50:43.407289Z","shell.execute_reply":"2024-12-31T04:50:44.166019Z"},"papermill":{"duration":0.694824,"end_time":"2024-12-30T16:12:56.971115","exception":false,"start_time":"2024-12-30T16:12:56.276291","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Low Correlation Between Features:**\n   - The majority of the features exhibit very low correlation values (close to 0), indicating that they are weakly related or independent.\n   - This suggests minimal multicollinearity, which is beneficial for model training as it reduces the risk of redundancy in predictive features.\n\n**Key Observations:**\n   - **Credit Score vs Annual Income:** Shows a slightly negative correlation (-0.20), suggesting that individuals with lower income might have slightly higher credit scores or vice versa.\n   - **Premium Amount vs Previous Claims:** The correlation value is positive (0.05), indicating that individuals with higher previous claims tend to have slightly higher premium amounts.\n   - **Health Score vs Premium Amount:** A weak positive correlation (0.01) suggests minimal influence of health score on premium amount.\n\n**Premium Amount Correlation:**\n   - Most features show minimal correlation with the target variable (`Premium Amount`), highlighting the potential for non-linear relationships that might require advanced modeling techniques (e.g., tree-based algorithms).","metadata":{"papermill":{"duration":0.044336,"end_time":"2024-12-30T16:12:57.061972","exception":false,"start_time":"2024-12-30T16:12:57.017636","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"## Categorical Feature Analysis","metadata":{"papermill":{"duration":0.044196,"end_time":"2024-12-30T16:12:57.150885","exception":false,"start_time":"2024-12-30T16:12:57.106689","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Function to display barplot and pie chart for categorical columns\ndef plot_categorical_distribution(data, column_name):\n    plt.figure(figsize=(12, 4))\n    \n    # Bar plot for categorical distribution\n    plt.subplot(1, 2, 1)\n    sns.countplot(y=column_name, data=data, palette='Set2')\n    plt.title(f'Distribution of {column_name}', fontsize=12)\n    plt.xlabel('Count', fontsize=10)\n    plt.ylabel(column_name, fontsize=10)\n\n    ax = plt.gca()\n    for p in ax.patches:\n        count = int(p.get_width())\n        ax.annotate(f'{count}', \n                    (p.get_width() + 0.1, p.get_y() + p.get_height() / 2), \n                    ha='left', va='center', fontsize=10, color='black')\n    \n    sns.despine(left=True, bottom=True)\n    \n    # Pie chart for percentage distribution\n    plt.subplot(1, 2, 2)\n    data[column_name].value_counts().plot.pie(\n        autopct='%1.1f%%', \n        colors=sns.color_palette('Set2', data[column_name].nunique()), \n        startangle=90, \n        explode=[0.05] * data[column_name].nunique(), \n        shadow=True\n    )\n    plt.title(f'Percentage Distribution of {column_name}', fontsize=12)\n    plt.ylabel('')  \n\n    plt.tight_layout()\n    plt.show()\n\ncategorical_columns = ['Gender', 'Marital Status', 'Education Level', 'Occupation', 'Location', \n                'Policy Type', 'Customer Feedback', 'Smoking Status', 'Exercise Frequency', 'Property Type']\n\nfor column in categorical_columns:\n    plot_categorical_distribution(train_data, column)","metadata":{"execution":{"iopub.status.busy":"2024-12-31T04:50:44.168938Z","iopub.execute_input":"2024-12-31T04:50:44.169271Z","iopub.status.idle":"2024-12-31T04:50:53.478395Z","shell.execute_reply.started":"2024-12-31T04:50:44.169249Z","shell.execute_reply":"2024-12-31T04:50:53.477423Z"},"papermill":{"duration":10.310672,"end_time":"2024-12-30T16:13:07.506423","exception":false,"start_time":"2024-12-30T16:12:57.195751","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Function to calculate count and percentage of unique values\ndef unique_values_table(data, categorical_columns):\n    results = {}\n    for column in categorical_columns:\n        value_counts = data[column].value_counts()\n        percentages = (value_counts / len(data)) * 100\n        results[column] = pd.DataFrame({\n            'Value': value_counts.index,\n            'Count': value_counts.values,\n            'Percentage (%)': percentages.values\n        })\n    return results\n\n# Specify categorical columns\ncategorical_columns = ['Gender', 'Marital Status', 'Education Level', 'Occupation', 'Location', \n                       'Policy Type', 'Customer Feedback', 'Smoking Status', \n                       'Exercise Frequency', 'Property Type']\n\n# Get unique value tables for each categorical column\nunique_values_results = unique_values_table(train_data, categorical_columns)\n\nfor column in categorical_columns[:10]:  \n    print(f\"Unique Values for {column}:\\n\")\n    display(unique_values_results[column])\n","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-12-31T04:50:53.479343Z","iopub.execute_input":"2024-12-31T04:50:53.479648Z","iopub.status.idle":"2024-12-31T04:50:54.249163Z","shell.execute_reply.started":"2024-12-31T04:50:53.479615Z","shell.execute_reply":"2024-12-31T04:50:54.248298Z"},"papermill":{"duration":0.841833,"end_time":"2024-12-30T16:13:08.408582","exception":false,"start_time":"2024-12-30T16:13:07.566749","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"#### Observational Insights:\n\n1. **Gender Distribution:**\n   - The dataset is balanced with a nearly equal proportion of males (50.21%) and females (49.79%).\n\n2. **Marital Status:**\n   - The three primary categories (Single, Married, Divorced) are distributed almost equally, each accounting for about 33%.\n   - A small percentage (1.54%) of data is categorized as `Unknown`, which could be imputed or treated based on the analysis.\n\n3. **Education Level:**\n   - Education levels are uniformly distributed among Bachelor's, Master's, PhD, and High School, with percentages ranging from 24% to 25%.\n   - No major outlier category exists in education, making this feature consistent.\n\n4. **Occupation:**\n   - A significant portion (29.84%) of data falls under the `Unknown` category, indicating missing values that were replaced during preprocessing.\n   - The remaining categories (Employed, Self-Employed, Unemployed) are evenly distributed around 23%.\n\n5. **Location:**\n   - The dataset shows an almost equal split between `Suburban`, `Rural`, and `Urban` areas, each contributing about 33% of the data.\n\n6. **Policy Type:**\n   - The types of policies (`Premium`, `Comprehensive`, `Basic`) are evenly distributed, with each type accounting for approximately one-third of the data.\n\n7. **Customer Feedback:**\n   - Feedback categories (`Average`, `Poor`, `Good`) account for around 31% each, with an `Unknown` category representing 6.48%.\n   - The presence of `Unknown` suggests gaps in feedback collection that might need attention.\n\n8. **Smoking Status:**\n   - A balanced distribution exists between `Yes` (50.16%) and `No` (49.84%) categories, making it a useful feature for potential analysis.\n\n9. **Exercise Frequency:**\n   - The dataset is evenly split across all four categories (`Weekly`, `Monthly`, `Rarely`, `Daily`), each contributing about 25%.\n\n10. **Property Type:**\n    - The three property types (`House`, `Apartment`, `Condo`) are evenly distributed, each accounting for approximately 33%.\n\n#### Key Observations:\n- Many categorical features have balanced distributions, which can be advantageous for model training as it reduces bias.\n- Certain features, such as `Occupation`, `Customer Feedback`, and `Marital Status`, contain `Unknown` categories due to imputed missing values. Their treatment depends on the modeling approach.\n- Most features exhibit uniform distribution across their categories, suggesting no dominance of a single class, which could enhance feature diversity in the model.","metadata":{"papermill":{"duration":0.062114,"end_time":"2024-12-30T16:13:08.532972","exception":false,"start_time":"2024-12-30T16:13:08.470858","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"## Categorical Features vs Premium Amount","metadata":{"papermill":{"duration":0.061576,"end_time":"2024-12-30T16:13:08.655384","exception":false,"start_time":"2024-12-30T16:13:08.593808","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# List of categorical columns\ncategorical_columns = [\n    'Gender', 'Marital Status', 'Education Level', 'Occupation', 'Location', \n    'Policy Type', 'Customer Feedback', 'Smoking Status', 'Exercise Frequency', 'Property Type'\n]\n\n# Loop through each categorical feature to display summary statistics and box plot\nfor column in categorical_columns:\n    # Calculate summary statistics grouped by the categorical column\n    stats = train_data.groupby(column)['Premium Amount'].agg(['mean', 'median', 'count'])\n    \n    # Display summary statistics\n    print(f\"\\nSummary Statistics for Premium Amount by {column}:\")\n    print(stats)\n    \n    # Plot box plot\n    plt.figure(figsize=(8, 4))\n    sns.boxplot(data=train_data, x=column, y='Premium Amount', palette='viridis')\n    plt.title(f'Premium Amount by {column}', fontsize=12)\n    plt.xlabel(column, fontsize=11)\n    plt.ylabel('Premium Amount', fontsize=11)\n    plt.xticks(rotation=45)\n    plt.tight_layout()\n    plt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-12-31T04:50:54.250047Z","iopub.execute_input":"2024-12-31T04:50:54.250363Z","iopub.status.idle":"2024-12-31T04:51:02.238666Z","shell.execute_reply.started":"2024-12-31T04:50:54.250333Z","shell.execute_reply":"2024-12-31T04:51:02.237669Z"},"papermill":{"duration":8.9391,"end_time":"2024-12-30T16:13:17.655018","exception":false,"start_time":"2024-12-30T16:13:08.715918","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <span style=\"color:transparent;\">Data Preprocessing</span>\n\n<div style=\"border-radius: 15px; border: 2px solid #6A1B9A; padding: 10px; background: linear-gradient(135deg, #9C27B0, #4CAF50); text-align: center; box-shadow: 0px 4px 8px rgba(0, 0, 0, 0.5);\">\n    <h1 style=\"color: #ffffff; text-shadow: 2px 2px 4px rgba(0, 0, 0, 0.7); font-weight: bold; margin-bottom: 5px; font-size: 28px; font-family: 'Roboto', sans-serif;\">\n        Data Preprocessing\n    </h1>\n</div>\n\n<!-- Include Google Fonts for a modern font -->\n<link href=\"https://fonts.googleapis.com/css2?family=Roboto:wght@700&display=swap\" rel=\"stylesheet\">\n","metadata":{"papermill":{"duration":0.070007,"end_time":"2024-12-30T16:13:17.79345","exception":false,"start_time":"2024-12-30T16:13:17.723443","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"## Converting Date Columns to Epoch Time","metadata":{"papermill":{"duration":0.069763,"end_time":"2024-12-30T16:13:17.932158","exception":false,"start_time":"2024-12-30T16:13:17.862395","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Retrieve columns with 'object' data type\ndatetime_columns = train_data.select_dtypes(include=['object']).columns\n\nfor col in datetime_columns:\n    try:\n        # Convert the column to datetime format\n        train_data[col] = pd.to_datetime(train_data[col], errors='raise')\n        test_data[col] = pd.to_datetime(test_data[col], errors='raise')\n        \n        # Convert datetime to epoch time (in seconds)\n        train_data[col] = train_data[col].astype(np.int64) / 10**9\n        test_data[col] = test_data[col].astype(np.int64) / 10**9\n\n        print(f\"Converted '{col}' to epoch time.\")\n    except Exception as e:\n        print(f\"Skipping column '{col}' due to: {e}\")\n","metadata":{"execution":{"iopub.status.busy":"2024-12-31T04:51:02.239513Z","iopub.execute_input":"2024-12-31T04:51:02.239967Z","iopub.status.idle":"2024-12-31T04:51:03.539767Z","shell.execute_reply.started":"2024-12-31T04:51:02.239933Z","shell.execute_reply":"2024-12-31T04:51:03.53893Z"},"papermill":{"duration":1.380901,"end_time":"2024-12-30T16:13:19.378931","exception":false,"start_time":"2024-12-30T16:13:17.99803","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Check the first few rows and data type of the 'Policy Start Date' column\ncolumn_name = 'Policy Start Date'\n\n# Display the first few rows\nprint(f\"Sample data for '{column_name}':\\n\", train_data[column_name].head())\n\n# Display the data type of the column\nprint(f\"Data type of '{column_name}' in train_data: {train_data[column_name].dtype}\")\n\n# Repeat for test_data\nprint(f\"Sample data for '{column_name}' in test_data:\\n\", test_data[column_name].head())\nprint(f\"Data type of '{column_name}' in test_data: {test_data[column_name].dtype}\")\n","metadata":{"execution":{"iopub.status.busy":"2024-12-31T04:51:03.540739Z","iopub.execute_input":"2024-12-31T04:51:03.541066Z","iopub.status.idle":"2024-12-31T04:51:03.548834Z","shell.execute_reply.started":"2024-12-31T04:51:03.54103Z","shell.execute_reply":"2024-12-31T04:51:03.548141Z"},"papermill":{"duration":0.076635,"end_time":"2024-12-30T16:13:19.522839","exception":false,"start_time":"2024-12-30T16:13:19.446204","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"- Successfully converts the 'Policy Start Date' column to datetime and then to epoch time.\n- The converted Policy Start Date values (e.g., 1.703345e+09) are epoch times in seconds.\n- The float64 data type indicates that the values are now numerical.","metadata":{"papermill":{"duration":0.067662,"end_time":"2024-12-30T16:13:19.660095","exception":false,"start_time":"2024-12-30T16:13:19.592433","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"## Label Encoding Categorical Features","metadata":{"papermill":{"duration":0.067724,"end_time":"2024-12-30T16:13:19.796212","exception":false,"start_time":"2024-12-30T16:13:19.728488","status":"completed"},"tags":[]}},{"cell_type":"code","source":"def identify_non_numerical_features(data, dataset_name):\n    non_numerical_features = data.select_dtypes(include=['object'])\n    print(f\"Non-Numerical Features and Unique Values in {dataset_name} dataset:\")\n    for column in non_numerical_features.columns:\n        unique_values = non_numerical_features[column].unique()\n        print(f\"\\n{column}: {unique_values}\")\n\n# Apply the function to training and test datasets\nidentify_non_numerical_features(train_data, \"Training\")\nprint(\"\\n\")\nidentify_non_numerical_features(test_data, \"Test\")","metadata":{"execution":{"iopub.status.busy":"2024-12-31T04:51:03.549789Z","iopub.execute_input":"2024-12-31T04:51:03.550087Z","iopub.status.idle":"2024-12-31T04:51:05.035356Z","shell.execute_reply.started":"2024-12-31T04:51:03.550054Z","shell.execute_reply":"2024-12-31T04:51:05.034613Z"},"papermill":{"duration":1.763545,"end_time":"2024-12-30T16:13:21.628107","exception":false,"start_time":"2024-12-30T16:13:19.864562","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Define the encoding strategies for specific features\nbinary_features = ['Gender', 'Smoking Status']\nordinal_features = {\n    'Exercise Frequency': ['Rarely', 'Monthly', 'Weekly', 'Daily']\n}\nnominal_features = ['Marital Status', 'Education Level', 'Occupation', \n                    'Location', 'Policy Type', 'Customer Feedback', 'Property Type']\n\n# Binary Encoding for binary features\nle = LabelEncoder()\nfor feature in binary_features:\n    train_data[feature] = le.fit_transform(train_data[feature])\n    test_data[feature] = le.transform(test_data[feature])\n\n# Ordinal Encoding for ordered features\nfor feature, order in ordinal_features.items():\n    oe = OrdinalEncoder(categories=[order])\n    train_data[feature] = oe.fit_transform(train_data[[feature]]).flatten()  # Flatten to 1D\n    test_data[feature] = oe.transform(test_data[[feature]]).flatten()       # Flatten to 1D\n\n# One-Hot Encoding for nominal features\ntrain_data = pd.get_dummies(train_data, columns=nominal_features, drop_first=True)\ntest_data = pd.get_dummies(test_data, columns=nominal_features, drop_first=True)\n","metadata":{"execution":{"iopub.status.busy":"2024-12-31T04:51:05.036157Z","iopub.execute_input":"2024-12-31T04:51:05.036377Z","iopub.status.idle":"2024-12-31T04:51:07.138003Z","shell.execute_reply.started":"2024-12-31T04:51:05.036351Z","shell.execute_reply":"2024-12-31T04:51:07.137301Z"},"papermill":{"duration":2.413709,"end_time":"2024-12-30T16:13:24.109838","exception":false,"start_time":"2024-12-30T16:13:21.696129","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Verifying Data Types Across Datasets","metadata":{"papermill":{"duration":0.070248,"end_time":"2024-12-30T16:13:24.396225","exception":false,"start_time":"2024-12-30T16:13:24.325977","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Create data type tables for train_data and test_data\ntrain_data_types = pd.DataFrame({\n    'Column Name': train_data.columns,\n    'Train Data Type': train_data.dtypes\n})\n\ntest_data_types = pd.DataFrame({\n    'Column Name': test_data.columns,\n    'Test Data Type': test_data.dtypes\n})\n\n# Merge the two tables for comparison\ndata_types_comparison = pd.merge(\n    train_data_types, \n    test_data_types, \n    on='Column Name', \n    how='outer'\n)\n\n# Display the data types comparison table\nprint(\"Data Types Comparison of Train and Test Datasets:\\n\")\ndisplay(data_types_comparison)\n","metadata":{"execution":{"iopub.status.busy":"2024-12-31T04:51:07.138816Z","iopub.execute_input":"2024-12-31T04:51:07.139115Z","iopub.status.idle":"2024-12-31T04:51:07.156711Z","shell.execute_reply.started":"2024-12-31T04:51:07.139075Z","shell.execute_reply":"2024-12-31T04:51:07.155857Z"},"papermill":{"duration":0.09384,"end_time":"2024-12-30T16:13:24.561711","exception":false,"start_time":"2024-12-30T16:13:24.467871","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Normalization of Numerical Features","metadata":{"papermill":{"duration":0.066894,"end_time":"2024-12-30T16:13:24.696926","exception":false,"start_time":"2024-12-30T16:13:24.630032","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Select numerical columns\nnumerical_columns = train_data.select_dtypes(include=['float64']).columns\nnumerical_columns = numerical_columns[numerical_columns != target_column]\n\n# Applying Normalization\nscaler = StandardScaler()\ntrain_data[numerical_columns] = scaler.fit_transform(train_data[numerical_columns])\ntest_data[numerical_columns] = scaler.transform(test_data[numerical_columns])\n","metadata":{"execution":{"iopub.status.busy":"2024-12-31T04:51:07.157739Z","iopub.execute_input":"2024-12-31T04:51:07.158051Z","iopub.status.idle":"2024-12-31T04:51:07.476236Z","shell.execute_reply.started":"2024-12-31T04:51:07.158019Z","shell.execute_reply":"2024-12-31T04:51:07.47559Z"},"papermill":{"duration":0.404768,"end_time":"2024-12-30T16:13:25.170511","exception":false,"start_time":"2024-12-30T16:13:24.765743","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"- The StandardScaler is used to standardize numerical features by removing the mean and scaling to unit variance.\n- The scaler is fit on the train_data to calculate the mean and standard deviation and then applied to both train_data and test_data.\n\n#### Why This Step is Important:\n- Normalization ensures that all numerical features contribute equally to the model, preventing features with larger magnitudes (e.g., `Annual Income`) from dominating those with smaller values (e.g., `Vehicle Age`).","metadata":{"papermill":{"duration":0.06807,"end_time":"2024-12-30T16:13:25.308325","exception":false,"start_time":"2024-12-30T16:13:25.240255","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"## Distribution Analysis of Preprocessed Features","metadata":{"papermill":{"duration":0.069609,"end_time":"2024-12-30T16:13:25.448274","exception":false,"start_time":"2024-12-30T16:13:25.378665","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Identify numeric columns only (excluding boolean columns)\nnumeric_columns = train_data.select_dtypes(include=[np.number]).columns\n\n# Calculate the number of rows and columns needed\nnum_features = len(numeric_columns)\nnum_cols = 4\nnum_rows = math.ceil(num_features / num_cols)\n\n# Create subplots\nfig, axes = plt.subplots(nrows=num_rows, ncols=num_cols, figsize=(16, num_rows * 3))\nviridis_cmap = cm.get_cmap('viridis', len(numeric_columns))\n\n# Plot each numeric column\nfor i, column in enumerate(numeric_columns):\n    ax = axes.flatten()[i]\n    train_data[column].hist(\n        ax=ax, \n        bins=20, \n        color=viridis_cmap(i / len(numeric_columns)),  \n        edgecolor='black', \n        linewidth=0.5\n    )\n    ax.set_title(column, fontsize=9)\n    ax.tick_params(axis='both', which='major', labelsize=6)\n    ax.grid(True, linestyle='--', alpha=0.6)  \n\n# Remove empty subplots if any\nfor j in range(i + 1, len(axes.flatten())):\n    fig.delaxes(axes.flatten()[j])\n\nplt.suptitle('Dataset Feature Distributions (train_data)', fontsize=11)\nplt.tight_layout()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-12-31T04:51:07.477016Z","iopub.execute_input":"2024-12-31T04:51:07.47725Z","iopub.status.idle":"2024-12-31T04:51:10.212579Z","shell.execute_reply.started":"2024-12-31T04:51:07.477231Z","shell.execute_reply":"2024-12-31T04:51:10.21168Z"},"papermill":{"duration":3.065083,"end_time":"2024-12-30T16:13:28.582735","exception":false,"start_time":"2024-12-30T16:13:25.517652","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Skewness Reduction with Log Transformation","metadata":{"papermill":{"duration":0.072067,"end_time":"2024-12-30T16:13:28.778219","exception":false,"start_time":"2024-12-30T16:13:28.706152","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Define the continuous columns\ncontinuous_columns_train = ['Annual Income', 'Premium Amount']  \ncontinuous_columns_test = ['Annual Income']  \n\n# Calculate skewness for the specified continuous columns\ntrain_skewness = train_data[continuous_columns_train].apply(skew)\ntest_skewness = test_data[continuous_columns_test].apply(skew)\n\n# Display results\nprint(\"Skewness for Training Dataset:\\n\")\ndisplay(train_skewness)\n\nprint(\"\\nSkewness for Test Dataset:\\n\")\ndisplay(test_skewness)","metadata":{"execution":{"iopub.status.busy":"2024-12-31T04:51:10.21335Z","iopub.execute_input":"2024-12-31T04:51:10.213597Z","iopub.status.idle":"2024-12-31T04:51:10.272757Z","shell.execute_reply.started":"2024-12-31T04:51:10.213575Z","shell.execute_reply":"2024-12-31T04:51:10.27204Z"},"papermill":{"duration":0.149733,"end_time":"2024-12-30T16:13:29.002123","exception":false,"start_time":"2024-12-30T16:13:28.85239","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Log-transform skewed features\ntrain_data['Annual Income'] = np.log1p(train_data['Annual Income'])\ntest_data['Annual Income'] = np.log1p(test_data['Annual Income'])","metadata":{"execution":{"iopub.status.busy":"2024-12-31T04:51:10.273613Z","iopub.execute_input":"2024-12-31T04:51:10.273949Z","iopub.status.idle":"2024-12-31T04:51:10.284592Z","shell.execute_reply.started":"2024-12-31T04:51:10.273917Z","shell.execute_reply":"2024-12-31T04:51:10.283746Z"},"papermill":{"duration":0.089755,"end_time":"2024-12-30T16:13:29.164338","exception":false,"start_time":"2024-12-30T16:13:29.074583","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Select only numeric columns from the training data\nnumeric_data = train_data.select_dtypes(include=['number'])\n\n# Add a new column for the log-transformed values of target variable\nnumeric_data['Log_Transformed_Premium'] = np.log1p(train_data[target_column])","metadata":{"execution":{"iopub.status.busy":"2024-12-31T04:51:10.285387Z","iopub.execute_input":"2024-12-31T04:51:10.285698Z","iopub.status.idle":"2024-12-31T04:51:10.422919Z","shell.execute_reply.started":"2024-12-31T04:51:10.285669Z","shell.execute_reply":"2024-12-31T04:51:10.422229Z"},"papermill":{"duration":0.224547,"end_time":"2024-12-30T16:13:29.467148","exception":false,"start_time":"2024-12-30T16:13:29.242601","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"#### **Explanation and Insights:**\n\n#### Skewness Analysis:\n1. **Training Dataset Skewness:**\n   - `Annual Income`: A skewness value of **1.522952** indicates moderate positive skewness, meaning that the distribution is tailing to the right, with a significant number of smaller income values and fewer larger ones.\n   - `Premium Amount`: A skewness value of **1.240914** also shows positive skewness, indicating that most premium values are clustered at the lower range with fewer higher premiums.\n\n2. **Test Dataset Skewness:**\n   - `Annual Income`: The skewness value of **1.516915** mirrors the training data, confirming that both datasets share similar distributions for this feature.\n\n#### Log Transformation:\n- Applying a log transformation to `Annual Income` reduces its skewness and compresses the range of larger values. This step makes the distribution closer to normal, benefiting algorithms sensitive to skewed data.\n  \n#### Adding Log of Target (`Log_Transformed_Premium`):\n- The addition of a `Log_Transformed_Premium` column in the training dataset creates a normalized target variable that minimizes the impact of extreme outliers, leading to better model performance.","metadata":{"papermill":{"duration":0.07636,"end_time":"2024-12-30T16:13:29.619999","exception":false,"start_time":"2024-12-30T16:13:29.543639","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"## Correlation Analysis of Preprocessed Data","metadata":{"papermill":{"duration":0.074315,"end_time":"2024-12-30T16:13:29.768509","exception":false,"start_time":"2024-12-30T16:13:29.694194","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Compute the correlation matrix\ntrain_corr_matrix = numeric_data.corr()\n\n# Create the heatmap\nplt.figure(figsize=(12, 9))\nsns.heatmap(\n    train_corr_matrix, \n    annot=True, \n    cmap='viridis',   \n    vmax=1, \n    vmin=-1,\n    annot_kws={\"size\": 8}, \n    fmt=\".3f\",\n    linewidths=0.5  \n)\n\n# Customize x and y tick labels\nplt.xticks(rotation=80, fontsize=9)\nplt.yticks(fontsize=9)\n\n# Add title and layout adjustments\nplt.title(\"Correlation Heatmap of the Train Data\", fontsize=11)\nplt.tight_layout()\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-12-31T04:51:10.423709Z","iopub.execute_input":"2024-12-31T04:51:10.424016Z","iopub.status.idle":"2024-12-31T04:51:11.941534Z","shell.execute_reply.started":"2024-12-31T04:51:10.423984Z","shell.execute_reply":"2024-12-31T04:51:11.940696Z"},"papermill":{"duration":1.818423,"end_time":"2024-12-30T16:13:31.661608","exception":false,"start_time":"2024-12-30T16:13:29.843185","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"#### Difference Between `Premium Amount` and `Log_Transformed_Premium` with Other Features:\n\n1. **Correlation Strength:**\n   - The correlation between `Log_Transformed_Premium` and other features tends to be slightly weaker compared to the correlation of the original `Premium Amount` with the same features. This is because the log transformation compresses the range of the `Premium Amount`, reducing the influence of extreme values (outliers).\n\n2. **Normalization of Relationships:**\n   - The log-transformed `Log_Transformed_Premium` shows more stable and normalized relationships with features like `Health Score` and `Customer Feedback`, which might indicate that these relationships are nonlinear in the raw `Premium Amount`.\n\n3. **Reduction in Skewness Impact:**\n   - In features like `Annual Income`, where the original correlation with `Premium Amount` is mildly negative, the correlation with `Log_Transformed_Premium` is further reduced. This suggests that the transformation reduces the skewed influence of high `Premium Amount` values.\n\n4. **Stronger Predictive Relationships:**\n   - The correlation between `Log_Transformed_Premium` and the original target (`Premium Amount`) is very high, indicating that the transformation retains most of the predictive information. However, the log-transformed target smooths out extreme values and better aligns with features like `Health Score` and `Previous Claims`.\n\n5. **Impact of High-Variance Features:**\n   - Features with high variability, such as `Credit Score` or `Annual Income`, exhibit more stable correlations with `Log_Transformed_Premium`, making them potentially more predictable in models using the transformed target.\n\n6. **Feature Contribution Differences:**\n   - The relationship between categorical features (e.g., `Location`, `Policy Type`) and the target remains nearly the same for `Premium Amount` and `Log_Transformed_Premium`, indicating that these features are less influenced by the target transformation.\n\n7. **Higher Correlation with `Log_Transformed_Premium` for Log-Normal Relationships:**\n   - Features like `Health Score` and `Policy Start Date` show slightly stronger positive correlations with `Log_Transformed_Premium` compared to `Premium Amount`, suggesting that these features have log-normal relationships with the target.\n\n#### Key Observations:\n- The log-transformation of the target variable (`Log_Transformed_Premium`) smooths the relationships between the target and input features, particularly for features influenced by outliers (e.g., `Annual Income` and `Previous Claims`).\n- It reduces the magnitude of extreme correlations, creating a more stable predictive target while retaining the most critical relationships for predictive modeling.\n- Features showing slight differences in correlation strength post-transformation indicate the utility of `Log_Transformed_Premium` for handling nonlinear relationships effectively.","metadata":{"papermill":{"duration":0.076949,"end_time":"2024-12-30T16:13:31.818913","exception":false,"start_time":"2024-12-30T16:13:31.741964","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"## Separating Features and Target","metadata":{"papermill":{"duration":0.078778,"end_time":"2024-12-30T16:13:31.976861","exception":false,"start_time":"2024-12-30T16:13:31.898083","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Separate Features and Target\nX_train = train_data.drop([target_column], axis=1)  # Features\ny_train = train_data[target_column]                   # Target variable","metadata":{"execution":{"iopub.status.busy":"2024-12-31T04:51:11.942603Z","iopub.execute_input":"2024-12-31T04:51:11.942924Z","iopub.status.idle":"2024-12-31T04:51:12.006185Z","shell.execute_reply.started":"2024-12-31T04:51:11.942892Z","shell.execute_reply":"2024-12-31T04:51:12.005526Z"},"papermill":{"duration":0.160009,"end_time":"2024-12-30T16:13:32.21523","exception":false,"start_time":"2024-12-30T16:13:32.055221","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Applying Log Transformation to the target variable","metadata":{"papermill":{"duration":0.084722,"end_time":"2024-12-30T16:13:32.389158","exception":false,"start_time":"2024-12-30T16:13:32.304436","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Apply log transformation to the target variable\ny_train_log = np.log1p(y_train)  # log1p is used for log(1 + x)","metadata":{"execution":{"iopub.status.busy":"2024-12-31T04:51:12.006873Z","iopub.execute_input":"2024-12-31T04:51:12.007099Z","iopub.status.idle":"2024-12-31T04:51:12.015916Z","shell.execute_reply.started":"2024-12-31T04:51:12.007078Z","shell.execute_reply":"2024-12-31T04:51:12.01512Z"},"papermill":{"duration":0.090854,"end_time":"2024-12-30T16:13:32.558349","exception":false,"start_time":"2024-12-30T16:13:32.467495","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Custom colormap using viridis\nviridis_cmap = cm.get_cmap(\"viridis\")\n\n# Select two colors from the colormap\ncolor1 = viridis_cmap(0.5)  \ncolor2 = viridis_cmap(0.3)  \n\n# Plot with the custom colors\nplt.figure(figsize=(9, 4))\n\n# Plot original target distribution\nplt.subplot(1, 2, 1)\nsns.histplot(y_train, kde=True, bins=30, color=color1)\nplt.title(f'Histogram of Target: {target_column} (y)', fontsize=11)\nplt.xlabel(f'{target_column} (y)', fontsize=10)\nplt.ylabel('Frequency', fontsize=10)\nplt.tick_params(axis='both', which='major', labelsize=7)\nplt.grid(True, linestyle='--', alpha=0.6)\n\n# Log-transformed target distribution\nplt.subplot(1, 2, 2)\nsns.histplot(y_train_log, kde=True, bins=30, color=color2)\nplt.title('Histogram of log(y + 1)', fontsize=11)\nplt.xlabel('log(y + 1)', fontsize=10)\nplt.ylabel('Frequency', fontsize=10)\nplt.tick_params(axis='both', which='major', labelsize=7)\nplt.grid(True, linestyle='--', alpha=0.6)\n\nplt.tight_layout()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-12-31T04:51:12.016791Z","iopub.execute_input":"2024-12-31T04:51:12.017047Z","iopub.status.idle":"2024-12-31T04:51:21.52423Z","shell.execute_reply.started":"2024-12-31T04:51:12.017027Z","shell.execute_reply":"2024-12-31T04:51:21.523399Z"},"papermill":{"duration":9.708654,"end_time":"2024-12-30T16:13:42.347992","exception":false,"start_time":"2024-12-30T16:13:32.639338","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"#### Histogram of `log(y + 1)`\n1. **Transformed Distribution:**\n   - After applying the logarithmic transformation (`log(y + 1)`), the distribution becomes **approximately normal**, reducing the skewness observed in the original data.\n   - This normalization ensures that most values fall within a manageable range for machine learning models, improving their performance.\n\n2. **Symmetry and Centralization:**\n   - The log transformation centers the distribution, with a peak around the range of `log(y + 1)` values between 6 and 7.\n\n\nBy compressing the scale of high premium values, the transformation mitigates the influence of outliers, reducing their weight in model training. A normal-like distribution aligns better with assumptions made by many regression models, potentially leading to more accurate predictions.","metadata":{"papermill":{"duration":0.079796,"end_time":"2024-12-30T16:13:42.508398","exception":false,"start_time":"2024-12-30T16:13:42.428602","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"### Removing Whitespaces in Feature Names","metadata":{"papermill":{"duration":0.08096,"end_time":"2024-12-30T16:13:42.67031","exception":false,"start_time":"2024-12-30T16:13:42.58935","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Remove whitespaces in feature names for X_train and test_data\nX_train.columns = X_train.columns.str.replace(' ', '_', regex=True)\ntest_data.columns = test_data.columns.str.replace(' ', '_', regex=True)\n\n# Verify the updated column names\nprint(\"Updated Feature Names in X_train:\")\nprint(X_train.columns)\n\nprint(\"\\nUpdated Feature Names in test_data:\")\nprint(test_data.columns)\n","metadata":{"execution":{"iopub.status.busy":"2024-12-31T04:51:21.52511Z","iopub.execute_input":"2024-12-31T04:51:21.525403Z","iopub.status.idle":"2024-12-31T04:51:21.532185Z","shell.execute_reply.started":"2024-12-31T04:51:21.525367Z","shell.execute_reply":"2024-12-31T04:51:21.531338Z"},"papermill":{"duration":0.088067,"end_time":"2024-12-30T16:13:42.838664","exception":false,"start_time":"2024-12-30T16:13:42.750597","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Validating Target and Transformed Target Data Types","metadata":{"papermill":{"duration":0.078302,"end_time":"2024-12-30T16:13:42.996533","exception":false,"start_time":"2024-12-30T16:13:42.918231","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Display data types\nprint(\"Feature Data Types (X_train):\")\nprint(X_train.dtypes)\n\nprint(\"\\nTarget Data Type (y_train):\")\nprint(y_train.dtypes)\n\nprint(\"\\nLog-Transformed Target Data Type (y_train_log):\")\nprint(y_train_log.dtypes)\n","metadata":{"execution":{"iopub.status.busy":"2024-12-31T04:51:21.533052Z","iopub.execute_input":"2024-12-31T04:51:21.533294Z","iopub.status.idle":"2024-12-31T04:51:21.551781Z","shell.execute_reply.started":"2024-12-31T04:51:21.533274Z","shell.execute_reply":"2024-12-31T04:51:21.551095Z"},"papermill":{"duration":0.088908,"end_time":"2024-12-30T16:13:43.164737","exception":false,"start_time":"2024-12-30T16:13:43.075829","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <span style=\"color:transparent;\">Model Training</span>\n\n<div style=\"border-radius: 15px; border: 2px solid #6A1B9A; padding: 10px; background: linear-gradient(135deg, #9C27B0, #4CAF50); text-align: center; box-shadow: 0px 4px 8px rgba(0, 0, 0, 0.5);\">\n    <h1 style=\"color: #ffffff; text-shadow: 2px 2px 4px rgba(0, 0, 0, 0.7); font-weight: bold; margin-bottom: 5px; font-size: 28px; font-family: 'Roboto', sans-serif;\">\n        Model Training\n    </h1>\n</div>\n\n<!-- Include Google Fonts for a modern font -->\n<link href=\"https://fonts.googleapis.com/css2?family=Roboto:wght@700&display=swap\" rel=\"stylesheet\">\n","metadata":{"papermill":{"duration":0.078961,"end_time":"2024-12-30T16:13:43.323728","exception":false,"start_time":"2024-12-30T16:13:43.244767","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"## Model Initialization","metadata":{}},{"cell_type":"code","source":"# Model Initialization\nlgb_gbdt = lgb.LGBMRegressor(\n    boosting_type='gbdt',\n    random_state=42,\n    learning_rate=0.030362233382902903,\n    n_estimators=998,\n    max_depth=9,\n    num_leaves=208,\n    min_child_samples=11,\n    subsample=0.8184667361186249,\n    colsample_bytree=0.8616459477375787,\n    reg_alpha=0.33080029457188864,\n    reg_lambda=0.20736962602335904,\n    objective='regression',\n    metric='rmse',\n    device='gpu',\n    verbose=-1\n)\n\nlgb_goss = lgb.LGBMRegressor(\n    boosting_type='goss',\n    random_state=42,\n    learning_rate=0.030362233382902903,\n    n_estimators=998,\n    max_depth=9,\n    num_leaves=208,\n    min_child_samples=11,\n    subsample=0.8184667361186249,\n    colsample_bytree=0.8616459477375787,\n    reg_alpha=0.33080029457188864,\n    reg_lambda=0.20736962602335904,\n    objective='regression',\n    metric='rmse',\n    device='gpu',\n    verbose=-1\n)\n","metadata":{"execution":{"iopub.status.busy":"2024-12-31T04:51:21.552567Z","iopub.execute_input":"2024-12-31T04:51:21.552795Z","iopub.status.idle":"2024-12-31T04:51:21.565012Z","shell.execute_reply.started":"2024-12-31T04:51:21.552769Z","shell.execute_reply":"2024-12-31T04:51:21.56431Z"},"papermill":{"duration":0.088665,"end_time":"2024-12-30T16:13:43.49265","exception":false,"start_time":"2024-12-30T16:13:43.403985","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"- Both **lgb_gbdt** and **lgb_goss** models are correctly initialized with appropriate hyperparameters, including GPU usage via device='gpu'.","metadata":{}},{"cell_type":"markdown","source":"## OOF Predictions","metadata":{}},{"cell_type":"code","source":"# Define the number of folds for OOF\nn_splits = 5\nkf = KFold(n_splits=n_splits, shuffle=True, random_state=42)\n\n# Initialize arrays to store OOF predictions and test predictions\noof_predictions_gbdt = np.zeros(len(X_train))\noof_predictions_goss = np.zeros(len(X_train))\n\ntest_predictions_gbdt = np.zeros((len(test_data), n_splits))\ntest_predictions_goss = np.zeros((len(test_data), n_splits))\n\n# Store RMSLE for each fold\nfold_rmsle_gbdt = []\nfold_rmsle_goss = []\n\n# OOF Training for LightGBM (GBDT) and LightGBM (GOSS)\nfor fold, (train_idx, val_idx) in enumerate(kf.split(X_train)):\n    print(f\"Training Fold {fold + 1}/{n_splits}...\")\n    \n    # Split the data into training and validation sets\n    X_train_fold, X_val_fold = X_train.iloc[train_idx], X_train.iloc[val_idx]\n    y_train_fold, y_val_fold = y_train_log.iloc[train_idx], y_train_log.iloc[val_idx]\n    \n    # LightGBM (GBDT)\n    lgb_gbdt.fit(\n        X_train_fold, y_train_fold,\n        eval_set=[(X_val_fold, y_val_fold)],\n        callbacks=[\n            lgb.early_stopping(stopping_rounds=50),  # Increased to 50\n            lgb.log_evaluation(period=10)\n        ]\n    )\n    oof_predictions_gbdt[val_idx] = lgb_gbdt.predict(X_val_fold)\n    test_predictions_gbdt[:, fold] = lgb_gbdt.predict(test_data)\n    fold_rmsle_gbdt.append(mean_squared_log_error(y_val_fold, oof_predictions_gbdt[val_idx]) ** 0.5)\n    \n    # LightGBM (GOSS)\n    lgb_goss.fit(\n        X_train_fold, y_train_fold,\n        eval_set=[(X_val_fold, y_val_fold)],\n        callbacks=[\n            lgb.early_stopping(stopping_rounds=50),  # Increased to 50\n            lgb.log_evaluation(period=10)\n        ]\n    )\n    oof_predictions_goss[val_idx] = lgb_goss.predict(X_val_fold)\n    test_predictions_goss[:, fold] = lgb_goss.predict(test_data)\n    fold_rmsle_goss.append(mean_squared_log_error(y_val_fold, oof_predictions_goss[val_idx]) ** 0.5)\n\n# Compute average RMSLE for each model\navg_rmsle_gbdt = np.mean(fold_rmsle_gbdt)\navg_rmsle_goss = np.mean(fold_rmsle_goss)\n\nprint(\"Average RMSLE (GBDT):\", avg_rmsle_gbdt)\nprint(\"Average RMSLE (GOSS):\", avg_rmsle_goss)","metadata":{"execution":{"iopub.status.busy":"2024-12-31T04:51:21.565751Z","iopub.execute_input":"2024-12-31T04:51:21.566Z","iopub.status.idle":"2024-12-31T04:56:14.673896Z","shell.execute_reply.started":"2024-12-31T04:51:21.565968Z","shell.execute_reply":"2024-12-31T04:56:14.67306Z"},"papermill":{"duration":317.389874,"end_time":"2024-12-30T16:19:00.962564","exception":false,"start_time":"2024-12-30T16:13:43.57269","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Compute weights based on RMSLE\n# Lower RMSLE => Higher Weight\ntotal_weight = 1 / avg_rmsle_gbdt + 1 / avg_rmsle_goss\nweight_gbdt = (1 / avg_rmsle_gbdt) / total_weight\nweight_goss = (1 / avg_rmsle_goss) / total_weight\n\nprint(\"Weight for GBDT:\", weight_gbdt)\nprint(\"Weight for GOSS:\", weight_goss)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-31T04:56:14.67472Z","iopub.execute_input":"2024-12-31T04:56:14.674956Z","iopub.status.idle":"2024-12-31T04:56:14.67974Z","shell.execute_reply.started":"2024-12-31T04:56:14.674935Z","shell.execute_reply":"2024-12-31T04:56:14.67888Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Compute weighted average of predictions\nfinal_test_predictions = (\n    weight_gbdt * test_predictions_gbdt.mean(axis=1) +\n    weight_goss * test_predictions_goss.mean(axis=1)\n)\n\n# Exponentiate the final predictions (log scale to original scale)\nfinal_test_predictions = np.expm1(final_test_predictions)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-31T04:56:14.680544Z","iopub.execute_input":"2024-12-31T04:56:14.680777Z","iopub.status.idle":"2024-12-31T04:56:14.731242Z","shell.execute_reply.started":"2024-12-31T04:56:14.680757Z","shell.execute_reply":"2024-12-31T04:56:14.730328Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Feature Importance\n\n- Feature importance is calculated for both GBDT and GOSS models.\n- Combined importance is averaged and plotted for the top 10 features, providing insights into the key drivers for predictions.","metadata":{}},{"cell_type":"code","source":"# Combine feature importance for both GBDT and GOSS models\nfeature_importance_gbdt = pd.DataFrame({\n    'Feature': X_train.columns,\n    'GBDT Importance': lgb_gbdt.feature_importances_,\n})\n\nfeature_importance_goss = pd.DataFrame({\n    'Feature': X_train.columns,\n    'GOSS Importance': lgb_goss.feature_importances_,\n})\n\n# Merge the feature importance for comparison\ncombined_feature_importance = pd.merge(\n    feature_importance_gbdt,\n    feature_importance_goss,\n    on='Feature',\n    how='inner'\n)\n\n# Sort features by average importance\ncombined_feature_importance['Avg Importance'] = (\n    combined_feature_importance['GBDT Importance'] +\n    combined_feature_importance['GOSS Importance']\n) / 2\n\ncombined_feature_importance = combined_feature_importance.sort_values(\n    by='Avg Importance', ascending=False\n)\n\n# Plot the feature importance\nplt.figure(figsize=(12, 8))\nsns.barplot(\n    x='Avg Importance', \n    y='Feature', \n    data=combined_feature_importance.head(10),  # Top 10 features\n    palette='viridis'\n)\nplt.title('Top 10 Features - Average Importance (GBDT & GOSS)', fontsize=16)\nplt.xlabel('Average Importance Score', fontsize=12)\nplt.ylabel('Features', fontsize=12)\nplt.grid(axis='x', linestyle='--', alpha=0.6)\nplt.tight_layout()\nplt.show()\n\n# Display top 10 features\nprint(\"Top 10 Features by Average Importance:\\n\")\ndisplay(combined_feature_importance.head(10))\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-31T04:56:14.732071Z","iopub.execute_input":"2024-12-31T04:56:14.732291Z","iopub.status.idle":"2024-12-31T04:56:14.995306Z","shell.execute_reply.started":"2024-12-31T04:56:14.732273Z","shell.execute_reply":"2024-12-31T04:56:14.994544Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Prediction Distribution\n\n- Histograms compare the distributions of the true values (y_train_log) and OOF predictions (oof_predictions_gbdt, oof_predictions_goss).\n- This visualizes how well the models capture the target variable's distribution.","metadata":{}},{"cell_type":"code","source":"# Calculate average predictions for test data\navg_test_predictions = (\n    weight_gbdt * test_predictions_gbdt.mean(axis=1) +\n    weight_goss * test_predictions_goss.mean(axis=1)\n)\n\n# Exponentiate predictions (log scale to original scale)\navg_test_predictions_exp = np.expm1(avg_test_predictions)\n\n# Visualization of Prediction Distributions\nviridis_cmap = cm.get_cmap(\"viridis\", 3)\n\nplt.figure(figsize=(12, 6))\n\n# Plot true values (Train Data)\nplt.hist(\n    y_train_log, bins=30, color=viridis_cmap(0), alpha=0.6, edgecolor=\"black\", label=\"True Values (Train)\"\n)\n\n# Plot predicted values (OOF Predictions)\nplt.hist(\n    oof_predictions_gbdt, bins=30, color=viridis_cmap(0.5), alpha=0.6, edgecolor=\"black\", label=\"Predicted Values (GBDT)\"\n)\n\nplt.hist(\n    oof_predictions_goss, bins=30, color=viridis_cmap(0.7), alpha=0.6, edgecolor=\"black\", label=\"Predicted Values (GOSS)\"\n)\n\n# Add titles and labels\nplt.title(\"Prediction Distributions - Train and OOF Predictions\", fontsize=16)\nplt.xlabel(\"Log-transformed Premium Amount (log(y + 1))\", fontsize=12)\nplt.ylabel(\"Frequency\", fontsize=12)\nplt.legend(fontsize=10)\nplt.grid(True, linestyle=\"--\", alpha=0.6)\n\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-31T04:56:14.996183Z","iopub.execute_input":"2024-12-31T04:56:14.996434Z","iopub.status.idle":"2024-12-31T04:56:15.568321Z","shell.execute_reply.started":"2024-12-31T04:56:14.996412Z","shell.execute_reply":"2024-12-31T04:56:15.56714Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <span style=\"color:transparent;\">Creating the Submission File</span>\n\n<div style=\"border-radius: 15px; border: 2px solid #6A1B9A; padding: 10px; background: linear-gradient(135deg, #9C27B0, #4CAF50); text-align: center; box-shadow: 0px 4px 8px rgba(0, 0, 0, 0.5);\">\n    <h1 style=\"color: #ffffff; text-shadow: 2px 2px 4px rgba(0, 0, 0, 0.7); font-weight: bold; margin-bottom: 5px; font-size: 28px; font-family: 'Roboto', sans-serif;\">\n        Creating the Submission File\n    </h1>\n</div>\n\n<!-- Include Google Fonts for a modern font -->\n<link href=\"https://fonts.googleapis.com/css2?family=Roboto:wght@700&display=swap\" rel=\"stylesheet\">\n","metadata":{"papermill":{"duration":0.088134,"end_time":"2024-12-30T16:19:01.137725","exception":false,"start_time":"2024-12-30T16:19:01.049591","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Prepare the submission file\nsubmission = pd.DataFrame({\"id\": test_data.index, target_column: final_test_predictions})\nsubmission.to_csv(\"ensemble_oof_submission.csv\", index=False)\nprint(\"Submission file created successfully!\")","metadata":{"execution":{"iopub.status.busy":"2024-12-31T04:56:15.569133Z","iopub.execute_input":"2024-12-31T04:56:15.569365Z","iopub.status.idle":"2024-12-31T04:56:16.899799Z","shell.execute_reply.started":"2024-12-31T04:56:15.569345Z","shell.execute_reply":"2024-12-31T04:56:16.899059Z"},"papermill":{"duration":1.479894,"end_time":"2024-12-30T16:19:02.706938","exception":false,"start_time":"2024-12-30T16:19:01.227044","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(submission.head(10))","metadata":{"execution":{"iopub.status.busy":"2024-12-31T04:56:16.900525Z","iopub.execute_input":"2024-12-31T04:56:16.900727Z","iopub.status.idle":"2024-12-31T04:56:16.906469Z","shell.execute_reply.started":"2024-12-31T04:56:16.900708Z","shell.execute_reply":"2024-12-31T04:56:16.905683Z"},"papermill":{"duration":0.100477,"end_time":"2024-12-30T16:19:02.903099","exception":false,"start_time":"2024-12-30T16:19:02.802622","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<div style=\"border-radius: 15px; border: 2px solid #6A1B9A; padding: 20px; background: linear-gradient(135deg, #9C27B0, #4CAF50); text-align: center; box-shadow: 0px 4px 8px rgba(0, 0, 0, 0.5);\">\n    <h1 style=\"color: #ffffff; text-shadow: 2px 2px 4px rgba(0, 0, 0, 0.7); font-weight: bold; margin-bottom: 10px; font-size: 28px; font-family: 'Roboto', sans-serif;\">\n        🙏 Thanks for Reading! 🚀\n    </h1>\n    <p style=\"color: #ffffff; font-size: 18px; text-align: center;\">\n        If you found this helpful, please upvote and share your thoughts!\n    </p>\n    <p style=\"color: #ffffff; font-size: 18px; text-align: center;\">\n        Happy Coding! 🙌😊\n    </p>\n</div>\n\n<!-- Include Google Fonts for a modern font -->\n<link href=\"https://fonts.googleapis.com/css2?family=Roboto:wght@700&display=swap\" rel=\"stylesheet\">\n","metadata":{"papermill":{"duration":0.088854,"end_time":"2024-12-30T16:19:03.085433","exception":false,"start_time":"2024-12-30T16:19:02.996579","status":"completed"},"tags":[]}}]}