{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":84896,"databundleVersionId":10305135,"sourceType":"competition"}],"dockerImageVersionId":30786,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<div style=\"font-family: 'Poppins'; font-weight: bold; letter-spacing: 0px; color: #FFFFFF; font-size: 300%; text-align: left; padding: 15px; background: #0A0F29; border: 8px solid #00FFFF; border-radius: 15px; box-shadow: 5px 5px 20px rgba(0, 0, 0, 0.5);\">\n    Predicting Insurance Premiums<br>\n</div>","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"background-color:#0A0F29; font-family:'Poppins', bold; color:#E0F7FA; font-size:140%; text-align:center; border: 2px solid #00FFFF; border-radius:15px; padding: 15px; box-shadow: 5px 5px 20px rgba(0, 0, 0, 0.5); font-weight: bold; letter-spacing: 1px; text-transform: uppercase;\">Introduction</div>","metadata":{}},{"cell_type":"markdown","source":"The 2024 Kaggle Playground Series - Season 4, Episode 12 presents a new challenge for data enthusiasts: predicting insurance premiums using a synthetic dataset. This competition offers participants the opportunity to hone their machine learning and feature engineering skills through a practical regression task.","metadata":{}},{"cell_type":"markdown","source":"## <div style=\"background-color:#0A0F29; font-family:'Poppins', bold; color:#E0F7FA; font-size:100%; text-align:center; border: 2px solid #0A0F29; border-radius:10px; padding: 10px; box-shadow: 5px 5px 20px rgba(0, 0, 0, 0.5); font-weight: bold; letter-spacing: 1px; text-transform: uppercase;\">Challenge Overview 📚</div>","metadata":{}},{"cell_type":"markdown","source":"- **Goal**: Predict insurance premiums (target variable: Premium Amount) using a set of predictor features.\n- **Competition Format**: Participants will work with a synthetically generated dataset modeled after real-world insurance data.\n- **Timeline**:\n    - Start Date: December 1, 2024\n    - Final Submission Deadline: December 31, 2024 (11:59 PM UTC)","metadata":{}},{"cell_type":"markdown","source":"## <div style=\"background-color:#0A0F29; font-family:'Poppins', bold; color:#E0F7FA; font-size:100%; text-align:center; border: 2px solid #0A0F29; border-radius:10px; padding: 10px; box-shadow: 5px 5px 20px rgba(0, 0, 0, 0.5); font-weight: bold; letter-spacing: 1px; text-transform: uppercase;\">Dataset Description 📊 </div>","metadata":{}},{"cell_type":"markdown","source":"- **Source**: The dataset is synthetic but based on a deep learning model trained on real-world insurance data.\n- **Files**:\n    - **train.csv**: Includes training data with features and the continuous target (Premium Amount).\n    - **test.csv**: Contains test data where predictions for Premium Amount are required.\n    - **sample_submission.csv**: A template submission file in the required format.\n\n- **Key Features**:\n    - Variables emulate real-world insurance data distributions but are modified to preserve confidentiality.\n    - Incorporating the original dataset may provide insights for better model performance.","metadata":{}},{"cell_type":"markdown","source":"## <div style=\"background-color:#0A0F29; font-family:'Poppins', bold; color:#E0F7FA; font-size:100%; text-align:center; border: 2px solid #0A0F29; border-radius:10px; padding: 10px; box-shadow: 5px 5px 20px rgba(0, 0, 0, 0.5); font-weight: bold; letter-spacing: 1px; text-transform: uppercase;\">Evaluation Metric</div>","metadata":{}},{"cell_type":"markdown","source":"Submissions are assessed using **Root Mean Squared Logarithmic Error (RMSLE)**:\n\n$$\n\\text{RMSLE} = \\sqrt{\\frac{1}{n} \\sum_{i=1}^n \\left(\\log(1 + \\text{pred}_i) - \\log(1 + \\text{actual}_i)\\right)^2}\n$$\n\nThe goal is to minimize the RMSLE value for predictions.","metadata":{}},{"cell_type":"markdown","source":"## <div style=\"background-color:#0A0F29; font-family:'Poppins', bold; color:#E0F7FA; font-size:100%; text-align:center; border: 2px solid #0A0F29; border-radius:10px; padding: 10px; box-shadow: 5px 5px 20px rgba(0, 0, 0, 0.5); font-weight: bold; letter-spacing: 1px; text-transform: uppercase;\">Submission Guidelines</div>","metadata":{}},{"cell_type":"markdown","source":"- Predict the `Premium Amount` for each row in the test set.\n- Ensure submissions include a header and follow this format:\n\n| id       | Premium Amount |\n|----------|----------------|\n| 1200000  | 1102.545       |\n| 1200001  | 1102.545       |\n| 1200002  | 1102.545       |","metadata":{"execution":{"iopub.status.busy":"2024-12-01T10:01:13.638274Z","iopub.execute_input":"2024-12-01T10:01:13.638684Z","iopub.status.idle":"2024-12-01T10:01:13.669005Z","shell.execute_reply.started":"2024-12-01T10:01:13.638632Z","shell.execute_reply":"2024-12-01T10:01:13.667388Z"}}},{"cell_type":"markdown","source":"# <div style=\"background-color:#0A0F29; font-family:'Poppins', bold; color:#E0F7FA; font-size:140%; text-align:center; border: 2px solid #00FFFF; border-radius:15px; padding: 15px; box-shadow: 5px 5px 20px rgba(0, 0, 0, 0.5); font-weight: bold; letter-spacing: 1px; text-transform: uppercase;\">Automated EDA</div>","metadata":{}},{"cell_type":"code","source":"# Install general-purpose data science and profiling libraries\n!pip install pandas==1.5.3 numba==0.58.1 visions==0.7.5 ydata-profiling==4.7.0 > /dev/null 2>&1\n\n# Install advanced machine learning frameworks\n!pip install catboost > /dev/null 2>&1\n!pip install ray==2.10.0 autogluon.tabular > /dev/null 2>&1\n\n# Install Optuna and its integrations\n!pip install optuna-integration[sklearn] > /dev/null 2>&1\n\n# Install LangChain and OpenAI-specific components\n!pip install langchain-core langchain-openai > /dev/null 2>&1\n\n# Install data visualization libraries\n!pip install sweetviz > /dev/null 2>&1","metadata":{"trusted":true,"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Standard Libraries\nimport json\nimport logging\nimport warnings\nfrom itertools import product\nfrom dataclasses import dataclass\nimport tempfile\n\n# Data Manipulation and Visualization\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport sweetviz as sv\nimport plotly.express as px\nimport plotly.graph_objects as go\nimport matplotlib.colors as mcolors\nfrom IPython.display import Markdown, display, IFrame\n\n# Data Profiling\nfrom ydata_profiling import ProfileReport\n\n# Feature Engineering\nfrom sklearn.preprocessing import StandardScaler, OneHotEncoder, KBinsDiscretizer\nfrom sklearn.compose import ColumnTransformer\nfrom sklearn.impute import SimpleImputer\nfrom sklearn.pipeline import Pipeline\nfrom featuretools import dfs, EntitySet\n\n# Model Selection and Validation\nfrom sklearn.model_selection import (\n    train_test_split,\n    StratifiedKFold,\n    cross_val_score,\n    RandomizedSearchCV\n)\n\n# Metrics and Scoring\nfrom sklearn.metrics import (\n    accuracy_score,\n    precision_score,\n    recall_score,\n    f1_score,\n    make_scorer\n)\n\n# Scikit-Learn Classifiers\nfrom sklearn.linear_model import LogisticRegression, SGDClassifier, RidgeClassifier\nfrom sklearn.svm import SVC\nfrom sklearn.tree import DecisionTreeClassifier\nfrom sklearn.naive_bayes import GaussianNB\nfrom sklearn.neighbors import KNeighborsClassifier\nfrom sklearn.ensemble import (\n    RandomForestClassifier,\n    GradientBoostingClassifier,\n    AdaBoostClassifier,\n    VotingClassifier\n)\n\n# External Libraries Classifiers\nfrom catboost import CatBoostClassifier\nfrom xgboost import XGBClassifier\nfrom lightgbm import LGBMClassifier, log_evaluation\n\n# Optuna for Hyperparameter Tuning\nimport optuna\nfrom optuna.integration import OptunaSearchCV\n\n# LangChain for LLM\nfrom langchain_core.prompts import ChatPromptTemplate\nfrom langchain_core.output_parsers import StrOutputParser\nfrom langchain_openai import ChatOpenAI\n\nfrom autogluon.tabular import TabularPredictor\nfrom autogluon.core.metrics import make_scorer\n\nimport pandas as pd\nimport numpy as np\nimport plotly.express as px\nimport plotly.graph_objects as go\nimport squarify\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nfrom dataclasses import dataclass\nfrom typing import List, Optional, Tuple, Dict\n\n# Suppress Warnings\nwarnings.filterwarnings(\"ignore\")\nwarnings.filterwarnings('ignore', category=RuntimeWarning, message=\"underflow encountered*\")","metadata":{"trusted":true,"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from kaggle_secrets import UserSecretsClient\nuser_secrets = UserSecretsClient()\nOPENAI_API_KEY = user_secrets.get_secret(\"openai_key\")","metadata":{"trusted":true,"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Define the LLM model using LangChain\nmodel = ChatOpenAI(\n    model='gpt-4o-2024-05-13',\n    temperature=0,\n    api_key=OPENAI_API_KEY\n)","metadata":{"trusted":true,"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"TIME_LIMIT=3600 * 2","metadata":{"trusted":true,"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Function to generate EDA summary\ndef eda_summary(df):\n    summary = {}\n    \n    # General Info\n    summary['general'] = {\n        'num_rows': df.shape[0],\n        'num_columns': df.shape[1],\n        'num_missing_values': df.isnull().sum().sum(),\n        'percent_missing_values': df.isnull().mean().mean() * 100\n    }\n    \n    # Column Data Types\n    summary['data_types'] = df.dtypes.to_dict()\n    \n    # Missing Value Summary (per column)\n    summary['missing_values'] = (\n        df.isnull()\n        .sum()\n        .to_frame(name='missing_count')\n        .assign(percent_missing=lambda x: (x['missing_count'] / df.shape[0]) * 100)\n        .to_dict(orient='index')\n    )\n    \n    # Numerical Summary (Mean, Median, Std, Min, Max)\n    describe_df = df.describe()\n    numerical_columns = ['mean', '50%', 'std', 'min', 'max']\n    available_columns = [col for col in numerical_columns if col in describe_df.columns]\n    summary['numerical_summary'] = (\n        describe_df[available_columns]\n        .rename(columns={'50%': 'median'})\n        .to_dict(orient='index')\n    )\n    \n    # Unique Counts for Categorical Columns\n    summary['categorical_summary'] = (\n        df.select_dtypes(include=['object', 'category'])\n        .nunique()\n        .to_frame(name='unique_counts')\n        .to_dict(orient='index')\n    )\n    \n    # Correlations\n    try:\n        summary['correlations'] = df.corr(numeric_only=True).to_dict()\n    except ValueError:\n        summary['correlations'] = \"Unable to calculate correlations due to data type issues.\"\n    \n    # Outlier Count based on IQR\n    outlier_summary = {}\n    for column in df.select_dtypes(include=[np.number]).columns:\n        Q1 = df[column].quantile(0.25)\n        Q3 = df[column].quantile(0.75)\n        IQR = Q3 - Q1\n        outliers = df[(df[column] < (Q1 - 1.5 * IQR)) | (df[column] > (Q3 + 1.5 * IQR))]\n        outlier_summary[column] = {\n            'outlier_count': outliers.shape[0],\n            'percent_outliers': (outliers.shape[0] / df.shape[0]) * 100\n        }\n    summary['outlier_summary'] = outlier_summary\n\n    return summary","metadata":{"trusted":true,"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Define DataConfig class\n@dataclass\nclass DataConfig:\n    target_column: str\n    id_column: str\n    train_path: str\n    test_path: str\n    random_seed: int = 42\n    colors: List[str] = None\n\n    def __post_init__(self):\n        if self.colors is None:\n            self.colors = px.colors.qualitative.Vivid  # Default to a vivid color scheme\n\n# Define EDA class\nclass EDA:\n    def __init__(self, config: DataConfig):\n        self.config = config\n        self.df_train = None\n        self.df_test = None\n        self.numerical_columns = []\n        self.categorical_columns = []\n        self.colors = config.colors\n\n    def load_data(self):\n        \"\"\"Load train and test data.\"\"\"\n        self.df_train = pd.read_csv(self.config.train_path)\n        self.df_test = pd.read_csv(self.config.test_path)\n        return self.df_train, self.df_test\n\n    def classify_columns(self):\n        \"\"\"Classify columns into numerical and categorical.\"\"\"\n        # Drop the ID and target columns for classification\n        df = self.df_train.drop([self.config.id_column, self.config.target_column], axis=1)\n        \n        # Classify columns and remove duplicates\n        self.numerical_columns = list(set(df.select_dtypes(include=['float64', 'int64']).columns.tolist()))\n        self.categorical_columns = list(set(df.select_dtypes(include=['object']).columns.tolist()))\n        \n        # Print the classification report\n        print(\"\\n\" + \"=\" * 50)\n        print(\"🔍 COLUMN CLASSIFICATION REPORT\")\n        print(\"=\" * 50)\n        \n        # Print numerical columns\n        print(\"\\n📊 Numerical Columns:\")\n        if self.numerical_columns:\n            for col in self.numerical_columns:\n                print(f\"  - {col}\")\n        else:\n            print(\"  None\")\n        \n        # Print categorical columns\n        print(\"\\n📋 Categorical Columns:\")\n        if self.categorical_columns:\n            for col in self.categorical_columns:\n                print(f\"  - {col}\")\n        else:\n            print(\"  None\")\n        \n        print(\"=\" * 50)\n\n    def basic_data_analysis(self):\n        \"\"\"Perform and print a basic data analysis report.\"\"\"\n        analysis = {\n            'shape': self.df_train.shape,\n            'duplicates': self.df_train.duplicated().sum(),\n            'missing_values': self.df_train.isnull().sum().to_dict(),\n            'dtypes': self.df_train.dtypes.to_dict()\n        }\n        \n        # Determine if the target is continuous or categorical\n        target_data = self.df_train[self.config.target_column]\n        if pd.api.types.is_numeric_dtype(target_data):\n            analysis['target_summary'] = {\n                'mean': target_data.mean(),\n                'median': target_data.median(),\n                'std': target_data.std(),\n                'min': target_data.min(),\n                'max': target_data.max()\n            }\n        else:\n            analysis['target_distribution'] = target_data.value_counts().to_dict()\n        \n        # Print the analysis report\n        print(\"\\n\" + \"=\" * 50)\n        print(\"📊 BASIC DATA ANALYSIS REPORT\")\n        print(\"=\" * 50)\n        print(f\"Rows: {analysis['shape'][0]:,}\")\n        print(f\"Columns: {analysis['shape'][1]}\")\n        print(f\"Duplicates: {analysis['duplicates']:,}\")\n        print(\"Missing Values:\")\n        for col, count in analysis['missing_values'].items():\n            if count > 0:\n                print(f\"  {col}: {count:,}\")\n        \n        # Print target summary\n        print(\"\\nTarget Summary:\")\n        if 'target_summary' in analysis:\n            print(\"  Continuous Target Variable\")\n            for stat, value in analysis['target_summary'].items():\n                print(f\"    {stat.capitalize()}: {value:.2f}\")\n        elif 'target_distribution' in analysis:\n            print(\"  Categorical Target Variable\")\n            total = sum(analysis['target_distribution'].values())\n            for target_class, count in analysis['target_distribution'].items():\n                percentage = (count / total) * 100\n                print(f\"    {target_class}: {count:,} ({percentage:.2f}%)\")\n        \n        print(\"=\" * 50)\n        return analysis\n\n    def visualize_missing_values(self):\n        \"\"\"Visualize missing values.\"\"\"\n        missing_values = self.df_train.isnull().mean() * 100\n        missing_values = missing_values[missing_values > 0].sort_values()\n        fig = go.Figure(go.Bar(\n            x=missing_values.values,\n            y=missing_values.index,\n            orientation='h',\n            marker_color=self.colors[0]\n        ))\n        fig.update_layout(title=\"Missing Value Percentage\", xaxis_title=\"Percentage\", yaxis_title=\"Feature\")\n        fig.show(renderer='iframe')\n\n    def plot_numerical_distributions(self, columns: List[str] = None):\n        \"\"\"Plot numerical column distributions with robust error handling.\"\"\"\n        # Use the provided column list or default to all numerical columns\n        columns_to_plot = columns if columns else list(set(self.numerical_columns))\n        \n        for col in columns_to_plot:\n            try:\n                assert col in self.df_train.columns, f\"Column '{col}' is not found in the dataset.\"\n                # Create a new figure for each column\n                fig = px.histogram(self.df_train, x=col, marginal=\"box\", color_discrete_sequence=[self.colors[1]])\n                fig.update_layout(title=f\"Distribution of {col}\", xaxis_title=col, yaxis_title=\"Count\")\n                fig.show(renderer='iframe')  # Explicit rendering of each figure\n            except AssertionError as e:\n                print(f\"AssertionError: {e}\")\n            except Exception as e:\n                print(f\"An error occurred while plotting '{col}': {e}\")\n\n    def plot_categorical_distributions(self, columns: List[str] = None, top_n: int = 20):\n        \"\"\"Plot categorical column distributions with robust error handling.\"\"\"\n        # Use the provided column list or default to all categorical columns\n        columns_to_plot = columns if columns else list(set(self.categorical_columns))\n        \n        for col in columns_to_plot:\n            try:\n                assert col in self.df_train.columns, f\"Column '{col}' is not found in the dataset.\"\n                # Create a new figure for each column\n                value_counts = self.df_train[col].value_counts().nlargest(top_n)\n                fig = px.bar(x=value_counts.index, y=value_counts.values, color_discrete_sequence=[self.colors[2]])\n                fig.update_layout(title=f\"Distribution of {col} (Top {top_n})\", xaxis_title=col, yaxis_title=\"Count\")\n                fig.show(renderer='iframe')  # Explicit rendering of each figure\n            except AssertionError as e:\n                print(f\"AssertionError: {e}\")\n            except Exception as e:\n                print(f\"An error occurred while plotting '{col}': {e}\")\n\n    def plot_treemap(self, categorical_column: str, top_n: int = 20):\n        \"\"\"Plot a treemap for a categorical column using Plotly.\"\"\"\n        \n        # Get the top categories\n        value_counts = self.df_train[categorical_column].value_counts().nlargest(top_n)\n    \n        # Create a DataFrame for the treemap\n        treemap_data = pd.DataFrame({\n            categorical_column: value_counts.index,\n            'Count': value_counts.values\n        })\n    \n        # Generate the treemap using Plotly\n        fig = px.treemap(\n            treemap_data,\n            path=[categorical_column],\n            values='Count',\n            color='Count',\n            color_continuous_scale='Viridis',\n            title=f\"Treemap of {categorical_column} (Top {top_n})\"\n        )\n    \n        # Display the treemap\n        fig.show(renderer='iframe')\n\n    def plot_sankey(self, categorical_column: str, top_n: int = 20):\n        \"\"\"Create a Sankey diagram for a categorical column and the target variable.\"\"\"\n        top_categories = self.df_train[categorical_column].value_counts().nlargest(top_n).index\n        filtered_data = self.df_train[self.df_train[categorical_column].isin(top_categories)]\n        sankey_data = filtered_data.groupby([categorical_column, self.config.target_column]).size().reset_index(name='Count')\n\n        all_labels = top_categories.tolist() + [f\"{self.config.target_column}: {v}\" for v in self.df_train[self.config.target_column].unique()]\n        source_indices = []\n        target_indices = []\n        values = []\n\n        for i, cat in enumerate(top_categories):\n            for j, target in enumerate(self.df_train[self.config.target_column].unique()):\n                count = sankey_data[(sankey_data[categorical_column] == cat) & (sankey_data[self.config.target_column] == target)]['Count']\n                if not count.empty:\n                    source_indices.append(i)\n                    target_indices.append(len(top_categories) + j)\n                    values.append(count.values[0])\n\n        node_colors = self.colors[:len(all_labels)]\n        fig = go.Figure(data=[go.Sankey(\n            node=dict(pad=10, thickness=20, label=all_labels, color=node_colors),\n            link=dict(source=source_indices, target=target_indices, value=values)\n        )])\n        fig.update_layout(title=f\"Sankey Diagram: {categorical_column} vs {self.config.target_column}\")\n        fig.show(renderer='iframe')\n\n    def correlation_matrix(self):\n        \"\"\"Plot a correlation matrix.\"\"\"\n        corr = self.df_train.corr(numeric_only=True)\n        fig = px.imshow(corr, text_auto=True, color_continuous_scale='RdBu')\n        fig.update_layout(title=\"Correlation Matrix\")\n        fig.show(renderer='iframe')","metadata":{"trusted":true,"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Define DataConfig class\n@dataclass\nclass DataConfigSeaborn:\n    target_column: str\n    id_column: str\n    train_path: str\n    test_path: str\n    random_seed: int = 42\n    colors: List[str] = None\n\n    def __post_init__(self):\n        if self.colors is None:\n            self.colors = sns.color_palette(\"deep\")\n\n# Define EDA class\nclass EDASeaborn:\n    def __init__(self, config: DataConfigSeaborn):\n        self.config = config\n        self.df_train = None\n        self.df_test = None\n        self.numerical_columns = []\n        self.categorical_columns = []\n        self.colors = config.colors\n\n    def load_data(self):\n        \"\"\"Load train and test data.\"\"\"\n        self.df_train = pd.read_csv(self.config.train_path)\n        self.df_test = pd.read_csv(self.config.test_path)\n        return self.df_train, self.df_test\n\n    def classify_columns(self):\n        \"\"\"Classify columns into numerical and categorical.\"\"\"\n        # Drop the ID and target columns for classification\n        df = self.df_train.drop([self.config.id_column, self.config.target_column], axis=1)\n        \n        # Classify columns and remove duplicates\n        self.numerical_columns = list(set(df.select_dtypes(include=['float64', 'int64']).columns.tolist()))\n        self.categorical_columns = list(set(df.select_dtypes(include=['object']).columns.tolist()))\n        \n        # Print the classification report\n        print(\"\\n\" + \"=\" * 50)\n        print(\"🔍 COLUMN CLASSIFICATION REPORT\")\n        print(\"=\" * 50)\n        \n        # Print numerical columns\n        print(\"\\n📊 Numerical Columns:\")\n        if self.numerical_columns:\n            for col in self.numerical_columns:\n                print(f\"  - {col}\")\n        else:\n            print(\"  None\")\n        \n        # Print categorical columns\n        print(\"\\n📋 Categorical Columns:\")\n        if self.categorical_columns:\n            for col in self.categorical_columns:\n                print(f\"  - {col}\")\n        else:\n            print(\"  None\")\n        \n        print(\"=\" * 50)\n\n    def basic_data_analysis(self):\n        \"\"\"Perform basic data analysis.\"\"\"\n        analysis = {\n            'shape': self.df_train.shape,\n            'duplicates': self.df_train.duplicated().sum(),\n            'missing_values': self.df_train.isnull().sum(),\n            'dtypes': self.df_train.dtypes\n        }\n        return analysis\n\n    def visualize_missing_values(self):\n        \"\"\"Visualize missing values.\"\"\"\n        missing_values = self.df_train.isnull().mean() * 100\n        missing_values = missing_values[missing_values > 0].sort_values()\n\n        plt.figure(figsize=(10, 6))\n        sns.barplot(x=missing_values.values, y=missing_values.index, palette=self.colors)\n        plt.title(\"Missing Value Percentage\")\n        plt.xlabel(\"Percentage\")\n        plt.ylabel(\"Feature\")\n        plt.show()\n\n    def plot_numerical_distributions(self, columns: List[str] = None):\n        \"\"\"Plot numerical column distributions.\"\"\"\n        columns_to_plot = columns if columns else self.numerical_columns\n\n        for col in columns_to_plot:\n            plt.figure(figsize=(10, 6))\n            sns.histplot(self.df_train[col], kde=True, color=self.colors[0])\n            plt.title(f\"Distribution of {col}\")\n            plt.xlabel(col)\n            plt.ylabel(\"Count\")\n            plt.show()\n\n    def plot_categorical_distributions(self, columns: List[str] = None, top_n: int = 20):\n        \"\"\"Plot categorical column distributions.\"\"\"\n        columns_to_plot = columns if columns else self.categorical_columns\n\n        for col in columns_to_plot:\n            value_counts = self.df_train[col].value_counts().nlargest(top_n)\n\n            plt.figure(figsize=(10, 6))\n            sns.barplot(x=value_counts.values, y=value_counts.index, palette=self.colors)\n            plt.title(f\"Distribution of {col} (Top {top_n})\")\n            plt.xlabel(\"Count\")\n            plt.ylabel(col)\n            plt.show()\n\n    def plot_treemap(self, categorical_column: str, top_n: int = 20):\n        \"\"\"Plot a treemap for a categorical column.\"\"\"\n        value_counts = self.df_train[categorical_column].value_counts().nlargest(top_n)\n\n        plt.figure(figsize=(10, 6))\n        sns.barplot(x=value_counts.values, y=value_counts.index, palette=self.colors)\n        plt.title(f\"Treemap of {categorical_column} (Top {top_n})\")\n        plt.xlabel(\"Count\")\n        plt.ylabel(categorical_column)\n        plt.show()\n\n    def correlation_matrix(self):\n        \"\"\"Plot a correlation matrix.\"\"\"\n        corr = self.df_train.corr(numeric_only=True)\n\n        plt.figure(figsize=(12, 8))\n        sns.heatmap(corr, annot=True, cmap=\"RdBu\", fmt=\".2f\")\n        plt.title(\"Correlation Matrix\")\n        plt.show()","metadata":{"trusted":true,"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Configuration\nconfig = DataConfig(\n    target_column=\"Premium Amount\",\n    id_column=\"id\",\n    train_path=\"/kaggle/input/playground-series-s4e12/train.csv\",\n    test_path=\"/kaggle/input/playground-series-s4e12/test.csv\"\n)\n\n# Initialize the EDA class\n#eda = EDASeaborn(config)\neda = EDA(config) # Plotly\n\n# Load data\ntrain_data, test_data = eda.load_data()\n\nprint('Train dataset: ')\ndisplay(train_data)\nprint('Test dataset: ')\ndisplay(test_data)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"eda.classify_columns()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-01T13:44:01.390605Z","iopub.execute_input":"2024-12-01T13:44:01.390958Z","iopub.status.idle":"2024-12-01T13:44:01.829499Z","shell.execute_reply.started":"2024-12-01T13:44:01.390924Z","shell.execute_reply":"2024-12-01T13:44:01.828266Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(\"Categorical Columns:\", eda.categorical_columns)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-01T13:44:01.831174Z","iopub.execute_input":"2024-12-01T13:44:01.83159Z","iopub.status.idle":"2024-12-01T13:44:01.837928Z","shell.execute_reply.started":"2024-12-01T13:44:01.831554Z","shell.execute_reply":"2024-12-01T13:44:01.836542Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(\"Numerical Columns:\", eda.numerical_columns)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-01T13:44:01.839526Z","iopub.execute_input":"2024-12-01T13:44:01.840446Z","iopub.status.idle":"2024-12-01T13:44:01.851011Z","shell.execute_reply.started":"2024-12-01T13:44:01.840409Z","shell.execute_reply":"2024-12-01T13:44:01.849582Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Perform EDA\neda.basic_data_analysis()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-01T13:44:01.852722Z","iopub.execute_input":"2024-12-01T13:44:01.853144Z","iopub.status.idle":"2024-12-01T13:44:04.50325Z","shell.execute_reply.started":"2024-12-01T13:44:01.853099Z","shell.execute_reply":"2024-12-01T13:44:04.502088Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# LLM automated EDA\n\n# Generate the summary\nsummary = eda_summary(train_data)\nsummary_json = json.dumps(summary, indent=4, default=str)\n\n# Define the prompt template for LangChain\ntemplate = \"\"\"Provide an analysis of the following EDA summary, The aim of this dataset and EDA is to understand how several variables influence depression.\nUltimately the aim is to build a classification model to predict depression:\n{context}\n\nKey insights and observations:\n\"\"\"\n\nprompt = ChatPromptTemplate.from_template(template)\n\n# Create a chain to pass the summary to the model\nchain = prompt | model | StrOutputParser()\n\n# Invoke the chain to analyze the EDA summary\nresult = chain.invoke(summary_json)\n\n# Print the result\ndisplay(Markdown(result))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-01T13:44:04.504928Z","iopub.execute_input":"2024-12-01T13:44:04.505423Z","iopub.status.idle":"2024-12-01T13:44:17.754216Z","shell.execute_reply.started":"2024-12-01T13:44:04.505369Z","shell.execute_reply":"2024-12-01T13:44:17.752921Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"eda.visualize_missing_values()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-01T13:44:17.758749Z","iopub.execute_input":"2024-12-01T13:44:17.75917Z","iopub.status.idle":"2024-12-01T13:44:18.547816Z","shell.execute_reply.started":"2024-12-01T13:44:17.759132Z","shell.execute_reply":"2024-12-01T13:44:18.546245Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"eda.plot_numerical_distributions(columns=[\"Vehicle Age\"])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-01T13:44:18.549704Z","iopub.execute_input":"2024-12-01T13:44:18.550225Z","iopub.status.idle":"2024-12-01T13:44:19.068211Z","shell.execute_reply.started":"2024-12-01T13:44:18.55016Z","shell.execute_reply":"2024-12-01T13:44:19.066651Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"eda.plot_numerical_distributions(columns=[\"Number of Dependents\"])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-01T13:44:19.070023Z","iopub.execute_input":"2024-12-01T13:44:19.070572Z","iopub.status.idle":"2024-12-01T13:44:19.528231Z","shell.execute_reply.started":"2024-12-01T13:44:19.07051Z","shell.execute_reply":"2024-12-01T13:44:19.526624Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"eda.plot_numerical_distributions(columns=[\"Previous Claims\"])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-01T13:44:19.530123Z","iopub.execute_input":"2024-12-01T13:44:19.530752Z","iopub.status.idle":"2024-12-01T13:44:19.959887Z","shell.execute_reply.started":"2024-12-01T13:44:19.530695Z","shell.execute_reply":"2024-12-01T13:44:19.958607Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"eda.plot_numerical_distributions(columns=[\"Insurance Duration\"])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-01T13:44:19.96142Z","iopub.execute_input":"2024-12-01T13:44:19.961822Z","iopub.status.idle":"2024-12-01T13:44:20.568066Z","shell.execute_reply.started":"2024-12-01T13:44:19.961787Z","shell.execute_reply":"2024-12-01T13:44:20.566733Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"eda.plot_numerical_distributions(columns=[\"Credit Score\"])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-01T13:44:20.569718Z","iopub.execute_input":"2024-12-01T13:44:20.570214Z","iopub.status.idle":"2024-12-01T13:44:21.077092Z","shell.execute_reply.started":"2024-12-01T13:44:20.570162Z","shell.execute_reply":"2024-12-01T13:44:21.074861Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"eda.plot_numerical_distributions(columns=[\"Annual Income\"])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-01T13:44:21.078717Z","iopub.execute_input":"2024-12-01T13:44:21.079196Z","iopub.status.idle":"2024-12-01T13:44:21.643759Z","shell.execute_reply.started":"2024-12-01T13:44:21.079144Z","shell.execute_reply":"2024-12-01T13:44:21.642256Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"eda.plot_numerical_distributions(columns=[\"Age\"])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-01T13:44:21.644856Z","iopub.execute_input":"2024-12-01T13:44:21.645174Z","iopub.status.idle":"2024-12-01T13:44:22.13725Z","shell.execute_reply.started":"2024-12-01T13:44:21.645144Z","shell.execute_reply":"2024-12-01T13:44:22.135828Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"eda.plot_numerical_distributions(columns=[\"Health Score\"])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-01T13:44:22.139037Z","iopub.execute_input":"2024-12-01T13:44:22.13958Z","iopub.status.idle":"2024-12-01T13:44:23.155689Z","shell.execute_reply.started":"2024-12-01T13:44:22.139526Z","shell.execute_reply":"2024-12-01T13:44:23.154384Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"eda.plot_categorical_distributions(columns=[\"Occupation\"])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-01T13:44:23.157174Z","iopub.execute_input":"2024-12-01T13:44:23.15763Z","iopub.status.idle":"2024-12-01T13:44:23.324318Z","shell.execute_reply.started":"2024-12-01T13:44:23.157596Z","shell.execute_reply":"2024-12-01T13:44:23.323194Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"eda.plot_categorical_distributions(columns=[\"Marital Status\"])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-01T13:44:23.325624Z","iopub.execute_input":"2024-12-01T13:44:23.325929Z","iopub.status.idle":"2024-12-01T13:44:23.506666Z","shell.execute_reply.started":"2024-12-01T13:44:23.3259Z","shell.execute_reply":"2024-12-01T13:44:23.505306Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"eda.plot_categorical_distributions(columns=[\"Policy Start Date\"])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-01T13:44:23.508053Z","iopub.execute_input":"2024-12-01T13:44:23.508486Z","iopub.status.idle":"2024-12-01T13:44:24.3516Z","shell.execute_reply.started":"2024-12-01T13:44:23.508445Z","shell.execute_reply":"2024-12-01T13:44:24.349911Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"eda.plot_categorical_distributions(columns=[\"Policy Type\"])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-01T13:44:24.352924Z","iopub.execute_input":"2024-12-01T13:44:24.353274Z","iopub.status.idle":"2024-12-01T13:44:24.528228Z","shell.execute_reply.started":"2024-12-01T13:44:24.353239Z","shell.execute_reply":"2024-12-01T13:44:24.526717Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"eda.plot_categorical_distributions(columns=[\"Exercise Frequency\"])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-01T13:44:24.530225Z","iopub.execute_input":"2024-12-01T13:44:24.530759Z","iopub.status.idle":"2024-12-01T13:44:24.704709Z","shell.execute_reply.started":"2024-12-01T13:44:24.53071Z","shell.execute_reply":"2024-12-01T13:44:24.703333Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"eda.plot_categorical_distributions(columns=[\"Property Type\"])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-01T13:44:24.706029Z","iopub.execute_input":"2024-12-01T13:44:24.706412Z","iopub.status.idle":"2024-12-01T13:44:24.886192Z","shell.execute_reply.started":"2024-12-01T13:44:24.706377Z","shell.execute_reply":"2024-12-01T13:44:24.884773Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"eda.plot_categorical_distributions(columns=[\"Location\"])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-01T13:44:24.887665Z","iopub.execute_input":"2024-12-01T13:44:24.888135Z","iopub.status.idle":"2024-12-01T13:44:25.065426Z","shell.execute_reply.started":"2024-12-01T13:44:24.888084Z","shell.execute_reply":"2024-12-01T13:44:25.064038Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"eda.plot_categorical_distributions(columns=[\"Gender\"])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-01T13:44:25.066729Z","iopub.execute_input":"2024-12-01T13:44:25.06705Z","iopub.status.idle":"2024-12-01T13:44:25.23953Z","shell.execute_reply.started":"2024-12-01T13:44:25.06702Z","shell.execute_reply":"2024-12-01T13:44:25.238212Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"eda.plot_categorical_distributions(columns=[\"Education Level\"])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-01T13:44:25.24118Z","iopub.execute_input":"2024-12-01T13:44:25.2417Z","iopub.status.idle":"2024-12-01T13:44:25.412448Z","shell.execute_reply.started":"2024-12-01T13:44:25.24165Z","shell.execute_reply":"2024-12-01T13:44:25.411007Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"eda.plot_categorical_distributions(columns=[\"Smoking Status\"])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-01T13:44:25.414098Z","iopub.execute_input":"2024-12-01T13:44:25.415429Z","iopub.status.idle":"2024-12-01T13:44:25.583496Z","shell.execute_reply.started":"2024-12-01T13:44:25.415381Z","shell.execute_reply":"2024-12-01T13:44:25.582137Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"eda.plot_categorical_distributions(columns=[\"Customer Feedback\"])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-01T13:44:25.584922Z","iopub.execute_input":"2024-12-01T13:44:25.585331Z","iopub.status.idle":"2024-12-01T13:44:25.752445Z","shell.execute_reply.started":"2024-12-01T13:44:25.585262Z","shell.execute_reply":"2024-12-01T13:44:25.751157Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"eda.plot_treemap('Occupation')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-01T13:44:25.754353Z","iopub.execute_input":"2024-12-01T13:44:25.754865Z","iopub.status.idle":"2024-12-01T13:44:25.92978Z","shell.execute_reply.started":"2024-12-01T13:44:25.754815Z","shell.execute_reply":"2024-12-01T13:44:25.928499Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"eda.correlation_matrix()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-01T13:44:25.93145Z","iopub.execute_input":"2024-12-01T13:44:25.931885Z","iopub.status.idle":"2024-12-01T13:44:26.429135Z","shell.execute_reply.started":"2024-12-01T13:44:25.931841Z","shell.execute_reply":"2024-12-01T13:44:26.427928Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Sweetviz Report Train vs Test\ntrain_data = eda.df_train\ntest_data = eda.df_test\ntarget_variable = \"Premium Amount\"\nsweetviz_report = sv.compare([train_data, \"Train\"], [test_data, \"Test\"], target_feat=target_variable)\n#sweetviz_report.show_html(filepath=\"Sweetviz_Report.html\", open_browser=False)\n#display(IFrame(src=\"Sweetviz_Report.html\", width=1000, height=600))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-01T13:44:26.431391Z","iopub.execute_input":"2024-12-01T13:44:26.431887Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# YData Profiling Train\ntrain_profile = ProfileReport(train_data, title=\"Train Data Profile Report\", explorative=True)\ntrain_profile_path = \"Train_Profile_Report.html\"\ntrain_profile.to_file(train_profile_path)\n#display(IFrame(src=train_profile_path, width=1000, height=600))","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# YData Profiling Test\ntest_profile = ProfileReport(test_data, title=\"Test Data Profile Report\", explorative=True)\ntest_profile_path = \"Test_Profile_Report.html\"\ntest_profile.to_file(test_profile_path)\n#display(IFrame(src=test_profile_path, width=1000, height=600))","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <div style=\"background-color:#0A0F29; font-family:'Poppins', bold; color:#E0F7FA; font-size:140%; text-align:center; border: 2px solid #00FFFF; border-radius:15px; padding: 15px; box-shadow: 5px 5px 20px rgba(0, 0, 0, 0.5); font-weight: bold; letter-spacing: 1px; text-transform: uppercase;\">AutoGluon Baseline Model</div>","metadata":{}},{"cell_type":"code","source":"# Paths\ntrain_path = config.train_path\ntest_path = config.test_path\nsubmission_template_path = \"/kaggle/input/playground-series-s4e12/sample_submission.csv\"\n\n# Load datasets\ntrain_data = pd.read_csv(train_path)\ntest_data = pd.read_csv(test_path)\n\n# RMSLE Metric Definition\ndef rmsle(y_true, y_pred):\n    y_true = np.maximum(y_true, 0)  # Ensure non-negative values\n    y_pred = np.maximum(y_pred, 0)\n    return np.sqrt(np.mean(np.square(np.log1p(y_true) - np.log1p(y_pred))))\n\nrmsle_metric = make_scorer(\n    name='RMSLE',\n    score_func=rmsle,\n    greater_is_better=False,\n    needs_proba=False\n)\n\n# Define target and features\nTARGET = \"Premium Amount\"\nID_COLUMN = \"id\"\n\n# Exclude 'id' and the target column from the features\nfeatures = [col for col in train_data.columns if col not in [TARGET, ID_COLUMN]]\n\n# Define a persistent directory for model files\nmodel_dir = \"model_output\"\n\n# Define and train the predictor\npredictor = TabularPredictor(\n    label=TARGET,  # Target column\n    eval_metric=rmsle_metric,  # Evaluation metric\n    problem_type=\"regression\",  # Problem type: regression\n    path=model_dir  # Persistent directory for model files\n).fit(\n    train_data[features + [TARGET]],  # Include only features and target during training\n    presets='best_quality',  # Preset for best quality models\n    time_limit=TIME_LIMIT,  # Training time limit\n    verbosity=0,  # Verbosity level for detailed logs\n    ag_args_fit={'num_gpus': 1},  # Use 1 GPU for all models\n    excluded_model_types=['FASTAI', 'KNN']  # Exclude FASTAI and KNN models\n)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Generate and display the fit summary\nresults = predictor.fit_summary()\n\n# Predict on the test data (use only the features, excluding 'id')\npredictions = predictor.predict(test_data[features])\n\n# Display the leaderboard\nleaderboard = predictor.leaderboard(extra_info=True)  # Shows performance of all trained models\nprint(\"Leaderboard:\")\ndisplay(leaderboard)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Display and optionally plot feature importance\nfeature_importances = predictor.feature_importance(train_data)\nprint(\"Feature Importances:\")\ndisplay(feature_importances)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Reset the index to make the features a column\nfeature_importances_reset = feature_importances.reset_index()\nfeature_importances_reset.columns = ['Feature', 'Importance', 'StdDev', 'P-Value', 'N', 'P99_High', 'P99_Low']\n\n# Plot feature importances\nfeature_importances_reset.set_index('Feature')['Importance'].plot(kind='barh', figsize=(10, 8))\nplt.title(\"Feature Importances\")\nplt.xlabel(\"Importance Score\")\nplt.ylabel(\"Features\")\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Prepare Submission\nsubmission = pd.read_csv(submission_template_path)\nsubmission[TARGET] = predictions\nsubmission.to_csv(\"submission.csv\", index=False)\n\nprint(\"Submission file created: 'submission.csv'\")","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <div style=\"background-color:#0A0F29; font-family:'Poppins', bold; color:#E0F7FA; font-size:140%; text-align:center; border: 2px solid #00FFFF; border-radius:15px; padding: 15px; box-shadow: 5px 5px 20px rgba(0, 0, 0, 0.5); font-weight: bold; letter-spacing: 1px; text-transform: uppercase;\">References</div>","metadata":{}},{"cell_type":"markdown","source":"- https://www.kaggle.com/code/stpeteishii/solution-to-a-plotly-graph-cannot-be-displayed","metadata":{}}]}