{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":84896,"databundleVersionId":10305135,"sourceType":"competition"}],"dockerImageVersionId":30804,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2024-12-11T19:01:40.359552Z","iopub.execute_input":"2024-12-11T19:01:40.360154Z","iopub.status.idle":"2024-12-11T19:01:41.499543Z","shell.execute_reply.started":"2024-12-11T19:01:40.3601Z","shell.execute_reply":"2024-12-11T19:01:41.49819Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nfrom scipy.stats import shapiro, kstest, kurtosis, skew\nfrom sklearn.preprocessing import MinMaxScaler, RobustScaler, StandardScaler, Normalizer\nimport seaborn as sns\nfrom sklearn.impute import KNNImputer","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-11T19:01:41.501476Z","iopub.execute_input":"2024-12-11T19:01:41.502174Z","iopub.status.idle":"2024-12-11T19:01:43.60747Z","shell.execute_reply.started":"2024-12-11T19:01:41.502119Z","shell.execute_reply":"2024-12-11T19:01:43.606369Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 1. EDA","metadata":{}},{"cell_type":"markdown","source":"## 1.1 Helper Functions","metadata":{}},{"cell_type":"markdown","source":"### 1.1.1. Analyze Feature Function","metadata":{}},{"cell_type":"markdown","source":"The `analyze_feature` function provides a detailed analysis of a specified feature (column) in a DataFrame. It returns key statistics about the feature, which can be helpful in understanding its data quality and characteristics.\n\n#### Key Metrics Analyzed:\n1. **Total Number of Values**: This metric shows the total number of entries (rows) in the feature, including null values.\n2. **Number of Null Values**: The function counts how many entries in the feature are null or missing.\n3. **Percentage of Null Values**: It calculates the proportion of null values relative to the total number of values in the feature, giving an indication of the completeness of the data.\n4. **Data Type**: This reveals the data type of the feature (e.g., integer, float, object), helping to understand the nature of the data.\n5. **Number of Unique Values**: The function calculates how many distinct values are present in the feature, which can be useful for understanding the diversity of the data.\n\n#### Usage:\nThis function is useful for performing an initial exploratory data analysis (EDA) on a feature in a dataset. By calling this function, users can quickly get insights into missing data, data types, and the general distribution of values within a feature.\n\nThe function outputs a dictionary containing these metrics, which can be further used for data preprocessing, cleaning, or detailed reporting.\n","metadata":{}},{"cell_type":"code","source":"def analyze_feature(data, feature_name):\n    \"\"\"\n    Analyzes a specified feature in a DataFrame for null values, data type, \n    total values, and unique values.\n\n    Parameters:\n    data (pd.DataFrame): The DataFrame containing the data.\n    feature_name (str): The name of the feature (column) to analyze.\n\n    Returns:\n    dict: A dictionary containing the following metrics for the feature:\n          - Number of null values\n          - Percentage of null values\n          - Data type of the feature\n          - Total number of values\n          - Number of unique values\n    \"\"\"\n    # Calculate the number of null rows in the selected feature\n    null_count = data[feature_name].isnull().sum()\n\n    # Calculate the total number of values in the feature\n    total_count = data[feature_name].shape[0]\n\n    # Calculate the percentage of null values\n    null_percentage = (null_count / total_count) * 100 if total_count > 0 else 0\n\n    # Get the data type of the feature\n    data_type = data[feature_name].dtype\n\n    # Calculate the number of unique values in the feature\n    unique_count = data[feature_name].nunique()\n\n    # Prepare results in a dictionary\n    results = {\n        'Total number of values': total_count,\n        'Number of null values': null_count,\n        'Percentage of null values': round(null_percentage, 2),\n        'Data type': str(data_type),\n        'Number of unique values': unique_count\n    }\n    \n    return results","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-11T19:01:45.97983Z","iopub.execute_input":"2024-12-11T19:01:45.981011Z","iopub.status.idle":"2024-12-11T19:01:45.988215Z","shell.execute_reply.started":"2024-12-11T19:01:45.980967Z","shell.execute_reply":"2024-12-11T19:01:45.986931Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 1.1.2. Check Normality Function","metadata":{}},{"cell_type":"markdown","source":"The `check_normality` function is an enhanced tool for analyzing the normality and distribution of a specific numeric feature in a DataFrame. It provides statistical insights and performs normality tests to assess how closely the feature follows a normal distribution.\n\n#### Key Metrics and Tests Performed:\n1. **Mean**: The average value of the feature, which provides a central measure of the data.\n2. **Median**: The middle value of the feature when sorted, offering a measure of central tendency less affected by outliers.\n3. **Variance**: Measures the spread of the data around the mean.\n4. **Standard Deviation**: The square root of variance, showing the average deviation from the mean.\n5. **Skewness**: Indicates the asymmetry of the distribution. Positive skew means a right tail, while negative skew indicates a left tail.\n6. **Kurtosis**: Measures the \"tailedness\" of the distribution. High kurtosis indicates heavy tails, while low kurtosis suggests light tails.\n7. **Shapiro-Wilk Test**: A statistical test to assess if the data is normally distributed. The p-value helps determine if the data significantly deviates from normality.\n8. **Kolmogorov-Smirnov (KS) Test**: A test that compares the feature's distribution to a normal distribution using the mean and standard deviation as parameters.\n\n#### Visualizations:\n- **Histogram**: A plot showing the frequency distribution of the feature, providing a visual check for normality.\n- **Box Plot**: A plot showing the spread of the feature's data and identifying potential outliers.\n\n#### Usage:\nThis function is useful for performing a thorough exploratory analysis of a numeric feature’s distribution. It calculates several statistical measures and conducts normality tests, providing a comprehensive view of the feature’s distribution. Additionally, it visualizes the data using a histogram and a box plot, helping users better understand the shape and spread of the feature.\n\nThe function returns a dictionary with the calculated metrics and test results, which can be used for further analysis or reporting.\n","metadata":{}},{"cell_type":"code","source":"def check_normality(data, feature_name):\n    \"\"\"\n    Enhanced function to analyze the normality and distribution of a feature in a DataFrame.\n\n    Parameters:\n    data (pd.DataFrame): The DataFrame containing the data.\n    feature_name (str): The name of the numeric feature to analyze.\n\n    Returns:\n    dict: A dictionary containing various statistics and test results.\n    \"\"\"\n    # Drop NaN values for the feature\n    feature_data = data[feature_name].dropna()\n\n    # Calculate basic statistics\n    mean = feature_data.mean()\n    median = feature_data.median()\n    variance = feature_data.var()\n    std_dev = feature_data.std()\n\n    # Calculate skewness and kurtosis\n    skewness = skew(feature_data)\n    kurtosis_value = kurtosis(feature_data)\n\n    # Perform Shapiro-Wilk Test\n    shapiro_stat, shapiro_p = shapiro(feature_data)\n\n    # Perform Kolmogorov-Smirnov Test against normal distribution\n    ks_stat, ks_p = kstest(feature_data, 'norm', args=(mean, std_dev))\n\n    # Prepare the result dictionary\n    results = {\n        'Mean': mean,\n        'Median': median,\n        'Variance': variance,\n        'Standard Deviation': std_dev,\n        'Skewness': skewness,\n        'Kurtosis': kurtosis_value,\n        'Shapiro-Wilk Statistic': shapiro_stat,\n        'Shapiro-Wilk p-value': shapiro_p,\n        'KS Statistic': ks_stat,\n        'KS p-value': ks_p,\n    }\n\n    # Plot the histogram and boxplot\n    plt.figure(figsize=(10, 5))\n\n    plt.subplot(1, 2, 1)\n    feature_data.hist(bins=30, edgecolor='k')\n    plt.title(f'Histogram of {feature_name}')\n    plt.xlabel(feature_name)\n    plt.ylabel('Frequency')\n\n    plt.subplot(1, 2, 2)\n    data.boxplot(column=feature_name)\n    plt.title(f'Box Plot of {feature_name}')\n\n    plt.tight_layout()\n    plt.show()\n\n    print(f\"Analysis for {feature_name}:\\n\", results)\n    return results","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-11T19:01:47.352443Z","iopub.execute_input":"2024-12-11T19:01:47.352832Z","iopub.status.idle":"2024-12-11T19:01:47.362082Z","shell.execute_reply.started":"2024-12-11T19:01:47.352779Z","shell.execute_reply":"2024-12-11T19:01:47.360841Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 1.1.3. Fill Missing Values KMM Feature Function","metadata":{}},{"cell_type":"markdown","source":"The `fill_missing_values_knn_feature` function uses the K-Nearest Neighbors (KNN) algorithm to fill missing values in a specific feature (column) of a DataFrame. It allows users to impute missing data in a targeted manner, focusing only on the feature of interest.\n\n#### Key Features:\n1. **K-Nearest Neighbors Algorithm**: This imputation method uses the values of the closest neighbors (rows with similar data points) to estimate the missing values in the selected feature. The number of neighbors to use can be adjusted using the `n_neighbors` parameter.\n2. **Targeted Imputation**: Instead of filling missing values across the entire dataset, this function isolates the feature (column) specified by the user, performing the imputation only on that feature.\n3. **Data Integrity**: The function operates on a copy of the original DataFrame to preserve the integrity of the input data, avoiding unintended modifications.\n\n#### Parameters:\n- **`data`**: The DataFrame containing the dataset with missing values.\n- **`feature_name`**: The name of the feature (column) where missing values should be filled.\n- **`n_neighbors`**: The number of neighbors to consider for the imputation (default value is 5).\n\n#### Usage:\nThis function is especially useful when there are missing values in a specific feature and you want to impute those values using the KNN algorithm based on the other data in the dataset. It provides a quick and efficient way to handle missing data without modifying the entire dataset.\n\nThe function returns the original DataFrame with the missing values in the specified feature filled, while the rest of the dataset remains unchanged.\n\n#### Example Output:\nAfter filling the missing values, the function will return the DataFrame with the updated feature. The output DataFrame will have no missing values in the specified feature, allowing for further analysis or modeling.\n","metadata":{}},{"cell_type":"code","source":"def fill_missing_values_knn_feature(data, feature_name, n_neighbors=5):\n    \"\"\"\n    Fills missing values in a specific feature (column) using the KNN algorithm.\n    \n    Parameters:\n    data (pd.DataFrame): The DataFrame containing the data with missing values.\n    feature_name (str): The name of the feature (column) whose missing values should be filled.\n    n_neighbors (int): The number of neighbors to use for imputation (default is 5).\n    \n    Returns:\n    pd.DataFrame: The DataFrame with missing values in the specified feature filled.\n    \"\"\"\n    # Create a copy of the data to avoid modifying the original DataFrame\n    data_copy = data.copy()\n\n    # Isolate the feature column\n    feature_data = data_copy[[feature_name]]\n    \n    # Initialize the KNN imputer\n    imputer = KNNImputer(n_neighbors=n_neighbors)\n\n    # Perform imputation for the selected feature\n    feature_data_filled = imputer.fit_transform(feature_data)\n\n    # Update the feature column in the original DataFrame with the filled values\n    data_copy[feature_name] = feature_data_filled\n\n    return data_copy\n\n# Example usage\n# data = pd.read_csv('data.csv')  # Load your data here\n# feature_name = 'Feature_X'  # Specify the feature you want to fill\n# filled_data = fill_missing_values_knn_feature(data, feature_name)\n# print(filled_data.head())\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-11T19:01:48.405845Z","iopub.execute_input":"2024-12-11T19:01:48.406225Z","iopub.status.idle":"2024-12-11T19:01:48.412466Z","shell.execute_reply.started":"2024-12-11T19:01:48.406193Z","shell.execute_reply":"2024-12-11T19:01:48.411327Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 1.1.4. EDA Pipeline","metadata":{}},{"cell_type":"markdown","source":"This section outlines a data processing pipeline designed to facilitate the exploration, loading, and analysis of a dataset in an efficient and user-friendly manner. The pipeline consists of multiple functions, each performing a specific task to assist in understanding and preparing data for further analysis or machine learning modeling.\n\n---\n\n#### 1. **Data Exploration and Visualization**\n   - **Purpose**: The pipeline provides essential exploratory data analysis (EDA) features that give insights into the structure, content, and quality of the dataset.\n   - **Key Features**:\n     - Visualizes basic statistics of the dataset.\n     - Displays the first few rows of the DataFrame for quick inspection.\n     - Checks for null values and provides a summary of missing data.\n     - Offers summary statistics for numerical features, such as mean, median, standard deviation, etc.\n     - Lists the number of unique values for each feature.\n\n   The `analyze_data` function ensures that a dataset is thoroughly explored by showing key aspects of its structure, missing values, and statistical distributions.\n\n---\n\n#### 2. **Dataset Loading and Validation**\n   - **Purpose**: This part of the pipeline ensures the dataset is correctly loaded from a specified file path, validating the file type and handling common errors.\n   - **Key Features**:\n     - Prompts the user to input a file path if not provided.\n     - Validates that the file is in `.csv` format, raising an error if the file format is incorrect.\n     - Handles common file errors, such as a file not being found or invalid file paths.\n   \n   The `load_dataset` function simplifies the process of loading a dataset by checking the file format and loading it into a pandas DataFrame for subsequent analysis.\n\n---\n\n#### 3. **Section Header Display for Reporting**\n   - **Purpose**: This function organizes the output into readable sections by clearly marking different analysis parts, improving the overall readability of reports.\n   - **Key Features**:\n     - Displays a section title, centered in a line of equal signs (`=`).\n     - Optionally, includes a description for additional context under the title.\n     - Can be used to structure various sections of the analysis, making it easier for users to navigate through the report.\n\n   The `print_section_header` function adds structure and clarity when displaying outputs, helping the user follow the flow of analysis.\n\n---\n\n#### 4. **Feature Selection for Target Variable**\n   - **Purpose**: In supervised learning, selecting the target feature is a crucial step. This function assists the user in selecting the appropriate target variable from the dataset.\n   - **Key Features**:\n     - Prompts the user to select the target feature from the dataset.\n     - Ensures that the user can easily designate which column they wish to predict.\n\n   The `select_target_feature` function plays an important role in the machine learning pipeline by identifying the feature that will be the focus of predictive modeling.\n\n---\n\n### Pipeline Workflow Summary\n\n1. **Loading and Validation**: The pipeline begins by loading a dataset from a specified path using the `load_dataset` function. The dataset must be in CSV format, and appropriate error handling ensures a smooth loading process.\n   \n2. **Data Exploration**: The pipeline proceeds with the `analyze_data` function, which performs comprehensive exploratory data analysis. This step includes visualizing basic statistics, detecting missing values, and understanding the distribution of the dataset.\n\n3. **Feature Selection**: Once the data is analyzed, the user can select the target feature for further predictive analysis using the `select_target_feature` function.\n\n4. **Section Reporting**: The `print_section_header` function is used throughout the pipeline to organize and present the results of each step in a readable manner.\n\n---\n\n### Conclusion\n\nThis pipeline provides a structured and efficient approach for data analysis and machine learning preprocessing. By automating key tasks like dataset loading, exploration, and feature selection, it enables users to quickly get started with data analysis and modeling, minimizing the need for manual intervention. The pipeline ensures that data is thoroughly understood, cleaned, and ready for modeling, facilitating better decision-making in subsequent stages of analysis.\n","metadata":{}},{"cell_type":"code","source":"def print_section_header(title, description=\"\"):\n    line = \"=\" * 50\n    print(line)\n    print(title.center(50))\n    print(line)\n    if description:\n        print(description)\n\ndef load_dataset(file_path=None):\n    # If the file path is not specified, the user is prompted to enter it\n    if file_path is None:\n        file_path = input(\"Please enter the file path: \")\n\n    # Check if the file extension is correct\n    if not file_path.endswith('.csv'):\n        raise ValueError(\"Invalid file format. Please provide a CSV file.\")\n    \n    try:\n        # Read the CSV file\n        data = pd.read_csv(file_path)\n        print(\"Dataset loaded successfully!\")\n        return data\n    except FileNotFoundError:\n        print(\"File not found. Please check the file path.\")\n    except Exception as e:\n        # General error message for other potential errors\n        print(f'An error occurred: {e}')\n\ndef analyze_data(data):\n    \"\"\"\n    Analyzes a given pandas DataFrame.\n\n    Parameters:\n    - data (pandas.DataFrame): The DataFrame to analyze.\n\n    Shows:\n    - First few rows of the DataFrame.\n    - DataFrame info.\n    - List of features with null values.\n    - Summary statistics for numerical features.\n    - Number of unique values in each feature.\n    \"\"\"\n\n    # 1. Check if the input is a pandas DataFrame\n    if not isinstance(data, pd.DataFrame):\n        raise ValueError(\"The input is not a pandas DataFrame. Please provide a valid DataFrame.\")\n    \n\n    # 2. Display the first few rows of the DataFrame\n    print_section_header(\"2A)First Few Rows of the DataFrame\")\n    print(data.head())\n\n    # 3. Display the DataFrame information\n    print_section_header(\"2B)DataFrame Info\")\n    data.info()\n\n    # 4. List features with null values\n    print_section_header(\"2C)Features with Null Values\")\n    null_features = data.columns[data.isnull().any()].tolist()\n    if null_features:\n        for feature in null_features:\n            null_count = data[feature].isnull().sum()\n            print(f'- {feature}: {null_count} null values')\n    else:\n        print(\"No features with null values.\")\n    \n    # 5. Display summary statistics for numerical features\n    print_section_header(\"2D)Summary Statistics for Numerical Features\")\n    print(data.describe())\n\n    # 6. Show the number of unique values in each feature\n    print_section_header(\"2E)Number of Unique Values in Each Feature\")\n    for column in data.columns:\n        unique_values = data[column].nunique()\n        print(f'- {column}: {unique_values} unique values')\n\n\n\ndef select_target_feature(data):\n    \"\"\"\n    Prompts the user to select a target feature from the DataFrame.\n\n    Parameters:\n    - data (pandas.DataFrame): The DataFrame from which to select a target feature.\n\n    Returns:\n    - str: The name of the selected target feature.\n    \"\"\"\n    if not isinstance(data, pd.DataFrame):\n        raise ValueError(\"The input is not a pandas DataFrame. Please provide a valid DataFrame.\")\n\n    available_features = data.columns.tolist()\n    print(\"Available features:\")\n    for i, feature in enumerate(available_features, 1):\n        print(f'{i}. {feature}')\n\n    while True:\n        try:\n            choice = int(input(\"Please enter the number of the target feature: \"))\n            if 1 <= choice <= len(available_features):\n                target_feature = available_features[choice - 1]\n                print(f\"Target feature '{target_feature}' selected successfully.\")\n                return target_feature\n            else:\n                print(\"Invalid number. Please enter a number from the list.\")\n        except ValueError:\n            print(\"Invalid input. Please enter a number.\")\n\ndef categorize_features(data):\n    \"\"\"\n    Categorizes features into numeric, categorical, boolean, and date/time features.\n\n    Parameters:\n    - data (pandas.DataFrame): The DataFrame whose features to categorize.\n\n    Returns:\n    - dict: Dictionary with feature categories.\n    \"\"\"\n    if not isinstance(data, pd.DataFrame):\n        raise ValueError(\"The input is not a pandas DataFrame. Please provide a valid DataFrame.\")\n\n    numeric_features = data.select_dtypes(include=['number']).columns.tolist()\n    categorical_features = data.select_dtypes(include=['object', 'category']).columns.tolist()\n    bool_features = data.select_dtypes(include=['bool']).columns.tolist()\n    datetime_features = data.select_dtypes(include=['datetime']).columns.tolist()\n\n    categorized_features = {\n        'Numeric': numeric_features,\n        'Categorical': categorical_features,\n        'Boolean': bool_features,\n        'DateTime': datetime_features\n    }\n\n    for category, features in categorized_features.items():\n        print_section_header(f'{category} Features')\n        if features:\n            for feature in features:\n                print(f'\\u2022 {feature}')  # • symbol\n        else:\n            print(\"None\")\n\n\ndef plot_histograms(data):\n    \"\"\"\n    Plots histograms for numeric features in the DataFrame,\n    excluding 'id' column. Arranges plots in a grid of three columns per row.\n    \"\"\"\n    # Exclude 'id' from numerical features\n    numeric_features = data.select_dtypes(include=['number']).columns.tolist()\n    if 'id' in numeric_features:\n        numeric_features.remove('id')\n\n    # Configure subplot grid\n    num_cols = 3\n    num_rows = (len(numeric_features) + num_cols - 1) // num_cols\n    fig, axes = plt.subplots(num_rows, num_cols, figsize=(15, 5 * num_rows))\n\n    # Plot histograms\n    for i, feature in enumerate(numeric_features):\n        row, col = divmod(i, num_cols)\n        ax = axes[row, col]\n        \n        # Calculate bins based on unique values\n        unique_values = data[feature].dropna().unique()\n        num_bins = min(len(unique_values), 50)  # Limit to a reasonable number\n        \n        sns.histplot(data[feature], bins=num_bins, ax=ax)\n        ax.set_title(f'Histogram of {feature}')\n\n    # Turn off any unused subplots\n    for j in range(i + 1, num_rows * num_cols):\n        fig.delaxes(axes.flatten()[j])\n\n    plt.tight_layout()\n    plt.show()\n\ndef plot_target_distribution(data, target_feature):\n    \"\"\"\n    Plots the distribution of the target feature, and displays the proportion of each value.\n\n    Parameters:\n    - data (pandas.DataFrame): The DataFrame containing the target feature.\n    - target_feature (str): The name of the target feature.\n    \"\"\"\n    # Check unique values to decide on plot type\n    unique_values = sorted(data[target_feature].unique())\n    \n    plt.figure(figsize=(10, 6))\n\n    if data[target_feature].dtype in ['int64', 'float64'] and len(unique_values) > 10:\n        sns.histplot(data[target_feature], kde=True)\n        plt.title(f'Distribution of Numerical Target Feature: {target_feature}')\n    else:\n        sns.countplot(x=target_feature, data=data, order=unique_values)\n        plt.title(f'Distribution of Categorical Target Feature: {target_feature}')\n        plt.xticks(rotation=45)\n\n        # Calculate and annotate proportions\n        total_count = len(data)\n        for i, value in enumerate(unique_values):\n            count = (data[target_feature] == value).sum()\n            proportion = count / total_count\n            plt.annotate(f'{proportion:.2%}', xy=(i, count), xytext=(0, 5),\n                         textcoords='offset points', ha='center', va='bottom')\n\n    plt.xlabel(f'{target_feature}')\n    plt.ylabel(\"Frequency\")\n    plt.xlim(min(unique_values) - 0.5, max(unique_values) + 0.5)\n    plt.show()\n\ndef plot_correlation_matrix(data):\n    \"\"\"\n    Plots the correlation matrix, excluding 'Name' and 'id' columns if they exist.\n    \"\"\"\n    # Exclude specific columns if they exist\n    columns_to_exclude = ['Name', 'id']\n    numeric_data = data.select_dtypes(include=['number']).copy()\n\n    # Drop specified columns if they are present\n    for col in columns_to_exclude:\n        if col in numeric_data.columns:\n            numeric_data.drop(columns=[col], inplace=True)\n\n    # Calculate and plot the correlation matrix\n    correlation_matrix = numeric_data.corr()\n    plt.figure(figsize=(12, 10))\n    sns.heatmap(correlation_matrix, annot=True, fmt=\".2f\", cmap='coolwarm', cbar=True)\n    plt.title(\"Correlation Matrix Heatmap\")\n    plt.xticks(rotation=45)\n    plt.yticks(rotation=0)\n    plt.tight_layout()\n    plt.show()\n\ndef pipeline(data=None, target_feature=None):\n    \"\"\"\n    Executes the full data analysis pipeline step-by-step.\n    \n    Parameters:\n    - data (pandas.DataFrame or str): The DataFrame to analyze or a file path to load the dataset. \n      If not provided, the user will be prompted to load a dataset.\n    - target_feature (str): The target feature to predict. If not provided, the user will be prompted to select it.\n    \"\"\"\n    # Step 1: Load the dataset if data is a file path or None\n    if isinstance(data, str):\n        # If a string is passed, treat it as a file path and load the dataset\n        print_section_header(\"Step 1: Load the Dataset\", \"Loading the dataset from the provided file path.\")\n        data = load_dataset(data)\n    elif data is None:\n        # If no data is provided, prompt the user to load a dataset\n        print_section_header(\"Step 1: Load the Dataset\", \"Loading the dataset for analysis.\")\n        data = load_dataset()\n\n    if data is not None:\n        # Step 2: Analyze the dataset\n        print_section_header(\"Step 2: Analyze the Dataset\", \"Analyzing data structure and missing values.\")\n        analyze_data(data)\n\n        # Step 3: Select the target feature if not provided\n        if target_feature is None:\n            print_section_header(\"Step 3: Select the Target Feature\", \"Select the feature to predict.\")\n            target_feature = select_target_feature(data)\n\n        # Step 4: Categorize features\n        print_section_header(\"Step 4: Categorize Features\", \"Categorizing features by type.\")\n        categorize_features(data)\n        \n        # Step 5: Visualize data\n        print_section_header(\"Step 5: Visualize Data\", \"Creating visualizations for exploratory data analysis.\")\n        plot_histograms(data)\n        \n        # Step 6: Target Feature Distribution\n        print_section_header(\"Step 6: Target Feature Distribution\", \"Visualizing the distribution of the target feature.\")\n        plot_target_distribution(data, target_feature)\n        \n        # Step 7: Correlation Matrix\n        print_section_header(\"Step 7: Correlation Matrix\", \"Analyzing feature correlations.\")\n        plot_correlation_matrix(data)\n\n        print_section_header(\"Pipeline Completed\", \"Pipeline executed successfully.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-11T19:02:09.774066Z","iopub.execute_input":"2024-12-11T19:02:09.774428Z","iopub.status.idle":"2024-12-11T19:02:09.804502Z","shell.execute_reply.started":"2024-12-11T19:02:09.774397Z","shell.execute_reply":"2024-12-11T19:02:09.803321Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 1.2. EDA","metadata":{}},{"cell_type":"code","source":"pipeline(data ='/kaggle/input/playground-series-s4e12/train.csv' ,target_feature='Premium Amount')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-11T19:02:31.265565Z","iopub.execute_input":"2024-12-11T19:02:31.265948Z","iopub.status.idle":"2024-12-11T19:02:58.669557Z","shell.execute_reply.started":"2024-12-11T19:02:31.265915Z","shell.execute_reply":"2024-12-11T19:02:58.668363Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 1.3. Dataset Overview","metadata":{}},{"cell_type":"markdown","source":"### Dataset Overview\n\nThis dataset comprises various features related to insurance policyholders. Each feature's description and its characteristics are outlined below:\n\n#### 1. `id` (int64)\n- **Description**: A unique identifier for each entry in the dataset.\n- **Non-Null Count**: 1,200,000\n- **Remarks**: Integer values used for indexing, with no missing entries.\n\n#### 2. `Age` (float64)\n- **Description**: Represents the age of the policyholder.\n- **Non-Null Count**: 1,181,295\n- **Remarks**: Contains missing values to account for.\n\n#### 3. `Gender` (object)\n- **Description**: Gender of the policyholder, typically 'Male' or 'Female'.\n- **Non-Null Count**: 1,200,000\n- **Remarks**: All entries are complete.\n\n#### 4. `Annual Income` (float64)\n- **Description**: The annual income of the policyholder, likely in dollars.\n- **Non-Null Count**: 1,155,051\n- **Remarks**: Requires imputation for missing values.\n\n#### 5. `Marital Status` (object)\n- **Description**: The marital status of the policyholder, with categories such as 'Married', 'Single', etc.\n- **Non-Null Count**: 1,181,471\n- **Remarks**: Some entries missing.\n\n#### 6. `Number of Dependents` (float64)\n- **Description**: The number of individuals financially dependent on the policyholder.\n- **Non-Null Count**: 1,090,328\n- **Remarks**: Contains missing values.\n\n#### 7. `Education Level` (object)\n- **Description**: The highest level of education attained by the policyholder, e.g., 'Bachelor's', 'Master's'.\n- **Non-Null Count**: 1,200,000\n- **Remarks**: Complete data.\n\n#### 8. `Occupation` (object)\n- **Description**: The occupation of the policyholder, such as 'Self-Employed', 'Engineer', etc.\n- **Non-Null Count**: 841,925\n- **Remarks**: Requires imputation for missing values.\n\n#### 9. `Health Score` (float64)\n- **Description**: A numeric score indicating health metrics, potentially standardized.\n- **Non-Null Count**: 1,125,924\n- **Remarks**: Further details and potential imputations needed for missing values.\n\n#### 10. `Location` (object)\n- **Description**: Describes the living area of the policyholder, like 'Urban', 'Suburban', or 'Rural'.\n- **Non-Null Count**: 1,200,000\n- **Remarks**: Fully populated.\n\n#### 11. `Policy Type` (object)\n- **Description**: The type of insurance policy held by the policyholder.\n- **Non-Null Count**: 1,200,000\n- **Remarks**: No missing entries.\n\n#### 12. `Previous Claims` (float64)\n- **Description**: Number of prior insurance claims made by the policyholder.\n- **Non-Null Count**: 835,971\n- **Remarks**: Requires further handling for null values.\n\n#### 13. `Vehicle Age` (float64)\n- **Description**: The age of the vehicle associated with the policy.\n- **Non-Null Count**: 1,199,994\n- **Remarks**: Nearly complete, minor missing values.\n\n#### 14. `Credit Score` (float64)\n- **Description**: A numeric score representing the creditworthiness of the policyholder.\n- **Non-Null Count**: 1,062,118\n- **Remarks**: Contains significant missing data.\n\n#### 15. `Insurance Duration` (float64)\n- **Description**: Duration of the insurance policy.\n- **Non-Null Count**: 1,199,999\n- **Remarks**: Almost complete.\n\n#### 16. `Policy Start Date` (object)\n- **Description**: The date when the insurance policy commenced.\n- **Non-Null Count**: 1,200,000\n- **Remarks**: No null entries.\n\n#### 17. `Customer Feedback` (object)\n- **Description**: Feedback or rating provided by the customer regarding services.\n- **Non-Null Count**: 1,122,176\n- **Remarks**: Some entries are missing.\n\n#### 18. `Smoking Status` (object)\n- **Description**: Indicates if the policyholder smokes, typically 'Yes' or 'No'.\n- **Non-Null Count**: 1,200,000\n- **Remarks**: Fully covered data.\n\n#### 19. `Exercise Frequency` (object)\n- **Description**: Frequency of exercise by the policyholder, e.g., 'Daily', 'Weekly', etc.\n- **Non-Null Count**: 1,200,000\n- **Remarks**: Complete entries.\n\n#### 20. `Property Type` (object)\n- **Description**: Type of property owned or lived in by the policyholder.\n- **Non-Null Count**: 1,200,000\n- **Remarks**: All entries accounted for.\n\n#### 21. `Premium Amount` (float64)\n- **Description**: The premium cost for the insurance policy.\n- **Non-Null Count**: 1,200,000\n- **Remarks**: Fully populated.\n\n### Notes\n\n- **Data Handling Needed**: Several features require imputation for missing values to prepare the dataset for analysis or modeling.\n- **Data Types**: The dataset includes both categorical (`object` type) and continuous (`float64` type) variables, which may necessitate different preprocessing approaches.\n\n","metadata":{}},{"cell_type":"markdown","source":"# 2. Future Analysis and Cleaning","metadata":{}},{"cell_type":"markdown","source":"## 2.1 Loading Datasets","metadata":{}},{"cell_type":"code","source":"# Read the training data\n# The CSV file is located at the specified path and is loaded into a DataFrame\ntrain_data = pd.read_csv('/kaggle/input/playground-series-s4e12/train.csv')\n\n# Read the test data\n# The test set is also read from a CSV file and loaded into a DataFrame\ntest_data = pd.read_csv('/kaggle/input/playground-series-s4e12/test.csv')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 2.2 ID","metadata":{}},{"cell_type":"code","source":"# Drop the 'id' column from the training data and create a new DataFrame\ndata_id_train = train_data.drop(columns=['id'])\n\n# Keep the 'id' column from the test data\nid_keep = test_data['id']  # Capture the 'id' values for later use\n\n# Drop the 'id' column from the test data and create a new DataFrame\ndata_id_test = test_data.drop(columns=['id'])","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 2.3. Age","metadata":{}},{"cell_type":"code","source":"analyze_feature(data_id_train, 'Age')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"analyze_feature(data_id_test, 'Age')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Drop rows where the 'Age' feature has NaN values\ndata_age_train = data_id_train.dropna(subset=['Age'])","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"analyze_feature(data_age_train, 'Age')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Calculate the median of the 'Age' column in the original DataFrame\nmedian_age = data_id_test['Age'].median()  # Find the median, ignoring NaN values\n\n# Create a copy of the original DataFrame to work with\ndata_age_test = data_id_test.copy()  # Duplicate the DataFrame to keep the original intact\n\n# Fill NaN values in the 'Age' column of the copied DataFrame using the median\ndata_age_test['Age'] = data_age_test['Age'].fillna(median_age)  \n# Replace NaN values with the calculated median","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"analyze_feature(data_age_test, 'Age')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"result = check_normality(data_age_train, 'Age')\nprint(result)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Analysis Results Interpretation\n\n- **Mean and Median**: The mean and median values are almost the same (41.15 and 41), suggesting a symmetric distribution. However, this alone is not a definitive indicator.\n\n- **Variance and Standard Deviation**: These metrics indicate the spread of the data. A standard deviation of 13.54 could be considered reasonable for age data.\n\n- **Skewness**: A value around -0.012 indicates a slightly negative skew but is nearly symmetric. Skewness close to 0 is often interpreted as a symmetric distribution.\n\n- **Kurtosis**: A value of -1.19 suggests the data distribution is flatter (platykurtic) and has fewer outliers than a normal distribution.\n\n- **Shapiro-Wilk and KS Tests**: The very small p-values (much less than 0.05) from both tests strongly suggest that the data does not follow a normal distribution. The Shapiro-Wilk p-value, in particular, is quite low.\n\n### Conclusions and Recommendations\n\nThese results indicate that the \"Age\" data is not normally distributed and has weak outlier presence. Given this:\n\n- **Min-Max Scaling**: Ideal for scaling non-normally distributed data, especially when you want to scale the data to a [0, 1] range.\n\n- **RobustScaler**: Use this if you want to minimize the impact of skewness and outliers. It scales the data using the median and IQR (interquartile range).\n\nThese scaling methods can help bring the data into a more balanced format, potentially improving the performance of algorithms.","metadata":{}},{"cell_type":"code","source":"# Min-Max Scaler\nmin_max_scaler = MinMaxScaler()\ndata_age_train['Age'] = min_max_scaler.fit_transform(data_age_train[['Age']])\ndata_age_test['Age'] = min_max_scaler.fit_transform(data_age_test[['Age']])","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 2.4. Annual Income","metadata":{}},{"cell_type":"code","source":"analyze_feature(data_age_train, 'Annual Income')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"analyze_feature(data_age_test, 'Annual Income')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"result = check_normality(data_age_train, 'Annual Income')\nprint(result)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Report on Missing Value Imputation for \"Annual Income\"\n\n#### Overview\n\nThe \"Annual Income\" feature has been analyzed to determine the most appropriate method for imputing missing values. Several statistical metrics and distribution tests were considered to make an informed decision.\n\n#### Analysis Results\n\n- **Mean and Median**: The mean is 32,739.94 while the median is 23,906.5. The significant difference between these two measures indicates a right-skewed distribution.\n\n- **Skewness**: The skewness value of 1.47 confirms the presence of right skewness in the data. This suggests that there are many large values affecting the distribution.\n\n- **Kurtosis**: A kurtosis of 1.80 shows that the distribution is flatter than a normal distribution (platykurtic), with fewer outliers.\n\n- **Normality Tests**: The Shapiro-Wilk and Kolmogorov-Smirnov tests both yield p-values significantly less than 0.05, indicating that the data does not follow a normal distribution.\n\n#### Recommendations\n\nGiven the analysis results, the use of **median imputation** for filling missing values in the \"Annual Income\" feature is recommended. Here’s why:\n\n1. **Robustness to Outliers**: The median is robust to outliers and skews, making it a reliable measure of central tendency for skewed distributions. In the case of \"Annual Income\", where many values are large, using the median will minimize the impact of these outliers.\n\n2. **Discrepancy Between Mean and Median**: The large discrepancy between the mean and median suggests that the mean is influenced by extreme values. Hence, relying on the median for imputation helps mitigate this issue.\n\n3. **Simplicity and Effectiveness**: Median imputation is straightforward to implement and effective for datasets with non-normal distributions.\n\n4. **Consistent Central Tendency**: By using the median, the imputed values are consistent with the central tendency representative of the majority of data points, excluding outliers.\n\n#### Conclusion\n\nThe right-skewed distribution and the presence of potentially influential outliers make median imputation the most prudent choice for addressing missing values in the \"Annual Income\" feature. This approach ensures a balanced representation of central tendency without the distortion from extreme values.\n","metadata":{}},{"cell_type":"code","source":"# Calculate the median of the 'Annual Income' column in the training data\nmedian_income_train = data_age_train['Annual Income'].median()  # Computes the median to fill missing values in the training set\n\n# Calculate the median of the 'Annual Income' column in the test data\nmedian_income_test = data_age_test['Annual Income'].median()  # Computes the median to fill missing values in the test set\n\n# Create a copy of the original training DataFrame\ndata_ai_train = data_age_train.copy()  # Makes a copy to preserve the original training data\n\n# Create a copy of the original test DataFrame\ndata_ai_test = data_age_test.copy()  # Makes a copy to preserve the original test data\n\n# Fill NaN values in 'Annual Income' with the median in the training data\ndata_ai_train['Annual Income'] = data_ai_train['Annual Income'].fillna(median_income_train)\n# Fills missing values using the median calculated from the training set\n\n# Fill NaN values in 'Annual Income' with the median in the test data\ndata_ai_test['Annual Income'] = data_ai_test['Annual Income'].fillna(median_income_test)\n# Fills missing values using the median calculated from the test set","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"analyze_feature(data_ai_train, 'Annual Income')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"analyze_feature(data_ai_test, 'Annual Income')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Recommended Scaling Method for \"Annual Income\"\n\nBased on the analysis of the \"Annual Income\" feature, considering factors like skewness and deviation from normal distribution, the following scaling methods are recommended:\n\n#### 1. RobustScaler\n- **Why This Method?**: The analysis indicates that the \"Annual Income\" data is skewed and has some outliers. RobustScaler uses the median and interquartile range (IQR) to scale the data, making it resistant to the influence of outliers.\n- **Use Case**: Robust approaches are ideal when there is significant skewness and the presence of outliers in the data distribution.\n\n#### 2. Min-Max Scaling\n- **Why This Method?**: If the goal is to normalize the data into a [0, 1] range, Min-Max scaling can be utilized. However, this method can be sensitive to the effects of outliers which might expand the range.\n- **Use Case**: Preferred when standard scaling is needed and when data needs to be constrained within a specific range.\n\n### Conclusion and Recommendation\n\nConsidering the skewness in the data distribution and the presence of outliers, **RobustScaler** is the most suitable choice for the \"Annual Income\" feature. This method minimizes the impact of outliers and skewness while preserving the central tendency of the data. As a result, your dataset will provide more consistent and reliable outcomes when used in machine learning models.\n\nIf you have any further questions or need additional recommendations, feel free to reach out!","metadata":{}},{"cell_type":"code","source":"# Initialize RobustScaler\nscaler = RobustScaler()\n\n# Apply RobustScaler to the 'Annual Income' column in the training set\ndata_ai_train['Annual Income'] = scaler.fit_transform(data_ai_train[['Annual Income']])\n\n# Apply the same scaler to the test set to ensure consistency\ndata_ai_test['Annual Income'] = scaler.transform(data_ai_test[['Annual Income']])","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 2.5. Health Score","metadata":{}},{"cell_type":"code","source":"analyze_feature(data_ai_train, 'Health Score')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"analyze_feature(data_ai_test, 'Health Score')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"result = check_normality(data_ai_train, 'Health Score')\nprint(result)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Report on Missing Value Imputation for \"Health Score\"\n\n#### Overview\n\nThis report evaluates the \"Health Score\" feature to determine the most suitable method for filling missing values. Based on statistical analysis, we recommend using median imputation. Below are the detailed reasons for this choice.\n\n#### Analysis Summary\n\n- **Mean and Median**: The mean is 25.61, and the median is 24.57, indicating that the data is almost symmetric. The closeness of these two measures suggests that the data is not heavily skewed.\n\n- **Skewness**: A skewness value of 0.282 indicates slight right skewness but is close to zero, showcasing near symmetry in the distribution.\n\n- **Kurtosis**: With a kurtosis of -0.78, the distribution is flatter than a normal distribution (platykurtic) with fewer outliers.\n\n- **Normality Tests**: The Shapiro-Wilk and Kolmogorov-Smirnov tests both report p-values significantly below 0.05, confirming that the data doesn’t follow a normal distribution strictly. However, the deviation is not extreme.\n\n#### Recommendation\n\n**Median Imputation**: \n\n- **Rationale**: Given the nearly symmetric distribution and the low impact of skewness, median provides a robust measure of central tendency. This method effectively reduces the influence of potential outliers and captures the central value of the dataset.\n\n- **Benefits**:\n  - **Robustness to Outliers**: Median is less affected by extreme values, ensuring a stable central tendency among the data.\n  - **Central Tendency Representation**: It aligns well with the near-symmetry and distribution characteristics observed in the \"Health Score\" data.\n \n#### Conclusion\n\nBased on the statistical evaluation of \"Health Score\", median imputation is the most appropriate method for handling missing values. It addresses the nearly symmetric nature of the data and provides reliable central value capturing, particularly in the presence of minimal skewness.\n\nBy choosing median imputation, we ensure that the imputation process enhances data consistency and reliability, which is crucial for subsequent analyses or model training.\n","metadata":{}},{"cell_type":"code","source":"# Calculate the median of the 'Health Score' column in the training data\nmedian_income_train = data_ai_train['Health Score'].median()  # Computes the median to fill missing values in the training set\n\n# Calculate the median of the 'Health Score' column in the test data\nmedian_income_test = data_ai_test['Health Score'].median()  # Computes the median to fill missing values in the test set\n\n# Create a copy of the original training DanltaFrame\ndata_hs_train = data_ai_train.copy()  # Makes a copy to preserve the original training data\n\n# Create a copy of the original test DataFrame\ndata_hs_test = data_ai_test.copy()  # Makes a copy to preserve the original test data\n\n# Fill NaN values in 'Health Score' with the median in the training data\ndata_hs_train['Health Score'] = data_hs_train['Health Score'].fillna(median_income_train)\n# Fills missing values using the median calculated from the training set\n\n# Fill NaN values in 'Health Score' with the median in the test data\ndata_hs_test['Health Score'] = data_hs_test['Health Score'].fillna(median_income_test)\n# Fills missing values using the median calculated from the test set","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Report on Scaling Method Recommendation for \"Health Score\"\n\n#### Overview\n\nThis report evaluates the scaling methods for the \"Health Score\" feature, considering its statistical analysis. Based on the distribution characteristics and analysis results, StandardScaler is recommended for scaling the \"Health Score\". Below are the detailed reasons for this choice.\n\n#### Analysis Summary\n\n- **Mean and Median**: The mean of 25.61 and median of 24.57 are close, indicating that the data is nearly symmetric. This symmetry suggests that the \"Health Score\" distribution approximates a normal shape.\n\n- **Skewness**: With a skewness value of 0.282, the data shows slight right skewness but remains near zero, reinforcing the near-normal distribution assumption.\n\n- **Kurtosis**: The kurtosis value of -0.78 indicates a flatter distribution (platykurtic) but with limited outliers, suitable for standard normalization.\n\n- **Normality Tests**: The Shapiro-Wilk and KS tests report p-values far below 0.05, indicating the data does not strictly follow a normal distribution. However, the deviation is not extreme, supporting a near-normal assumption for practical purposes.\n\n#### Recommendation\n\n**StandardScaler**:\n\n- **Rationale**: Given the symmetry in the data and proximity to a normal distribution, StandardScaler is well-suited for this feature. It transforms the data by centering it around zero and scaling it to unit variance, which supports many machine learning models that assume normally distributed data.\n\n- **Benefits**:\n  - **Alignment with Model Assumptions**: Suitable for linear models or algorithms like SVM or logistic regression that assume input features are normally distributed.\n  - **Mean Centering and Unit Variance**: Facilitates convergence during optimization by standardizing feature scales to a similar range.\n\n#### Conclusion\n\nThe close alignment of mean and median and the low skewness suggest that the \"Health Score\" is poised for effective scaling with StandardScaler. This choice ensures consistent scaling across features, potentially improving model performance and stability in scenarios where normal distribution assumptions hold important.\n\nBy opting for StandardScaler, the processed \"Health Score\" data becomes more amenable to various analytical tasks and machine learning models, fostering improved results and insights.\n","metadata":{}},{"cell_type":"code","source":"# Initialize StandardScaler\nscaler = StandardScaler()\n\n# Fit and transform the 'Health Score' in the training set\ndata_hs_train['Health Score'] = scaler.fit_transform(data_hs_train[['Health Score']])\n\n# Transform the 'Health Score' in the test set\ndata_hs_test['Health Score'] = scaler.transform(data_hs_test[['Health Score']])","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 2.6. Credit Score","metadata":{}},{"cell_type":"code","source":"analyze_feature(data_hs_train, 'Credit Score')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"analyze_feature(data_hs_test, 'Credit Score')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"result = check_normality(data_hs_train, 'Credit Score')\nprint(result)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Report on Missing Value Imputation for \"Credit Score\"\n\n#### Overview\n\nThis report evaluates the \"Credit Score\" feature to determine the most effective method for filling missing values. Given the statistical characteristics observed in the dataset, median imputation is recommended. Below, the rationale for this choice is detailed.\n\n#### Analysis Summary\n\n- **Mean and Median**: The mean is 592.99, and the median is 595.00. The proximity of these values suggests a symmetric distribution, indicating that the median is a reliable measure of central tendency.\n\n- **Skewness**: The skewness value is -0.114, indicating a slightly left-skewed distribution. However, this skewness is minimal and suggests a roughly symmetric dataset.\n\n- **Kurtosis**: With a kurtosis value of -1.09, the data distribution is flatter than a normal distribution, displaying a platykurtic nature with fewer outliers.\n\n- **Normality Tests**: Both the Shapiro-Wilk and Kolmogorov-Smirnov tests yield p-values significantly below 0.05, indicating that the data does not strictly follow a normal distribution, but the symmetry observed justifies using central measures like median for imputation.\n\n#### Recommendation\n\n**Median Imputation**:\n\n- **Rationale**: Given the near-symmetry and low skewness in the distribution, the median is an ideal choice for imputation. It effectively captures the central tendency of the data while minimizing the impact of outliers or skewness.\n\n- **Benefits**:\n  - **Resilience to Outliers**: Median is less sensitive to extreme values, making it a robust choice for datasets where the mean might be misleading.\n  - **Consistency with Data Symmetry**: Aligns well with the observed symmetric nature of the data, ensuring imputed values are representative of the true data distribution.\n\n#### Conclusion\n\nBased on the statistical assessment of the \"Credit Score\" data, median imputation is the most suitable strategy for addressing missing values. This approach leverages the symmetric nature of the data, ensuring accurate representation of the dataset while maintaining robustness against outliers, thereby supporting improved analytical outcomes and model performance.\n","metadata":{}},{"cell_type":"code","source":"# Calculate the median of the 'Credit Score' column in the training data\nmedian_income_train = data_hs_train['Credit Score'].median()  # Computes the median to fill missing values in the training set\n\n# Calculate the median of the 'Credit Score' column in the test data\nmedian_income_test = data_hs_test['Credit Score'].median()  # Computes the median to fill missing values in the test set\n\n# Create a copy of the original training DanltaFrame\ndata_cs_train = data_hs_train.copy()  # Makes a copy to preserve the original training data\n\n# Create a copy of the original test DataFrame\ndata_cs_test = data_ai_test.copy()  # Makes a copy to preserve the original test data\n\n# Fill NaN values in 'Credit Score' with the median in the training data\ndata_cs_train['Credit Score'] = data_cs_train['Credit Score'].fillna(median_income_train)\n# Fills missing values using the median calculated from the training set\n\n# Fill NaN values in 'Credit Score' with the median in the test data\ndata_cs_test['Credit Score'] = data_cs_test['Credit Score'].fillna(median_income_test)\n# Fills missing values using the median calculated from the test set","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"analyze_feature(data_cs_train, 'Credit Score')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"analyze_feature(data_cs_test, 'Credit Score')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Report on Scaling Method Recommendation for \"Credit Score\"\n\n#### Overview\n\nThis report assesses the scaling methods for the \"Credit Score\" feature, considering its statistical analysis. Based on the data characteristics observed, StandardScaler is recommended for scaling. The rationale for this recommendation is provided below.\n\n#### Analysis Summary\n\n- **Mean and Median**: The mean of 592.99 and median of 595.00 are closely aligned, indicating that the data distribution is almost symmetric.\n  \n- **Skewness**: A skewness value of -0.114 signifies a minor left skew, implying that the overall distribution is nearly symmetric.\n\n- **Kurtosis**: With a kurtosis of -1.09, the distribution is flatter than a normal curve (platykurtic), suggesting fewer extreme outliers.\n\n- **Normality Tests**: Despite the Shapiro-Wilk and KS tests indicating non-normal distributions with very low p-values, the minimal skewness suggests a near-normal or symmetric distribution.\n\n#### Recommendation\n\n**StandardScaler**:\n\n- **Rationale**: Given the close proximity of mean and median, low skewness, and reasonable variance, applying StandardScaler ensures that the data is centered around zero and has a unit variance. This is particularly suitable for models that benefit from normally distributed data.\n\n- **Benefits**:\n  - **Handles Variability**: StandardScaler effectively manages variations in scale among features, ensuring that none disproportionately influences the model.\n  - **Optimal for Symmetric Data**: Enhances consistency across feature scales by standardizing them to have the same importance in a model.\n  - **Supports Many Models**: Well-suited for algorithms like SVM or logistic regression, which benefit from data that resembles a normal distribution.\n\n#### Conclusion\n\nBased on the statistical evaluation of \"Credit Score\", StandardScaler is the most appropriate method for scaling. It aligns with the symmetric nature and low skewness observed in the data, ensuring a balanced feature set for analytical modeling.\n\nBy choosing StandardScaler, you enhance the consistency of the data representation and potentially improve the robustness and performance of any machine learning models applied.","metadata":{}},{"cell_type":"code","source":"# Initialize StandardScaler\nscaler = StandardScaler()\n\n# Fit and transform the 'Credit Score' in the training set\ndata_cs_train['Credit Score'] = scaler.fit_transform(data_cs_train[['Credit Score']])\n\n# Transform the 'Credit Score' in the test set\ndata_cs_test['Credit Score'] = scaler.transform(data_cs_test[['Credit Score']])","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"result = check_normality(data_cs_train, 'Credit Score')\nprint(result)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 2.7. Vehicle Age","metadata":{}},{"cell_type":"code","source":"analyze_feature(data_cs_train, 'Vehicle Age')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"analyze_feature(data_cs_test, 'Vehicle Age')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"result = check_normality(data_cs_train, 'Vehicle Age')\nprint(result)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Calculate the median of the 'Vehicle Age' column in the test data\nmedian_income_test = data_cs_test['Vehicle Age'].median()  # Computes the median to fill missing values in the test set\n\n# Create a copy of the original training DanltaFrame\ndata_vc_train = data_cs_train.copy()  # Makes a copy to preserve the original training data\n\n# Create a copy of the original test DataFrame\ndata_vc_test = data_cs_test.copy()  # Makes a copy to preserve the original test data\n\n# Fill NaN values in 'Vehicle Age' with the median in the test data\ndata_vc_test['Vehicle Age'] = data_vc_test['Vehicle Age'].fillna(median_income_test)\n# Fills missing values using the median calculated from the test set\n\n# Drop rows where the 'Vehicle Age' feature has NaN values\ndata_vc_train = data_vc_train.dropna(subset=['Vehicle Age'])","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Initialize StandardScaler\nscaler = StandardScaler()\n\n# Fit and transform the 'Vehicle Age' in the training set\ndata_vc_train['Vehicle Age'] = scaler.fit_transform(data_vc_train[['Vehicle Age']])\n\n# Transform the 'Vehicle Age' in the test set\ndata_vc_test['Vehicle Age'] = scaler.transform(data_vc_test[['Vehicle Age']])","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 2.8. Gender","metadata":{}},{"cell_type":"code","source":"analyze_feature(data_vc_train, 'Gender')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"analyze_feature(data_vc_test, 'Gender')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Create a copy of the original training DanltaFrame\ndata_g_train = data_vc_train.copy()  # Makes a copy to preserve the original training data\n\n# Create a copy of the original test DataFrame\ndata_g_test = data_vc_test.copy()  # Makes a copy to preserve the original test data\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# İki değerli değişkeni 0 ve 1 ile değiştirin\ndata_g_train['Gender'] = data_g_train['Gender'].replace({'Male': 1, 'Female': 0})\ndata_g_test['Gender'] = data_g_test['Gender'].replace({'Male': 1, 'Female': 0})","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 2.9. Education Level","metadata":{}},{"cell_type":"code","source":"analyze_feature(data_g_train, 'Education Level')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"analyze_feature(data_g_test, 'Education Level')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Create a copy of the original training DanltaFrame\ndata_el_train = data_g_train.copy()  # Makes a copy to preserve the original training data\n\n# Create a copy of the original test DataFrame\ndata_el_test = data_g_test.copy()  # Makes a copy to preserve the original test data\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Apply one-hot encoding\ndata_el_train = pd.get_dummies(data_el_train, columns=['Education Level'])\ndata_el_test = pd.get_dummies(data_el_test, columns=['Education Level'])","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 2.10. Marital Status","metadata":{}},{"cell_type":"code","source":"analyze_feature(data_el_train, 'Marital Status')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"analyze_feature(data_el_test, 'Marital Status')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"data_el_train['Marital Status'].unique()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Create a copy of the original training DanltaFrame\ndata_ms_train = data_el_train.copy()  # Makes a copy to preserve the original training data\n\n# Create a copy of the original test DataFrame\ndata_ms_test = data_el_test.copy()  # Makes a copy to preserve the original test data","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Drop rows where the 'Marital Status' feature has NaN values\ndata_ms_train = data_ms_train.dropna(subset=['Marital Status'])","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"analyze_feature(data_ms_train, 'Marital Status')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(data_ms_test['Marital Status'].value_counts(normalize=True))","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Report on Missing Value Imputation for \"Marital Status\"\n\n#### Overview\n\nThis report outlines the strategy employed for filling missing values in the \"Marital Status\" feature, leveraging the dataset's inherent distribution properties. Given the balanced nature of categories, a random imputation approach based on observed category proportions was selected.\n\n#### Data Analysis\n\n- **Category Distribution**:\n  - **Single**: 33.47%\n  - **Married**: 33.37%\n  - **Divorced**: 33.16%\n\nThe distribution analysis revealed near-equivalent proportions across all categories, indicating a well-balanced dataset with no distinct dominant category.\n\n#### Imputation Strategy: Random Imputation Based on Distribution\n\n- **Rationale**: The similarity in category proportions makes random imputation an excellent strategy to preserve this balance while filling missing values. This approach minimizes the risk of introducing bias that could skew model insights or analysis outcomes.\n\n#### Outcome\n\n- **Distribution Preservation**: This method successfully maintained the original distribution of categories across the dataset, facilitating a balanced representation post-imputation.\n- **Integrity of Analysis**: By retaining proportional integrity, subsequent data analysis and modeling efforts can proceed without the distortion that unevenly imputed data might impose.\n\n#### Conclusion\n\nRandom imputation based on category distribution was selected as the optimal strategy for handling missing values in the \"Marital Status\" feature. This method respects the dataset's intrinsic balance and supports robust analytical and modeling pursuits.","metadata":{}},{"cell_type":"code","source":"# Assign missing values based on category proportions\ndata_ms_test['Marital Status'] = data_ms_test['Marital Status'].apply(\n    lambda x: np.random.choice(['Single', 'Married', 'Divorced'],\n                               p=[0.3347, 0.3337, 0.3316]) if pd.isna(x) else x\n)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"analyze_feature(data_ms_test, 'Marital Status')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Apply one-hot encoding\ndata_ms_train = pd.get_dummies(data_ms_train, columns=['Marital Status'])\ndata_ms_test = pd.get_dummies(data_ms_test, columns=['Marital Status'])","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 2.11. Property Type","metadata":{}},{"cell_type":"code","source":"analyze_feature(data_ms_train, 'Property Type')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"analyze_feature(data_ms_test, 'Property Type')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Create a copy of the original training DataFrame\ndata_pt_train = data_ms_train.copy()  # Makes a copy to preserve the original training data\n\n# Create a copy of the original test DataFrame\ndata_pt_test = data_ms_test.copy()  # Makes a copy to preserve the original test data","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Apply one-hot encoding\ndata_pt_train = pd.get_dummies(data_pt_train, columns=['Property Type'])\ndata_pt_test = pd.get_dummies(data_pt_test, columns=['Property Type'])","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 2.12. Policy Type","metadata":{}},{"cell_type":"code","source":"analyze_feature(data_pt_train, 'Policy Type')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"analyze_feature(data_pt_test, 'Policy Type')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Create a copy of the original training DataFrame\ndata_pot_train = data_pt_train.copy()  # Makes a copy to preserve the original training data\n\n# Create a copy of the original test DataFrame\ndata_pot_test = data_pt_test.copy()  # Makes a copy to preserve the original test data","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Apply one-hot encoding\ndata_pot_train = pd.get_dummies(data_pot_train, columns=['Policy Type'])\ndata_pot_test = pd.get_dummies(data_pot_test, columns=['Policy Type'])","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 2.13. Number of Dependents","metadata":{}},{"cell_type":"code","source":"analyze_feature(data_pot_train, 'Number of Dependents')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"analyze_feature(data_pot_test, 'Number of Dependents')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Create a copy of the original training DataFrame\ndata_nod_train = data_pot_train.copy()  # Makes a copy to preserve the original training data\n\n# Create a copy of the original test DataFrame\ndata_nod_test = data_pot_test.copy()  # Makes a copy to preserve the original test data","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(data_nod_train['Number of Dependents'].value_counts(normalize=True))","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Replace NaN values with 'Unknown'\ndata_nod_train['Number of Dependents'].fillna('Unknown', inplace=True)\ndata_nod_test['Number of Dependents'].fillna('Unknown', inplace=True)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Apply one-hot encoding\ndata_nod_train = pd.get_dummies(data_nod_train, columns=['Number of Dependents'])\ndata_nod_test = pd.get_dummies(data_nod_test, columns=['Number of Dependents'])","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 2.14. Occupation","metadata":{}},{"cell_type":"code","source":"analyze_feature(data_nod_train, 'Occupation')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"analyze_feature(data_nod_test, 'Occupation')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"data_nod_train['Occupation'].unique()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Create a copy of the original training DataFrame\ndata_ocp_train = data_nod_train.copy()  # Makes a copy to preserve the original training data\n\n# Create a copy of the original test DataFrame\ndata_ocp_test = data_nod_test.copy()  # Makes a copy to preserve the original test data","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Replace NaN values with 'Unknown'\ndata_ocp_train['Occupation'].fillna('Unknown', inplace=True)\ndata_ocp_test['Occupation'].fillna('Unknown', inplace=True)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Apply one-hot encoding\ndata_ocp_train = pd.get_dummies(data_ocp_train, columns=['Occupation'])\ndata_ocp_test = pd.get_dummies(data_ocp_test, columns=['Occupation'])","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 2.15. Location","metadata":{}},{"cell_type":"code","source":"analyze_feature(data_ocp_train, 'Location')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"analyze_feature(data_ocp_test, 'Location')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"data_ocp_train['Location'].unique()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Create a copy of the original training DataFrame\ndata_lc_train = data_ocp_train.copy()  # Makes a copy to preserve the original training data\n\n# Create a copy of the original test DataFrame\ndata_lc_test = data_ocp_test.copy()  # Makes a copy to preserve the original test data","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Apply one-hot encoding\ndata_lc_train = pd.get_dummies(data_lc_train, columns=['Location'])\ndata_lc_test = pd.get_dummies(data_lc_test, columns=['Location'])","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 2.16. Smoking Status","metadata":{}},{"cell_type":"code","source":"analyze_feature(data_lc_train, 'Smoking Status')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"analyze_feature(data_lc_test, 'Smoking Status')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Create a copy of the original training DataFrame\ndata_sm_train = data_lc_train.copy()  # Makes a copy to preserve the original training data\n\n# Create a copy of the original test DataFrame\ndata_sm_test = data_lc_test.copy()  # Makes a copy to preserve the original test data","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"data_lc_test['Smoking Status'].unique()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# İki değerli değişkeni 0 ve 1 ile değiştirin\ndata_sm_train['Smoking Status'] = data_sm_train['Smoking Status'].replace({'Yes': 1, 'No': 0})\ndata_sm_test['Smoking Status'] = data_sm_test['Smoking Status'].replace({'Yes': 1, 'No': 0})","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 2.17. Insurance Duration","metadata":{}},{"cell_type":"code","source":"analyze_feature(data_sm_train, 'Insurance Duration')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"analyze_feature(data_sm_test, 'Insurance Duration')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"data_sm_train['Insurance Duration'].unique()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Create a copy of the original training DataFrame\ndata_ind_train = data_sm_train.copy()  # Makes a copy to preserve the original training data\n\n# Create a copy of the original test DataFrame\ndata_ind_test = data_sm_test.copy()  # Makes a copy to preserve the original test data","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Drop rows where the 'Marital Status' feature has NaN values\ndata_ind_train = data_ind_train.dropna(subset=['Insurance Duration'])","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Calculate the median of the 'Insurance Duration' column in the test data\nmedian_income_test = data_ind_test['Insurance Duration'].median()  # Computes the median to fill missing values in the test set\n\n# Fill NaN values in 'Vehicle Age' with the median in the test data\ndata_ind_test['Insurance Duration'] = data_ind_test['Insurance Duration'].fillna(median_income_test)\n# Fills missing values using the median calculated from the test set","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"analyze_feature(data_ind_test, 'Insurance Duration')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(data_ind_train['Insurance Duration'].value_counts(normalize=True))","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Rationale for Using a Standard Scaler for \"Insurance Duration\"\n\n#### Overview\n\nThe \"Insurance Duration\" feature displays a fairly uniform distribution across its categories, with each value representing a significant portion of the dataset. Given this balanced distribution, using a standard scaler for preprocessing is a suitable approach.\n\n#### Key Justifications\n\n- **Uniform Distribution**:\n  - The proportions of each insurance duration value are closely aligned, indicating that the feature is well-distributed across possible durations. \n  - This uniformity suggests that variance and mean are appropriate statistical measures for this data.\n\n- **Normalization Consistency**:\n  - Applying a standard scaler will center the data around a mean of 0 and scale it to unit variance, which is beneficial for many machine learning algorithms that assume normally distributed input features.\n  \n- **Enhanced Model Performance**:\n  - Normalizing numerical features to a standard scale often leads to improved convergence rates during training for algorithms like gradient descent, resulting in potentially more efficient and accurate models.\n\nBy implementing standard scaling, you ensure that \"Insurance Duration\" is aligned with other scaled features, maintaining consistency throughout the dataset and supporting robust model training.","metadata":{}},{"cell_type":"code","source":"# Initialize StandardScaler\nscaler = StandardScaler()\n\n# Fit and transform the 'Insurance Duration' in the training set\ndata_ind_train['Insurance Duration'] = scaler.fit_transform(data_ind_train[['Insurance Duration']])\n\n# Transform the 'Insurance Duration' in the test set\ndata_ind_test['Insurance Duration'] = scaler.transform(data_ind_test[['Insurance Duration']])","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 2.18. Policy Start Date","metadata":{}},{"cell_type":"code","source":"data_ind_train['Policy Start Date'].head()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Create a copy of the original training DataFrame\ndata_psd_train = data_ind_train.copy()  # Makes a copy to preserve the original training data\n\n# Create a copy of the original test DataFrame\ndata_psd_test = data_ind_test.copy()  # Makes a copy to preserve the original test data","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Convert the 'Policy Start Date' column in the training data to datetime format\ndata_psd_train['Policy Start Date'] = pd.to_datetime(data_psd_train['Policy Start Date'])  \n# This conversion allows for easier date manipulation and extraction of date components\n\n# Convert the 'Policy Start Date' column in the test data to datetime format\ndata_psd_test['Policy Start Date'] = pd.to_datetime(data_psd_test['Policy Start Date'])  \n# Standardizing the date format ensures consistency and enables temporal analysis in test data","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"analyze_feature(data_psd_train, 'Policy Start Date')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"analyze_feature(data_psd_test, 'Policy Start Date')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"data_psd_train['Year'] = data_psd_train['Policy Start Date'].dt.year\ndata_psd_train['Month'] = data_psd_train['Policy Start Date'].dt.month\ndata_psd_train['Day'] = data_psd_train['Policy Start Date'].dt.day\ndata_psd_train['DayOfWeek'] = data_psd_train['Policy Start Date'].dt.dayofweek","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"data_psd_test['Year'] = data_psd_test['Policy Start Date'].dt.year\ndata_psd_test['Month'] = data_psd_test['Policy Start Date'].dt.month\ndata_psd_test['Day'] = data_psd_test['Policy Start Date'].dt.day\ndata_psd_test['DayOfWeek'] = data_psd_test['Policy Start Date'].dt.dayofweek","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Month için dönüşüm\ndata_psd_train['Month_sin'] = np.sin(2 * np.pi * data_psd_train['Month'] / 12)\ndata_psd_train['Month_cos'] = np.cos(2 * np.pi * data_psd_train['Month'] / 12)\n\n# Day için dönüşüm (ayın günleri, 1-31 arası döngüsel)\ndata_psd_train['Day_sin'] = np.sin(2 * np.pi * data_psd_train['Day'] / 31)\ndata_psd_train['Day_cos'] = np.cos(2 * np.pi * data_psd_train['Day'] / 31)\n\n# DayOfWeek için dönüşüm (haftanın günleri, 0-6 arası döngüsel)\ndata_psd_train['DayOfWeek_sin'] = np.sin(2 * np.pi * data_psd_train['DayOfWeek'] / 7)\ndata_psd_train['DayOfWeek_cos'] = np.cos(2 * np.pi * data_psd_train['DayOfWeek'] / 7)\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Month için dönüşüm\ndata_psd_test['Month_sin'] = np.sin(2 * np.pi * data_psd_test['Month'] / 12)\ndata_psd_test['Month_cos'] = np.cos(2 * np.pi * data_psd_test['Month'] / 12)\n\n# Day için dönüşüm (ayın günleri, 1-31 arası döngüsel)\ndata_psd_test['Day_sin'] = np.sin(2 * np.pi * data_psd_test['Day'] / 31)\ndata_psd_test['Day_cos'] = np.cos(2 * np.pi * data_psd_test['Day'] / 31)\n\n# DayOfWeek için dönüşüm (haftanın günleri, 0-6 arası döngüsel)\ndata_psd_test['DayOfWeek_sin'] = np.sin(2 * np.pi * data_psd_test['DayOfWeek'] / 7)\ndata_psd_test['DayOfWeek_cos'] = np.cos(2 * np.pi * data_psd_test['DayOfWeek'] / 7)\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"data_psd_train.info()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"data_psd_train = data_psd_train.drop(['Policy Start Date', 'Year', 'Month', 'Day', 'DayOfWeek'], axis=1)\ndata_psd_test = data_psd_test.drop(['Policy Start Date', 'Year', 'Month', 'Day', 'DayOfWeek'], axis=1)\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"data_psd_train.info()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"#### Feature Engineering: Encoding Temporal Data\n\nTemporal data often contains cyclical patterns that are crucial for machine learning models to understand. This report outlines the steps taken to preprocess and encode the `Policy Start Date` column in the dataset, ensuring the features accurately capture the cyclical nature of time-based components.\n\n---\n\n##### 1. Converting `Policy Start Date` to Datetime Format\nThe `Policy Start Date` column was initially stored as a raw text or non-datetime format. It was converted to a standardized datetime type to facilitate easier extraction of date components and ensure consistency across the dataset.\n\n---\n\n##### 2. Extracting Date Components\nFrom the standardized `Policy Start Date`, we extracted key temporal features:\n- **Year**: Represents the calendar year.\n- **Month**: Represents the calendar month (1–12).\n- **Day**: Represents the day of the month (1–31).\n- **Day of the Week**: Represents the day of the week (0–6, where 0 is Monday).\n\nThese components provide granular details about the temporal data.\n\n---\n\n##### 3. Encoding Temporal Features with Cyclical Transformations\nTo account for the cyclical nature of time (e.g., December is closer to January than to June), sine and cosine transformations were applied to the extracted features:\n- **Month**: Encoded to reflect the annual cycle.\n- **Day**: Encoded to reflect the monthly cycle.\n- **Day of the Week**: Encoded to reflect the weekly cycle.\n\nThis transformation ensures that the temporal relationships are preserved and correctly interpreted by machine learning models, especially those that rely on Euclidean distance or linear relationships.\n\n---\n\n##### 4. Removing Redundant Features\nAfter deriving the encoded features, the original columns (`Policy Start Date`, `Year`, `Month`, `Day`, and `Day of the Week`) were removed from the dataset. This step reduces redundancy and potential multicollinearity, leaving only the meaningful cyclical representations.\n\n---\n\n##### Benefits of This Approach\n- **Improved Feature Representation**: Captures the periodic nature of temporal data effectively.\n- **Reduced Redundancy**: Eliminates unnecessary columns while preserving essential information.\n- **Enhanced Model Interpretability**: Allows machine learning algorithms to better understand and leverage temporal patterns in the data.\n\nThis preprocessing step ensures the dataset is clean, efficient, and ready for modeling with temporal awareness.\n","metadata":{}},{"cell_type":"markdown","source":"## 2.19. Customer Feedback","metadata":{}},{"cell_type":"code","source":"analyze_feature(data_psd_train, 'Customer Feedback')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"analyze_feature(data_psd_test, 'Customer Feedback')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"data_psd_train['Customer Feedback'].unique()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(data_psd_train['Customer Feedback'].value_counts(normalize=True))","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Create a copy of the original training DataFrame\ndata_cf_train = data_psd_train.copy()  # Makes a copy to preserve the original training data\n\n# Create a copy of the original test DataFrame\ndata_cf_test = data_psd_test.copy()  # Makes a copy to preserve the original test data","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Replace NaN values with 'Unknown'\ndata_cf_train['Customer Feedback'].fillna('Unknown', inplace=True)\ndata_cf_test['Customer Feedback'].fillna('Unknown', inplace=True)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Apply one-hot encoding\ndata_cf_train = pd.get_dummies(data_cf_train, columns=['Customer Feedback'])\ndata_cf_test = pd.get_dummies(data_cf_test, columns=['Customer Feedback'])","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 2.20. Exercise Frequency","metadata":{}},{"cell_type":"code","source":"analyze_feature(data_cf_train, 'Exercise Frequency')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"analyze_feature(data_cf_test, 'Exercise Frequency')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"data_cf_train['Exercise Frequency'].unique()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Create a copy of the original training DataFrame\ndata_exf_train = data_cf_train.copy()  # Makes a copy to preserve the original training data\n\n# Create a copy of the original test DataFrame\ndata_exf_test = data_cf_test.copy()  # Makes a copy to preserve the original test data","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Apply one-hot encoding\ndata_exf_train = pd.get_dummies(data_exf_train, columns=['Exercise Frequency'])\ndata_exf_test = pd.get_dummies(data_exf_test, columns=['Exercise Frequency'])","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 2.21. Previous Claims","metadata":{}},{"cell_type":"markdown","source":"### Variation 1 / Random Choice - Scalar","metadata":{}},{"cell_type":"code","source":"analyze_feature(data_cf_train, 'Previous Claims')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"analyze_feature(data_cf_test, 'Previous Claims')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(data_cf_train['Previous Claims'].value_counts(normalize=True))","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Create a copy of the original training DataFrame\ndata_pcl_train = data_exf_train.copy()  # Makes a copy to preserve the original training data\n\n# Create a copy of the original test DataFrame\ndata_pcl_test = data_exf_test.copy()  # Makes a copy to preserve the original test data","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Previous Claims sütunundaki eksik değerleri doldurmak için oranları belirtiyoruz\nproportions = {\n    0.0: 0.365706,\n    1.0: 0.360090,\n    2.0: 0.200233,\n    3.0: 0.058445,\n    4.0: 0.012677,\n    5.0: 0.002410,\n    6.0: 0.000358,\n    7.0: 0.000070,\n    8.0: 0.000010,\n    9.0: 0.000001\n}\n\n# Oranları normalize ediyoruz (toplamları 1 olmalı, fakat zaten uygun görünüyor)\ncategories = list(proportions.keys())\nprobabilities = list(proportions.values())\n\n# Eksik değerleri rastgele dolduruyoruz\ndata_pcl_train['Previous Claims'] = data_pcl_train['Previous Claims'].apply(\n    lambda x: np.random.choice(categories, p=probabilities) if pd.isna(x) else x\n)\n\n# Eksik değerleri rastgele dolduruyoruz\ndata_pcl_test['Previous Claims'] = data_pcl_test['Previous Claims'].apply(\n    lambda x: np.random.choice(categories, p=probabilities) if pd.isna(x) else x\n)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"result = check_normality(data_pcl_train, 'Previous Claims')\nprint(result)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(data_pcl_train['Previous Claims'].value_counts(normalize=True))","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Initialize RobustScaler\nscaler = RobustScaler()\n\n# Apply RobustScaler to the 'Annual Income' column in the training set\ndata_pcl_train['Previous Claims'] = scaler.fit_transform(data_pcl_train[['Previous Claims']])\n\n# Apply the same scaler to the test set to ensure consistency\ndata_pcl_test['Previous Claims'] = scaler.transform(data_pcl_test[['Previous Claims']])","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Variation 2 / Unknown - One Hot Encoding","metadata":{}},{"cell_type":"code","source":"# Create a copy of the original training DataFrame for variation \ndata_pcl1_train = data_exf_train.copy()  # Makes a copy to preserve the original training data\n\n# Create a copy of the original test DataFrame for variation\ndata_pcl1_test = data_exf_test.copy()  # Makes a copy to preserve the original test data","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Replace NaN values with 'Unknown'\ndata_pcl1_train['Previous Claims'].fillna('Unknown', inplace=True)\ndata_pcl1_test['Previous Claims'].fillna('Unknown', inplace=True)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Apply one-hot encoding\ndata_pcl1_train = pd.get_dummies(data_pcl1_train, columns=['Previous Claims'])\ndata_pcl1_test = pd.get_dummies(data_pcl1_test, columns=['Previous Claims'])","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 3. Creating CSV Files","metadata":{}},{"cell_type":"code","source":"# Writing the DataFrame to CSV files\ndata_pcl1_train.to_csv(\"cleaned/ddata_train_1.csv\", index=False)  # Save the training dataset 1 to a CSV file\ndata_pcl_train.to_csv(\"cleaned/ddata_train_2.csv\", index=False)  # Save the training dataset 2 to a CSV file\n\n# Adding 'id' column from the original train_data to the test sets\ndata_pcl1_test['id'] = train_data['id']  # Add 'id' column to test set 1 from the training data\ndata_pcl_test['id'] = train_data['id']  # Add 'id' column to test set 2 from the training data\n\n# If you want to save the test set as well, you can do so by:\ndata_pcl1_test.to_csv(\"cleaned/ddata_test_1.csv\", index=False)  # Save the test dataset 1 to a CSV file\ndata_pcl_test.to_csv(\"cleaned/ddata_test_2.csv\", index=False)  # Save the test dataset 2 to a CSV file","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}