{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":84896,"databundleVersionId":10305135,"sourceType":"competition"}],"dockerImageVersionId":30804,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<center><h1 style=\"color:white; background-color:teal; padding:20px; font-family:Courier New, monospace; border-radius:8px\">Insurance Dataset</h1></center>\n","metadata":{}},{"cell_type":"code","source":"from IPython.display import HTML,display\nHTML('<center><img src=\"https://static.investindia.gov.in/s3fs-public/2019-05/Insurance1.jpg\" alt=\"Premium insurence\" width=\"1300\"></center>')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T13:35:04.359963Z","iopub.execute_input":"2024-12-13T13:35:04.360386Z","iopub.status.idle":"2024-12-13T13:35:04.390143Z","shell.execute_reply.started":"2024-12-13T13:35:04.360329Z","shell.execute_reply":"2024-12-13T13:35:04.388932Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<center><h1 style=\"color:white; background-color:teal; padding:20px; font-family:Courier New, monospace; border-radius:8px\">Importing Packages</h1></center>\n","metadata":{"execution":{"iopub.status.busy":"2024-12-11T11:56:16.014814Z","iopub.execute_input":"2024-12-11T11:56:16.015203Z","iopub.status.idle":"2024-12-11T11:56:16.021495Z","shell.execute_reply.started":"2024-12-11T11:56:16.015168Z","shell.execute_reply":"2024-12-11T11:56:16.020088Z"}}},{"cell_type":"code","source":"import pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport plotly.express as px\nimport warnings\nwarnings.filterwarnings(\"ignore\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T13:35:04.392199Z","iopub.execute_input":"2024-12-13T13:35:04.392649Z","iopub.status.idle":"2024-12-13T13:35:06.338084Z","shell.execute_reply.started":"2024-12-13T13:35:04.392599Z","shell.execute_reply":"2024-12-13T13:35:06.337091Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df_test = pd.read_csv('/kaggle/input/playground-series-s4e12/test.csv')\ndf_train = pd.read_csv('/kaggle/input/playground-series-s4e12/train.csv', index_col=\"id\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T13:35:06.340153Z","iopub.execute_input":"2024-12-13T13:35:06.34068Z","iopub.status.idle":"2024-12-13T13:35:17.031313Z","shell.execute_reply.started":"2024-12-13T13:35:06.34064Z","shell.execute_reply":"2024-12-13T13:35:17.02961Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<center><h1 style=\"color:white; background-color:teal; padding:20px; font-family:Courier New, monospace; border-radius:8px\">Exploratory Data Analysis(EDA)</h1></center>\n","metadata":{}},{"cell_type":"markdown","source":"###  Feature Understanding","metadata":{}},{"cell_type":"markdown","source":"1. **Age:** Age of the insured individual (Numerical)\n2. **Gender:** Gender of the insured individual (Categorical: Male, Female)\n3. **Annual Income:** Annual income of the insured individual (Numerical, skewed)\n4. **Marital Status:** Marital status of the insured individual (Categorical: Single, Married, Divorced)\n5. **Number of Dependents:** Number of dependents (Numerical, with missing values)\n6. **Education Level:** Highest education level attained (Categorical: High School, Bachelor's, Master's, PhD)\n7. **Occupation:** Occupation of the insured individual (Categorical: Employed, Self-Employed, Unemployed)\n8. **Health Score:** A score representing the health status (Numerical, skewed)\n9. **Location:** Type of location (Categorical: Urban, Suburban, Rural)\n10. **Policy Type:** Type of insurance policy (Categorical: Basic, Comprehensive, Premium)\n11. **Previous Claims:** Number of previous claims made (Numerical, with outliers)\n12. **Vehicle Age:** Age of the vehicle insured (Numerical)\n13. **Credit Score:** Credit score of the insured individual (Numerical, with missing values)\n14. **Insurance Duration:** Duration of the insurance policy (Numerical, in years)\n15. **Premium Amount:** Target variable representing the insurance premium amount (Numerical, skewed)\n16. **Policy Start Date:** Start date of the insurance policy (Text, improperly formatted)\n17. **Customer Feedback:** Short feedback comments from customers (Text)\n18. **Smoking Status:** Smoking status of the insured individual (Categorical: Yes, No)\n19. **Exercise Frequency:** Frequency of exercise (Categorical: Daily, Weekly, Monthly, Rarely)\n20. **Property Type:** Type of property owned (Categorical: House, Apartment, Condo)","metadata":{"execution":{"iopub.status.busy":"2024-12-11T11:58:15.580504Z","iopub.execute_input":"2024-12-11T11:58:15.5812Z","iopub.status.idle":"2024-12-11T11:58:15.590736Z","shell.execute_reply.started":"2024-12-11T11:58:15.581151Z","shell.execute_reply":"2024-12-11T11:58:15.589365Z"}}},{"cell_type":"code","source":"df_train.info()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T13:35:17.032728Z","iopub.execute_input":"2024-12-13T13:35:17.033245Z","iopub.status.idle":"2024-12-13T13:35:17.732247Z","shell.execute_reply.started":"2024-12-13T13:35:17.033185Z","shell.execute_reply":"2024-12-13T13:35:17.730876Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df_train.describe().T","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T13:35:17.734413Z","iopub.execute_input":"2024-12-13T13:35:17.734783Z","iopub.status.idle":"2024-12-13T13:35:18.434512Z","shell.execute_reply.started":"2024-12-13T13:35:17.734745Z","shell.execute_reply":"2024-12-13T13:35:18.433187Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df_train.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T13:35:18.435712Z","iopub.execute_input":"2024-12-13T13:35:18.436068Z","iopub.status.idle":"2024-12-13T13:35:18.460878Z","shell.execute_reply.started":"2024-12-13T13:35:18.436033Z","shell.execute_reply":"2024-12-13T13:35:18.459358Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df_train.isnull().sum()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T13:35:18.46257Z","iopub.execute_input":"2024-12-13T13:35:18.463062Z","iopub.status.idle":"2024-12-13T13:35:19.109761Z","shell.execute_reply.started":"2024-12-13T13:35:18.463009Z","shell.execute_reply":"2024-12-13T13:35:19.108467Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"for col in df_train:\n    print(f'{col}:{df_train[col].nunique()}')\n    \nprint('------------------ unique names ----------------------------')\n\nfor col in df_train:\n    print(f'{col}:{df_train[col].unique()}')\n    print('--------------- next feature ------------------')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T13:35:19.111172Z","iopub.execute_input":"2024-12-13T13:35:19.111529Z","iopub.status.idle":"2024-12-13T13:35:21.470925Z","shell.execute_reply.started":"2024-12-13T13:35:19.111492Z","shell.execute_reply":"2024-12-13T13:35:21.469371Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Display original column names\nprint(\"Original Columns:\", df_train.columns)\n\n# Create a new list of column names using a loop\nnew_columns = []\nfor col in df_train.columns:\n    # Modify each column name\n    new_col = col.replace(' ', '_')\n    new_columns.append(new_col)\n\n# Assign the new column names to the DataFrame\ndf_train.columns = new_columns\nprint(f'Rename Columns:{df_train.columns}')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T13:35:21.472351Z","iopub.execute_input":"2024-12-13T13:35:21.472668Z","iopub.status.idle":"2024-12-13T13:35:21.482893Z","shell.execute_reply.started":"2024-12-13T13:35:21.472637Z","shell.execute_reply":"2024-12-13T13:35:21.480895Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"categorical_columns = []\nnumerical_columns = []\n\n# Iterate through columns and their data types\nfor column, dtype in df_train.dtypes.items():\n    if dtype == \"object\":\n        categorical_columns.append(column) \n    elif dtype == \"float64\":\n        numerical_columns.append(column)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T13:35:21.485616Z","iopub.execute_input":"2024-12-13T13:35:21.486056Z","iopub.status.idle":"2024-12-13T13:35:21.523879Z","shell.execute_reply.started":"2024-12-13T13:35:21.486008Z","shell.execute_reply":"2024-12-13T13:35:21.522338Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<div style=\"background-color: #ee82ee; padding: 10px;\">\n    <h3> &#128521;Note:- </h3>\n    \n<h4> &#128313; I tried to explore the data without filling NaN values to gain insights and understand the differences in data distribution with and without filling NaN values.</h4>\n\n<h4>&#128313;You can check the insights gained below after filling NaN values.</h4>\n</div>","metadata":{}},{"cell_type":"markdown","source":"<center><h1 style=\"color:white; background-color:teal; padding:20px; font-family:Courier New, monospace; border-radius:8px\">Univariate  Analysis</h1></center>\n","metadata":{}},{"cell_type":"markdown","source":"for plot in numerical_columns:\n    plt.figure(figsize=(17,8))\n    p = plt.hist(df_train[plot], bins=30, edgecolor='black')\n    plt.title(f\"Histogram of {plot}\")\n    plt.xlabel(plot)\n    plt.ylabel(\"Frequency\")\n    \n    # Loop through the patches to add labels\n    for rect in p[2]:\n        height = rect.get_height()\n        plt.text(rect.get_x() + rect.get_width() / 2, height, str(int(height)),\n                 ha='center', va='bottom', fontsize=7)  # Position the label at the top of each bar\n\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2024-12-12T12:08:56.420969Z","iopub.execute_input":"2024-12-12T12:08:56.421364Z","iopub.status.idle":"2024-12-12T12:09:00.102126Z","shell.execute_reply.started":"2024-12-12T12:08:56.421331Z","shell.execute_reply":"2024-12-12T12:09:00.100993Z"},"_kg_hide-input":true}},{"cell_type":"markdown","source":"<div style=\"background-color: #ee82ee; padding: 10px;\">\n    <h3> Numerical Features</h3>\n</div>","metadata":{}},{"cell_type":"code","source":"def univariate_numerical_plot(column_name):\n    plt.figure(figsize=(30, 12))\n    # Create a bar plot\n    plot = sns.countplot(x=column_name, data=df_train)\n\n    plt.xlabel(column_name)\n    plt.title(f\"Univariate Analysis of {column_name}\")\n\n    # Loop through the patches to add labels\n    for rect in plot.patches:  # Use `plot.patches` to access bars\n        height = rect.get_height()\n        plt.text(rect.get_x() + rect.get_width() / 2, height, str(int(height)),\n                 ha='center', va='bottom', fontsize=10)  # Position the label at the top of each bar\n\n    plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T13:35:21.528307Z","iopub.execute_input":"2024-12-13T13:35:21.5288Z","iopub.status.idle":"2024-12-13T13:35:21.53988Z","shell.execute_reply.started":"2024-12-13T13:35:21.528761Z","shell.execute_reply":"2024-12-13T13:35:21.537434Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"univariate_numerical_plot('Age')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T13:35:21.54162Z","iopub.execute_input":"2024-12-13T13:35:21.542056Z","iopub.status.idle":"2024-12-13T13:35:22.466541Z","shell.execute_reply.started":"2024-12-13T13:35:21.542015Z","shell.execute_reply":"2024-12-13T13:35:22.465407Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<div style=\"background-color: #69cdcb; padding: 10px;\">\n    <h3> &#128521;Insights gained </h3>\n    \n<h4> &#128313; In this dataset, the total number of people aged in the range 18-30 is below ~25k. </h4>\r\n<h4> &#128313; The number of people aged 31-60 is in the range of ~25k to ~26k. </h4>\r\n\n\n</div>","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(17,8))\np = plt.hist(df_train['Annual_Income'], bins=30, edgecolor='black')\n# p = sns.countplot(x = df_train['Annual_Income'])\nplt.title(f\"Histogram of {'Annual_Income'}\")\nplt.xlabel('Annual_Income')\nplt.ylabel(\"Frequency\")\n\n# Loop through the patches to add labels\nfor rect in p[2]:\n    height = rect.get_height()\n    plt.text(rect.get_x() + rect.get_width() / 2, height, str(int(height)),\n             ha='center', va='bottom', fontsize=7)  # Position the label at the top of each bar\n\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T13:35:22.467763Z","iopub.execute_input":"2024-12-13T13:35:22.468132Z","iopub.status.idle":"2024-12-13T13:35:22.899485Z","shell.execute_reply.started":"2024-12-13T13:35:22.468096Z","shell.execute_reply":"2024-12-13T13:35:22.89823Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"univariate_numerical_plot('Number_of_Dependents')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T13:35:22.901189Z","iopub.execute_input":"2024-12-13T13:35:22.901618Z","iopub.status.idle":"2024-12-13T13:35:23.391725Z","shell.execute_reply.started":"2024-12-13T13:35:22.901579Z","shell.execute_reply":"2024-12-13T13:35:23.389949Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<div style=\"background-color: #69cdcb; padding: 10px;\">\n    <h3> &#128521;Insights gained </h3>\n    \n<h4> &#128313; In this plot Higest of the dependents is 3 and 4 . </h4>\n<h4> &#128313; Most of the all categories are ~ to comapre with other categories. </h4>\n\n</div>","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(17,8))\np = plt.hist(df_train['Health_Score'], bins=30, edgecolor='black')\nplt.title(f\"Histogram of {'Health_Score'}\")\nplt.xlabel('Health_Score')\nplt.ylabel(\"Frequency\")\n\n# Loop through the patches to add labels\nfor rect in p[2]:\n    height = rect.get_height()\n    plt.text(rect.get_x() + rect.get_width() / 2, height, str(int(height)),\n             ha='center', va='bottom', fontsize=7)  # Position the label at the top of each bar\n\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T13:35:23.393638Z","iopub.execute_input":"2024-12-13T13:35:23.394401Z","iopub.status.idle":"2024-12-13T13:35:23.935538Z","shell.execute_reply.started":"2024-12-13T13:35:23.394359Z","shell.execute_reply":"2024-12-13T13:35:23.934399Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<div style=\"background-color: #69cdcb; padding: 10px;\">\n    <h3> &#128521;Insights gained </h3>\n    \n<h4> &#128313; Majority of the people health score is 10 to 35. </h4>\n<h4>&#128313; In this plot outliers health score is 58. </h4>\n\n</div>","metadata":{}},{"cell_type":"code","source":"univariate_numerical_plot('Previous_Claims')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T13:35:23.936598Z","iopub.execute_input":"2024-12-13T13:35:23.93693Z","iopub.status.idle":"2024-12-13T13:35:24.406601Z","shell.execute_reply.started":"2024-12-13T13:35:23.936895Z","shell.execute_reply":"2024-12-13T13:35:24.405426Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<div style=\"background-color: #69cdcb; padding: 10px;\">\n    <h3> &#128521;Insights gained </h3>\n    \n<h4>&#128313; In this dataset, the majority of people don't have previous claims.</h4>\r\n<h4>&#128313; Most people with previous claims have 1-2 claims.</h4>\r\n<h4>&#128313; Fewer people have 3-4 previous claims.</h4\r\n\n\n</div>","metadata":{}},{"cell_type":"code","source":"univariate_numerical_plot('Vehicle_Age')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T13:35:24.408096Z","iopub.execute_input":"2024-12-13T13:35:24.408444Z","iopub.status.idle":"2024-12-13T13:35:25.021803Z","shell.execute_reply.started":"2024-12-13T13:35:24.408408Z","shell.execute_reply":"2024-12-13T13:35:25.020484Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plt.figure(figsize=(17,8))\np = plt.hist(df_train['Credit_Score'], bins=30, edgecolor='black')\n# p = sns.countplot(x = df_train['Annual_Income'])\nplt.title(f\"Histogram of {'Credit_Score'}\")\nplt.xlabel('Credit_Score')\nplt.ylabel(\"Frequency\")\n\n# Loop through the patches to add labels\nfor rect in p[2]:\n    height = rect.get_height()\n    plt.text(rect.get_x() + rect.get_width() / 2, height, str(int(height)),\n             ha='center', va='bottom', fontsize=7)  # Position the label at the top of each bar\n\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T13:35:25.02362Z","iopub.execute_input":"2024-12-13T13:35:25.024068Z","iopub.status.idle":"2024-12-13T13:35:25.395268Z","shell.execute_reply.started":"2024-12-13T13:35:25.024028Z","shell.execute_reply":"2024-12-13T13:35:25.394052Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"univariate_numerical_plot('Insurance_Duration')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T13:35:25.396597Z","iopub.execute_input":"2024-12-13T13:35:25.397108Z","iopub.status.idle":"2024-12-13T13:35:25.84168Z","shell.execute_reply.started":"2024-12-13T13:35:25.397055Z","shell.execute_reply":"2024-12-13T13:35:25.840281Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<div style=\"background-color: #ee82ee; padding: 10px;\">\n    <h3> Target Column</h3>\n</div>","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(17,8))\np = plt.hist(df_train['Premium_Amount'], bins=30, edgecolor='black')\n# p = sns.countplot(x = df_train['Annual_Income'])\nplt.title(f\"Histogram of {'Premium_Amount'}\")\nplt.xlabel('Premium_Amount')\nplt.ylabel(\"Frequency\")\n\n# Loop through the patches to add labels\nfor rect in p[2]:\n    height = rect.get_height()\n    plt.text(rect.get_x() + rect.get_width() / 2, height, str(int(height)),\n             ha='center', va='bottom', fontsize=7)  # Position the label at the top of each bar\n\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T13:35:25.843113Z","iopub.execute_input":"2024-12-13T13:35:25.843478Z","iopub.status.idle":"2024-12-13T13:35:26.329116Z","shell.execute_reply.started":"2024-12-13T13:35:25.843442Z","shell.execute_reply":"2024-12-13T13:35:26.327885Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<div style=\"background-color: #ee82ee; padding: 10px;\">\n    <h3> Categorical Features</h3>\n</div>","metadata":{}},{"cell_type":"code","source":"def univariate_categorical_plot(column_name):\n    plt.figure(figsize=(12, 7))\n    # Create a count plot\n    plot = sns.countplot(x=column_name, data=df_train)\n\n    plt.xlabel(column_name)\n    plt.title(f\"Univariate Analysis of {column_name}\")\n\n    # Loop through the patches to add labels\n    for rect in plot.patches:  # Use `plot.patches` to access bars\n        height = rect.get_height()\n        plt.text(rect.get_x() + rect.get_width() / 2, height, str(int(height)),\n                 ha='center', va='bottom', fontsize=15)  # Position the label at the top of each bar\n\n    plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T13:35:26.330452Z","iopub.execute_input":"2024-12-13T13:35:26.330762Z","iopub.status.idle":"2024-12-13T13:35:26.339633Z","shell.execute_reply.started":"2024-12-13T13:35:26.33073Z","shell.execute_reply":"2024-12-13T13:35:26.338421Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"univariate_categorical_plot('Gender')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T13:35:26.341365Z","iopub.execute_input":"2024-12-13T13:35:26.342237Z","iopub.status.idle":"2024-12-13T13:35:27.243606Z","shell.execute_reply.started":"2024-12-13T13:35:26.342179Z","shell.execute_reply":"2024-12-13T13:35:27.242397Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"univariate_categorical_plot('Marital_Status')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T13:35:27.245039Z","iopub.execute_input":"2024-12-13T13:35:27.245425Z","iopub.status.idle":"2024-12-13T13:35:28.257773Z","shell.execute_reply.started":"2024-12-13T13:35:27.245385Z","shell.execute_reply":"2024-12-13T13:35:28.256185Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"univariate_categorical_plot('Education_Level')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T13:35:28.259485Z","iopub.execute_input":"2024-12-13T13:35:28.259828Z","iopub.status.idle":"2024-12-13T13:35:29.139388Z","shell.execute_reply.started":"2024-12-13T13:35:28.259791Z","shell.execute_reply":"2024-12-13T13:35:29.138182Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"univariate_categorical_plot('Occupation')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T13:35:29.140578Z","iopub.execute_input":"2024-12-13T13:35:29.140937Z","iopub.status.idle":"2024-12-13T13:35:30.026125Z","shell.execute_reply.started":"2024-12-13T13:35:29.140896Z","shell.execute_reply":"2024-12-13T13:35:30.024932Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"univariate_categorical_plot('Location')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T13:35:30.027614Z","iopub.execute_input":"2024-12-13T13:35:30.028088Z","iopub.status.idle":"2024-12-13T13:35:30.945367Z","shell.execute_reply.started":"2024-12-13T13:35:30.028033Z","shell.execute_reply":"2024-12-13T13:35:30.944135Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"univariate_categorical_plot('Policy_Type')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T13:35:30.946751Z","iopub.execute_input":"2024-12-13T13:35:30.947154Z","iopub.status.idle":"2024-12-13T13:35:31.889681Z","shell.execute_reply.started":"2024-12-13T13:35:30.947107Z","shell.execute_reply":"2024-12-13T13:35:31.887744Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"univariate_categorical_plot('Customer_Feedback')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T13:35:31.891739Z","iopub.execute_input":"2024-12-13T13:35:31.892217Z","iopub.status.idle":"2024-12-13T13:35:32.819335Z","shell.execute_reply.started":"2024-12-13T13:35:31.892176Z","shell.execute_reply":"2024-12-13T13:35:32.818096Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"univariate_categorical_plot('Smoking_Status')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T13:35:32.824585Z","iopub.execute_input":"2024-12-13T13:35:32.824958Z","iopub.status.idle":"2024-12-13T13:35:33.66008Z","shell.execute_reply.started":"2024-12-13T13:35:32.824922Z","shell.execute_reply":"2024-12-13T13:35:33.658615Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"univariate_categorical_plot('Exercise_Frequency')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T13:35:33.661883Z","iopub.execute_input":"2024-12-13T13:35:33.662409Z","iopub.status.idle":"2024-12-13T13:35:34.574763Z","shell.execute_reply.started":"2024-12-13T13:35:33.662347Z","shell.execute_reply":"2024-12-13T13:35:34.573626Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"univariate_categorical_plot('Property_Type')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T13:35:34.576172Z","iopub.execute_input":"2024-12-13T13:35:34.576517Z","iopub.status.idle":"2024-12-13T13:35:35.540984Z","shell.execute_reply.started":"2024-12-13T13:35:34.576474Z","shell.execute_reply":"2024-12-13T13:35:35.539798Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<center><h1 style=\"color:white; background-color:teal; padding:20px; font-family:Courier New, monospace; border-radius:8px\">BiVariate Analysis</h1></center>\n","metadata":{}},{"cell_type":"code","source":"def categorical_plots(X):\n    plt.figure(figsize=(14, 7))\n    plot = sns.violinplot(x=df_train[X], y=df_train['Premium_Amount'])\n    plt.title(f'Scatter Plot of Premium Amount by {X}')\n    plt.show()\n    return","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T13:35:35.542563Z","iopub.execute_input":"2024-12-13T13:35:35.543038Z","iopub.status.idle":"2024-12-13T13:35:35.549709Z","shell.execute_reply.started":"2024-12-13T13:35:35.542987Z","shell.execute_reply":"2024-12-13T13:35:35.548533Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"categorical_plots('Number_of_Dependents')\nprint(df_train['Number_of_Dependents'].value_counts())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T13:35:35.551195Z","iopub.execute_input":"2024-12-13T13:35:35.551572Z","iopub.status.idle":"2024-12-13T13:35:38.807441Z","shell.execute_reply.started":"2024-12-13T13:35:35.551535Z","shell.execute_reply":"2024-12-13T13:35:38.80634Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"categorical_plots('Previous_Claims')\nprint(df_train['Previous_Claims'].value_counts())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T13:35:38.809025Z","iopub.execute_input":"2024-12-13T13:35:38.809368Z","iopub.status.idle":"2024-12-13T13:35:41.383333Z","shell.execute_reply.started":"2024-12-13T13:35:38.80933Z","shell.execute_reply":"2024-12-13T13:35:41.38229Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df_train[df_train['Previous_Claims'] == 8]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T13:35:41.384781Z","iopub.execute_input":"2024-12-13T13:35:41.385113Z","iopub.status.idle":"2024-12-13T13:35:41.415335Z","shell.execute_reply.started":"2024-12-13T13:35:41.38508Z","shell.execute_reply":"2024-12-13T13:35:41.414204Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"categorical_plots('Insurance_Duration')\nprint(df_train['Insurance_Duration'].value_counts())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T13:35:41.416449Z","iopub.execute_input":"2024-12-13T13:35:41.416734Z","iopub.status.idle":"2024-12-13T13:35:44.835704Z","shell.execute_reply.started":"2024-12-13T13:35:41.416704Z","shell.execute_reply":"2024-12-13T13:35:44.834256Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"categorical_plots('Gender')\ndf_train['Gender'].value_counts()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T13:35:44.838066Z","iopub.execute_input":"2024-12-13T13:35:44.838493Z","iopub.status.idle":"2024-12-13T13:35:48.227783Z","shell.execute_reply.started":"2024-12-13T13:35:44.838455Z","shell.execute_reply":"2024-12-13T13:35:48.226638Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"categorical_plots('Marital_Status')\nprint(df_train['Marital_Status'].value_counts())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T13:35:48.229132Z","iopub.execute_input":"2024-12-13T13:35:48.229448Z","iopub.status.idle":"2024-12-13T13:35:51.679669Z","shell.execute_reply.started":"2024-12-13T13:35:48.229416Z","shell.execute_reply":"2024-12-13T13:35:51.678577Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"categorical_plots('Education_Level')\nprint(df_train['Education_Level'].value_counts())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T13:35:51.681032Z","iopub.execute_input":"2024-12-13T13:35:51.681467Z","iopub.status.idle":"2024-12-13T13:35:55.257305Z","shell.execute_reply.started":"2024-12-13T13:35:51.681415Z","shell.execute_reply":"2024-12-13T13:35:55.255335Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"categorical_plots('Occupation')\nprint(df_train['Occupation'].value_counts())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T13:35:55.258807Z","iopub.execute_input":"2024-12-13T13:35:55.259175Z","iopub.status.idle":"2024-12-13T13:35:58.054767Z","shell.execute_reply.started":"2024-12-13T13:35:55.259136Z","shell.execute_reply":"2024-12-13T13:35:58.053419Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"categorical_plots('Location')\nprint(df_train['Location'].value_counts())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T13:35:58.056386Z","iopub.execute_input":"2024-12-13T13:35:58.056865Z","iopub.status.idle":"2024-12-13T13:36:01.546924Z","shell.execute_reply.started":"2024-12-13T13:35:58.056794Z","shell.execute_reply":"2024-12-13T13:36:01.545791Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"categorical_plots('Policy_Type')\nprint(df_train['Policy_Type'].value_counts())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T13:36:01.548201Z","iopub.execute_input":"2024-12-13T13:36:01.548547Z","iopub.status.idle":"2024-12-13T13:36:05.064623Z","shell.execute_reply.started":"2024-12-13T13:36:01.548511Z","shell.execute_reply":"2024-12-13T13:36:05.063527Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"categorical_plots('Customer_Feedback')\nprint(df_train['Customer_Feedback'].value_counts())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T13:36:05.066413Z","iopub.execute_input":"2024-12-13T13:36:05.066896Z","iopub.status.idle":"2024-12-13T13:36:08.545809Z","shell.execute_reply.started":"2024-12-13T13:36:05.066823Z","shell.execute_reply":"2024-12-13T13:36:08.544575Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"categorical_plots('Smoking_Status')\nprint(df_train['Smoking_Status'].value_counts())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T13:36:08.547226Z","iopub.execute_input":"2024-12-13T13:36:08.547568Z","iopub.status.idle":"2024-12-13T13:36:11.976989Z","shell.execute_reply.started":"2024-12-13T13:36:08.547531Z","shell.execute_reply":"2024-12-13T13:36:11.975942Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"categorical_plots('Exercise_Frequency')\nprint(df_train['Exercise_Frequency'].value_counts())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T13:36:11.978479Z","iopub.execute_input":"2024-12-13T13:36:11.978945Z","iopub.status.idle":"2024-12-13T13:36:15.521045Z","shell.execute_reply.started":"2024-12-13T13:36:11.978883Z","shell.execute_reply":"2024-12-13T13:36:15.519892Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"categorical_plots('Property_Type')\nprint(df_train['Property_Type'].value_counts())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T13:36:15.522285Z","iopub.execute_input":"2024-12-13T13:36:15.522595Z","iopub.status.idle":"2024-12-13T13:36:19.05265Z","shell.execute_reply.started":"2024-12-13T13:36:15.522563Z","shell.execute_reply":"2024-12-13T13:36:19.051312Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<div style=\"background-color: #69cdcb; padding: 10px;\">\n    <h3> &#128521;Insights gained </h3>\n    \n<h4> &#128313; For categorical features, when compared with the target feature, each category has different types. Most categories show similar distributions across the target feature.\n</h4>","metadata":{}},{"cell_type":"markdown","source":"for plot in categorical_columns:\n    plt.figure(figsize=(14, 7))\n    sns.violinplot(x=df_train[plot], y=df_train['Premium Amount'])\n    plt.title(f'Scatter Plot of Premium Amount by {plot}')\n    plt.show()","metadata":{"_kg_hide-input":true}},{"cell_type":"markdown","source":"<center><h1 style=\"color:white; background-color:teal; padding:20px; font-family:Courier New, monospace; border-radius:8px\">Data Cleaning</h1></center>\n","metadata":{"_kg_hide-input":true}},{"cell_type":"markdown","source":"#  change the data type \ndf_train['Policy_Start_Date'].astype('string') # before change the data type is object but i changed to text datatype.\n","metadata":{"_kg_hide-input":true}},{"cell_type":"markdown","source":"### Filling the Null Values","metadata":{"_kg_hide-input":true}},{"cell_type":"markdown","source":"for i in numerical_columns:\n    if i == 'id':\n        pass\n    else:\n        mean_value = df_train[i].mean()\n        df_train[i].fillna(value=mean_value, inplace=True) \n        print(f'{i}:{mean_value}')\n    print(\"completed\")\n        # break","metadata":{"_kg_hide-input":true}},{"cell_type":"markdown","source":"df_train.isnull().sum()","metadata":{"_kg_hide-input":true}},{"cell_type":"markdown","source":"### Hot Coding for Categorical Data ","metadata":{"_kg_hide-input":true}},{"cell_type":"markdown","source":"df_train.replace({\n    'Marital Status': {'Single': 0, 'Married': 1, 'Divorced': 2},\n    'Occupation': {'Self-Employed': 0, 'Employed': 1, 'Unemployed': 2},\n    'Customer Feedback': {'Poor': 0, 'Average': 1, 'Good': 2}\n}, inplace=True)","metadata":{"_kg_hide-input":true}},{"cell_type":"markdown","source":"# Fill NaN values with a specific integer (e.g., 0)\ndf_train['Marital Status'].fillna(0, inplace=True)\ndf_train['Occupation'].fillna(0, inplace=True)\ndf_train['Customer Feedback'].fillna(0, inplace=True)","metadata":{"_kg_hide-input":true}},{"cell_type":"markdown","source":"df_train['Marital Status'] = df_train['Marital Status'].astype(int)\ndf_train['Occupation'] = df_train['Occupation'].astype(int)\ndf_train['Customer Feedback'] = df_train['Customer Feedback'].astype(int)","metadata":{"_kg_hide-input":true}},{"cell_type":"markdown","source":"df_train['Marital Status'].mean()","metadata":{"_kg_hide-input":true}},{"cell_type":"markdown","source":"df_train[df_train['Age'] <= 19]","metadata":{"_kg_hide-input":true}}]}