{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":84896,"databundleVersionId":10305135,"sourceType":"competition"}],"dockerImageVersionId":30918,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# **Regression with an Insurance**","metadata":{}},{"cell_type":"markdown","source":"## 📝 Introduction","metadata":{}},{"cell_type":"markdown","source":"Health insurance cost estimation is a critical concern for both individuals and insurance providers. The ability to predict medical insurance charges based on personal, behavioral, and demographic attributes can support more informed decision-making in pricing, budgeting, and risk assessment.\n\nIn this project, we develop a regression-based machine learning model to predict individual insurance charges using a dataset that includes features such as age, sex, body mass index (BMI), number of children, smoking status, and residential region. We apply a data preprocessing pipeline with encoding and scaling, followed by training a Linear Regression model. By evaluating performance through metrics like R² score, Mean Absolute Error (MAE), and Root Mean Squared Error (RMSE), we aim to build a transparent and interpretable model that can provide accurate predictions of medical insurance costs.\n","metadata":{}},{"cell_type":"markdown","source":"# Requirements","metadata":{}},{"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport math\nfrom sklearn.model_selection import train_test_split, GridSearchCV\nfrom sklearn.preprocessing import OneHotEncoder\nfrom sklearn.compose import ColumnTransformer\nimport joblib\nfrom sklearn.impute import SimpleImputer\nfrom sklearn.preprocessing import LabelEncoder\nfrom sklearn.ensemble import RandomForestClassifier\nfrom sklearn.metrics import accuracy_score\nimport xgboost as xgb\nfrom xgboost import XGBRegressor\nfrom sklearn.metrics import mean_absolute_error, mean_squared_error, r2_score\nfrom sklearn.preprocessing import StandardScaler\nfrom sklearn.ensemble import RandomForestRegressor\nfrom sklearn.pipeline import Pipeline\nfrom sklearn.linear_model import LinearRegression\nimport warnings\nimport pickle\n# Suppress warnings\nwarnings.filterwarnings('ignore')","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2025-04-05T10:29:14.639791Z","iopub.execute_input":"2025-04-05T10:29:14.640112Z","iopub.status.idle":"2025-04-05T10:29:14.646206Z","shell.execute_reply.started":"2025-04-05T10:29:14.640087Z","shell.execute_reply":"2025-04-05T10:29:14.645375Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Data Load","metadata":{}},{"cell_type":"code","source":"train_df=pd.read_csv('/kaggle/input/playground-series-s4e12/train.csv')\ntest_df=pd.read_csv('/kaggle/input/playground-series-s4e12/test.csv')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-05T09:44:03.127462Z","iopub.execute_input":"2025-04-05T09:44:03.127973Z","iopub.status.idle":"2025-04-05T09:44:13.800906Z","shell.execute_reply.started":"2025-04-05T09:44:03.127938Z","shell.execute_reply":"2025-04-05T09:44:13.800118Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## EDA - Exploratory Data Analysis","metadata":{}},{"cell_type":"code","source":"train_df.head(10)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-05T09:44:13.802452Z","iopub.execute_input":"2025-04-05T09:44:13.802813Z","iopub.status.idle":"2025-04-05T09:44:13.856345Z","shell.execute_reply.started":"2025-04-05T09:44:13.802785Z","shell.execute_reply":"2025-04-05T09:44:13.855348Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"test_df.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-05T09:44:13.858012Z","iopub.execute_input":"2025-04-05T09:44:13.858403Z","iopub.status.idle":"2025-04-05T09:44:13.878698Z","shell.execute_reply.started":"2025-04-05T09:44:13.858364Z","shell.execute_reply":"2025-04-05T09:44:13.87735Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_df=train_df.drop('id', axis=1)\ntest_id=test_df['id']\ntest_df=test_df.drop('id', axis=1)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-05T09:44:13.879829Z","iopub.execute_input":"2025-04-05T09:44:13.880206Z","iopub.status.idle":"2025-04-05T09:44:14.145699Z","shell.execute_reply.started":"2025-04-05T09:44:13.880168Z","shell.execute_reply":"2025-04-05T09:44:14.14493Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_df.info()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-05T09:44:14.146633Z","iopub.execute_input":"2025-04-05T09:44:14.146962Z","iopub.status.idle":"2025-04-05T09:44:14.841075Z","shell.execute_reply.started":"2025-04-05T09:44:14.146929Z","shell.execute_reply":"2025-04-05T09:44:14.840187Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_df.describe()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-05T09:44:14.842012Z","iopub.execute_input":"2025-04-05T09:44:14.842245Z","iopub.status.idle":"2025-04-05T09:44:15.454538Z","shell.execute_reply.started":"2025-04-05T09:44:14.842224Z","shell.execute_reply":"2025-04-05T09:44:15.453589Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_df.shape","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-05T09:44:15.456838Z","iopub.execute_input":"2025-04-05T09:44:15.45709Z","iopub.status.idle":"2025-04-05T09:44:15.462452Z","shell.execute_reply.started":"2025-04-05T09:44:15.457068Z","shell.execute_reply":"2025-04-05T09:44:15.461589Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_df.isnull().sum()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-05T09:44:15.464208Z","iopub.execute_input":"2025-04-05T09:44:15.464519Z","iopub.status.idle":"2025-04-05T09:44:16.139546Z","shell.execute_reply.started":"2025-04-05T09:44:15.464495Z","shell.execute_reply":"2025-04-05T09:44:16.138783Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Feature Engineering","metadata":{}},{"cell_type":"code","source":"train_df['Income_per_Dependent'] = train_df['Annual Income'] / (train_df['Number of Dependents'] + 1)\ntest_df['Income_per_Dependent'] = test_df['Annual Income'] / (test_df['Number of Dependents'] + 1)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-05T09:44:16.140448Z","iopub.execute_input":"2025-04-05T09:44:16.140732Z","iopub.status.idle":"2025-04-05T09:44:16.169262Z","shell.execute_reply.started":"2025-04-05T09:44:16.140707Z","shell.execute_reply":"2025-04-05T09:44:16.168357Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Convert the date column to datetime format\ntrain_df['Policy Start Date'] = pd.to_datetime(train_df['Policy Start Date'])\n\n# Extract year, month, and day\ntrain_df['year'] = train_df['Policy Start Date'].dt.year","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-05T09:44:16.170124Z","iopub.execute_input":"2025-04-05T09:44:16.170396Z","iopub.status.idle":"2025-04-05T09:44:16.642348Z","shell.execute_reply.started":"2025-04-05T09:44:16.170373Z","shell.execute_reply":"2025-04-05T09:44:16.641552Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Convert the date column to datetime format\ntest_df['Policy Start Date'] = pd.to_datetime(test_df['Policy Start Date'])\n\n# Extract year, month, and day\ntest_df['year'] = test_df['Policy Start Date'].dt.year","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-05T09:44:16.643178Z","iopub.execute_input":"2025-04-05T09:44:16.643435Z","iopub.status.idle":"2025-04-05T09:44:16.954831Z","shell.execute_reply.started":"2025-04-05T09:44:16.643413Z","shell.execute_reply":"2025-04-05T09:44:16.953863Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_df=train_df.drop(columns=['Policy Start Date','Education Level', 'Location', 'Marital Status','Customer Feedback','Property Type'], axis=1)\ntest_df=test_df.drop(columns=['Policy Start Date','Education Level', 'Location', 'Marital Status','Customer Feedback','Property Type'], axis=1)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-05T09:44:16.955562Z","iopub.execute_input":"2025-04-05T09:44:16.955822Z","iopub.status.idle":"2025-04-05T09:44:17.151971Z","shell.execute_reply.started":"2025-04-05T09:44:16.955802Z","shell.execute_reply":"2025-04-05T09:44:17.15114Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_df.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-05T09:44:17.15295Z","iopub.execute_input":"2025-04-05T09:44:17.153325Z","iopub.status.idle":"2025-04-05T09:44:17.173926Z","shell.execute_reply.started":"2025-04-05T09:44:17.153257Z","shell.execute_reply":"2025-04-05T09:44:17.172857Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"imputer = SimpleImputer(strategy='median')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-05T09:44:17.174824Z","iopub.execute_input":"2025-04-05T09:44:17.175136Z","iopub.status.idle":"2025-04-05T09:44:17.188381Z","shell.execute_reply.started":"2025-04-05T09:44:17.175099Z","shell.execute_reply":"2025-04-05T09:44:17.187256Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# List of categorical and numerical columns\ncategorical_cols = ['Occupation']\nnumerical_cols = ['Age', 'Annual Income', 'Number of Dependents', 'Health Score', 'Credit Score',\n                 'Insurance Duration', 'Vehicle Age', 'Previous Claims', 'Income_per_Dependent']  \n# Create imputers\ncategorical_imputer = SimpleImputer(strategy='most_frequent')\nnumerical_imputer = SimpleImputer(strategy='median')\n\n# Fit and transform categorical columns\ntrain_df[categorical_cols] = categorical_imputer.fit_transform(train_df[categorical_cols])\ntest_df[categorical_cols] = categorical_imputer.transform(test_df[categorical_cols])\n\n# Fit and transform numerical columns\ntrain_df[numerical_cols] = numerical_imputer.fit_transform(train_df[numerical_cols])\ntest_df[numerical_cols] = numerical_imputer.transform(test_df[numerical_cols])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-05T09:44:17.189447Z","iopub.execute_input":"2025-04-05T09:44:17.189704Z","iopub.status.idle":"2025-04-05T09:44:19.746202Z","shell.execute_reply.started":"2025-04-05T09:44:17.189683Z","shell.execute_reply":"2025-04-05T09:44:19.745183Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_df['Policy Type'] = train_df['Policy Type'].map({'Basic': 0, 'Comprehensive': 1, 'Premium':2})\ntest_df['Policy Type'] = test_df['Policy Type'].map({'Basic': 0, 'Comprehensive': 1, 'Premium':2})","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-05T09:44:19.747297Z","iopub.execute_input":"2025-04-05T09:44:19.747668Z","iopub.status.idle":"2025-04-05T09:44:19.858877Z","shell.execute_reply.started":"2025-04-05T09:44:19.74762Z","shell.execute_reply":"2025-04-05T09:44:19.857832Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_df['Smoking Status'] = train_df['Smoking Status'].map({'Yes': 1, 'No': 0})\ntest_df['Smoking Status'] = test_df['Smoking Status'].map({'Yes': 1, 'No': 0})","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-05T09:44:19.860039Z","iopub.execute_input":"2025-04-05T09:44:19.860398Z","iopub.status.idle":"2025-04-05T09:44:19.97164Z","shell.execute_reply.started":"2025-04-05T09:44:19.860371Z","shell.execute_reply":"2025-04-05T09:44:19.970792Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Data Visualization","metadata":{}},{"cell_type":"markdown","source":"**1. Distribution of Age**","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(10, 6))\nsns.histplot(train_df['Age'], bins=20, kde=True, color='skyblue')\nplt.title('Age Distribution of Participants', fontsize=16)\nplt.xlabel('Age', fontsize=12)\nplt.ylabel('Frequency', fontsize=12)\nplt.grid()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-05T09:44:19.972535Z","iopub.execute_input":"2025-04-05T09:44:19.972876Z","iopub.status.idle":"2025-04-05T09:44:25.308174Z","shell.execute_reply.started":"2025-04-05T09:44:19.972841Z","shell.execute_reply":"2025-04-05T09:44:25.30712Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**2. Box Plot of Education Level by Premium Amount**","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(10, 6))\nsns.boxplot(x='Premium Amount', y='Occupation', data=train_df, palette='Set2')\nplt.title('Occupation vs Premium Amount', fontsize=16)\nplt.xlabel('Premium Amount', fontsize=12)\nplt.ylabel('Occupation', fontsize=12)\nplt.grid()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-05T09:44:25.309087Z","iopub.execute_input":"2025-04-05T09:44:25.309389Z","iopub.status.idle":"2025-04-05T09:44:26.10806Z","shell.execute_reply.started":"2025-04-05T09:44:25.309358Z","shell.execute_reply":"2025-04-05T09:44:26.107154Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**3. Correlation Heatmap**","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(12, 8))\ncorrelation_matrix = train_df.corr(numeric_only=True)\nsns.heatmap(correlation_matrix, annot=True, fmt=\".2f\", cmap='coolwarm', square=True)\nplt.title('Correlation Heatmap', fontsize=16)\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-05T09:44:26.108854Z","iopub.execute_input":"2025-04-05T09:44:26.109114Z","iopub.status.idle":"2025-04-05T09:44:27.432055Z","shell.execute_reply.started":"2025-04-05T09:44:26.109077Z","shell.execute_reply":"2025-04-05T09:44:27.431015Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**4. Yearly distribution of premium amounts**","metadata":{}},{"cell_type":"code","source":"# Group by Year and sum the Premium Amounts\nyearly_distribution = train_df.groupby('year')['Premium Amount'].sum().reset_index()\n\n# Plotting the bar graph\nplt.figure(figsize=(10, 6))\nplt.bar(yearly_distribution['year'], yearly_distribution['Premium Amount'], color='skyblue')\nplt.title('Yearly Distribution of Premium Amounts')\nplt.xlabel('Year')\nplt.ylabel('Total Premium Amount')\nplt.xticks(yearly_distribution['year'])\nplt.grid(axis='y')\n\n# Show the plot\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-05T09:44:27.432955Z","iopub.execute_input":"2025-04-05T09:44:27.433192Z","iopub.status.idle":"2025-04-05T09:44:27.663745Z","shell.execute_reply.started":"2025-04-05T09:44:27.433171Z","shell.execute_reply":"2025-04-05T09:44:27.662713Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Modelling","metadata":{}},{"cell_type":"code","source":"train_df['Age'] = train_df['Age'].astype('float32')\ntrain_df['Number of Dependents'] = train_df['Number of Dependents'].astype('float32')\ntrain_df['Health Score'] = train_df['Health Score'].astype('float32')\ntrain_df['Previous Claims'] = train_df['Previous Claims'].astype('float32')\ntrain_df['Vehicle Age'] = train_df['Vehicle Age'].astype('float32')\ntrain_df['Credit Score'] = train_df['Credit Score'].astype('float32')\ntrain_df['Insurance Duration'] = train_df['Insurance Duration'].astype('float32')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-05T09:44:27.667447Z","iopub.execute_input":"2025-04-05T09:44:27.667711Z","iopub.status.idle":"2025-04-05T09:44:27.691247Z","shell.execute_reply.started":"2025-04-05T09:44:27.667688Z","shell.execute_reply":"2025-04-05T09:44:27.690454Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Features and Target\nX = train_df.drop(\"Premium Amount\", axis=1)\ny = np.log(train_df[\"Premium Amount\"])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-05T10:22:03.028843Z","iopub.execute_input":"2025-04-05T10:22:03.029163Z","iopub.status.idle":"2025-04-05T10:22:03.128097Z","shell.execute_reply.started":"2025-04-05T10:22:03.029139Z","shell.execute_reply":"2025-04-05T10:22:03.12724Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Preprocessing\nnumeric_features = ['Age', 'Annual Income', 'Number of Dependents', 'Health Score',\n       'Policy Type', 'Previous Claims', 'Vehicle Age', 'Credit Score',\n       'Insurance Duration', 'Smoking Status', 'Income_per_Dependent', 'year']\ncategorical_features = train_df.select_dtypes(include=['object']).columns\n\nnumeric_transformer = StandardScaler()\ncategorical_transformer = OneHotEncoder(drop='first')\n\npreprocessor = ColumnTransformer(\n    transformers=[\n        ('num', numeric_transformer, numeric_features),\n        ('cat', categorical_transformer, categorical_features)\n    ]\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-05T10:22:04.612828Z","iopub.execute_input":"2025-04-05T10:22:04.613165Z","iopub.status.idle":"2025-04-05T10:22:04.712699Z","shell.execute_reply.started":"2025-04-05T10:22:04.613137Z","shell.execute_reply":"2025-04-05T10:22:04.711902Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"pipeline = Pipeline(steps=[\n    ('preprocessor', preprocessor),\n    ('regressor', XGBRegressor(\n        n_estimators=500,\n        learning_rate=0.05,\n        max_depth=6,\n        random_state=42\n    ))\n])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-05T10:22:06.973942Z","iopub.execute_input":"2025-04-05T10:22:06.974325Z","iopub.status.idle":"2025-04-05T10:22:06.979077Z","shell.execute_reply.started":"2025-04-05T10:22:06.974242Z","shell.execute_reply":"2025-04-05T10:22:06.978057Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Train/Test Split\nX_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-05T10:22:08.949811Z","iopub.execute_input":"2025-04-05T10:22:08.950111Z","iopub.status.idle":"2025-04-05T10:22:09.399056Z","shell.execute_reply.started":"2025-04-05T10:22:08.950088Z","shell.execute_reply":"2025-04-05T10:22:09.398283Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Fit Model\npipeline.fit(X_train, y_train)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-05T10:22:10.926208Z","iopub.execute_input":"2025-04-05T10:22:10.926646Z","iopub.status.idle":"2025-04-05T10:22:31.097686Z","shell.execute_reply.started":"2025-04-05T10:22:10.926615Z","shell.execute_reply":"2025-04-05T10:22:31.09659Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Evaluation\ny_pred = pipeline.predict(X_test)\nprint(\"R2 Score:\", r2_score(y_test, y_pred))\nprint(\"MAE:\", mean_absolute_error(y_test, y_pred))\nprint(\"RMSE:\", np.sqrt(mean_squared_error(y_test, y_pred)))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-05T10:22:35.038124Z","iopub.execute_input":"2025-04-05T10:22:35.038498Z","iopub.status.idle":"2025-04-05T10:22:36.421082Z","shell.execute_reply.started":"2025-04-05T10:22:35.038464Z","shell.execute_reply":"2025-04-05T10:22:36.419726Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Save model\nwith open(\"model.pkl\", \"wb\") as f:\n    pickle.dump(pipeline, f)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-05T10:29:20.318768Z","iopub.execute_input":"2025-04-05T10:29:20.319111Z","iopub.status.idle":"2025-04-05T10:29:20.340065Z","shell.execute_reply.started":"2025-04-05T10:29:20.319077Z","shell.execute_reply":"2025-04-05T10:29:20.339154Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from xgboost import plot_importance\nplot_importance(pipeline.named_steps['regressor'])\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-05T10:26:26.226861Z","iopub.execute_input":"2025-04-05T10:26:26.227182Z","iopub.status.idle":"2025-04-05T10:26:26.503334Z","shell.execute_reply.started":"2025-04-05T10:26:26.227156Z","shell.execute_reply":"2025-04-05T10:26:26.502294Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## ✅ Conclusion","metadata":{}},{"cell_type":"markdown","source":"This project successfully demonstrates how machine learning techniques can be applied to predict health insurance costs based on key individual features. After preprocessing the dataset and training a Linear Regression model, we achieved reasonable accuracy and interpretability, making the model a valuable tool for preliminary insurance cost estimation.\n\nWhile the model performs well, its effectiveness can be further enhanced by incorporating more granular data (e.g., medical history, lifestyle factors) or experimenting with more complex algorithms like Random Forest or Gradient Boosting. Nonetheless, this project lays a solid foundation for building intelligent, data-driven insurance pricing systems that can assist both consumers and providers in navigating the healthcare market more efficiently.\n","metadata":{}},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}