{"metadata":{"kernelspec":{"name":"ir","display_name":"R","language":"R"},"language_info":{"name":"R","codemirror_mode":"r","pygments_lexer":"r","mimetype":"text/x-r-source","file_extension":".r","version":"4.4.0"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":84896,"databundleVersionId":10305135,"sourceType":"competition"}],"dockerImageVersionId":30749,"isInternetEnabled":true,"language":"r","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# PGS4E12 - Regression with an insurance dataset\nThe goal of this Month is to predict insurance premium amount from. In this notebook, I'll take a first look at the data and perform a quick visual EDA.\n\nComments and feedback are appreciated. If you find any of this useful, please use as you see fit!\n\n\n## Load packages and data","metadata":{}},{"cell_type":"code","source":"# Load libraries\nlibrary(dplyr, quiet = TRUE)\nlibrary(ggplot2, quiet = TRUE)\nlibrary(gridExtra, quiet = TRUE)\nlibrary(tidyr, quiet = TRUE)\nlibrary(reshape2, quiet = TRUE)\nlibrary(lubridate, quiet = TRUE)\nlibrary(DataExplorer, quiet = TRUE)\nlibrary(corrplot, quiet = TRUE)","metadata":{"_uuid":"051d70d956493feee0c6d64651c6a088724dca2a","_execution_state":"idle","trusted":true,"execution":{"iopub.status.busy":"2024-12-09T06:59:45.856401Z","iopub.execute_input":"2024-12-09T06:59:45.858566Z","iopub.status.idle":"2024-12-09T06:59:45.899398Z","shell.execute_reply":"2024-12-09T06:59:45.896749Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# load data\ntrain_data <- read.csv(\"/kaggle/input/playground-series-s4e12/train.csv\")\ntest_data <- read.csv(\"/kaggle/input/playground-series-s4e12/test.csv\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-09T06:59:45.904728Z","iopub.execute_input":"2024-12-09T06:59:45.907503Z","iopub.status.idle":"2024-12-09T07:00:06.50745Z","shell.execute_reply":"2024-12-09T07:00:06.504928Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Data overview","metadata":{}},{"cell_type":"code","source":"summary(train_data)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-09T07:00:06.514407Z","iopub.execute_input":"2024-12-09T07:00:06.516473Z","iopub.status.idle":"2024-12-09T07:00:07.057996Z","shell.execute_reply":"2024-12-09T07:00:07.054885Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"summary(test_data)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-09T07:00:07.062273Z","iopub.execute_input":"2024-12-09T07:00:07.064249Z","iopub.status.idle":"2024-12-09T07:00:07.404731Z","shell.execute_reply":"2024-12-09T07:00:07.400676Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Missing values\nplot missing values by columns:","metadata":{}},{"cell_type":"code","source":"DataExplorer::plot_missing(train_data)\nDataExplorer::plot_missing(test_data)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-09T07:00:07.408915Z","iopub.execute_input":"2024-12-09T07:00:07.410942Z","iopub.status.idle":"2024-12-09T07:00:09.668916Z","shell.execute_reply":"2024-12-09T07:00:09.666324Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The training data contains 1200000 rows and 21 columns. Both the train and test data contain a substantial amount of missing values in the varaibles 'Age', 'Annual.Income', 'Number.of.Dependents' , 'Health.Score', 'Previous.Claims' and 'Credit.Score'.","metadata":{}},{"cell_type":"markdown","source":"## Barplots of discrete variables","metadata":{}},{"cell_type":"code","source":"# frequency distribution of all discrete variables\nplot_bar(train_data)\nplot_bar(test_data)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-09T07:00:09.672575Z","iopub.execute_input":"2024-12-09T07:00:09.674358Z","iopub.status.idle":"2024-12-09T07:00:17.160275Z","shell.execute_reply":"2024-12-09T07:00:17.157766Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"discrete variables in test and training data have very similar frequency distributions.","metadata":{}},{"cell_type":"markdown","source":"## Histograms of numerical variables","metadata":{}},{"cell_type":"code","source":"plot_histogram(train_data)\nplot_histogram(test_data)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-09T07:00:17.165287Z","iopub.execute_input":"2024-12-09T07:00:17.167037Z","iopub.status.idle":"2024-12-09T07:00:41.426816Z","shell.execute_reply":"2024-12-09T07:00:41.424337Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Histogram of outcome variable","metadata":{}},{"cell_type":"code","source":"# histogram of outcome variable \noptions(repr.plot.width = 8, repr.plot.height = 8)\n\nggplot(train_data, aes(x = Premium.Amount)) + \n    geom_histogram(col = \"white\", binwidth = 100) +\n    theme_bw()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-09T07:00:41.431093Z","iopub.execute_input":"2024-12-09T07:00:41.432923Z","iopub.status.idle":"2024-12-09T07:00:42.199642Z","shell.execute_reply":"2024-12-09T07:00:42.197094Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Again, the histograms for the test and training data look very similar. The predictor variable is right skewed.","metadata":{}},{"cell_type":"markdown","source":"## Correlation plot\n","metadata":{}},{"cell_type":"code","source":"plot_correlation(\n  train_data,\n  type = c(\"all\", \"discrete\", \"continuous\"),\n  maxcat = 20L,\n  cor_args = list(),\n  geom_text_args = list(),\n  title = NULL,\n  ggtheme = theme_gray(),\n  theme_config = list(legend.position = \"bottom\", axis.text.x = element_text(angle = 90))\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-09T07:00:42.204302Z","iopub.execute_input":"2024-12-09T07:00:42.206078Z","iopub.status.idle":"2024-12-09T07:01:27.354432Z","shell.execute_reply":"2024-12-09T07:01:27.351203Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"No notable linear relationships between variables.","metadata":{}},{"cell_type":"markdown","source":"#### That's it for now folks! Comments and feedback are always appreciated!\n\n#### Happy competition!","metadata":{}}]}