{"metadata":{"kernelspec":{"name":"ir","display_name":"R","language":"R"},"language_info":{"name":"R","codemirror_mode":"r","pygments_lexer":"r","mimetype":"text/x-r-source","file_extension":".r","version":"4.0.5"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This R environment comes with many helpful analytics packages installed\n# It is defined by the kaggle/rstats Docker image: https://github.com/kaggle/docker-rstats\n# For example, here's a helpful package to load\n\nlibrary(tidyverse) # metapackage of all tidyverse packages\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nlist.files(path = \"../input\")\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"051d70d956493feee0c6d64651c6a088724dca2a","_execution_state":"idle","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Source - \n- https://cran.r-project.org/web/packages/vtable/vignettes/vtablefunction.html","metadata":{}},{"cell_type":"markdown","source":"### Source \n- https://www.rdocumentation.org/packages/psych/versions/2.3.6\n- https://www.rdocumentation.org/packages/vtable/versions/1.4.4/topics/vtable\n- https://cran.r-project.org/web/packages/arsenal/index.html\n- https://github.com/mayoverse/arsenal\n- https://www.rdocumentation.org/packages/arsenal/versions/3.6.3\n- https://cran.r-project.org/web/packages/table1/vignettes/table1-examples.html\n- https://cran.r-project.org/web/packages/egg/vignettes/Ecosystem.html","metadata":{}},{"cell_type":"code","source":"library(readr)\ndf <- read_csv(\"/kaggle/input/rsna-2023-abdominal-trauma-detection/train.csv\")\nhead(df)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Calculate descriptive statistics\nmean <- mean(df$any_injury)\nmedian <- median(df$any_injury)\nmode <- mode(df$any_injury)\nstd_dev <- sd(df$any_injury)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Print the descriptive statistics\nprint(mean)\nprint(median)\nprint(mode)\nprint(std_dev)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Load the plotly and psych packages\nlibrary(plotly)\nlibrary(psych)\n\n# Calculate the summary statistics of the dataframe\nstats <- df %>% describe()\n\n# Create an interactive plot of the statistics\np <- plot_ly(stats, x = names(stats), y = stats, type = \"bar\")\n\n# Show the plot\np","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Correlation Plot\n- melt the correlation matrix using the melt() function. Covert the matrix into a data frame with three columns: Var1, Var2, and value. The Var1 and Var2 columns contain the names of the variables, and the value column contains the correlation coefficient between the two variables.\n- Then interactive correlation plot using the ggplot() function. The geom_tile() layer creates a heatmap of the correlation coefficients. The scale_fill_gradient() layer specifies the color scale for the heatmap. The labs() layer adds titles and labels to the plot. The theme() layer adjusts the appearance of the plot.","metadata":{}},{"cell_type":"code","source":"library(reshape2)\n# Create a correlation matrix\ncorr_mat <- cor(df)\n\n# Melt the correlation matrix\nmelted_corr_mat <- melt(corr_mat)\n\n# Create an interactive correlation plot\nggplot(data = melted_corr_mat, aes(x = Var1, y = Var2, fill = value)) +\n  geom_tile() +\n  scale_fill_gradient(low = \"white\", high = \"red\") +\n  labs(title = \"Correlation\", x = \"Variable 1\", y = \"Variable 2\") +\n  theme(axis.text.x = element_text(angle = 45, hjust = 1))\n\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"library(readr)\n\n# Define the names of the CSV files\nfilenames <- c(\"/kaggle/input/rsna-2023-abdominal-trauma-detection/image_level_labels.csv\", \"/kaggle/input/rsna-2023-abdominal-trauma-detection/train.csv\", \"/kaggle/input/rsna-2023-abdominal-trauma-detection/train_series_meta.csv\")\n\n# Read the CSV files into data frames\ndata1 <- read_csv(filenames[1])\ndata2 <- read_csv(filenames[2])\ndata3 <- read_csv(filenames[3])\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Merge the dataframes using`merge()` function\ndf_merged <- merge(data1, data2, by = \"patient_id\")\ndf_merged <- merge(df_merged, data3, by = \"patient_id\")\n\nprint(head(df_merged))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(colnames(df_merged))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Check the data type of each column\nprint(str(df_merged))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_merged <- df_merged[, !(colnames(df_merged) %in% c(\"injury_name\"))]","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_merged <- na.omit(df_merged)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"correlation_matrix <- cor(df_merged)\n\n# Create a correlation plot\nplot(correlation_matrix, main=\"Correlation Plot\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"install.packages(\"sjlabelled\")\ninstall.packages(\"vtable\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# ?vtable","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"library(sjlabelled)\nget_label(df_merged)\n\ndata_labels <- enframe(get_label(df_merged))\ndata_labels","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"descriptive_statistics <- df_merged %>% describe() %>% as_tibble() %>% select(\"n\",\"min\",\"max\",\"mean\",\"median\")\nstats_table <- cbind(data_labels,descriptive_statistics)\nstats_table","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"library(arsenal)\ntable <- tableby(any_injury ~ ., data=df_merged)\nsummary(table, text=TRUE)\nlabels(table)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"barplot(df_merged$extravasation_healthy, df_merged$spleen_healthy, xlab=\"Percentage\", ylab=\"Proportion\")\nggplot(df_merged, aes(x=extravasation_healthy, y=spleen_healthy)) + geom_bar(stat=\"identity\") + \n  labs(x=\"Percentage\", y=\"Proportion\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"library(\"gridExtra\")\nggplot1 <- ggplot(df_merged, aes(x = kidney_healthy, y = kidney_low)) + geom_bar(stat='identity')\n\nggplot2 <- ggplot(df_merged, aes(x = spleen_high, y = spleen_high)) + geom_bar(stat='identity')\n\ngrid.arrange(ggplot1, ggplot2, ncol = 2)\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# library(table1)\n# # Display the dataframe with inferential statistics\n# table1(df_merged, \n#         render.continuous = function(x) {\n#           with(stats.apply.rounding(stats.default(x), digits=2), c(\"\", \"Mean (SD)\"=sprintf(\"%s (± %s)\", MEAN, SD)))\n#         },\n#         render.categorical = function(x) {\n#           c(\"\", sapply(stats.default(x), function(y) with(y, sprintf(\"%d (%0.0f %%)\", FREQ, PCT))))\n#         })\n            \n            ","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}