R语言中基于行计数的GP转诊数据ggplot可视化问题
Hey there! Let's work through your two ggplot visualization tasks using your GP referral data. I'll use tidyverse tools (dplyr for data wrangling + ggplot2 for plotting) since they play nicely with tibbles, and the code will scale perfectly to your full 3000+ row dataset.
First, let's make sure we have your sample data loaded (you can skip this part if your full dataset is already in your environment):
# Load sample data sample_data <- structure(list(LAST_NAME_GP = c("NOORDHOF", "ONBEKEND", "RAHIMTOOLA", "HIEMSTRA", "VIS", "OLDENBURG", "SLACHTER", "NOORDHOF", "VOSKUILEN", "STEVENS", "COMANS", "HIJMERING", "PHILIPS", "VIS", "LOUTER"), INSTITUTION = c("OPVOEDPOLI B.V.", "PARLAN", "PARLAN", "PARLAN", "OPVOEDPOLI B.V.", "TRIVERSUM", "ALKMAARSE PSYCHOLOGENPRAKTIJK", "TRIVERSUM", "STICHTING KRAM", "TRIVERSUM", "TRIVERSUM", "TRIVERSUM", "OPVOEDPOLI B.V.", "TRIVERSUM", "ELINE BIESHEUVEL" )), row.names = c(NA, -15L), class = c("tbl_df", "tbl", "data.frame" ))
This task focuses on counting how many times each GP appears in the dataset (each row = one referral). We'll first summarize the data, then plot it as a clean bar chart.
# Load required packages library(tidyverse) # Step 1: Calculate total referrals per GP gp_total_referrals <- sample_data %>% count(LAST_NAME_GP, name = "total_referrals") %>% # Sort by total referrals (descending) so busiest GPs are first arrange(desc(total_referrals)) # Step 2: Plot with ggplot ggplot(gp_total_referrals, aes(x = reorder(LAST_NAME_GP, total_referrals), y = total_referrals)) + geom_bar(stat = "identity", fill = "#2c3e50") + # Flip axes to make long GP names easier to read coord_flip() + labs( title = "Total Referrals by GP", x = "GP Last Name", y = "Number of Referrals" ) + theme_minimal()
Key notes for your full dataset:
- If you have hundreds of GPs, the plot might get cluttered. You can filter to show only the top N busiest GPs with
slice_max(total_referrals, n = 20)after thearrange()step. - The
reorder()function ensures bars are sorted by referral count, not alphabetical order—this makes trends way easier to spot.
Here we need to count referrals per GP-institution pair. I'll show two common visualization options, plus a pro tip for handling large datasets.
Option 1: Grouped Bar Chart (per GP, split by institution)
This lets you compare referral counts across institutions for each GP:
# Step 1: Calculate referrals per GP + institution combination gp_institution_referrals <- sample_data %>% count(LAST_NAME_GP, INSTITUTION, name = "referral_count") %>% # Sort to keep the busiest GP-institution pairs at the top arrange(desc(referral_count)) # Step 2: Plot grouped bars ggplot(gp_institution_referrals, aes(x = reorder(LAST_NAME_GP, referral_count), y = referral_count, fill = INSTITUTION)) + geom_bar(stat = "identity", position = "dodge") + coord_flip() + labs( title = "Referrals by GP and Target Institution", x = "GP Last Name", y = "Number of Referrals", fill = "Institution" ) + theme_minimal() + # Move legend to bottom to avoid cluttering the plot theme(legend.position = "bottom")
Option 2: Stacked Bar Chart (each bar is a GP, segments are institutions)
This is perfect for quickly seeing the breakdown of institutions per GP:
ggplot(gp_institution_referrals, aes(x = reorder(LAST_NAME_GP, referral_count), y = referral_count, fill = INSTITUTION)) + geom_bar(stat = "identity") + coord_flip() + labs( title = "Referral Breakdown by GP and Institution", x = "GP Last Name", y = "Number of Referrals", fill = "Institution" ) + theme_minimal()
Pro tip for large datasets:
If you have tons of institutions, the legend will get messy. You can use facet_wrap(~ INSTITUTION) to create separate plots for each institution, or filter to only show the top 5 most-referred institutions with slice_max(referral_count, n = 5, by = INSTITUTION).
内容的提问来源于stack exchange,提问作者Hannie

