使用ggplot绘制多过程数据:含日期、测量值及上下限可视化需求
Perfect, let's walk through how to get this done cleanly with R and ggplot2. Your dataset size is totally manageable, and breaking this into two key steps—date cleaning and plotting—will make it straightforward.
Step 1: Prep Your Data (Date & Factor Handling)
First, we need to make sure your date column is recognized as a date type (not just text) and that your process types are treated as categorical factors (this helps with grouping/coloring in ggplot). We'll use tidyverse for data wrangling and lubridate for easy date parsing:
# Load required packages library(tidyverse) library(lubridate) # Clean the data df_clean <- df %>% # Convert date column to date format—adjust ymd() to match your actual date structure! # Use mdy() if dates are like "MM/DD/YYYY", dmy() for "DD/MM/YYYY" mutate( date = ymd(date), process_types = as.factor(process_types) )
Double-check the date conversion worked with head(df_clean$date)—you should see dates formatted as YYYY-MM-DD instead of strings.
Step 2: Build the ggplot Visualization
Now we'll create a plot that shows measured values over time for each process, with upper/lower limits included. Since you have 14 process types, using faceting (one plot per process) will keep things readable instead of cramming all lines into one chart.
Option 1: If Upper/Lower Limits Vary by Date
If each row has a unique upper/lower limit tied to that date and process, use this code:
ggplot(df_clean, aes(x = date)) + # Add upper limit (gray dashed line) geom_line(aes(y = upper_limit), color = "gray50", linetype = "dashed") + # Add lower limit (gray dashed line) geom_line(aes(y = lower_limit), color = "gray50", linetype = "dashed") + # Add measured value line (colored by process type) geom_line(aes(y = measured_value, color = process_types), linewidth = 1) + # Facet by process type—adjust ncol to fit your screen facet_wrap(~process_types, ncol = 4) + # Customize labels and theme labs( title = "Measured Values vs. Control Limits by Process", x = "Date", y = "Measurement Value", color = "Process Type" ) + # Format date axis to avoid overlapping text scale_x_date(date_labels = "%Y-%m-%d", date_breaks = "1 month") + theme_minimal() + theme( axis.text.x = element_text(angle = 45, hjust = 1), plot.title = element_text(hjust = 0.5, size = 14) )
Option 2: If Upper/Lower Limits Are Fixed per Process
If each process has a single static upper/lower limit (not changing with date), we'll first extract those fixed values, then use geom_hline to add horizontal lines to each facet:
# Extract fixed limits per process process_limits <- df_clean %>% group_by(process_types) %>% summarise( lower_limit = unique(lower_limit), upper_limit = unique(upper_limit) ) # Build the plot ggplot(df_clean, aes(x = date, y = measured_value, color = process_types)) + geom_line(linewidth = 1) + # Add fixed lower limit to each facet geom_hline(data = process_limits, aes(yintercept = lower_limit), color = "gray50", linetype = "dashed") + # Add fixed upper limit to each facet geom_hline(data = process_limits, aes(yintercept = upper_limit), color = "gray50", linetype = "dashed") + facet_wrap(~process_types, ncol = 4) + labs( title = "Measured Values vs. Fixed Control Limits by Process", x = "Date", y = "Measurement Value", color = "Process Type" ) + scale_x_date(date_labels = "%Y-%m-%d", date_breaks = "1 month") + theme_minimal() + theme( axis.text.x = element_text(angle = 45, hjust = 1), plot.title = element_text(hjust = 0.5, size = 14) )
Quick Tips
- If your date axis still looks cluttered, adjust
date_breaks(e.g., use"2 months"instead of"1 month"). - If you prefer all processes on one single chart instead of facets, remove the
facet_wrap()line—just note that 14 colors might get busy, so you could usescale_color_viridis_d()for more distinguishable hues. - Always check for missing values with
sum(is.na(df_clean))—missing dates or measurements can mess up the line plot.
内容的提问来源于stack exchange,提问作者Jenny

