R语言累积图绘制方法及适配统计检验选型咨询
Hey there! Let’s walk through exactly how to create those three cumulative plots you’re after in R, plus cover which statistical tests fit best for your analysis. I’ll use the tidyverse and ggplot2 packages since they’re the most intuitive for this kind of work—make sure you’ve got them installed first with install.packages(c("tidyverse")) if you haven’t already.
This is essentially an empirical cumulative distribution function (ECDF) plot, which shows what percentage of your stop events fall below a given Q value. To highlight the 80th percentile (the Q value where 80% of stops occur), we can add reference lines for clarity.
Code Example:
library(tidyverse) # Replace `your_data` with your actual dataset name ggplot(your_data, aes(x = Q)) + stat_ecdf(geom = "step", color = "#2c3e50", linewidth = 1) + # Add 80th percentile reference lines geom_vline(xintercept = quantile(your_data$Q, 0.8, na.rm = TRUE), color = "#e74c3c", linetype = "dashed", linewidth = 1) + geom_hline(yintercept = 0.8, color = "#e74c3c", linetype = "dashed", linewidth = 1) + # Annotate the 80% threshold value annotate("text", x = quantile(your_data$Q, 0.8, na.rm = TRUE) + 0.5, y = 0.8, label = paste0("80% of stops at Q = ", round(quantile(your_data$Q, 0.8, na.rm = TRUE), 2)), hjust = 0) + labs(title = "Cumulative Distribution of Q Values for Stop Events", x = "Q (mg/l)", y = "Cumulative % of Stop Events") + theme_minimal()
Note: The na.rm = TRUE argument handles any missing values in your Q column.
First, we need to filter your data to only include records where Q ≥ 1mg/l, then sort by date and calculate the running total of stops over time.
Code Example:
# Prepare data: filter threshold, sort by date, compute cumulative stops cumulative_stops_over_threshold <- your_data %>% filter(Q >= 1) %>% arrange(Date) %>% mutate(cumulative_stops = cumsum(Stops)) # Convert Date to proper date format if needed if (!inherits(your_data$Date, "Date")) { cumulative_stops_over_threshold$Date <- as.Date(cumulative_stops_over_threshold$Date) } # Plot ggplot(cumulative_stops_over_threshold, aes(x = Date, y = cumulative_stops)) + geom_line(color = "#2980b9", linewidth = 1) + geom_point(size = 2, color = "#2980b9") + labs(title = "Cumulative Stop Events Since Q Exceeded 1mg/l", x = "Date", y = "Cumulative Number of Stops") + theme_minimal() + theme(axis.text.x = element_text(angle = 45, hjust = 1))
This is the simplest cumulative plot—we just sort all your data by date and compute the running total of stops across the full dataset.
Code Example:
# Prepare data: sort by date, compute cumulative stops cumulative_stops_all <- your_data %>% arrange(Date) %>% mutate(cumulative_stops = cumsum(Stops)) # Convert Date to proper date format if needed if (!inherits(your_data$Date, "Date")) { cumulative_stops_all$Date <- as.Date(cumulative_stops_all$Date) } # Plot ggplot(cumulative_stops_all, aes(x = Date, y = cumulative_stops)) + geom_line(color = "#27ae60", linewidth = 1) + labs(title = "Cumulative Stop Events Over Time", x = "Date", y = "Cumulative Number of Stops") + theme_minimal() + theme(axis.text.x = element_text(angle = 45, hjust = 1))
Here are the best tests based on what you’re trying to investigate:
- Relationship between Q and Stop Events: Use the Spearman Rank Correlation Test (
cor.test(your_data$Q, your_data$Stops, method = "spearman")). This works better than Pearson correlation because it handles non-linear relationships, which are common in cumulative count data. - Trend in Cumulative Stops Over Time: To test if stop frequency is increasing/decreasing significantly over time, use a Poisson Regression (since stops are count data):
glm(Stops ~ Date, data = your_data, family = poisson()). If you notice overdispersion (common with count data), switch to a Negative Binomial Regression using theMASSpackage:glm.nb(Stops ~ Date, data = your_data). - Difference in Stop Rates Before/After Q = 1mg/l: If you want to compare stop frequency before vs. after the 1mg/l threshold, use a Chi-Squared Test (group data into "Q <1" and "Q >=1", count stops in each group) or Fisher’s Exact Test if your sample sizes are small.
内容的提问来源于stack exchange,提问作者ac93

