R语言:创建分组mean_rp列及解决区域均值时序图异常问题
R语言新手问题:新增均值列+修复区域时序图异常
需求说明
- 在数据框中新增
mean_rp列:计算每个month+region组合的rp均值,再乘以对应行的year数值
- 在数据框中新增
- 已生成
mean_rp列,但绘制不同region的mean_rp时序图时线条异常,原代码仅替换了code为region、rp为mean_rp,希望每个区域显示一条对应颜色的线条
- 已生成
附用户提供的数据与代码
数据结构
structure(list(month = c("JAN", "FEV", "MAR", "ABR", "MAI", "JUN", "JUL", "AGO", "SET", "OUT", "NOV", "DEZ", "JAN", "FEV"), year = c(2004L, 2004L, 2004L, 2004L, 2004L, 2004L, 2004L, 2004L, 2004L, 2004L, 2004L, 2004L, 2004L, 2004L), code = c("AL", "AL", "AL", "AL", "AL", "AL", "AL", "AL", "AL", "AL", "AL", "AL", "AM", "AM"), region = c("NE", "NE", "NE", "NE", "NE", "NE", "NE", "NE", "NE", "NE", "NE", "NE", "N", "N"), month_num = c(1L, 2L, 3L, 4L, 5L, 6L, 7L, 8L, 9L, 10L, 11L, 12L, 1L, 2L), rp = c(151.351257324219, 150.433822631836, 145.326431274414, 144.790817260742, 139.691024780273, 138.706481933594, 137.455856323242, 145.046249389648, 136.064834594727, 135.468658447266, 134.540267944336, 137.561904907227, 142.4482421875, 141.584777832031)), row.names = c(NA, 14L), class = "data.frame")
原单州可视化代码
dc$date <- as.Date(dc$date) plot(subset(dc, code == "RS")$date, subset(dc, code == "RS")$rp, type = "l", xlab = "Time", ylab = "Real price of cement", ylim=c(0,400)) points(subset(dc, code == "MG")$date, subset(dc, code == "MG")$rp, type = "l",col="red") points(subset(dc, code == "MT")$date, subset(dc, code == "MT")$rp, type = "l",col="yellow") points(subset(dc, code == "PE")$date, subset(dc, code == "PE")$rp, type = "l",col="blue")
#Instruments: Shades of grey dc$date <- as.Date(dc$date) points(subset(dc, code == "RS")$date, subset(dc, code == "RS")$rip_diesel/20, type = "l", col="darkgrey") points(subset(dc, code == "RS")$date, subset(dc, code == "RS")$rip_fueloil/10, type = "l", col="darkgrey") points(subset(dc, code == "RS")$date, subset(dc, code == "RS")$rip_vegcoal,type = "l", col="darkgrey") legend("topleft",legend = c("Price of Cement - RS", "Price of Cement - MG","Price of Cement - MT", "Price of Cement - PE", "Price of Diesel (/20)","Price of Fuel Oil (/10)","Price of Veg. Oil"),lty = c(1,1,1,1,1,1,1),col = c("black", "yellow", "blue", "dark grey", "grey", "lightgrey"),inset = c(0.15,0))
异常的区域均值可视化代码
plot(subset(dc, region == "N")$date, subset(dc, region == "N")$mean_rp, type = "l", xlab = "Time", ylab = "Real price of cement", ylim=c(0,300)) points(subset(dc, region == "NE")$date, subset(dc, region == "NE")$mean_rp, type = "l", col="red") points(subset(dc, region == "MW")$date, subset(dc, region == "MW")$mean_rp, type = "l", col="yellow") points(subset(dc, region == "S")$date, subset(dc, region == "S")$mean_rp, type = "l", col="blue") points(subset(dc, region == "SE")$date, subset(dc, region == "SE")$mean_rp, type = "l",col="green")
解决方案
1. 新增mean_rp列
方法1:用dplyr包(新手友好,语法直观)
# 首次使用需安装dplyr install.packages("dplyr") library(dplyr) # 分组计算月度区域均值,再乘以年份得到mean_rp dc <- dc %>% group_by(month, region) %>% mutate(rp_month_mean = mean(rp, na.rm = TRUE)) %>% ungroup() %>% mutate(mean_rp = rp_month_mean * year) %>% select(-rp_month_mean) # 移除中间计算列(可选)
方法2:用base R
# 计算每个month+region的rp均值 rp_month_region_mean <- aggregate(rp ~ month + region, data = dc, FUN = mean, na.rm = TRUE) names(rp_month_region_mean)[3] <- "rp_mean" # 合并回原数据框并计算mean_rp dc <- merge(dc, rp_month_region_mean, by = c("month", "region")) dc$mean_rp <- dc$rp_mean * dc$year dc$rp_mean <- NULL # 移除中间列
2. 修复时序图异常
异常原因:原代码针对单个code(州)绘图,每个州的时间序列唯一且连续;但按region绘图时,同一区域下有多个州,直接subset会得到重复时间点的mean_rp值,导致线条乱跳。
方法1:用ggplot2(推荐,易维护)
install.packages("ggplot2") library(ggplot2) # 先确保每个区域每个时间点只有一个值 dc_agg <- dc %>% group_by(region, date) %>% summarise(mean_rp = first(mean_rp)) %>% ungroup() # 绘制时序图 ggplot(dc_agg, aes(x = date, y = mean_rp, color = region)) + geom_line() + labs(x = "时间", y = "水泥实际价格", title = "各区域mean_rp时序图") + scale_color_manual(values = c("N" = "black", "NE" = "red", "MW" = "yellow", "S" = "blue", "SE" = "green")) + theme_minimal()
方法2:用base R
# 按区域和日期排序,去重得到唯一的区域-时间组合 dc_sorted <- dc[order(dc$region, dc$date), ] dc_unique <- dc_sorted[!duplicated(dc_sorted[, c("region", "date")]), ] # 初始化绘图框架 plot(NA, xlim = range(dc_unique$date), ylim = c(0, 300), xlab = "时间", ylab = "水泥实际价格") # 定义区域与颜色对应关系 regions <- c("N", "NE", "MW", "S", "SE") colors <- c("black", "red", "yellow", "blue", "green") # 循环绘制每个区域的线条 for(i in seq_along(regions)){ sub_data <- subset(dc_unique, region == regions[i]) lines(sub_data$date, sub_data$mean_rp, col = colors[i]) } # 添加图例 legend("topleft", legend = regions, col = colors, lty = 1, inset = c(0.15, 0))
内容的提问来源于stack exchange,提问作者emilyinr
相关产品推荐
相关产品推荐

