使用dplyr按WELL_CATEG分组求均值仅返回单个值的问题
问题描述
我在R中使用dplyr计算分组条件均值时遇到异常:想按WELL_CATEG字段分组,计算每个因子水平下TCPNoZeros的均值(命名为TCP_MEAN),但执行代码后仅返回一个全局均值,而非预期的多分组结果。
执行的代码:
#mean_by_group df2 <- df1000 %>% group_by(WELL_CATEG) %>% summarise(TCP_MEAN = mean(TCPNoZeros, na.rm=T))
实际输出:
> df2 TCP_MEAN 1 0.193232
已知WELL_CATEG包含多个水平:"Domestic"、"Monitoring"、"Municipal"、"Water Supply, Other",应返回4组结果。数据集样本如下(完整数据含NA值):
structure(list(WELL_ID = c("5700542-001", "5700551-001", "5700552-001", "5700554-001", "5700571-011", "5700571-012", "5700575-001", "T0604700079-NW-2", "T0604700079-MW-9", "T0604700079-NW-3", "T0604700079-NW-4", "T0604700079-NW-9", "T0604700079-NW-8", "T0604700079-NW-6", "S4-TUSK-TLA02", "DM-U-03", "RED-06", "RICE-20", "RICE-09", "MODFP-03", "5700652-010", "USGS-361422119431201", "USGS-363645119420702", "USGS-364112119352701", "USGS-364258119380201", "USGS-363418119384203", "USGS-382718121224901", "USGS-372205120433801" ), WELL_CATEG = c("Domestic", "Domestic", "Domestic", "Domestic", "Domestic", "Domestic", "Domestic", "Monitoring", "Monitoring", "Monitoring", "Monitoring", "Monitoring", "Monitoring", "Monitoring", "Municipal", "Municipal", "Municipal", "Municipal", "Municipal", "Municipal", "Water Supply, Other", "Water Supply, Other", "Water Supply, Other", "Water Supply, Other", "Water Supply, Other", "Water Supply, Other", "Water Supply, Other", "Water Supply, Other"), Shape_Leng = c(6283.185307, 6283.185307, 6283.185307, 6283.185307, 6283.185307, 6283.185307, 6283.185307, 6283.185307, 6283.185307, 6283.185307, 6283.185307, 6283.185307, 6283.185307, 6283.185307, 6283.185307, 6283.185307, 6283.185307, 6283.185307, 6283.185307, 6283.185307, 6283.185307, 6283.185307, 6283.185307, 6283.185307, 6283.185307, 6283.185307, 6283.185307, 6283.185307), TCP_MEAN = c(0, 0, 0, 0, 0, 0, 0, 0.531914894, 0.535714286, 0.638297872, 0.714285714, 0.731707317, 0.731707317, 0.760869565, 0.006, 0.0625, 0.12, 0.18, 0.18, 0.18, 1, 2, 2, 2, 2, 4, 4, 6), TCPNoZeros = c(0.0025, 0.0025, 0.0025, 0.0025, 0.0025, 0.0025, 0.0025, 0.531914894, 0.535714286, 0.638297872, 0.714285714, 0.731707317, 0.731707317, 0.760869565, 0.006, 0.0625, 0.12, 0.18, 0.18, 0.18, 0.0025, 0.0025, 0.0025, 0.0025, 0.0025, 0.0025, 0.0025, 0.0025)), class = "data.frame", row.names = c(NA, -28L))
排查与解决步骤
1. 检查分组字段的有效性
先确认WELL_CATEG的类型和实际分组情况:
# 查看变量类型 class(df1000$WELL_CATEG) # 查看唯一分组值 unique(df1000$WELL_CATEG) # 统计各分组的样本量 table(df1000$WELL_CATEG)
如果unique()返回的结果符合预期,但table()显示只有一组,大概率是分组值存在隐形空格或特殊字符,导致分组被合并。可以用trimws()清理:
df1000$WELL_CATEG <- trimws(df1000$WELL_CATEG)
2. 确认dplyr版本与语法
旧版dplyr的summarise()默认会取消分组,但不会导致仅返回一个值。建议更新到最新版dplyr:
install.packages("dplyr") library(dplyr)
同时可以显式添加ungroup()(语法更清晰,不影响结果):
df2 <- df1000 %>% group_by(WELL_CATEG) %>% summarise(TCP_MEAN = mean(TCPNoZeros, na.rm = TRUE)) %>% ungroup()
3. 检查TCPNoZeros的变量类型
如果TCPNoZeros是列表列而非数值向量,mean()会计算全局均值。检查并转换类型:
# 查看变量类型 class(df1000$TCPNoZeros) # 若为列表,转换为数值向量 df1000$TCPNoZeros <- as.numeric(unlist(df1000$TCPNoZeros))
4. 用样本数据验证
用你提供的样本数据运行代码,正常会返回4组结果:
# 加载样本数据 sample_df <- structure(...) # 替换为你的样本数据结构 # 执行分组计算 sample_df %>% group_by(WELL_CATEG) %>% summarise(TCP_MEAN = mean(TCPNoZeros, na.rm = TRUE))
预期输出:
# A tibble: 4 × 2 WELL_CATEG TCP_MEAN <chr> <dbl> 1 Domestic 0.0025 2 Monitoring 0.654 3 Municipal 0.121 4 Water Supply, Other 0.0025
如果样本数据能得到正确结果,说明你的完整数据存在上述某类问题。
内容的提问来源于stack exchange,提问作者BHope
相关产品推荐
相关产品推荐

