You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用dplyr按WELL_CATEG分组求均值仅返回单个值的问题

问题描述

我在R中使用dplyr计算分组条件均值时遇到异常:想按WELL_CATEG字段分组,计算每个因子水平下TCPNoZeros的均值(命名为TCP_MEAN),但执行代码后仅返回一个全局均值,而非预期的多分组结果。

执行的代码:

#mean_by_group
df2 <- df1000 %>% group_by(WELL_CATEG) %>% summarise(TCP_MEAN = mean(TCPNoZeros, na.rm=T))

实际输出:

> df2
 TCP_MEAN
 1 0.193232

已知WELL_CATEG包含多个水平:"Domestic"、"Monitoring"、"Municipal"、"Water Supply, Other",应返回4组结果。数据集样本如下(完整数据含NA值):

structure(list(WELL_ID = c("5700542-001", "5700551-001", "5700552-001", 
"5700554-001", "5700571-011", "5700571-012", "5700575-001", "T0604700079-NW-2", 
"T0604700079-MW-9", "T0604700079-NW-3", "T0604700079-NW-4", "T0604700079-NW-9", 
"T0604700079-NW-8", "T0604700079-NW-6", "S4-TUSK-TLA02", "DM-U-03", 
"RED-06", "RICE-20", "RICE-09", "MODFP-03", "5700652-010", "USGS-361422119431201", 
"USGS-363645119420702", "USGS-364112119352701", "USGS-364258119380201", 
"USGS-363418119384203", "USGS-382718121224901", "USGS-372205120433801"
), WELL_CATEG = c("Domestic", "Domestic", "Domestic", "Domestic", 
"Domestic", "Domestic", "Domestic", "Monitoring", "Monitoring", 
"Monitoring", "Monitoring", "Monitoring", "Monitoring", "Monitoring", 
"Municipal", "Municipal", "Municipal", "Municipal", "Municipal", 
"Municipal", "Water Supply, Other", "Water Supply, Other", "Water Supply, Other", 
"Water Supply, Other", "Water Supply, Other", "Water Supply, Other", 
"Water Supply, Other", "Water Supply, Other"), Shape_Leng = c(6283.185307, 
6283.185307, 6283.185307, 6283.185307, 6283.185307, 6283.185307, 
6283.185307, 6283.185307, 6283.185307, 6283.185307, 6283.185307, 
6283.185307, 6283.185307, 6283.185307, 6283.185307, 6283.185307, 
6283.185307, 6283.185307, 6283.185307, 6283.185307, 6283.185307, 
6283.185307, 6283.185307, 6283.185307, 6283.185307, 6283.185307, 
6283.185307, 6283.185307), TCP_MEAN = c(0, 0, 0, 0, 0, 0, 0, 
0.531914894, 0.535714286, 0.638297872, 0.714285714, 0.731707317, 
0.731707317, 0.760869565, 0.006, 0.0625, 0.12, 0.18, 0.18, 0.18, 
1, 2, 2, 2, 2, 4, 4, 6), TCPNoZeros = c(0.0025, 0.0025, 0.0025, 
0.0025, 0.0025, 0.0025, 0.0025, 0.531914894, 0.535714286, 0.638297872, 
0.714285714, 0.731707317, 0.731707317, 0.760869565, 0.006, 0.0625, 
0.12, 0.18, 0.18, 0.18, 0.0025, 0.0025, 0.0025, 0.0025, 0.0025, 
0.0025, 0.0025, 0.0025)), class = "data.frame", row.names = c(NA, 
-28L))
排查与解决步骤

1. 检查分组字段的有效性

先确认WELL_CATEG的类型和实际分组情况:

# 查看变量类型
class(df1000$WELL_CATEG)
# 查看唯一分组值
unique(df1000$WELL_CATEG)
# 统计各分组的样本量
table(df1000$WELL_CATEG)

如果unique()返回的结果符合预期,但table()显示只有一组,大概率是分组值存在隐形空格或特殊字符,导致分组被合并。可以用trimws()清理:

df1000$WELL_CATEG <- trimws(df1000$WELL_CATEG)

2. 确认dplyr版本与语法

旧版dplyr的summarise()默认会取消分组,但不会导致仅返回一个值。建议更新到最新版dplyr:

install.packages("dplyr")
library(dplyr)

同时可以显式添加ungroup()(语法更清晰,不影响结果):

df2 <- df1000 %>% 
  group_by(WELL_CATEG) %>% 
  summarise(TCP_MEAN = mean(TCPNoZeros, na.rm = TRUE)) %>%
  ungroup()

3. 检查TCPNoZeros的变量类型

如果TCPNoZeros是列表列而非数值向量,mean()会计算全局均值。检查并转换类型:

# 查看变量类型
class(df1000$TCPNoZeros)
# 若为列表,转换为数值向量
df1000$TCPNoZeros <- as.numeric(unlist(df1000$TCPNoZeros))

4. 用样本数据验证

用你提供的样本数据运行代码,正常会返回4组结果:

# 加载样本数据
sample_df <- structure(...) # 替换为你的样本数据结构
# 执行分组计算
sample_df %>% 
  group_by(WELL_CATEG) %>% 
  summarise(TCP_MEAN = mean(TCPNoZeros, na.rm = TRUE))

预期输出:

# A tibble: 4 × 2
  WELL_CATEG          TCP_MEAN
  <chr>                  <dbl>
1 Domestic             0.0025 
2 Monitoring           0.654  
3 Municipal            0.121  
4 Water Supply, Other  0.0025 

如果样本数据能得到正确结果,说明你的完整数据存在上述某类问题。


内容的提问来源于stack exchange,提问作者BHope

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 09:55:07