R语言按井号、采样日期分组分析地下水DataFrame方法
R实现地下水检测数据分组判定方案
R完全支持类似Python的for循环遍历逻辑,且针对这类数千行规模的分组统计场景,还有更简洁高效的向量化实现方案,不需要手动拆分数据集后反复拼接,具体实现如下:
方案一:向量化分组计算(推荐,代码简洁性能好)
首先预处理列名,把带空格的原始字段名改成方便调用的形式,假设原始数据框名为gw_sample:
# 重命名列,对应原始四个字段:井号、采样日期、检测化合物、检测结果 colnames(gw_sample) <- c("well_id", "sample_date", "compound", "result")
dplyr实现
适合习惯tidyverse语法的场景:
library(dplyr) calc_res <- gw_sample %>% group_by(well_id, sample_date) %>% # 核心逻辑:统计分组内结果>0.2的去重化合物数,判断是否≥2 summarise( is_over = n_distinct(compound[result > 0.2]) >= 2, .groups = "drop" ) %>% # 拼接为要求的输出格式 mutate(output = paste0(well_id, " ", sample_date, " -> ", tolower(as.character(is_over))))
最终要求格式的结果存储在calc_res$output向量中,直接打印即可得到类似A 2020-01-01 -> true的结果,其余不符合条件的分组自动返回false。
data.table实现
适合数据量更大的场景,性能最优:
library(data.table) setDT(gw_sample) calc_res <- gw_sample[, .(is_over = uniqueN(compound[result > 0.2]) >= 2), by = .(well_id, sample_date) ][, output := paste0(well_id, " ", sample_date, " -> ", tolower(as.character(is_over)))]
方案二:显式for循环实现(和Python遍历逻辑一致)
如果需要写和Python逻辑完全对齐的循环写法也可以实现,步骤是先取全部分组组合,再逐组遍历判断:
# 提取所有不重复的井号+采样日期分组 all_groups <- unique(gw_sample[, c("well_id", "sample_date")]) output <- vector("character", nrow(all_groups)) # 逐组遍历判断 for (i in seq_along(output)) { cur_well <- all_groups$well_id[i] cur_date <- all_groups$sample_date[i] # 筛选当前分组的子集 sub_data <- gw_sample[gw_sample$well_id == cur_well & gw_sample$sample_date == cur_date, ] # 统计符合条件的去重化合物数量 valid_compound_cnt <- length(unique(sub_data$compound[sub_data$result > 0.2])) # 拼接结果 output[i] <- paste0(cur_well, " ", cur_date, " -> ", tolower(as.character(valid_compound_cnt >= 2))) }
数千行数据规模下,这个循环写法的运行速度和向量化方案没有明显感知差异。
之前split方案的补全实现
如果已经用split拆分了数据集,只需要加一层列表遍历就能解决输出格式问题:
# 原split拆分步骤 split_data <- split(gw_sample, list(gw_sample$well_id, gw_sample$sample_date), drop = TRUE) # 遍历拆分后的列表拼接结果 output <- sapply(names(split_data), function(group_name){ current_sub <- split_data[[group_name]] flag <- length(unique(current_sub$compound[current_sub$result > 0.2])) >= 2 # split默认用.连接分组名,替换回空格即可 paste0(gsub("\\.", " ", group_name), " -> ", tolower(as.character(flag))) })
注意点
- R原生逻辑值输出为大写
TRUE/FALSE,用tolower()转成小写即可和示例要求的true/false对齐 - 如果采样日期列是Date/POSIX时间类型,拼接时会自动转为标准日期字符串,不需要额外格式化
- 去重统计化合物数量时一定要加去重逻辑,避免同一化合物重复检测导致计数错误
内容的提问来源于stack exchange,提问作者Mailynn
相关产品推荐
相关产品推荐

