You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言按井号、采样日期分组分析地下水DataFrame方法

R实现地下水检测数据分组判定方案

R完全支持类似Python的for循环遍历逻辑,且针对这类数千行规模的分组统计场景,还有更简洁高效的向量化实现方案,不需要手动拆分数据集后反复拼接,具体实现如下:

方案一:向量化分组计算(推荐,代码简洁性能好)

首先预处理列名,把带空格的原始字段名改成方便调用的形式,假设原始数据框名为gw_sample:

# 重命名列,对应原始四个字段:井号、采样日期、检测化合物、检测结果
colnames(gw_sample) <- c("well_id", "sample_date", "compound", "result")

dplyr实现

适合习惯tidyverse语法的场景:

library(dplyr)

calc_res <- gw_sample %>%
  group_by(well_id, sample_date) %>%
  # 核心逻辑:统计分组内结果>0.2的去重化合物数,判断是否≥2
  summarise(
    is_over = n_distinct(compound[result > 0.2]) >= 2,
    .groups = "drop"
  ) %>%
  # 拼接为要求的输出格式
  mutate(output = paste0(well_id, " ", sample_date, " -> ", tolower(as.character(is_over))))

最终要求格式的结果存储在calc_res$output向量中,直接打印即可得到类似A 2020-01-01 -> true的结果,其余不符合条件的分组自动返回false。

data.table实现

适合数据量更大的场景,性能最优:

library(data.table)
setDT(gw_sample)

calc_res <- gw_sample[, 
  .(is_over = uniqueN(compound[result > 0.2]) >= 2), 
  by = .(well_id, sample_date)
][, output := paste0(well_id, " ", sample_date, " -> ", tolower(as.character(is_over)))]

方案二:显式for循环实现(和Python遍历逻辑一致)

如果需要写和Python逻辑完全对齐的循环写法也可以实现,步骤是先取全部分组组合,再逐组遍历判断:

# 提取所有不重复的井号+采样日期分组
all_groups <- unique(gw_sample[, c("well_id", "sample_date")])
output <- vector("character", nrow(all_groups))

# 逐组遍历判断
for (i in seq_along(output)) {
  cur_well <- all_groups$well_id[i]
  cur_date <- all_groups$sample_date[i]
  # 筛选当前分组的子集
  sub_data <- gw_sample[gw_sample$well_id == cur_well & gw_sample$sample_date == cur_date, ]
  # 统计符合条件的去重化合物数量
  valid_compound_cnt <- length(unique(sub_data$compound[sub_data$result > 0.2]))
  # 拼接结果
  output[i] <- paste0(cur_well, " ", cur_date, " -> ", tolower(as.character(valid_compound_cnt >= 2)))
}

数千行数据规模下,这个循环写法的运行速度和向量化方案没有明显感知差异。

之前split方案的补全实现

如果已经用split拆分了数据集,只需要加一层列表遍历就能解决输出格式问题:

# 原split拆分步骤
split_data <- split(gw_sample, list(gw_sample$well_id, gw_sample$sample_date), drop = TRUE)
# 遍历拆分后的列表拼接结果
output <- sapply(names(split_data), function(group_name){
  current_sub <- split_data[[group_name]]
  flag <- length(unique(current_sub$compound[current_sub$result > 0.2])) >= 2
  # split默认用.连接分组名,替换回空格即可
  paste0(gsub("\\.", " ", group_name), " -> ", tolower(as.character(flag)))
})

注意点

  • R原生逻辑值输出为大写TRUE/FALSE,用tolower()转成小写即可和示例要求的true/false对齐
  • 如果采样日期列是Date/POSIX时间类型,拼接时会自动转为标准日期字符串,不需要额外格式化
  • 去重统计化合物数量时一定要加去重逻辑,避免同一化合物重复检测导致计数错误

内容的提问来源于stack exchange,提问作者Mailynn

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.29 15:48:36