基于条件删除R数据框重复观测并新建richness列的实现方法
解决代码
以下实现完全匹配你的需求,1.27亿行大体积数据可以优先使用下方的优化方案,运算效率更高。
标准dplyr实现
# 加载依赖包 library(dplyr) result <- df %>% # 按地点+物种分组,处理同一地点下同一物种的多记录问题 group_by(location, species) %>% summarise( # 判断该物种在当前地点是否同时存在两种定植类型 has_both_type = n_distinct(establishment) == 2, # stress字段优先取natural行的非NA值,单种类型时直接取现有非NA值 stress = dplyr::if_else( has_both_type, stress[establishment == "natural" & !is.na(stress)][1], stress[!is.na(stress)][1] ), # maturity字段逻辑同stress maturity = dplyr::if_else( has_both_type, maturity[establishment == "natural" & !is.na(maturity)][1], maturity[!is.na(maturity)][1] ), # 存在两种类型时保留anthropogenic标识,否则保留原有标识 establishment = dplyr::if_else( has_both_type, "anthropogenic", first(establishment) ), .groups = "drop" ) %>% # 按地点分组统计物种总数,填充到组内所有行 group_by(location) %>% mutate(richness = n_distinct(species)) %>% ungroup()
大体积数据优化方案(适配1.27亿行规模)
用dtplyr实现dplyr语法转data.table执行,运算效率比标准dplyr高数倍:
# 加载依赖包 library(dplyr) library(dtplyr) library(data.table) result <- lazy_dt(df) %>% group_by(location, species) %>% summarise( has_both_type = n_distinct(establishment) == 2, stress = dplyr::if_else( has_both_type, stress[establishment == "natural" & !is.na(stress)][1], stress[!is.na(stress)][1] ), maturity = dplyr::if_else( has_both_type, maturity[establishment == "natural" & !is.na(maturity)][1], maturity[!is.na(maturity)][1] ), establishment = dplyr::if_else( has_both_type, "anthropogenic", first(establishment) ), .groups = "drop" ) %>% group_by(location) %>% mutate(richness = n_distinct(species)) %>% ungroup() %>% as.data.frame() # 运算结束后转为常规data.frame格式
说明
如果单个location+species分组下natural行存在多个非NA的stress/maturity值,代码默认取第一个出现的值,你可以根据需求调整取数逻辑(比如替换为取均值、最大值等)。所有因子类字段的层级会自动保留,输出结构和你给出的示例完全一致。
内容的提问来源于stack exchange,提问作者Jose
相关产品推荐
相关产品推荐

