R语言:对分组数据应用基于THB标记的Dist条件生成布尔值
解决按组判断Dist与组内THB最小Dist的布尔标记问题
首先确保加载dplyr包:
library(dplyr)
正确处理代码(方法一:分组内直接计算)
直接在分组操作中,计算每组内Remark == "THB"的Dist最小值,再进行布尔判断:
Profile_processed <- Profile %>% group_by(Profil) %>% mutate( # 提取当前组内Remark为"THB"的Dist值,计算最小值(忽略NA) min_thb_dist = min(ifelse(Remark == "THB", Dist, NA), na.rm = TRUE), # 核心判断:当前记录的Dist <= 组内THB的最小Dist则标记为TRUE P = Dist <= min_thb_dist ) %>% ungroup() # 可选,取消分组状态
正确处理代码(方法二:先汇总再合并)
先单独提取THB记录,按Profil分组计算最小Dist,再合并到原数据集进行判断:
# 预计算每个Profil对应的THB最小Dist thb_min_summary <- Profile %>% filter(Remark == "THB") %>% group_by(Profil) %>% summarise(min_thb_dist = min(Dist)) %>% ungroup() # 合并数据并生成标记 Profile_processed <- Profile %>% left_join(thb_min_summary, by = "Profil") %>% mutate(P = Dist <= min_thb_dist)
之前代码的问题分析
- 第一段代码错误:
- 误用
max而非需求中的min - 写法
Dist %>% filter(Remark == "THB")不符合dplyr语法,无法在组内正确筛选对应Remark的Dist值
- 误用
- 第二段代码错误:
Int$Dist是全局的子集数据,没有按Profil分组匹配,导致所有组都使用了全局的max值,不符合按组判断的需求
内容的提问来源于stack exchange,提问作者C. Guff
相关产品推荐
相关产品推荐

