如何简化DataFrame,标记checklist在另一列是否存在指定值
简化DataFrame:按checklist判断是否存在"X"值
需求:按checklist_id对原始DataFrame分组,判断每个分组的how_many_birds列中是否存在代表存在性的"X"值,最终生成仅包含checklist ID和判断结果(存在则为1,不存在则为0)的简化DataFrame。
示例数据
原始DataFrame(original_df)
checklist_id <- c(1,1,2,2,2,3,3,3,3) obs_id <- c("obs1", "obs2", "obs3", "obs4", "obs5", "obs6", "obs7", "obs8", "obs9") how_many_birds <- c('7','8','X','2','3','1','6','8','X') original_df <- data.frame(checklist_id, obs_id, how_many_birds)
输出:
checklist_id obs_id how_many_birds 1 1 obs1 7 2 1 obs2 8 3 2 obs3 X 4 2 obs4 2 5 2 obs5 3 6 3 obs6 1 7 3 obs7 6 8 3 obs8 8 9 3 obs9 X
目标DataFrame(goal_df)
checklist_id_goal <- c(1,2,3) at_least_one_x <- c(0,1,1) goal_df <- data.frame(checklist_id_goal, at_least_one_x)
输出:
checklist_id_goal at_least_one_x 1 1 0 2 2 1 3 3 1
解决方法
方法1:使用dplyr包分组聚合
library(dplyr) result_df <- original_df %>% group_by(checklist_id) %>% summarise(at_least_one_x = as.integer(any(how_many_birds == "X")), .groups = "drop")
方法2:使用base R实现
# 按checklist_id分组,判断每组是否包含X x_check <- tapply(original_df$how_many_birds, original_df$checklist_id, function(x) any(x == "X")) # 转换为目标格式的数据框 result_df <- data.frame( checklist_id = as.integer(names(x_check)), at_least_one_x = as.integer(x_check) )
两种方法生成的result_df与目标goal_df结构完全一致。
内容的提问来源于stack exchange,提问作者Rachael
相关产品推荐
相关产品推荐

