如何用dplyr筛选仅选normal的受访者及normal+负面情绪的受访者
解决方案
需求1:创建标识「仅选中normal」的变量
要筛选仅选中normal的受访者,需同时满足两个条件:
emotions_normal取值为1(选中normal)- 其余所有情绪列(愤怒、焦虑、紧张、压力、开心)的取值均为0(未选中)
用dplyr的mutate结合rowSums即可实现,通过计算除normal外所有列的总和,总和为0就代表其他情绪都没选:
emotion_chart <- emotion_chart %>% mutate(is_only_normal = (emotions_normal == 1) & (rowSums(select(., -emotions_normal)) == 0))
如果需要输出0/1格式而非逻辑值,将判断部分用as.integer()包裹即可:
emotion_chart <- emotion_chart %>% mutate(is_only_normal = as.integer((emotions_normal == 1) & (rowSums(select(., -emotions_normal)) == 0)))
需求2:创建标识「选中normal且至少一种负面情绪」的变量
先定义负面情绪列:emotions_angry、emotions_anxious、emotions_nervous、emotions_stressed,需满足:
emotions_normal取值为1- 上述负面情绪列中至少有一个取值为1
通过rowSums计算负面情绪列的总和,总和≥1就代表至少选中一种负面情绪:
# 定义负面情绪列名,方便后续调整范围 negative_emotions <- c("emotions_angry", "emotions_anxious", "emotions_nervous", "emotions_stressed") emotion_chart <- emotion_chart %>% mutate(is_normal_plus_negative = (emotions_normal == 1) & (rowSums(select(., all_of(negative_emotions))) >= 1))
同样,若需要0/1格式,用as.integer()包裹逻辑判断部分即可。
补充说明
rowSums适合快速处理多列二值变量的逐行组合判断select(., -emotions_normal)用于排除normal列,直接判断其余情绪的选中情况- 提前定义
negative_emotions向量,后续调整负面情绪范围时只需修改该向量,代码可维护性更高
内容的提问来源于stack exchange,提问作者debussyn
相关产品推荐
相关产品推荐

