基于字段特定值选择列:dplyr select(where)报错排查
问题
我正在基于调查数据创建表格,该调查要求受访者对活动各元素的满意度进行评分。若受访者对某活动元素表示不满意,需填写自由文本子问题说明不满原因。
目前我已将所有活动元素及其评分格式化为单独列,同时保留了通用的自由文本不满反馈字段,并过滤掉了该字段中的NA值:
# 当前输出 Response_ID | Reasons_for_Dissatisfaction | Rank_Food | Rank_Facilities | Rank_Content ---------------------------------------------------------------------------------------------------------- 1 | <Free-text feedback> | Somewhat dissatisfied | Very dissatisfied | Very satisfied ---------------------------------------------------------------------------------------------------------- 2 | <Free-text feedback> | Very dissatisfied | Somewhat dissatisfied | Very satisfied
我希望表格仅显示受访者“Somewhat dissatisfied”(有点不满意)或“Very dissatisfied”(非常不满意)的活动元素列,同时保留自由文本形式的评分原因列,目标输出如下:
# 目标输出 Response_ID | Reasons_for_Dissatisfaction | Rank_Food | Rank_Facilities ----------------------------------------------------------------------------------------- 1 | <Free-text feedback> | Somewhat dissatisfied | Very dissatisfied ----------------------------------------------------------------------------------------- 2 | <Free-text feedback> | Very dissatisfied | Somewhat dissatisfied
我需要代码能随数据集更新自动调整:若新增反馈中出现新的不满意活动元素,需自动将该元素列及对应记录加入表格,因此不能使用绝对列引用。我尝试用select(where())筛选仅包含“Somewhat dissatisfied”或“Very dissatisfied”值的列,同时保留“Response_ID”和“Reasons_for_Dissatisfaction”列,但运行报错:
# 示例代码 satisfactionrank <- data.frame( `Response_ID`=c(1,2), `Reasons_for_Dissatisfaction`=c("<Free-text feedback>","<Free-text feedback>"), `Rank_Food`=c("Somewhat dissatisfied","Very dissatisfied"), `Rank_Facilities`=c("Very dissatisfied","Somewhat dissatisfied"), `Rank_Content`=c("Very satisfied","Very satisfied") ) eventdissatisfaction <- satisfactionrank %>% select( `Response_ID`, `Reasons_for_Dissatisfaction`, where(contains("Rank") & any(. == "Somewhat dissatisfied") | any(. == "Very dissatisfied")) )
报错信息:
Error in `select()`: ℹ In argument: `where(...)`. Caused by error in `where()`: ! Can't convert `fn`, a logical vector, to a function. Run `rlang::last_trace()` to see where the error occurred.
请问我的代码哪里出错了?
解决方案
错误原因
你的代码存在两个核心问题:
where()使用逻辑错误:where()要求传入一个返回逻辑值的函数(或匿名函数)来判断列是否符合条件,但你直接将列名筛选函数contains("Rank")与列值判断逻辑混用,导致返回的是逻辑向量而非函数,触发类型转换错误。- 逻辑运算符优先级问题:
&的优先级高于|,原表达式会被错误解析,无法正确匹配你需要的条件组合。
修正后的代码
library(dplyr) satisfactionrank <- data.frame( `Response_ID`=c(1,2), `Reasons_for_Dissatisfaction`=c("<Free-text feedback>","<Free-text feedback>"), `Rank_Food`=c("Somewhat dissatisfied","Very dissatisfied"), `Rank_Facilities`=c("Very dissatisfied","Somewhat dissatisfied"), `Rank_Content`=c("Very satisfied","Very satisfied") ) eventdissatisfaction <- satisfactionrank %>% select( Response_ID, Reasons_for_Dissatisfaction, # 先匹配列名含"Rank"的列,再判断该列是否存在不满意评分 where(~ grepl("Rank", colnames(.)) & any(. %in% c("Somewhat dissatisfied", "Very dissatisfied"))) ) print(eventdissatisfaction)
代码说明
where(~ ...):用波浪线定义匿名函数,.指代当前列的所有值,colnames(.)指代当前列的列名。grepl("Rank", colnames(.)):精准筛选所有活动元素评分列(列名含"Rank")。any(. %in% c(...)):判断当前列中是否存在任意一个"有点不满意"或"非常不满意"的评分,存在则保留该列。- 逻辑分组:通过括号明确条件执行顺序,确保先筛选列名,再判断列值,同时用
%in%简化多值匹配的写法。
输出结果
Response_ID Reasons_for_Dissatisfaction Rank_Food Rank_Facilities 1 1 <Free-text feedback> Somewhat dissatisfied Very dissatisfied 2 2 <Free-text feedback> Very dissatisfied Somewhat dissatisfied
该代码可自动适配新增的Rank_*列:只要新列中存在不满意评分,就会被自动加入结果表格。
内容的提问来源于stack exchange,提问作者Mary Rachel
相关产品推荐
相关产品推荐

