使用R识别调查中的直线作答:检测受访者11题全选1或2
在R中检测分组受访者的全一致回答(全1或全2)
我的数据集包含11055条观测与12个变量,每个受访者由rid标识,对应11组观测。我需要检测是否存在受访者的11项回答全部为"1"或全部为"2"——需求类似检测单行内多列值一致的场景,但这里是按受访者分组后,检查其所有观测的目标列是否全为1或全为2。
附数据集结构:
combineddata <- structure(list( rid = c(12, 12, 12, 12, 12, 12), design_row = c(23, 24, 25, 26, 27, 28), scenario = c(1, 2, 3, 4, 5, 6), seq = c(6, 5, 3, 10, 11, 9), choice = c(2, 2, 2, 2, 1, 1), qol = c(4, 4, 1, 3, 3, 2), life = c(4, 4, 4, 4, 3, 4), benefit = c(1, 3, 1, 3, 3, 1), drug_a_cert = c(3, 1, 3, 3, 2, 4), drug_a_wait = c(4, 1, 4, 4, 3, 3), drug_b_cert = c(2, 2, 2, 2, 1, 1), drug_b_wait = c(1, 4, 1, 1, 1, 2)), row.names = c(NA, -6L), class = c("tbl_df", "tbl", "data.frame"))
实现方案
完全可以在R中实现,以下是两种高效的实现方式,支持生成标记变量、统计符合条件的受访者数量等需求:
1. 使用dplyr包(代码简洁易维护)
假设需要检测的回答列是choice(实际使用时替换为你的目标列):
library(dplyr) # 为每条观测添加标记列,标记所属受访者是否全选1或全选2 result <- combineddata %>% group_by(rid) %>% mutate( all_one = all(choice == 1), # 全为1的标记 all_two = all(choice == 2), # 全为2的标记 straight_line = all_one | all_two # 符合任一条件的总标记 ) %>% ungroup() # 统计符合条件的受访者总数 straight_count <- result %>% distinct(rid, straight_line) %>% filter(straight_line) %>% nrow() # 输出结果 cat("符合条件的受访者数量:", straight_count, "\n")
2. 使用Base R(无需额外安装包)
# 按rid分组检查每组是否全为1或全为2 group_check <- tapply(combineddata$choice, combineddata$rid, function(x) { all(x == 1) | all(x == 2) }) # 转换为数据框便于查看 group_result <- data.frame( rid = as.integer(names(group_check)), straight_line = as.logical(group_check) ) # 统计符合条件的数量 straight_count_base <- sum(group_result$straight_line) # 输出结果 cat("符合条件的受访者数量:", straight_count_base, "\n")
扩展说明
- 如果需要检测多列(比如多个回答变量),可以使用
across()函数扩展dplyr代码:result <- combineddata %>% group_by(rid) %>% mutate(across(c(choice, qol, benefit), ~all(.x == 1) | all(.x == 2), .names = "{col}_straight")) %>% ungroup() - 代码中的
choice列仅为示例,请根据实际需求替换为你要检测的具体回答列名。
内容的提问来源于stack exchange,提问作者Bob_123
相关产品推荐
相关产品推荐

