如何基于字符串向量中的两类关键词重编码生成新向量?
问题:提取字符串向量中的目标关键词并生成二分类向量
示例数据
data <- tibble::tibble( w = c("Strongly disagree", "Somewhat disagree", "Disagree", "Somewhat agree", "Strongly agree", "Agree"), x = c("Definitely true", "Probably true", "Somewhat false", "Definitely false", "Definitely true", "Definitely false"), y = c("Definitely not doing enough", "Definitely doing enough", "Possibly not doing enough", "Possibly doing enough", "Definitely not doing enough", "Somehat doing enough"), z = c("Very comfortable", "Comfortable", "Somewhat comfortable", "Very uncomfortable", "Somewhat uncomfortable", "Comfortable") )
需求说明
每列字符串包含两类特定关键词(如w列的agree/disagree),需要生成仅保留这些关键词的新向量,处理后w列预期结果:
[1] "Disagree" "Disagree" "Disagree" "Agree" "Agree" "Agree"
现有代码问题
原函数make_dicho在处理无前缀的单关键词字符串(如"Disagree")时返回NA,原因是正则表达式(?<=\s).+仅匹配空格后的内容,无空格时无法匹配:
# 原函数 make_dicho <- function(df = data, var) { df %>% dplyr::pull({{ var }}) %>% haven::as_factor() %>% stringr::str_extract("(?<=\\s).+") %>% stringr::str_to_sentence() } # 测试结果(存在NA) data %>% make_dicho(., w) # [1] "Disagree" "Disagree" NA "Agree" "Agree" NA
解决方案
修改函数逻辑,直接匹配目标关键词而非依赖空格,同时支持不区分大小写的匹配,确保所有包含关键词的字符串都能被正确提取:
改进后的通用函数
make_dicho <- function(var, pattern) { var %>% haven::as_factor() %>% # 不区分大小写匹配目标关键词 stringr::str_extract(stringr::regex(pattern, ignore_case = TRUE)) %>% # 统一格式为首字母大写 stringr::str_to_sentence() }
在dplyr::mutate中使用
针对不同列传入对应的关键词模式:
processed_data <- data %>% mutate( w_dicho = make_dicho(w, "agree|disagree"), x_dicho = make_dicho(x, "true|false"), y_dicho = make_dicho(y, "doing enough|not doing enough"), z_dicho = make_dicho(z, "comfortable|uncomfortable") )
测试验证
处理w列的结果:
processed_data %>% pull(w_dicho) # [1] "Disagree" "Disagree" "Disagree" "Agree" "Agree" "Agree"
所有单关键词的情况都能正确提取,不再返回NA。
内容的提问来源于stack exchange,提问作者jdenn0514
相关产品推荐
相关产品推荐

