使用mutate、case_when、%in%重编码长字符串变量部分匹配的问题
解决方案:无需拆分字符串,直接用子字符串匹配重编码
你之前的代码失效是因为%in%做的是精确全匹配,而你的post_caption是长字符串,显然不会和"dog"这类短字符串完全相等。不需要拆分变量,直接用字符串检测函数就能实现需求。
核心方法:用stringr::str_detect检测子字符串
结合dplyr::case_when,通过str_detect判断长字符串中是否包含目标子串,从而重编码post_type。
示例代码
先构造你的示例数据框:
library(dplyr) library(stringr) post_info <- tibble( post_type = c("type1", "type2", "type3", "type4"), post_caption = c("This post is about a dog", "This post is about a cat", "This post is about a walrus", "This post is about space") )
然后执行重编码:
post_info %>% mutate(post_type_r = case_when( # 检测是否包含任意动物关键词 str_detect(post_caption, regex("dog|cat|walrus", ignore_case = TRUE)) ~ "animal_post", # 检测是否包含任意其他类关键词 str_detect(post_caption, regex("space|rocks|trees", ignore_case = TRUE)) ~ "other_post", # 其他情况保留原post_type TRUE ~ post_type ))
进阶优化:匹配完整单词(避免误匹配)
如果需要确保匹配的是完整单词(比如不想把"doggy"识别成"dog"),可以加上单词边界\\b:
post_info %>% mutate(post_type_r = case_when( str_detect(post_caption, regex("\\bdog\\b|\\bcat\\b|\\bwalrus\\b", ignore_case = TRUE)) ~ "animal_post", str_detect(post_caption, regex("\\bspace\\b|\\brocks\\b|\\btrees\\b", ignore_case = TRUE)) ~ "other_post", TRUE ~ post_type ))
运行结果
执行后会得到新变量post_type_r,对应结果为:"animal_post", "animal_post", "animal_post", "other_post",完全符合需求。
内容的提问来源于stack exchange,提问作者RMRH
相关产品推荐
相关产品推荐

