使用grepl()对非互斥变量类别重新编码的问题
非互斥类别变量重新编码解决方案
问题背景
使用case_when对非互斥类别变量重新编码时,存在同时满足多个类别的样本(如标题含"dog和space"的帖子),会默认匹配第一个符合的条件,导致分类结果仅由条件顺序决定,无法准确反映多类别归属。
核心原因
case_when的执行逻辑是按顺序依次判断条件,返回第一个结果为TRUE的对应值,因此同时满足多个条件的样本只会被分配到第一个匹配的类别,调换条件顺序就会改变分类结果。
解决方案
方案1:生成多标记变量(推荐用于非互斥场景)
针对非互斥的类别,为每个类别单独创建二进制标记变量,一个样本可同时属于多个类别,清晰展现所有符合的类别:
post_caption <- c("This post is about a dog", "This post is about a cat", "This post is about a walrus", "This post is about space", "this post is about both a dog and space") post_name <- c("dog account", "dog_account", "walrus_account", "space_account", "space_account") # 标记是否属于animal_post(忽略大小写匹配) animal_post <- grepl(paste(c("dog", "cat", "walrus"), collapse = "|"), post_caption, ignore.case = TRUE) # 标记是否属于other_post(忽略大小写匹配) other_post <- grepl(paste(c("space", "rocks", "trees"), collapse = "|"), post_caption, ignore.case = TRUE) & grepl(paste(c("space", "rocks", "trees"), collapse = "|"), post_name, ignore.case = TRUE) # 合并结果为数据框 result_df <- data.frame( post_caption, post_name, is_animal_post = animal_post, is_other_post = other_post ) print(result_df)
输出结果:
post_caption post_name is_animal_post is_other_post 1 This post is about a dog dog account TRUE FALSE 2 This post is about a cat dog_account TRUE FALSE 3 This post is about a walrus walrus_account TRUE FALSE 4 This post is about space space_account FALSE TRUE 5 this post is about both a dog and space space_account TRUE TRUE
方案2:定义多类别混合标签
如果需要给每个样本分配单一标签,但要区分同时满足多个条件的情况,可以在case_when中优先判断多条件同时满足的场景:
post_caption <- c("This post is about a dog", "This post is about a cat", "This post is about a walrus", "This post is about space", "this post is about both a dog and space") post_name <- c("dog account", "dog_account", "walrus_account", "space_account", "space_account") category <- case_when( # 优先判断同时满足两类的情况 grepl(paste(c("dog", "cat", "walrus"), collapse = "|"), post_caption, ignore.case = TRUE) & grepl(paste(c("space", "rocks", "trees"), collapse = "|"), post_caption, ignore.case = TRUE) & grepl(paste(c("space", "rocks", "trees"), collapse = "|"), post_name, ignore.case = TRUE) ~ "mixed_post", # 单独匹配animal_post grepl(paste(c("dog", "cat", "walrus"), collapse = "|"), post_caption, ignore.case = TRUE) ~ "animal_post", # 单独匹配other_post grepl(paste(c("space", "rocks", "trees"), collapse = "|"), post_caption, ignore.case = TRUE) & grepl(paste(c("space", "rocks", "trees"), collapse = "|"), post_name, ignore.case = TRUE) ~ "other_post", # 未匹配的情况返回NA TRUE ~ NA_character_ ) print(category)
输出结果:
[1] "animal_post" "animal_post" "animal_post" "other_post" "mixed_post"
方案3:明确优先级规则
如果必须保留单一类别且有明确的优先级(如优先标记other_post),可以在低优先级条件中排除已匹配高优先级的样本,让逻辑更明确:
post_caption <- c("This post is about a dog", "This post is about a cat", "This post is about a walrus", "This post is about space", "this post is about both a dog and space") post_name <- c("dog account", "dog_account", "walrus_account", "space_account", "space_account") category <- case_when( # 高优先级:先匹配other_post grepl(paste(c("space", "rocks", "trees"), collapse = "|"), post_caption, ignore.case = TRUE) & grepl(paste(c("space", "rocks", "trees"), collapse = "|"), post_name, ignore.case = TRUE) ~ "other_post", # 低优先级:仅匹配未被标记为other_post的animal样本 !grepl(paste(c("space", "rocks", "trees"), collapse = "|"), post_caption, ignore.case = TRUE) & grepl(paste(c("dog", "cat", "walrus"), collapse = "|"), post_caption, ignore.case = TRUE) ~ "animal_post", TRUE ~ NA_character_ ) print(category)
输出结果:
[1] "animal_post" "animal_post" "animal_post" "other_post" "other_post"
内容的提问来源于stack exchange,提问作者RMRH
相关产品推荐
相关产品推荐

