如何基于source列匹配global_sources唯一值批量替换其NA值?
解决方案
步骤1:定义分类优先级
先从global_sources列提取所有已有的唯一分类,并将referral放在优先级末尾,确保其他分类优先匹配(比如google / referral会优先匹配google):
library(stringr) # 提取所有非NA的唯一分类 categories <- unique(df$global_sources[!is.na(df$global_sources)]) # 调整优先级:把referral移到最后 categories <- c(setdiff(categories, "referral"), "referral")
步骤2:批量替换NA值
遍历所有global_sources为NA的行,按优先级顺序检查source列是否包含分类,取第一个匹配的分类替换NA:
# 找到所有需要替换的行索引 na_rows <- which(is.na(df$global_sources)) # 逐个处理NA行 for(i in na_rows){ # 按优先级匹配分类 matched <- categories[str_detect(df$source[i], categories)] # 如果有匹配项,取第一个替换 if(length(matched) > 0){ df$global_sources[i] <- matched[1] } }
用dplyr简化实现
如果习惯用tidyverse语法,也可以用map_chr实现向量化操作,代码更简洁:
library(dplyr) library(purrr) df <- df %>% mutate( # 定义优先级分类列表 priority_cats = list(c(setdiff(unique(na.omit(global_sources)), "referral"), "referral")), # 替换NA值 global_sources = ifelse(is.na(global_sources), map_chr(source, ~{ match <- priority_cats[[1]][str_detect(.x, priority_cats[[1]])] ifelse(length(match) > 0, match[1], NA) }), global_sources) ) %>% select(-priority_cats)
最终结果
运行上述代码后,示例数据的global_sources列会被正确填充:
| source | global_sources |
|---|---|
| facebook / cpc | |
| googletest | |
| adwords.google | |
| organic | organic |
| source / referral | referral |
| google / source | |
| facebook / test |
内容的提问来源于stack exchange,提问作者Nicolás Rojas U
相关产品推荐
相关产品推荐

