使用str_detect批量统计数据框中多个关键词的出现频率
前提假设
先统一命名你的数据结构,可根据你的实际变量名替换:
- 存储Youtube标题的dataframe命名为
yt_df,标题列名为title - 71个关键词的独立列表命名为
keyword_list
方案1:tidyverse 简洁实现(推荐)
无需手动写循环,用purrr批量处理即可,代码容错性更高:
library(tidyverse) result <- map_df(keyword_list, function(kw){ # 统计当前关键词的匹配结果 match_res <- str_detect(yt_df$title, fixed(kw, ignore_case = T)) match_sum <- summary(match_res) # 构造单行统计结果 tibble( `Video Title` = kw, `Freq True` = as.numeric(match_sum["TRUE"]), `Freq False` = as.numeric(match_sum["FALSE"]) ) })
参数说明:fixed()用来避免关键词里的正则特殊字符导致匹配报错,ignore_case = T是忽略大小写匹配,如果不需要这两个规则可以直接删掉。
方案2:基础R for循环实现(匹配你提到的循环需求)
完全用基础R语法实现,无需加载额外包:
# 先初始化空结果表,提前分配内存运行效率更高 result <- data.frame( `Video Title` = character(length(keyword_list)), `Freq True` = numeric(length(keyword_list)), `Freq False` = numeric(length(keyword_list)), check.names = F ) # 循环遍历每个关键词完成统计 for (i in seq_along(keyword_list)) { kw <- keyword_list[i] match_res <- grepl(kw, yt_df$title, ignore.case = T) match_sum <- summary(match_res) result[i,1] <- kw result[i,2] <- as.numeric(match_sum["TRUE"]) result[i,3] <- as.numeric(match_sum["FALSE"]) }
结果校验
运行完成后可执行head(result)查看前几行格式是否符合要求,正常情况下71个关键词对应71行统计结果,每行的Freq True+Freq False的和恒等于7234,可用来快速校验统计是否正确。
内容的提问来源于stack exchange,提问作者Teebs
相关产品推荐
相关产品推荐

