如何在R的quanteda中同时统计单个词与词典词频?
Got it! 你想要把单个目标词的统计和自定义词组的频次统计合并成一步完成,用quanteda就能轻松实现,我给你两种靠谱的方法:
方法一:构建包含单个词和词组的统一词典
最简洁的方式是把需要单独统计的词和词组都放进同一个词典里,每个单独的词作为独立的词条,词组归为一个统一的类别,一次调用dfm()就能得到所有结果:
library("quanteda") txt <- "In west Philadelphia born and raised On the playground was where I spent most of my days Chillin' out maxin' relaxin' all cool And all shootin some b-ball outside of the school When a couple of guys who were up to no good Started making trouble in my neighborhood I got in one little fight and my mom got scared." # 构建合并后的词典:单独词各成一项,词组归为all_terms combined_dict <- dictionary( list( trouble = "trouble", fight = "fight", all_terms = c("chillin", "relaxin", "shootin") ) ) # 生成dfm并查看结果 result_dfm <- dfm(txt, dictionary = combined_dict) result_dfm
运行后输出:
Document-feature matrix of: 1 document, 3 features (0.0% sparse). features docs trouble fight all_terms text1 1 1 3
方法二:管道流中灵活合并统计
如果你更习惯用tokens的管道操作流程,可以先筛选出所有目标词,再把词组合并成一个统一特征:
library("quanteda") txt <- "In west Philadelphia born and raised On the playground was where I spent most of my days Chillin' out maxin' relaxin' all cool And all shootin some b-ball outside of the school When a couple of guys who were up to no good Started making trouble in my neighborhood I got in one little fight and my mom got scared." # 定义需要统计的单个词和词组 single_terms <- c("trouble", "fight") phrase_terms <- c("chillin", "relaxin", "shootin") # 一站式管道处理 result <- tokens(txt) %>% # 筛选出所有目标词(单个词+词组) tokens_select(pattern = c(single_terms, phrase_terms)) %>% # 转换为dfm dfm() %>% # 把词组替换为统一的all_terms标签 dfm_replace(pattern = phrase_terms, replacement = "all_terms") %>% # 压缩合并相同标签的频次 dfm_compress() result
这个方法的输出和方法一完全一致,适合需要在流程中动态调整词列表的场景。
你可以根据自己的使用习惯选择任意一种方法,都能实现你想要的“一步完成单个词+词组总频次统计”的需求~
内容的提问来源于stack exchange,提问作者Kasi
相关产品推荐
相关产品推荐

