使用quanteda生成统一数据框可视化词典术语频次的技术问询
问题描述
我正在分析数千篇报纸文章文本,打算构建医疗、税收、犯罪等议题词典,每个词典条目包含多个术语(比如医生、护士、医院等)。作为诊断环节,我想查看每个词典类别中占比最高的术语。目前我已经能单独打印各词典条目的top特征,但需要生成一个统一的数据框来做可视化,现有代码如下:
library(quanteda) # set path path_data <- system.file("extdata/", package = "readtext") # import csv file dat_inaug <- read.csv(paste0(path_data, "/csv/inaugCorpus.csv")) corp_inaug <- corpus(dat_inaug, text_field = "texts") corp_inaug %>% tokens(., remove_punct = T) %>% tokens_tolower() %>% tokens_select(., pattern=stopwords("en"), selection="remove")->tok # I have about eight or nine dictionaries dict<-dictionary(list(liberty=c("freedom", "free"), justice=c("justice", "law"))) # This produces a dfm of all the individual terms making up the dictionary tok %>% tokens_select(pattern=dict) %>% dfm() %>% topfeatures() # This produces the top features just making up the 'justice' dictionary entry tok %>% tokens_select(pattern=dict['justice']) %>% dfm() %>% topfeatures() # This gets me close to what I want, but I can't figure out how to collapse this now # to visualize which are the most frequent terms that are making up each dictionary category dict %>% map(., function(x) tokens_select(tok, pattern=x)) %>% map(., dfm) %>% map(., topfeatures)
解决方案
使用purrr::imap_dfr替代普通map,结合enframe将每个类别的top特征转换为带类别标识的数据框,最终合并成统一结构:
library(quanteda) library(purrr) library(tibble) library(dplyr) # 保留原有数据预处理代码 path_data <- system.file("extdata/", package = "readtext") dat_inaug <- read.csv(paste0(path_data, "/csv/inaugCorpus.csv")) corp_inaug <- corpus(dat_inaug, text_field = "texts") tok <- corp_inaug %>% tokens(remove_punct = TRUE) %>% tokens_tolower() %>% tokens_select(pattern = stopwords("en"), selection = "remove") # 定义词典 dict <- dictionary(list( liberty = c("freedom", "free"), justice = c("justice", "law") )) # 生成统一数据框 top_terms_df <- dict %>% imap_dfr(function(term_list, category_name) { # 处理单个类别术语 tok %>% tokens_select(pattern = term_list) %>% dfm() %>% topfeatures(n = Inf) %>% # n=Inf返回所有术语,可改为具体数值如10取Top10 enframe(name = "术语", value = "频次") %>% mutate(类别 = category_name) }) # 查看结果 print(top_terms_df)
关键说明:
imap_dfr:遍历词典时同时获取类别名和术语列表,并自动将结果按行合并成数据框enframe:将topfeatures返回的命名向量(术语:频次)转换为规范的数据框格式- 可通过修改
topfeatures(n = ...)控制每个类别返回的术语数量,方便后续可视化(比如ggplot2分组条形图)
内容的提问来源于stack exchange,提问作者spindoctor
相关产品推荐
相关产品推荐

