You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用quanteda生成统一数据框可视化词典术语频次的技术问询

问题描述

我正在分析数千篇报纸文章文本,打算构建医疗、税收、犯罪等议题词典,每个词典条目包含多个术语(比如医生、护士、医院等)。作为诊断环节,我想查看每个词典类别中占比最高的术语。目前我已经能单独打印各词典条目的top特征,但需要生成一个统一的数据框来做可视化,现有代码如下:

library(quanteda)
# set path
path_data <- system.file("extdata/", package = "readtext")

# import csv file
dat_inaug <- read.csv(paste0(path_data, "/csv/inaugCorpus.csv"))
corp_inaug <- corpus(dat_inaug, text_field = "texts") 
corp_inaug %>% 
  tokens(., remove_punct = T) %>% 
  tokens_tolower() %>% 
  tokens_select(., pattern=stopwords("en"), selection="remove")->tok

# I have about eight or nine dictionaries 
dict<-dictionary(list(liberty=c("freedom", "free"), 
                      justice=c("justice", "law")))
# This produces a dfm of all the individual terms making up the dictionary
tok %>% 
  tokens_select(pattern=dict) %>% 
  dfm() %>% 
  topfeatures()
  
# This produces the top features just making up the 'justice' dictionary entry
tok %>% 
  tokens_select(pattern=dict['justice']) %>% 
  dfm() %>% 
  topfeatures()
# This gets me close to what I want, but I can't figure out how to collapse this now 
# to visualize which are the most frequent terms that are making up each dictionary category

dict %>% 
  map(., function(x) tokens_select(tok, pattern=x)) %>% 
  map(., dfm) %>% 
  map(., topfeatures) 
解决方案

使用purrr::imap_dfr替代普通map,结合enframe将每个类别的top特征转换为带类别标识的数据框,最终合并成统一结构:

library(quanteda)
library(purrr)
library(tibble)
library(dplyr)

# 保留原有数据预处理代码
path_data <- system.file("extdata/", package = "readtext")
dat_inaug <- read.csv(paste0(path_data, "/csv/inaugCorpus.csv"))
corp_inaug <- corpus(dat_inaug, text_field = "texts") 
tok <- corp_inaug %>% 
  tokens(remove_punct = TRUE) %>% 
  tokens_tolower() %>% 
  tokens_select(pattern = stopwords("en"), selection = "remove")

# 定义词典
dict <- dictionary(list(
  liberty = c("freedom", "free"), 
  justice = c("justice", "law")
))

# 生成统一数据框
top_terms_df <- dict %>%
  imap_dfr(function(term_list, category_name) {
    # 处理单个类别术语
    tok %>%
      tokens_select(pattern = term_list) %>%
      dfm() %>%
      topfeatures(n = Inf) %>% # n=Inf返回所有术语,可改为具体数值如10取Top10
      enframe(name = "术语", value = "频次") %>%
      mutate(类别 = category_name)
  })

# 查看结果
print(top_terms_df)

关键说明:

  • imap_dfr:遍历词典时同时获取类别名和术语列表,并自动将结果按行合并成数据框
  • enframe:将topfeatures返回的命名向量(术语:频次)转换为规范的数据框格式
  • 可通过修改topfeatures(n = ...)控制每个类别返回的术语数量,方便后续可视化(比如ggplot2分组条形图)

内容的提问来源于stack exchange,提问作者spindoctor

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 00:25:18