You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用quanteda可视化KWIC结果的pre/post列及多词表达式

基于quanteda KWIC结果的上下文文本可视化方案

1. 提取并预处理pre/post列文本

首先需要把KWIC结果里的pre和post列文本合并,转换成可分析的格式,同时完成分词、去停用词等基础预处理:

library(quanteda)
library(wordcloud2)
library(ggplot2)

# 假设你的KWIC结果存储在kwic_result对象中
# 合并前后上下文文本
context_text <- c(kwic_result$pre, kwic_result$post)

# 构建语料库并做预处理
corpus_context <- corpus(context_text)
tokens_context <- tokens(corpus_context, remove_punct = TRUE, remove_numbers = TRUE) %>%
  tokens_tolower() %>%
  tokens_remove(stopwords("en")) # 中文文本替换为stopwords("zh"),也可自定义停用词

# 生成词频统计数据框
dfm_context <- dfm(tokens_context)
freq_df <- textstat_frequency(dfm_context)

2. 多词词云实现

通用词云

直接基于合并后的上下文文本生成词云:

wordcloud2(freq_df[, c("feature", "frequency")], size = 1.2, shape = "circle", backgroundColor = "#f8f9fa")

区分pre/post的分组词云

如果需要对比关键词前后的词汇差异,可以分别统计后用ggwordcloud做分组可视化:

library(ggwordcloud)

# 分别处理pre和post列
pre_tokens <- tokens(kwic_result$pre, remove_punct = TRUE, remove_numbers = TRUE) %>%
  tokens_tolower() %>%
  tokens_remove(stopwords("en"))
post_tokens <- tokens(kwic_result$post, remove_punct = TRUE, remove_numbers = TRUE) %>%
  tokens_tolower() %>%
  tokens_remove(stopwords("en"))

# 生成分组词频数据
pre_freq <- textstat_frequency(dfm(pre_tokens)) %>% mutate(source = "前上下文")
post_freq <- textstat_frequency(dfm(post_tokens)) %>% mutate(source = "后上下文")
combined_freq <- rbind(pre_freq, post_freq)

# 绘制分组词云
ggplot(combined_freq, aes(label = feature, size = frequency, color = source)) +
  geom_text_wordcloud(seed = 123) +
  scale_size_area(max_size = 15) +
  scale_color_manual(values = c("#3498db", "#e74c3c")) +
  theme_minimal()

3. 其他词频可视化图表

Top N高频词条形图

展示上下文里最常见的词汇:

# 取前20个高频词
top_20 <- head(freq_df, 20)

ggplot(top_20, aes(x = reorder(feature, frequency), y = frequency)) +
  geom_bar(stat = "identity", fill = "#2c3e50") +
  coord_flip() +
  labs(title = "上下文文本Top20高频词", x = "词汇", y = "词频") +
  theme_minimal()

pre/post词频对比条形图

直观对比关键词前后的高频词汇差异:

# 筛选前后上下文共有的Top15词汇
common_terms <- intersect(pre_freq$feature, post_freq$feature)[1:15]
pre_top <- pre_freq[pre_freq$feature %in% common_terms, ]
post_top <- post_freq[post_freq$feature %in% common_terms, ]

# 合并对比数据
compare_df <- merge(pre_top[, c("feature", "frequency")], post_top[, c("feature", "frequency")],
                    by = "feature", suffixes = c("_pre", "_post"))

# 绘制对比条形图
ggplot(compare_df, aes(x = reorder(feature, frequency_pre))) +
  geom_bar(aes(y = frequency_pre), stat = "identity", fill = "#3498db", alpha = 0.7) +
  geom_bar(aes(y = -frequency_post), stat = "identity", fill = "#e74c3c", alpha = 0.7) +
  coord_flip() +
  labs(title = "关键词前后上下文高频词对比", x = "词汇", y = "词频") +
  theme_minimal()

注意事项

  • 中文文本需替换停用词列表为stopwords("zh"),也可根据需求添加自定义停用词
  • 预处理阶段可按需加入词干提取(tokens_wordstem())进一步精简词汇
  • 词云和图表的样式可通过调整对应函数参数(如颜色、尺寸)优化展示效果

内容的提问来源于stack exchange,提问作者Juan Jose Echeverry De Mendoza

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 19:05:18