如何用quanteda可视化KWIC结果的pre/post列及多词表达式
基于quanteda KWIC结果的上下文文本可视化方案
1. 提取并预处理pre/post列文本
首先需要把KWIC结果里的pre和post列文本合并,转换成可分析的格式,同时完成分词、去停用词等基础预处理:
library(quanteda) library(wordcloud2) library(ggplot2) # 假设你的KWIC结果存储在kwic_result对象中 # 合并前后上下文文本 context_text <- c(kwic_result$pre, kwic_result$post) # 构建语料库并做预处理 corpus_context <- corpus(context_text) tokens_context <- tokens(corpus_context, remove_punct = TRUE, remove_numbers = TRUE) %>% tokens_tolower() %>% tokens_remove(stopwords("en")) # 中文文本替换为stopwords("zh"),也可自定义停用词 # 生成词频统计数据框 dfm_context <- dfm(tokens_context) freq_df <- textstat_frequency(dfm_context)
2. 多词词云实现
通用词云
直接基于合并后的上下文文本生成词云:
wordcloud2(freq_df[, c("feature", "frequency")], size = 1.2, shape = "circle", backgroundColor = "#f8f9fa")
区分pre/post的分组词云
如果需要对比关键词前后的词汇差异,可以分别统计后用ggwordcloud做分组可视化:
library(ggwordcloud) # 分别处理pre和post列 pre_tokens <- tokens(kwic_result$pre, remove_punct = TRUE, remove_numbers = TRUE) %>% tokens_tolower() %>% tokens_remove(stopwords("en")) post_tokens <- tokens(kwic_result$post, remove_punct = TRUE, remove_numbers = TRUE) %>% tokens_tolower() %>% tokens_remove(stopwords("en")) # 生成分组词频数据 pre_freq <- textstat_frequency(dfm(pre_tokens)) %>% mutate(source = "前上下文") post_freq <- textstat_frequency(dfm(post_tokens)) %>% mutate(source = "后上下文") combined_freq <- rbind(pre_freq, post_freq) # 绘制分组词云 ggplot(combined_freq, aes(label = feature, size = frequency, color = source)) + geom_text_wordcloud(seed = 123) + scale_size_area(max_size = 15) + scale_color_manual(values = c("#3498db", "#e74c3c")) + theme_minimal()
3. 其他词频可视化图表
Top N高频词条形图
展示上下文里最常见的词汇:
# 取前20个高频词 top_20 <- head(freq_df, 20) ggplot(top_20, aes(x = reorder(feature, frequency), y = frequency)) + geom_bar(stat = "identity", fill = "#2c3e50") + coord_flip() + labs(title = "上下文文本Top20高频词", x = "词汇", y = "词频") + theme_minimal()
pre/post词频对比条形图
直观对比关键词前后的高频词汇差异:
# 筛选前后上下文共有的Top15词汇 common_terms <- intersect(pre_freq$feature, post_freq$feature)[1:15] pre_top <- pre_freq[pre_freq$feature %in% common_terms, ] post_top <- post_freq[post_freq$feature %in% common_terms, ] # 合并对比数据 compare_df <- merge(pre_top[, c("feature", "frequency")], post_top[, c("feature", "frequency")], by = "feature", suffixes = c("_pre", "_post")) # 绘制对比条形图 ggplot(compare_df, aes(x = reorder(feature, frequency_pre))) + geom_bar(aes(y = frequency_pre), stat = "identity", fill = "#3498db", alpha = 0.7) + geom_bar(aes(y = -frequency_post), stat = "identity", fill = "#e74c3c", alpha = 0.7) + coord_flip() + labs(title = "关键词前后上下文高频词对比", x = "词汇", y = "词频") + theme_minimal()
注意事项
- 中文文本需替换停用词列表为
stopwords("zh"),也可根据需求添加自定义停用词 - 预处理阶段可按需加入词干提取(
tokens_wordstem())进一步精简词汇 - 词云和图表的样式可通过调整对应函数参数(如颜色、尺寸)优化展示效果
内容的提问来源于stack exchange,提问作者Juan Jose Echeverry De Mendoza
相关产品推荐
相关产品推荐

