You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用R的tidytext+ggplot2实现三组作者词比例的三分面对比图?

生成三组作者词频比例对比的分面图

原代码仅固定以简·奥斯汀为Y轴生成两组对比,要补充H.G.威尔斯与勃朗特姐妹的对比并整合为3列分面图,需修改数据处理逻辑与可视化映射,完整代码如下:

library(gutenbergr)
library(janeaustenr)
library(dplyr)
library(stringr)
library(tidytext)
library(tidyr)
library(scales)
library(ggplot2)

# 原有文本预处理代码保持不变
original_books <- austen_books() %>%
  group_by(book) %>%
  mutate(linenumber = row_number(),
         chapter = cumsum(str_detect(text, 
                                     regex("^chapter [\\divxlc]",
                                           ignore_case = TRUE)))) %>%
  ungroup()

tidy_books <- original_books %>%
  unnest_tokens(word, text) %>%
  anti_join(stop_words)

hgwells <- gutenberg_download(c(35, 36, 5230, 159))
tidy_hgwells <- hgwells %>%
  unnest_tokens(word, text) %>%
  anti_join(stop_words)

bronte <- gutenberg_download(c(1260, 768, 969, 9182, 767))
tidy_bronte <- bronte %>%
  unnest_tokens(word, text) %>%
  anti_join(stop_words)

# 修改词频比例计算逻辑,生成所有作者的宽格式数据
wide_proportions <- bind_rows(
  mutate(tidy_bronte, author = "Brontë Sisters"),
  mutate(tidy_hgwells, author = "H.G. Wells"), 
  mutate(tidy_books, author = "Jane Austen")
) %>% 
  mutate(word = str_extract(word, "[a-z']+")) %>%
  count(author, word) %>%
  group_by(author) %>%
  mutate(proportion = n / sum(n)) %>% 
  select(-n) %>% 
  pivot_wider(names_from = author, values_from = proportion) %>%
  drop_na()  # 移除存在缺失值的单词,避免绘图报错

# 生成所有不重复的两两对比组合
comparison_pairs <- expand_grid(
  y_author = c("Jane Austen", "H.G. Wells", "Brontë Sisters"),
  x_author = c("Jane Austen", "H.G. Wells", "Brontë Sisters")
) %>% 
  filter(y_author > x_author)  # 保留无序组合,避免重复对比

# 转换为适合分面绘图的长格式数据
plot_data <- comparison_pairs %>%
  rowwise() %>%
  mutate(
    plot_data = list(
      wide_proportions %>%
        select(word, y = all_of(y_author), x = all_of(x_author)) %>%
        mutate(diff = abs(y - x))
    )
  ) %>%
  unnest(plot_data) %>%
  mutate(facet_title = paste(y_author, "vs", x_author))

# 绘制3列分面图
ggplot(plot_data, aes(x = x, y = y, color = diff)) +
  geom_abline(color = "gray40", lty = 2) +
  geom_jitter(alpha = 0.1, size = 2.5, width = 0.3, height = 0.3) +
  geom_text(aes(label = word), check_overlap = TRUE, vjust = 1.5) +
  scale_x_log10(labels = percent_format()) +
  scale_y_log10(labels = percent_format()) +
  scale_color_gradient(limits = c(0, 0.001), 
                       low = "darkslategray4", high = "gray75") +
  facet_wrap(~facet_title, ncol = 3) +
  theme(legend.position = "none") +
  labs(x = NULL, y = NULL)

关键修改说明

  • 宽格式比例数据:先将所有作者的词频比例整理为宽格式,方便后续任意两两组合的提取
  • 对比组合生成:用expand_grid生成所有可能的作者对,再过滤掉重复/自对比的组合,保留3组有效对比
  • 动态映射适配:通过rowwise和unnest将每个对比对的x/y值单独提取,让ggplot能动态适配不同分面的坐标轴
  • 分面布局调整:将facet_wrap的ncol设为3,实现3列并列展示所有对比图

内容的提问来源于stack exchange,提问作者Dijkie85

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 06:25:27