You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用dplyr将treatment_implementation整合进词上下文分析数据集?

解决方法:合并word_concordances与main_df并分析MOM提及变化

1. 确认连接键

首先检查word_concordances和main_df的共同唯一标识列——quanteda的kwic()输出默认包含docname列,对应原始语料的文档ID/名称,你需要确认main_df中是否有匹配的列(比如同名的docname,或是document_id这类标识列)。

运行以下代码查看列名:

# 查看word_concordances的列
colnames(word_concordances)
# 查看main_df的列
colnames(main_df)

2. 合并数据集

用dplyr::left_join()将main_df中的treatment_implementation变量合并到word_concordances中,确保指定正确的连接键:

library(dplyr)

# 假设连接键是docname,按需替换为实际列名
merged_data <- word_concordances %>%
  left_join(
    main_df %>% select(docname, treatment_implementation), # 仅保留需要的列
    by = "docname" # 若两边列名不同,用c("word_con_col" = "main_df_col")格式指定
  )

3. 后续分析与可视化

合并完成后即可开展统计和可视化工作:

统计政策前后提及次数

# 统计政策实施前/后的MOM提及文档数(因你用了distinct去重,统计的是提及该术语的文档数量)
mention_stats <- merged_data %>%
  group_by(treatment_implementation) %>%
  summarise(mention_count = n()) %>%
  mutate(treatment_label = case_when(
    treatment_implementation == 0 ~ "政策实施前",
    treatment_implementation == 1 ~ "政策实施后"
  ))

可视化政策前后差异

library(ggplot2)

ggplot(mention_stats, aes(x = treatment_label, y = mention_count)) +
  geom_col(fill = "#4CAF50", width = 0.6) +
  labs(title = "MOM术语在政策实施前后的提及情况",
       x = "政策阶段",
       y = "提及文档数") +
  theme_light()

随时间的提及变化(若有时间变量)

如果main_df包含时间列(比如date),可以合并后分析时间趋势:

# 合并时间列
merged_data_with_time <- word_concordances %>%
  left_join(
    main_df %>% select(docname, treatment_implementation, date),
    by = "docname"
  ) %>%
  mutate(date = as.Date(date)) # 确保时间格式正确

# 按月聚合提及次数
time_trend <- merged_data_with_time %>%
  mutate(month = lubridate::floor_date(date, "month")) %>%
  group_by(month, treatment_implementation) %>%
  summarise(mention_count = n()) %>%
  ungroup() %>%
  mutate(treatment_label = case_when(
    treatment_implementation == 0 ~ "政策实施前",
    treatment_implementation == 1 ~ "政策实施后"
  ))

# 绘制时间趋势图
ggplot(time_trend, aes(x = month, y = mention_count, color = treatment_label)) +
  geom_line(linewidth = 1) +
  labs(title = "MOM术语随时间的提及变化",
       x = "时间",
       y = "提及文档数",
       color = "政策阶段") +
  theme_minimal()

注意事项

  • 若需要统计总提及次数而非提及文档数,需去掉代码中的distinct(word_concordances, post,.keep_all = TRUE)步骤。
  • 如果连接键不是docname,请替换为两个数据集中实际对应的唯一标识列(比如行号rowid)。

内容的提问来源于stack exchange,提问作者nesta1990

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.08 10:35:55