如何用dplyr将treatment_implementation整合进词上下文分析数据集?
解决方法:合并word_concordances与main_df并分析MOM提及变化
1. 确认连接键
首先检查word_concordances和main_df的共同唯一标识列——quanteda的kwic()输出默认包含docname列,对应原始语料的文档ID/名称,你需要确认main_df中是否有匹配的列(比如同名的docname,或是document_id这类标识列)。
运行以下代码查看列名:
# 查看word_concordances的列 colnames(word_concordances) # 查看main_df的列 colnames(main_df)
2. 合并数据集
用dplyr::left_join()将main_df中的treatment_implementation变量合并到word_concordances中,确保指定正确的连接键:
library(dplyr) # 假设连接键是docname,按需替换为实际列名 merged_data <- word_concordances %>% left_join( main_df %>% select(docname, treatment_implementation), # 仅保留需要的列 by = "docname" # 若两边列名不同,用c("word_con_col" = "main_df_col")格式指定 )
3. 后续分析与可视化
合并完成后即可开展统计和可视化工作:
统计政策前后提及次数
# 统计政策实施前/后的MOM提及文档数(因你用了distinct去重,统计的是提及该术语的文档数量) mention_stats <- merged_data %>% group_by(treatment_implementation) %>% summarise(mention_count = n()) %>% mutate(treatment_label = case_when( treatment_implementation == 0 ~ "政策实施前", treatment_implementation == 1 ~ "政策实施后" ))
可视化政策前后差异
library(ggplot2) ggplot(mention_stats, aes(x = treatment_label, y = mention_count)) + geom_col(fill = "#4CAF50", width = 0.6) + labs(title = "MOM术语在政策实施前后的提及情况", x = "政策阶段", y = "提及文档数") + theme_light()
随时间的提及变化(若有时间变量)
如果main_df包含时间列(比如date),可以合并后分析时间趋势:
# 合并时间列 merged_data_with_time <- word_concordances %>% left_join( main_df %>% select(docname, treatment_implementation, date), by = "docname" ) %>% mutate(date = as.Date(date)) # 确保时间格式正确 # 按月聚合提及次数 time_trend <- merged_data_with_time %>% mutate(month = lubridate::floor_date(date, "month")) %>% group_by(month, treatment_implementation) %>% summarise(mention_count = n()) %>% ungroup() %>% mutate(treatment_label = case_when( treatment_implementation == 0 ~ "政策实施前", treatment_implementation == 1 ~ "政策实施后" )) # 绘制时间趋势图 ggplot(time_trend, aes(x = month, y = mention_count, color = treatment_label)) + geom_line(linewidth = 1) + labs(title = "MOM术语随时间的提及变化", x = "时间", y = "提及文档数", color = "政策阶段") + theme_minimal()
注意事项
- 若需要统计总提及次数而非提及文档数,需去掉代码中的
distinct(word_concordances, post,.keep_all = TRUE)步骤。 - 如果连接键不是
docname,请替换为两个数据集中实际对应的唯一标识列(比如行号rowid)。
内容的提问来源于stack exchange,提问作者nesta1990
相关产品推荐
相关产品推荐

