如何自动化R脚本替换目标词汇 批量生成词频占比趋势图
R脚本自动化生成词汇占比趋势图方案
实现思路
将固定的绘图逻辑封装为自定义函数,把目标词、可选的坐标轴范围等作为可传入参数,所有原本硬编码目标词的位置都替换为参数变量,后续调用时仅需传入目标词汇即可自动完成全流程计算与绘图,无需反复修改核心代码。
完整可复用代码
# 自定义绘图函数,参数说明: # target_word:待分析的目标词汇 # ylim:y轴取值范围,默认和原脚本逻辑一致,可按需调整 plot_word_prop <- function(target_word, ylim = c(0, 0.005)) { # 计算目标词年度占比 dfm_target <- dfm(corpus_toks) %>% dfm_group(groups = "Year") %>% dfm_weight(scheme = "prop") %>% dfm_select(pattern = target_word) colnames(dfm_target)[1] <- target_word dfm_target2 <- convert(dfm_target, to = "data.frame") dfm_target2$doc_id <- as.numeric(dfm_target2$doc_id) # 数据精度处理 dfm_target2[[target_word]] <- round(dfm_target2[[target_word]], digits = 5) dfm_target3 <- reshape2::melt(dfm_target2, id.vars = "doc_id") options(scipen = 999) # 生成趋势图 target_plot <- ggplot(data = dfm_target3, aes(x = doc_id, y = value)) + geom_line(colour = "red") + scale_x_continuous(limits = c(1999, 2019), breaks = seq(1999, 2019, 1)) + scale_y_continuous(limits = ylim, breaks = seq(ylim[1], ylim[2], 0.001)) + labs(title = paste0(stringr::str_to_title(target_word), " by proportions"), x = "年份", y = "占比") + theme(axis.text.x = element_text(angle = 45, hjust = 1)) # 返回绘图对象供后续调整/保存 return(target_plot) }
使用示例
- 单个词汇绘图:直接传入目标词即可生成对应图表
# 生成research的占比趋势图 research_plot <- plot_word_prop("research") # 预览图表 print(research_plot) # 保存图表到本地 ggsave(paste0("research_占比趋势图.png"), research_plot, width = 10, height = 6)
- 批量生成多个词汇的图表:将所有待处理词汇存入向量后循环调用即可
# 定义需要批量分析的词汇列表 word_list <- c("research", "science", "technology", "education") # 批量生成图表并自动保存到本地,所有绘图结果存入plot_list列表 plot_list <- lapply(word_list, function(w) { p <- plot_word_prop(w) ggsave(paste0(w, "_占比趋势图.png"), p, width = 10, height = 6) return(p) }) # 查看列表中第二个词汇的趋势图 print(plot_list[[2]])
内容的提问来源于stack exchange,提问作者J.Doe
相关产品推荐
相关产品推荐

