并行化时含变量条件格式的flextable循环渲染RMarkdown失败
问题描述
使用循环从带条件格式flextable的Rmd模板生成docx报告时,将串行for循环改为foreach() %dopar%并行运行后失败,报错“找不到对象'boolean'”。但用%do%串行运行完全正常,问题根源是flextable的条件格式公式中引用了主脚本定义的boolean变量。
复现示例
主脚本
library(dplyr) library(foreach) library(doParallel) boolean <- TRUE tables <- list() for (n in 1:10) { tables[[n]] <- iris %>% slice_sample(n = 20) } number_of_cores <- parallel::detectCores() - 1 clusters <- parallel::makeCluster(number_of_cores) doParallel::registerDoParallel(clusters) foreach( x = 1:10, .export = c("boolean", "tables"), .packages = c("dplyr", "flextable", "officer"), .verbose = TRUE ) %dopar% { rmarkdown::render("Template.rmd", output_file = paste0("Iris Subset ", x)) }
报错的Rmd模板(变量在flextable公式内)
--- title: "Iris Data" output: word_document date: "" --- ```{r table} flextable(tables[[x]]) %>% bg(i = ~ (Species == "setosa" & boolean), bg = "yellow", part = "body") %>% bg(i = ~ (Species == "versicolor" & !boolean), bg = "skyblue", part = "body") %>% autofit()
正常运行的Rmd模板(变量在flextable外)
--- title: "Iris Data" output: word_document date: "" --- ```{r table} if (boolean) { flextable(tables[[x]]) %>% bg(i = ~ (Species == "setosa"), bg = "yellow", part = "body") %>% autofit() } else { flextable(tables[[x]]) %>% bg(i = ~ (Species == "versicolor"), bg = "skyblue", part = "body") %>% autofit() }
原因分析
flextable的bg()函数中i参数的公式(~开头)会在公式自身的专属环境中求值,而非当前工作环境。并行运行时,子进程环境与主进程完全隔离,即便通过.export传递了boolean,公式仍无法在自身环境中找到该变量;而串行运行时,公式环境与工作环境一致,因此能正常识别变量。
解决办法
方法1:将变量注入数据框
把boolean作为列添加到目标表格中,让公式直接引用数据框内部的列,彻底避免依赖外部变量:
# 修改Rmd模板中的表格代码 flextable(tables[[x]] %>% mutate(boolean = !!boolean)) %>% bg(i = ~ (Species == "setosa" & boolean), bg = "yellow", part = "body") %>% bg(i = ~ (Species == "versicolor" & !boolean), bg = "skyblue", part = "body") %>% autofit()
方法2:用rlang::inject强制注入变量
使用rlang的注入功能,将外部变量直接插入公式,强制公式在当前环境中求值:
# 修改Rmd模板中的表格代码 library(rlang) flextable(tables[[x]]) %>% bg(i = inject(~ (Species == "setosa" & !!boolean)), bg = "yellow", part = "body") %>% bg(i = inject(~ (Species == "versicolor" & !!(boolean))), bg = "skyblue", part = "body") %>% autofit()
方法3:提前计算逻辑向量,弃用公式
直接计算需要高亮的行索引,将索引值传给i参数,完全避开公式环境的问题:
# 修改Rmd模板中的表格代码 ft <- flextable(tables[[x]]) # 计算需要高亮的行位置 setosa_rows <- which(tables[[x]]$Species == "setosa" & boolean) versicolor_rows <- which(tables[[x]]$Species == "versicolor" & !boolean) ft %>% bg(i = setosa_rows, bg = "yellow", part = "body") %>% bg(i = versicolor_rows, bg = "skyblue", part = "body") %>% autofit()
额外优化:通过params传递变量(替代.export)
可以不用.export,改用rmarkdown::render的params参数传递变量,逻辑更清晰,但仍需结合上述方法解决公式环境问题:
# 修改主脚本的foreach部分 foreach( x = 1:10, .packages = c("dplyr", "flextable", "officer", "rmarkdown"), .verbose = TRUE ) %dopar% { rmarkdown::render( "Template.rmd", output_file = paste0("Iris Subset ", x), params = list(boolean = boolean, table_data = tables[[x]]) ) } # 对应的Rmd模板开头添加params定义 --- title: "Iris Data" output: word_document date: "" params: boolean: TRUE table_data: NULL --- # 表格代码结合方法1(注入列)修改 ```{r table} flextable(params$table_data %>% mutate(boolean = params$boolean)) %>% bg(i = ~ (Species == "setosa" & boolean), bg = "yellow", part = "body") %>% bg(i = ~ (Species == "versicolor" & !boolean), bg = "skyblue", part = "body") %>% autofit()
内容的提问来源于stack exchange,提问作者Samuel Reichler
相关产品推荐
相关产品推荐

