R正则查找第n次匹配项实现LaTeX表格按同尺寸矩阵规则着色
实现思路
不用硬写正则匹配第n个单元格,采用「提取内容-拆分单元格-逐位匹配着色-拼接回写」的逻辑,容错性远高于直接正则替换,适配绝大多数无合并单元格的tabular场景:
- 从输入的LaTeX代码中提取tabular环境内部的内容行
- 逐行拆分得到所有单元格内容,和同尺寸的着色矩阵一一对应
- 对非默认黑色的单元格套上
\textcolor{颜色}{内容}的LaTeX命令 - 把处理后的内容拼接回原LaTeX结构输出
完整可运行R代码
# 加载依赖包,仅需stringr,无其他复杂依赖 library(stringr) # ---------------------- 输入部分 ---------------------- # 输入的LaTeX表格代码,用raw字符串格式避免转义反斜杠 latex_table <- r"( \begin{table}[!htbp] \begin{tabular}{ccc} 1 & 2 & 3 \\ 4 & 5 & 6 \\ \end{tabular} \end{table} )" # 着色规则矩阵,维度必须和表格实际单元格数完全对应 color_mat <- matrix( c("red", "black", "black", "black", "black", "blue"), nrow = 2, byrow = TRUE ) # ----------------------------------------------------- # 提取tabular环境内的内容 tabular_content <- str_extract( latex_table, "(?<=\\begin\\{tabular\\}\\{.*?\\}\\n).*?(?=\\n\\end\\{tabular\\})", dotall = TRUE ) row_list <- str_split(tabular_content, "\\n")[[1]] |> Filter(\(x) x != "", x = _) # 逐行逐单元格处理着色 processed_rows <- c() for (i in seq_along(row_list)) { # 拆分单元格,去除行尾的\\和多余空白 cells <- str_remove(row_list[i], "\\\\\\\\$") |> str_split("&") |> unlist() |> str_trim() # 按着色矩阵处理当前行单元格 for (j in seq_along(cells)) { cur_color <- color_mat[i, j] if (cur_color != "black") { cells[j] <- sprintf("\\textcolor{%s}{%s}", cur_color, cells[j]) } } # 拼接回LaTeX行格式 processed_row <- paste(cells, collapse = " & ") |> paste0(" \\\\") processed_rows <- c(processed_rows, processed_row) } # 替换回原LaTeX代码,输出结果 final_latex <- str_replace( latex_table, "(?<=\\begin\\{tabular\\}\\{.*?\\}\\n).*?(?=\\n\\end\\{tabular\\})", paste(processed_rows, collapse = "\n") ) cat(final_latex)
输出结果
运行上述代码后输出的内容和预期完全一致:
\begin{table}[!htbp] \begin{tabular}{ccc} \textcolor{red}{1} & 2 & 3 \\ 4 & 5 & \textcolor{blue}{6} \\ \end{tabular} \end{table}
注意事项
- 如果表格包含合并单元格,仅需要调整着色矩阵的对应位置规则即可,整体处理逻辑不变
- 如果表格中存在
&字符本身(比如用作逻辑与符号,被包裹在$公式环境内),需要提前对这类特殊&做转义处理,再拆分单元格
内容的提问来源于stack exchange,提问作者Daisy Yu
相关产品推荐
相关产品推荐

