Python正则匹配LaTeX宏参数:支持括号配对与换行规则
提取LaTeX宏必填参数的Python实现(支持嵌套大括号与换行处理)
问题分析
现有函数存在三个核心缺陷:
- 无法处理大括号嵌套:原正则的
.*会贪婪匹配到文档末尾的大括号,而非目标宏的闭合括号 - 多行内容处理失效:
re.MULTILINE模式不能跨多行正确提取内容,且未按要求转换换行符 - 可选方括号参数处理不严谨:未彻底忽略方括号内的可选参数,容易出现匹配错位
解决方案
采用递归正则表达式处理嵌套大括号,配合针对性的换行转换逻辑,实现需求:
import re def tag(file_content, macro_name): # 构造正则:匹配宏名 + 跳过可选方括号参数 + 递归提取嵌套大括号内的内容 regex_pattern = re.compile( rf'\\{re.escape(macro_name)}\s*(?:\[[^\]]*\])?\s*{{((?:[^{{}}]|(?R))*)}}', re.DOTALL ) # 查找匹配内容 match_result = regex_pattern.search(file_content) if not match_result: return "" raw_content = match_result.group(1) # 按连续换行分割段落,保留原段落结构 paragraphs = re.split(r'\n{2,}', raw_content) processed_paragraphs = [] for para in paragraphs: # 段落内单个换行替换为空格,合并多余连续空格 cleaned_para = re.sub(r'\n', ' ', para).strip() cleaned_para = re.sub(r'\s+', ' ', cleaned_para) processed_paragraphs.append(cleaned_para) # 用单个换行连接处理后的段落 return '\n'.join(processed_paragraphs)
代码细节说明
- 递归正则逻辑:
(?:[^{}]|(?R))*表示匹配非大括号字符,或递归匹配完整的大括号结构,完美解决嵌套大括号的配对问题 - 可选参数忽略:
(?:\[[^\]]*\])?匹配并跳过宏名后的可选方括号参数(?:声明非捕获组,避免干扰结果提取) - 换行符转换:
- 先按连续两个及以上换行分割段落,保留原文档的段落分隔
- 每个段落内的单个换行替换为空格,同时合并多余的连续空格
- 最终用单个换行连接段落,符合需求中的格式要求
- 宏名转义:
re.escape(macro_name)处理宏名中可能存在的正则特殊字符,避免匹配出错
测试验证
针对提供的示例LaTeX代码,调用tag(file_content, 'abstract')将返回:
Problems of using {\it regular structures} of memory $\Sigma =\{\Sigma(z) \vert z\in M\}$ and $\tilde{\Sigma} = \{\tilde{\Sigma}(z) \vert z\in R\}$ to model the processes of synthesis of knowledge structures with the help of knowledge processing operations from special classes of such operations are considered. These structures define the formats of the memory subdomains used. Area formats match domain element structures (domains of definition and meaning of operations)...
内容的提问来源于stack exchange,提问作者Crosfield
相关产品推荐
相关产品推荐

