如何解决Rust Pest解析器无法识别LaTeX加粗格式的问题
如何修正Pest语法以正确解析LaTeX的加粗格式和纯文本?
我正尝试用Pest库构建一个基础LaTeX解析器,目前只需要处理行、加粗格式(\textbf{...})和纯文本,但在纯文本处理上遇到了问题。已假设纯文本不包含\和}字符,现有语法如下:
lines = { line ~ (NEWLINE ~ line)* } line = { token* } token = { text_bold | text_plain } text_bold = { "\\textbf{" ~ text_plain ~ "}" } text_plain = ${ inner ~ ("\\" | "}" | NEWLINE) } inner = @{ char* } char = { !("\\" | "}" | NEWLINE) ~ ANY } main = { SOI ~ lines ~ EOI }
测试时发现当前语法无法正确识别加粗标签,输入示例及错误解析结果如下:
输入:
Before \textbf{middle} after. New line
错误输出:
- lines > line
- token > text_plain > inner: "Before "
- token > text_plain > inner: "textbf{middle"
- token > text_plain > inner: " after."
- token > text_plain > inner: "New line"
我尝试将${ inner ~ ("\\" | "}" | NEWLINE) }替换为${ inner }但解析失败,添加&前缀也没用,请问该怎么修改语法才能正确识别行和加粗标签?
问题根源
当前语法有两个核心问题:
text_plain的定义强制要求文本以\、}或换行结尾,这不符合纯文本的实际场景(比如行尾的普通字符就无法匹配)。- 由于
text_plain会尝试匹配到\为止,导致\textbf{被拆分成纯文本,而不是触发text_bold的匹配——Pest的匹配优先级是尝试更长的匹配,但这里text_plain的规则先捕获了\之前的内容,剩下的textbf{middle又被当成纯文本。
修改后的正确语法
lines = { line ~ (NEWLINE ~ line)* } line = { token* } token = { text_bold | text_plain } text_bold = { "\\textbf{" ~ text_plain ~ "}" } text_plain = ${ (char)+ } char = { !("\\" | "}" | NEWLINE) ~ ANY } main = { SOI ~ lines ~ EOI }
关键修改说明
- 移除冗余规则:删掉了多余的
inner规则,直接用(char)+定义text_plain,表示纯文本是一个或多个连续的合法字符(非\、}、换行)。 - 修正纯文本匹配逻辑:不再强制要求纯文本以特定字符结尾,只要是符合要求的连续字符就构成纯文本。这样当遇到
\时,text_plain匹配失败,Pest会自动转而尝试匹配text_bold,从而正确识别\textbf{...}结构。
修改后,输入的解析结果会变为:
- lines > line
- token > text_plain: "Before "
- token > text_bold
- text_plain: "middle"
- token > text_plain: " after."
- lines > line
- token > text_plain: "New line"
内容的提问来源于stack exchange,提问作者Ito Pakito
相关产品推荐
相关产品推荐

