You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何解决Rust Pest解析器无法识别LaTeX加粗格式的问题

如何修正Pest语法以正确解析LaTeX的加粗格式和纯文本?

我正尝试用Pest库构建一个基础LaTeX解析器,目前只需要处理行、加粗格式(\textbf{...})和纯文本,但在纯文本处理上遇到了问题。已假设纯文本不包含\和}字符,现有语法如下:

lines = { line ~ (NEWLINE ~ line)* }
line = { token* }

token = { text_bold | text_plain }

text_bold = { "\\textbf{" ~ text_plain ~ "}" }
text_plain = ${ inner ~ ("\\" | "}" | NEWLINE) }
inner = @{ char* }
char = {
    !("\\" | "}" | NEWLINE) ~ ANY
}

main = {
  SOI ~
  lines ~
  EOI
}

测试时发现当前语法无法正确识别加粗标签,输入示例及错误解析结果如下:

输入:

Before \textbf{middle} after.
New line

错误输出:

  • lines > line
    • token > text_plain > inner: "Before "
    • token > text_plain > inner: "textbf{middle"
    • token > text_plain > inner: " after."
    • token > text_plain > inner: "New line"

我尝试将${ inner ~ ("\\" | "}" | NEWLINE) }替换为${ inner }但解析失败,添加&前缀也没用,请问该怎么修改语法才能正确识别行和加粗标签?


问题根源

当前语法有两个核心问题:

  1. text_plain的定义强制要求文本以\、}或换行结尾,这不符合纯文本的实际场景(比如行尾的普通字符就无法匹配)。
  2. 由于text_plain会尝试匹配到\为止,导致\textbf{被拆分成纯文本,而不是触发text_bold的匹配——Pest的匹配优先级是尝试更长的匹配,但这里text_plain的规则先捕获了\之前的内容,剩下的textbf{middle又被当成纯文本。

修改后的正确语法

lines = { line ~ (NEWLINE ~ line)* }
line = { token* }

token = { text_bold | text_plain }

text_bold = { "\\textbf{" ~ text_plain ~ "}" }
text_plain = ${ (char)+ }
char = { !("\\" | "}" | NEWLINE) ~ ANY }

main = {
  SOI ~
  lines ~
  EOI
}

关键修改说明

  1. 移除冗余规则:删掉了多余的inner规则,直接用(char)+定义text_plain,表示纯文本是一个或多个连续的合法字符(非\、}、换行)。
  2. 修正纯文本匹配逻辑:不再强制要求纯文本以特定字符结尾,只要是符合要求的连续字符就构成纯文本。这样当遇到\时,text_plain匹配失败,Pest会自动转而尝试匹配text_bold,从而正确识别\textbf{...}结构。

修改后,输入的解析结果会变为:

  • lines > line
    • token > text_plain: "Before "
    • token > text_bold
      • text_plain: "middle"
    • token > text_plain: " after."
  • lines > line
    • token > text_plain: "New line"

内容的提问来源于stack exchange,提问作者Ito Pakito

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 23:55:17