You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在TextX语法中解析并保留C风格单行注释至模型?

问题:TextX无法正确解析C/C++风格单行注释

我正尝试构建支持C/C++风格单行注释(例如//this is a comment)的TextX语法。测试阶段编写了如下语法规则:

Ledger:entries+=Entry;

Entry: Comment;

Comment:/\/\/.*$/;

但解析包含两行注释的文件:

//comment 1
//comment 2

时,出现错误:

textx.exceptions.TextXSyntaxError: file.txt:3:1: Expected Comment => 'omment 2 *'

解析器调试输出如下:

*** PARSING MODEL ***
>> Matching rule Model=Sequence at position 0 => *//comment
   >> Matching rule Ledger=Sequence in Model at position 0 => *//comment
      >> Matching rule __asgn_oneormore=OneOrMore[entries] in Ledger at position 0 => *//comment
         ?? Try match rule Comment=RegExMatch(\/\/.*$) in __asgn_oneormore at position 0 => *//comment
         ?? Try match rule Comment=RegExMatch(\/\/.*$) in __asgn_oneormore at position 0 => *//comment
' at 0 => '*//comment 1*'ent 1
         ?? Try match rule Comment=RegExMatch(\/\/.*$) in __asgn_oneormore at position 13 => omment 1 *//comment
' at 13 => 'omment 1 *//comment 2*'
         ?? Try match rule Comment=RegExMatch(\/\/.*$) in __asgn_oneormore at position 26 => omment 2 *
         -- NoMatch at 26
         -- NoMatch at 26
      <<- Not matched rule __asgn_oneormore=OneOrMore[entries] in __asgn_oneormore at position 0 => *//comment
   <<- Not matched rule Ledger=Sequence in Ledger at position 0 => *//comment
<<- Not matched rule Model=Sequence in Model at position 0 => *//comment

当为Entry规则添加FLOAT选项:

Ledger:entries+=Entry;

Entry: Comment | FLOAT;

Comment:/\/\/.*$/;

并在文件中添加数值:

//comment 1
//comment 2
7

此时不再抛出异常,但注释未被解析进模型,解析器调试输出如下:

>> Matching rule Model=Sequence at position 0 => *//comment
   >> Matching rule Ledger=Sequence in Model at position 0 => *//comment
      >> Matching rule __asgn_oneormore=OneOrMore[entries] in Ledger at position 0 => *//comment
         >> Matching rule Entry=OrderedChoice in __asgn_oneormore at position 0 => *//comment
            ?? Try match rule Comment=RegExMatch(\/\/.*$) in Entry at position 0 => *//comment
            ?? Try match rule Comment=RegExMatch(\/\/.*$) in Entry at position 0 => *//comment
' at 0 => '*//comment 1*'omment 1
            ?? Try match rule Comment=RegExMatch(\/\/.*$) in Entry at position 13 => omment 1 *//comment
' at 13 => 'omment 1 *//comment 2*'
            ?? Try match rule Comment=RegExMatch(\/\/.*$) in Entry at position 26 => omment 2 *7
            -- NoMatch at 26
            -- NoMatch at 26
            ?? Try match rule FLOAT=RegExMatch([+-]?(\d+(\.\d*)?|\.\d+)([eE][+-]?\d+)?(?<=[\w\.])(?![\w\.])) in Entry at position 0 => *//comment
            ++ Match '7' at 26 => 'omment 2 *7*'
         <<+ Matched rule Entry=OrderedChoice in Entry at position 27 => mment 2 7*
         >> Matching rule Entry=OrderedChoice in __asgn_oneormore at position 27 => mment 2 7*
            ?? Try match rule Comment=RegExMatch(\/\/.*$) in Entry at position 27 => mment 2 7*
            ?? Try match rule Comment=RegExMatch(\/\/.*$) in Entry at position 27 => mment 2 7*
            -- NoMatch at 27
            -- NoMatch at 27
            ?? Try match rule FLOAT=RegExMatch([+-]?(\d+(\.\d*)?|\.\d+)([eE][+-]?\d+)?(?<=[\w\.])(?![\w\.])) in Entry at position 27 => mment 2 7*
            -- NoMatch at 27
         <<- Not matched rule Entry=OrderedChoice in Entry at position 27 => mment 2 7*
      <<+ Matched rule __asgn_oneormore=OneOrMore[entries] in __asgn_oneormore at position 27 => mment 2 7*
   <<+ Matched rule Ledger=Sequence in Ledger at position 27 => mment 2 7*
   ?? Try match rule EOF in Model at position 27 => mment 2 7*
<<+ Matched rule Model=Sequence in Model at position 27 => mment 2 7*

问题原因

核心问题在于正则表达式/\/\/.*$/的行为:

  • $默认匹配整个输入的结尾,而非单行结尾。TextX默认不启用多行模式,导致第一个注释的正则会匹配从//开始到整个文件末尾的所有内容,包括后续的换行和第二行注释。
  • 当解析器尝试匹配第二个Comment时,指针已经跳到文件末尾附近,无法找到新的//开头,因此报错。
  • 添加FLOAT后,解析器跳过了无法匹配的注释内容,直接匹配到最后的数值,所以不报错但注释被忽略。

解决方案

修改Comment的正则表达式,明确匹配到单行末尾而非整个输入末尾,有两种可行方案:

方案1:启用多行模式

使用(?m)标记开启多行模式,让$匹配每行的结束位置:

Ledger: entries+=Entry;

Entry: Comment | FLOAT;

Comment: /(?m)\/\/.*$/;

方案2:直接匹配非换行字符

用[^\r\n]匹配//之后所有非换行的字符,直观限制在单行内:

Ledger: entries+=Entry;

Entry: Comment | FLOAT;

Comment: /\/\/[^\r\n]*/;

补充说明

  • 两种方案都能让每个注释仅匹配当前行的内容,解析器会自动处理空白字符(包括换行),匹配完单行注释后,会自动跳过换行符继续匹配下一个Entry。
  • 测试时,两行注释会被识别为两个独立的Comment实例,存入entries列表;包含注释和数值的文件,两者都会被正确解析并存入模型。

内容的提问来源于stack exchange,提问作者kabla002

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.18 06:05:55