如何在TextX语法中解析并保留C风格单行注释至模型?
问题:TextX无法正确解析C/C++风格单行注释
我正尝试构建支持C/C++风格单行注释(例如//this is a comment)的TextX语法。测试阶段编写了如下语法规则:
Ledger:entries+=Entry; Entry: Comment; Comment:/\/\/.*$/;
但解析包含两行注释的文件:
//comment 1 //comment 2
时,出现错误:
textx.exceptions.TextXSyntaxError: file.txt:3:1: Expected Comment => 'omment 2 *'
解析器调试输出如下:
*** PARSING MODEL *** >> Matching rule Model=Sequence at position 0 => *//comment >> Matching rule Ledger=Sequence in Model at position 0 => *//comment >> Matching rule __asgn_oneormore=OneOrMore[entries] in Ledger at position 0 => *//comment ?? Try match rule Comment=RegExMatch(\/\/.*$) in __asgn_oneormore at position 0 => *//comment ?? Try match rule Comment=RegExMatch(\/\/.*$) in __asgn_oneormore at position 0 => *//comment ' at 0 => '*//comment 1*'ent 1 ?? Try match rule Comment=RegExMatch(\/\/.*$) in __asgn_oneormore at position 13 => omment 1 *//comment ' at 13 => 'omment 1 *//comment 2*' ?? Try match rule Comment=RegExMatch(\/\/.*$) in __asgn_oneormore at position 26 => omment 2 * -- NoMatch at 26 -- NoMatch at 26 <<- Not matched rule __asgn_oneormore=OneOrMore[entries] in __asgn_oneormore at position 0 => *//comment <<- Not matched rule Ledger=Sequence in Ledger at position 0 => *//comment <<- Not matched rule Model=Sequence in Model at position 0 => *//comment
当为Entry规则添加FLOAT选项:
Ledger:entries+=Entry; Entry: Comment | FLOAT; Comment:/\/\/.*$/;
并在文件中添加数值:
//comment 1 //comment 2 7
此时不再抛出异常,但注释未被解析进模型,解析器调试输出如下:
>> Matching rule Model=Sequence at position 0 => *//comment >> Matching rule Ledger=Sequence in Model at position 0 => *//comment >> Matching rule __asgn_oneormore=OneOrMore[entries] in Ledger at position 0 => *//comment >> Matching rule Entry=OrderedChoice in __asgn_oneormore at position 0 => *//comment ?? Try match rule Comment=RegExMatch(\/\/.*$) in Entry at position 0 => *//comment ?? Try match rule Comment=RegExMatch(\/\/.*$) in Entry at position 0 => *//comment ' at 0 => '*//comment 1*'omment 1 ?? Try match rule Comment=RegExMatch(\/\/.*$) in Entry at position 13 => omment 1 *//comment ' at 13 => 'omment 1 *//comment 2*' ?? Try match rule Comment=RegExMatch(\/\/.*$) in Entry at position 26 => omment 2 *7 -- NoMatch at 26 -- NoMatch at 26 ?? Try match rule FLOAT=RegExMatch([+-]?(\d+(\.\d*)?|\.\d+)([eE][+-]?\d+)?(?<=[\w\.])(?![\w\.])) in Entry at position 0 => *//comment ++ Match '7' at 26 => 'omment 2 *7*' <<+ Matched rule Entry=OrderedChoice in Entry at position 27 => mment 2 7* >> Matching rule Entry=OrderedChoice in __asgn_oneormore at position 27 => mment 2 7* ?? Try match rule Comment=RegExMatch(\/\/.*$) in Entry at position 27 => mment 2 7* ?? Try match rule Comment=RegExMatch(\/\/.*$) in Entry at position 27 => mment 2 7* -- NoMatch at 27 -- NoMatch at 27 ?? Try match rule FLOAT=RegExMatch([+-]?(\d+(\.\d*)?|\.\d+)([eE][+-]?\d+)?(?<=[\w\.])(?![\w\.])) in Entry at position 27 => mment 2 7* -- NoMatch at 27 <<- Not matched rule Entry=OrderedChoice in Entry at position 27 => mment 2 7* <<+ Matched rule __asgn_oneormore=OneOrMore[entries] in __asgn_oneormore at position 27 => mment 2 7* <<+ Matched rule Ledger=Sequence in Ledger at position 27 => mment 2 7* ?? Try match rule EOF in Model at position 27 => mment 2 7* <<+ Matched rule Model=Sequence in Model at position 27 => mment 2 7*
问题原因
核心问题在于正则表达式/\/\/.*$/的行为:
$默认匹配整个输入的结尾,而非单行结尾。TextX默认不启用多行模式,导致第一个注释的正则会匹配从//开始到整个文件末尾的所有内容,包括后续的换行和第二行注释。- 当解析器尝试匹配第二个Comment时,指针已经跳到文件末尾附近,无法找到新的
//开头,因此报错。 - 添加FLOAT后,解析器跳过了无法匹配的注释内容,直接匹配到最后的数值,所以不报错但注释被忽略。
解决方案
修改Comment的正则表达式,明确匹配到单行末尾而非整个输入末尾,有两种可行方案:
方案1:启用多行模式
使用(?m)标记开启多行模式,让$匹配每行的结束位置:
Ledger: entries+=Entry; Entry: Comment | FLOAT; Comment: /(?m)\/\/.*$/;
方案2:直接匹配非换行字符
用[^\r\n]匹配//之后所有非换行的字符,直观限制在单行内:
Ledger: entries+=Entry; Entry: Comment | FLOAT; Comment: /\/\/[^\r\n]*/;
补充说明
- 两种方案都能让每个注释仅匹配当前行的内容,解析器会自动处理空白字符(包括换行),匹配完单行注释后,会自动跳过换行符继续匹配下一个Entry。
- 测试时,两行注释会被识别为两个独立的Comment实例,存入
entries列表;包含注释和数值的文件,两者都会被正确解析并存入模型。
内容的提问来源于stack exchange,提问作者kabla002
相关产品推荐
相关产品推荐

