ANTLR4解析Python风格注释报错:输入不匹配与多余输入问题
问题描述
我编写了名为ngc_grammar的ANTLR4语法文件,目标是解析Python风格的单行(#)和多行(""")注释,但测试时出现以下问题:
单行注释测试场景
# This is single line comment.\nme = "Test Name"\nt = test(me)\nreturn t测试失败# This is single line comment.测试通过me = "Test Name"\n# This is single line comment.\nt = test(me)\nreturn t测试通过
多行注释测试场景
"""This is comment.\nThis is comment.\n"""\nme = "Test Name"\nt = test(me)\nreturn t测试失败"""This is comment.\nThis is comment.\n"""测试通过me = "Test Name"\n"""This is comment.\nThis is comment.\n"""\nt = test(me)\nreturn t测试通过
报错信息
antlr4.error.Errors.ParseCancellationException: Line 2: 0 extraneous input 'me' expecting {, NEWLINE}
同时存在mismatched input '<EOF>'错误。
原始语法文件
grammar ngc_grammar; parse : (multi_statement | NEWLINE*) EOF ; multi_statement : (statement NEWLINE*)+ ; statement : return_statement | conditional_statement | function_invocation | assignment | for_loop ; conditional_statement : if_statement (NEWLINE+ elif_statement)* (NEWLINE+ else_statement)? ; if_statement : IF SPACE logical_expression ':' NEWLINE SPACE statement ; elif_statement : ELIF SPACE logical_expression ':' NEWLINE SPACE statement ; else_statement : ELSE ':' NEWLINE SPACE statement ; return_statement : RETURN SPACE (expr | logical_expression) ; logical_expression : '(' SPACE? logical_expression SPACE? ')' #BracketedLogicalExpr | NOT SPACE logical_expression #LogicalNotExpr | left=logical_expression SPACE operator=OR SPACE right=logical_expression #LogicalExpr | left=logical_expression SPACE operator=AND SPACE right=logical_expression #LogicalExpr | boolean_expression #LogicalBooleanExpr ; boolean_expression : left=expr SPACE? operator=(GT | GTE | LT | LTE | EQ | NE | IN | NOT_IN | IS | IS_NOT) SPACE? right=expr #ComparisonExpr | expr #BooleanFunctionExpr ; expr : term (SPACE? arith_expr)* ; arith_expr : operator=(PLUS|MINUS) SPACE? term ; term : factor (SPACE? factor_expr)* ; factor_expr : operator=(TIMES|DIV) SPACE? factor ; factor : operator=(PLUS|MINUS) SPACE? factor | value ; value : atom_expr | function_invocation ; atom_expr : atom trailer* #AtomExprAtom | '(' SPACE? expr SPACE? ')' #AtomExprBracket ; trailer : '(' parameters ')' # TrailerFunction | '[' string_atom ']' # TrailerIndex | '.' identifier # TrailerProp ; function_invocation : identifier '(' parameters ')' # FunctionInvocation | atom trailer+ # FunctionAccessor ; assignment : identifier SPACE? '=' SPACE? expr #VariableAssignment | variable_accessor '.' identifier SPACE? '=' SPACE? expr #PropAssignment ; parameters : (value SPACE? (',' SPACE? value)*)? ; array : '[' SPACE? (value SPACE? (',' SPACE? value)*)? SPACE? ']' ; for_loop : FOR SPACE var=identifier SPACE* (',' SPACE* index=identifier)? SPACE IN SPACE atom ':' NEWLINE SPACE assignment #forLoop ; atom : none_atom | boolean_atom | float_atom | integer_atom | string_atom | array | variable_accessor ; variable_accessor : identifier ; none_atom : NONE ; boolean_atom : TRUE | FALSE ; integer_atom : INTEGER ; float_atom : FLOAT ; string_atom : STRING_LITERAL ; identifier : NAME ; FOR: 'for'; AND: 'and'; ELIF: 'elif'; ELSE: 'else'; FALSE: 'False'; FLOAT: ( '0' | [1-9] [0-9]* ) '.' [0-9]+; IF: 'if'; IN: 'in'; IS: 'is'; IS_NOT: 'is not'; NOT_IN: 'not in'; INTEGER: ( '0' | [1-9] [0-9]* ); NEWLINE: ( '\r'? '\n' | '\r' | '\f' ); NONE: 'None'; NOT: 'not'; OR: 'or'; RETURN: 'return'; SPACE: [ \t]+; STRING_LITERAL : '"'.*?'"' | '\''.*?'\'' ; TRUE: 'True'; GT: '>'; GTE: '>='; LT: '<'; LTE: '<='; EQ: '=='; NE: '!='; PLUS: '+'; MINUS: '-'; TIMES: '*'; DIV: '/'; NAME: ID_START ID_CONTINUE*; SINGLELINECOMMENT: '#' ~[\r\n]* -> skip; MULTILINECOMMENT: ('"""') .*? (MULTILINECOMMENT | '"""') -> skip; fragment ID_START : '_' | [A-Z] | [a-z] ; fragment ID_CONTINUE : ID_START | [0-9] ;
解决方案
问题根源在于parse规则的结构限制、多行注释的递归匹配逻辑错误,以及字符串字面量的规则不符合Python语法。具体修改如下:
1. 调整parse规则
原规则要求开头只能是multi_statement或空换行,但注释被skip后相当于无Token,导致后续语句无法被正确识别。修改为允许任意数量的语句加换行:
parse : (statement NEWLINE*)* EOF ;
同时删除冗余的multi_statement规则,因为新的parse规则已经覆盖了多语句场景。
2. 修复多行注释规则
原规则的递归嵌套逻辑错误(Python本身不支持嵌套多行注释),简化为直接匹配"""包裹的内容:
MULTILINECOMMENT: '"""' .*? '"""' -> skip;
3. 修正字符串字面量规则
原规则用"代替双引号,无法识别Python风格的双引号字符串,替换为实际双引号:
STRING_LITERAL : '"' .*? '"' | '\'' .*? '\'' ;
修改后的完整语法文件
grammar ngc_grammar; parse : (statement NEWLINE*)* EOF ; statement : return_statement | conditional_statement | function_invocation | assignment | for_loop ; conditional_statement : if_statement (NEWLINE+ elif_statement)* (NEWLINE+ else_statement)? ; if_statement : IF SPACE logical_expression ':' NEWLINE SPACE statement ; elif_statement : ELIF SPACE logical_expression ':' NEWLINE SPACE statement ; else_statement : ELSE ':' NEWLINE SPACE statement ; return_statement : RETURN SPACE (expr | logical_expression) ; logical_expression : '(' SPACE? logical_expression SPACE? ')' #BracketedLogicalExpr | NOT SPACE logical_expression #LogicalNotExpr | left=logical_expression SPACE operator=OR SPACE right=logical_expression #LogicalExpr | left=logical_expression SPACE operator=AND SPACE right=logical_expression #LogicalExpr | boolean_expression #LogicalBooleanExpr ; boolean_expression : left=expr SPACE? operator=(GT | GTE | LT | LTE | EQ | NE | IN | NOT_IN | IS | IS_NOT) SPACE? right=expr #ComparisonExpr | expr #BooleanFunctionExpr ; expr : term (SPACE? arith_expr)* ; arith_expr : operator=(PLUS|MINUS) SPACE? term ; term : factor (SPACE? factor_expr)* ; factor_expr : operator=(TIMES|DIV) SPACE? factor ; factor : operator=(PLUS|MINUS) SPACE? factor | value ; value : atom_expr | function_invocation ; atom_expr : atom trailer* #AtomExprAtom | '(' SPACE? expr SPACE? ')' #AtomExprBracket ; trailer : '(' parameters ')' # TrailerFunction | '[' string_atom ']' # TrailerIndex | '.' identifier # TrailerProp ; function_invocation : identifier '(' parameters ')' # FunctionInvocation | atom trailer+ # FunctionAccessor ; assignment : identifier SPACE? '=' SPACE? expr #VariableAssignment | variable_accessor '.' identifier SPACE? '=' SPACE? expr #PropAssignment ; parameters : (value SPACE? (',' SPACE? value)*)? ; array : '[' SPACE? (value SPACE? (',' SPACE? value)*)? SPACE? ']' ; for_loop : FOR SPACE var=identifier SPACE* (',' SPACE* index=identifier)? SPACE IN SPACE atom ':' NEWLINE SPACE assignment #forLoop ; atom : none_atom | boolean_atom | float_atom | integer_atom | string_atom | array | variable_accessor ; variable_accessor : identifier ; none_atom : NONE ; boolean_atom : TRUE | FALSE ; integer_atom : INTEGER ; float_atom : FLOAT ; string_atom : STRING_LITERAL ; identifier : NAME ; FOR: 'for'; AND: 'and'; ELIF: 'elif'; ELSE: 'else'; FALSE: 'False'; FLOAT: ( '0' | [1-9] [0-9]* ) '.' [0-9]+; IF: 'if'; IN: 'in'; IS: 'is'; IS_NOT: 'is not'; NOT_IN: 'not in'; INTEGER: ( '0' | [1-9] [0-9]* ); NEWLINE: ( '\r'? '\n' | '\r' | '\f' ); NONE: 'None'; NOT: 'not'; OR: 'or'; RETURN: 'return'; SPACE: [ \t]+; STRING_LITERAL : '"' .*? '"' | '\'' .*? '\'' ; TRUE: 'True'; GT: '>'; GTE: '>='; LT: '<'; LTE: '<='; EQ: '=='; NE: '!='; PLUS: '+'; MINUS: '-'; TIMES: '*'; DIV: '/'; NAME: ID_START ID_CONTINUE*; SINGLELINECOMMENT: '#' ~[\r\n]* -> skip; MULTILINECOMMENT: '"""' .*? '"""' -> skip; fragment ID_START : '_' | [A-Z] | [a-z] ; fragment ID_CONTINUE : ID_START | [0-9] ;
修改后,无论注释位于代码开头、中间还是结尾,都能被正确跳过,后续代码也能被正常解析。
内容的提问来源于stack exchange,提问作者radhey_mishra
相关产品推荐
相关产品推荐

