You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Java与ANTLR4解析配置文件时无法获取Name字段的问题

问题解决:ANTLR4解析配置文件的语法错误与键值对识别问题

问题概述

使用Java+ANTLR4解析带注释、层级结构的配置文件时,出现以下问题:

  1. 语法错误提示:line 2:1 no viable alternative at input '[WORLD'
  2. 无法将Name=The Game ; needs a string识别为有效的键值对
  3. 解析树未正确体现层级嵌套结构

原始代码与输出

配置文件

; comment
[WORLD]

[game] ; is a child of WORLD
Name=The Game ; needs a string

[/game] ; end of game

[/WORLD] ; end of WORLD

原始GameData语法规则

grammar GameData;

// Whitespace and comments (ignored)
WS  : [ \t\r\n]+ -> skip;
COMMENT     : ';' .*? ('\r'? '\n' | EOF) -> skip;

WORLD   : '[Ww][Oo][Rr][Ll][Dd]';
GAME    : '[Gg][Aa][Mm][Ee]';


// Header and closing
header     : '[' WORLD ']' ;
closing    : '[' '/' WORLD ']' ;

// Section (game in this case)
game_section    : '[' GAME ']' ( keyValuePair )* ;

// Key-value pair
keyValuePair : ID '=' value ';' ;

// Value (string in this case)
value       : STRING ;

// Terminals
DOT      : '.';
ID       : [a-zA-Z]+;
STRING    : '"' .*? '"' ;

file        : (header|game_section)*;

Java解析代码

import org.antlr.v4.runtime.*;
import org.antlr.v4.runtime.tree.*;
import parser.*;
public class Main {
    public static void main(String[] args) throws Exception {
        CharStream input = CharStreams.fromFileName("test.bd");
        GameDataLexer lexer = new GameDataLexer(input);
        CommonTokenStream tokens = new CommonTokenStream(lexer);
        GameDataParser parser = new GameDataParser(tokens);
        ParseTree tree = parser.file();

        System.out.println("getChildCount "+tree.getChildCount());
        for(int i=0;i<tree.getChildCount();i++) 
            System.out.println("getChild "+tree.getChild(i).toString());
        System.out.println(tree.toStringTree(parser));
    }
}

原始解析输出

line 2:1 no viable alternative at input '[WORLD'
getChildCount 18
getChild [
getChild WORLD
getChild ]
getChild [
getChild game
getChild ]
getChild Name
getChild =
getChild The
getChild Game
getChild [
getChild /
getChild game
getChild ]
getChild [
getChild /
getChild WORLD
getChild ]
(file [ WORLD ] [ game ] Name = The Game [ / game ] [ / WORLD ])

问题根源分析

  1. 词法规则错误:WORLD和GAME的定义用单引号包裹字符组,导致规则匹配的是字面量字符串[Ww][Oo][Rr][Ll][Dd]而非大小写的WORLD/world。
  2. 键值对匹配失败:value仅定义为带双引号的STRING,但配置中的The Game无引号,且未定义无引号字符串规则,无法匹配keyValuePair。
  3. 层级结构未定义:语法规则未处理WORLD包含game_section的嵌套关系,file规则仅允许平级的header和game_section。
  4. 符号未显式定义:[、]、/、=等符号未定义词法规则,导致lexer拆分token时出现歧义。

修复后的解决方案

修复后的GameData语法规则

grammar GameData;

// 忽略空白字符与注释
WS          : [ \t\r\n]+ -> skip;
COMMENT     : ';' .*? ('\r'? '\n' | EOF) -> skip;

// 显式定义符号词法规则
LBRACK      : '[';
RBRACK      : ']';
SLASH       : '/';
EQ          : '=';

// 匹配大小写的WORLD和GAME
WORLD       : [Ww][Oo][Rr][Ll][Dd];
GAME        : [Gg][Aa][Mm][Ee];

// 终端规则
DOT         : '.';
ID          : [a-zA-Z]+;
STRING      : '"' .*? '"'; // 带引号的字符串
UNQUOTED_STRING : ~[;\r\n]+; // 无引号字符串,直到分号或换行

// 层级结构定义
file        : world_section EOF;

world_section : LBRACK WORLD RBRACK 
                (game_section)* 
                LBRACK SLASH WORLD RBRACK;

game_section : LBRACK GAME RBRACK 
               (keyValuePair)* 
               LBRACK SLASH GAME RBRACK;

// 键值对规则,支持带引号或无引号的value
keyValuePair : ID EQ (STRING | UNQUOTED_STRING) ';' ;

修复说明

  1. 修正词法规则:将WORLD/GAME改为字符组匹配,显式定义LBRACK/RBRACK等符号,消除token拆分歧义。
  2. 扩展value类型:新增UNQUOTED_STRING规则,支持无引号的字符串值,让keyValuePair能匹配配置中的Name=The Game ; ...。
  3. 定义嵌套结构:通过world_section包含game_section的规则,正确体现配置的层级关系。
  4. 完善file规则:让file仅包含一个完整的world_section,符合配置文件的结构。

修复后的测试结果

生成新的lexer和parser后运行Java代码,将不再出现语法错误,解析树会正确识别:

  • [WORLD]到[/WORLD]的完整world区块
  • [game]到[/game]的game区块
  • Name=The Game ; needs a string作为有效的keyValuePair

解析树输出示例:

(file (world_section [ WORLD ] (game_section [ game ] (keyValuePair Name = The Game ;) ) [ / WORLD ]) <EOF>)

内容的提问来源于stack exchange,提问作者Bafoeg

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.25 15:44:50