You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在自定义词法分析器中识别并输出无效词素及对应错误信息

自定义int/float词法分析器的无效词素识别修复方案

正在开发一款支持int/float输入的自定义词法分析器,已完成有效词素的识别与Token分配逻辑,但在识别无效词素并输出精准错误信息时遇到瓶颈。尝试过三版findInvalidLexemes方法,均存在逻辑漏洞:

  • 第一版错误假设数据类型为首个Token,无法处理多个无效Token的场景
  • 第二版将输入拆分为单个字符识别,未覆盖赋值语句的完整校验逻辑
  • 第三版错误预设数据类型后紧跟等号,正则校验规则过于严苛

期望实现的效果:

  • 输入int 23jordan=23; → 输出Invalid identifier 23jordan
  • 输入int x=; → 输出Invalid, Missing constant

现有有效Token识别代码如下:

private static void printLexemes(String input, String dataType) {
    Map<String, String> lexemes = new HashMap<>();
    lexemes.put(dataType, "Data_Type");

    String[] tokens = input.split("\\s+|(?<=[=,;])(?=[^\\s])|(?=[=,;])(?<=[^\\s])");

    boolean afterDataType = false;

    String currentToken = "";

    for (String token : tokens) {
        if (!token.isEmpty()) {
            if (!afterDataType) {
                if (token.equals(dataType)) {
                    afterDataType = true;
                } else {
                    lexemes.put(token, "invalid lexeme");
                }
            } else {
                if (token.matches("[a-zA-Z_][a-zA-Z0-9_]*")) {
                    lexemes.put(token, "IDENTIFIER");
                } else if (token.matches("\\d+(\\.\\d+)?")) {
                    if (dataType.equals("float")) {
                        float floatValue = Float.parseFloat(token);
                        currentToken = Float.toString(floatValue);
                        lexemes.put(currentToken, "Constant(F)");

                    } else {
                        currentToken = token;
                        lexemes.put(currentToken, "Constant(I)");

                    }
                } else if (token.equals(";")) {
                    lexemes.put(token, "Semi_Colon");
                } else if (token.equals(",")) {
                    lexemes.put(token, "Comma");
                } else if (token.equals("=")) {
                    lexemes.put(token, "Equal_Sign");
                } else {
                    lexemes.put(token, "invalid lexeme");
                }
            }
        }
    }
}

修复后的无效词素识别实现

核心思路是基于状态机跟踪语法规则,配合与有效Token识别一致的拆分逻辑,精准定位无效词素或缺失的语法元素:

import java.util.ArrayList;
import java.util.List;
import java.util.regex.Pattern;
import java.util.Map;
import java.util.HashMap;

public class Lexer {
    // 标识符正则:必须以字母/下划线开头,后续可跟字母/数字/下划线
    private static final Pattern IDENTIFIER_PATTERN = Pattern.compile("^[a-zA-Z_][a-zA-Z0-9_]*$");
    // 整数/浮点数正则
    private static final Pattern CONSTANT_PATTERN = Pattern.compile("^\\d+(\\.\\d+)?$");
    // 支持的基础数据类型
    private static final List<String> SUPPORTED_TYPES = List.of("int", "float");

    // 原有效Token识别方法保留
    private static void printLexemes(String input, String dataType) {
        Map<String, String> lexemes = new HashMap<>();
        lexemes.put(dataType, "Data_Type");

        String[] tokens = input.split("\\s+|(?<=[=,;])(?=[^\\s])|(?=[=,;])(?<=[^\\s])");

        boolean afterDataType = false;
        String currentToken = "";

        for (String token : tokens) {
            if (!token.isEmpty()) {
                if (!afterDataType) {
                    if (token.equals(dataType)) {
                        afterDataType = true;
                    } else {
                        lexemes.put(token, "invalid lexeme");
                    }
                } else {
                    if (token.matches("[a-zA-Z_][a-zA-Z0-9_]*")) {
                        lexemes.put(token, "IDENTIFIER");
                    } else if (token.matches("\\d+(\\.\\d+)?")) {
                        if (dataType.equals("float")) {
                            float floatValue = Float.parseFloat(token);
                            currentToken = Float.toString(floatValue);
                            lexemes.put(currentToken, "Constant(F)");
                        } else {
                            currentToken = token;
                            lexemes.put(currentToken, "Constant(I)");
                        }
                    } else if (token.equals(";")) {
                        lexemes.put(token, "Semi_Colon");
                    } else if (token.equals(",")) {
                        lexemes.put(token, "Comma");
                    } else if (token.equals("=")) {
                        lexemes.put(token, "Equal_Sign");
                    } else {
                        lexemes.put(token, "invalid lexeme");
                    }
                }
            }
        }
    }

    // 修复后的无效词素识别方法
    public static List<String> findInvalidLexemes(String input) {
        List<String> errors = new ArrayList<>();
        // 复用有效Token识别的拆分规则,确保Token一致性
        String[] tokens = input.split("\\s+|(?<=[=,;])(?=[^\\s])|(?=[=,;])(?<=[^\\s])");

        // 状态枚举:跟踪当前语法位置的预期Token类型
        enum ParseState {
            EXPECTING_TYPE,
            EXPECTING_IDENTIFIER,
            EXPECTING_EQUALS,
            EXPECTING_CONSTANT,
            EXPECTING_SEMICOLON
        }

        ParseState currentState = ParseState.EXPECTING_TYPE;
        String lastValidIdentifier = null;

        for (String token : tokens) {
            if (token.isBlank()) continue;

            switch (currentState) {
                case EXPECTING_TYPE:
                    if (SUPPORTED_TYPES.contains(token)) {
                        currentState = ParseState.EXPECTING_IDENTIFIER;
                    } else {
                        errors.add("Invalid data type: " + token);
                        // 即使类型错误,仍继续校验后续Token
                        currentState = ParseState.EXPECTING_IDENTIFIER;
                    }
                    break;

                case EXPECTING_IDENTIFIER:
                    if (token.equals(";")) {
                        errors.add("Missing identifier and declaration");
                        currentState = ParseState.EXPECTING_SEMICOLON;
                    } else if (IDENTIFIER_PATTERN.matcher(token).matches()) {
                        lastValidIdentifier = token;
                        currentState = ParseState.EXPECTING_EQUALS;
                    } else {
                        errors.add("Invalid identifier: " + token);
                        currentState = ParseState.EXPECTING_EQUALS;
                    }
                    break;

                case EXPECTING_EQUALS:
                    if (token.equals("=")) {
                        currentState = ParseState.EXPECTING_CONSTANT;
                    } else if (token.equals(";")) {
                        errors.add("Missing equal sign and constant for: " + lastValidIdentifier);
                        currentState = ParseState.EXPECTING_SEMICOLON;
                    } else {
                        errors.add("Unexpected token after identifier: " + token);
                    }
                    break;

                case EXPECTING_CONSTANT:
                    if (CONSTANT_PATTERN.matcher(token).matches()) {
                        currentState = ParseState.EXPECTING_SEMICOLON;
                    } else if (token.equals(";")) {
                        errors.add("Invalid, Missing constant");
                        currentState = ParseState.EXPECTING_SEMICOLON;
                    } else {
                        errors.add("Invalid constant: " + token);
                    }
                    break;

                case EXPECTING_SEMICOLON:
                    if (!token.equals(";")) {
                        errors.add("Unexpected token after declaration: " + token);
                    }
                    break;
            }
        }

        // 检查结束时的未完成状态
        switch (currentState) {
            case EXPECTING_TYPE:
                errors.add("Missing data type");
                break;
            case EXPECTING_IDENTIFIER:
                errors.add("Missing identifier");
                break;
            case EXPECTING_EQUALS:
                errors.add("Missing equal sign for: " + lastValidIdentifier);
                break;
            case EXPECTING_CONSTANT:
                errors.add("Invalid, Missing constant");
                break;
        }

        return errors;
    }

    // 测试示例
    public static void main(String[] args) {
        // 测试1:无效标识符
        System.out.println(findInvalidLexemes("int 23jordan=23;"));
        // 输出:[Invalid identifier: 23jordan]

        // 测试2:缺失常量
        System.out.println(findInvalidLexemes("int x=;"));
        // 输出:[Invalid, Missing constant]

        // 测试3:有效声明
        System.out.println(findInvalidLexemes("float y=3.14;"));
        // 输出:[]

        // 测试4:无效数据类型
        System.out.println(findInvalidLexemes("double z=5;"));
        // 输出:[Invalid data type: double]
    }
}

关键修复点说明

  1. 统一Token拆分逻辑:复用有效Token识别中的拆分正则,避免因拆分规则不一致导致的校验偏差。
  2. 状态机驱动校验:通过枚举ParseState跟踪当前语法位置,明确每个阶段预期的Token类型,精准判断无效或缺失元素。
  3. 精准错误提示:针对不同场景生成对应提示,比如无效标识符、缺失常量、数据类型错误等。
  4. 容错性处理:即使某个Token错误,仍继续校验后续内容,尽可能收集所有问题。

内容的提问来源于stack exchange,提问作者Newbie

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 09:10:01