You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

词法分析器未输出全部Token与Lexeme问题求助

词法分析器遗漏最后一个元素的问题

我的词法分析器在处理运算符或括号结尾的输入时,无法输出最后一个元素。比如输入"(47+12);",输出如下:

Next token is: LEFT_PAREN Next lexeme is (
Next token is: INT_LIT Next lexeme is 47
Next token is: ADD_OP Next lexeme is +
Next token is: INT_LIT Next lexeme is 12
Next token is: RIGHT_PAREN Next lexeme is )

可以看到遗漏了分号对应的SEMICOLON标记。以下是我的Java实现代码:

package lexPackage;
 
import java.util.*;
import java.io.*;

public class Lex {

    static final int MAX_LEXEME_LEN = 100;
    static Token charClass;                            // 与枚举对比识别字符类别
    static int lexLen;                                 // 当前词素长度
    static char lexeme[] = new char[MAX_LEXEME_LEN];   // 当前词素的字符数组
    static char nextChar;
    static Token nextToken;
    static int charIndex;

     // 标记与类别
    enum Token {
        INT_LIT,
        ADD_OP,
        LEFT_PAREN,
        RIGHT_PAREN,
        DIGIT,
        SEMICOLON,
        UNKNOWN,
    }

    public static void main(String[] args) throws IOException {

        String line;

        try {
            line = "(47+12);";
            charIndex = 0;
            
            if (getChar(line)) {
                // 在数组范围内执行词法分析
                while (charIndex < line.length()) {
                    lex(line);
                }
            }
        } catch (FileNotFoundException e) {
            System.out.println(e.toString());
        } catch (Exception e) {
            e.printStackTrace();
        }
        
    }
    
    private static Token lookup(char ch) {
        switch (ch) {
        case '(':
            addChar();
            nextToken = Token.LEFT_PAREN;
            break;
        case ')':
            addChar();
            nextToken = Token.RIGHT_PAREN;
            break;
        case '+':
            addChar();
            nextToken = Token.ADD_OP;
            break;
        case ';':
            addChar();
            nextToken = Token.SEMICOLON;
            break;
        }
        return nextToken;
    }

    /************* addChar - 将nextChar添加到词素的函数 *************/
    private static boolean addChar() {
        if (lexLen <= 98) {
            lexeme[lexLen++] = nextChar;
            lexeme[lexLen] = 0;
            return true;
        } else {
            System.out.println("Error - lexeme is too long \n");
            return false;
        }
    }

    /************* getChar - 获取行中下一个字符的函数 *************/
    private static boolean getChar(String ln) {
        if (charIndex >= ln.length()) {
            return false;
        }
        nextChar = ln.charAt(charIndex++);
        if (Character.isDigit(nextChar)) {
            charClass = Token.DIGIT;
        } else {
            charClass = Token.UNKNOWN;
        }
        return true;
    }

    /************* lex - 用于算术表达式的简单词法分析器 *************/
    public static Token lex(String ln) throws IOException {
        lexLen = 0;
        
        switch (charClass) {
        // 解析整数字面量
        case DIGIT:
            nextToken = Token.INT_LIT;
            addChar();
            
            if (getChar(ln)) {
                while (charClass == Token.DIGIT) {
                    addChar();
                    if (!getChar(ln)) {
                        break;
                    }
                }
                
                if (charClass == Token.UNKNOWN && charIndex == ln.length()) {
                    charIndex--;
                }
            }
            break;
            
         // 括号与运算符
        case UNKNOWN:
            lookup(nextChar);
            getChar(ln);
            break;
            
        default:
            nextToken = Token.UNKNOWN;
            break;
        }
        
        // 打印每个标记及其对应的词素
        System.out.printf("Next token is: %-12s Next lexeme is %s\n", String.valueOf(nextToken), String.valueOf(lexeme, 0, lexLen));
        return nextToken;
    }
}

问题原因

核心问题在于main方法的循环条件与lex方法的字符读取逻辑不匹配:

  1. 原循环使用while (charIndex < line.length()),当处理最后一个字符(分号)时,lex方法的UNKNOWN分支会调用getChar(ln),将charIndex增加到等于字符串长度,此时循环直接终止,导致最后一次lex的输出被跳过。
  2. 仅在DIGIT分支中处理了charIndex回退的情况,但对于UNKNOWN类型的最后一个字符,没有对应的逻辑保证其被处理。

修复方案

将main方法中的while循环改为do-while循环,确保无论charIndex是否达到长度,当前字符的lex处理都会先执行:

public static void main(String[] args) throws IOException {
    String line;
    try {
        line = "(47+12);";
        charIndex = 0;
        if (getChar(line)) {
            // 使用do-while确保最后一个字符被处理
            do {
                lex(line);
            } while (charIndex < line.length());
        }
    } catch (FileNotFoundException e) {
        System.out.println(e.toString());
    } catch (Exception e) {
        e.printStackTrace();
    }
}

修复效果

修改后运行代码,输入"(47+12);"会输出完整的标记:

Next token is: LEFT_PAREN Next lexeme is (
Next token is: INT_LIT Next lexeme is 47
Next token is: ADD_OP Next lexeme is +
Next token is: INT_LIT Next lexeme is 12
Next token is: RIGHT_PAREN Next lexeme is )
Next token is: SEMICOLON Next lexeme is ;


内容的提问来源于stack exchange,提问作者1597jony

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 22:30:32