You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Java词法分析器输出存在空行问题求助

解决Java词法分析器输出空行问题

我开发了一个Java词法分析器,目标是为每个符号输出对应的token,要求每行仅输出一个token且无空行,但当前输出存在空行。我尝试过使用trim()、replaceAll()方法解决,但都无效。

我的Java代码

import java.io.BufferedReader;
import java.io.FileReader;
import java.io.IOException;
import java.util.regex.Matcher;
import java.util.regex.Pattern;

public class Lexer {

    public static void Tokenize(String fileName) {
        BufferedReader reader = null;
        try {
            reader = new BufferedReader(new FileReader(fileName));
            String line = null;
            while ((line = reader.readLine()) != null) {
                line = removeComments(line);
                String[] tokens = tokenizeLine(line);
                for (String token : tokens) {
                    System.out.println(token);
                }
            }
        } catch (IOException e) {
            System.err.println("Error reading file: " + e.getMessage());
        } finally {
            try {
                if (reader != null) {
                    reader.close();
                }
            } catch (IOException e) {
                System.err.println("Error closing file: " + e.getMessage());
            }
        }
        System.out.println("SYNTAX ERROR: INVALID IDENTIFIER NAME");

    }

    private static String removeComments(String line) {
        // Remove inline comments
        line = line.replaceAll("//.*", "");
        // Remove block comments
        Pattern pattern = Pattern.compile("/\\*.*?\\*/", Pattern.DOTALL);
        Matcher matcher = pattern.matcher(line);
        return matcher.replaceAll("");
    }

    private static String[] tokenizeLine(String line) {

        String[] tokens = line.split("\\s+ |(?=[\\[\\](){}<>=,;+-/*%|&!])|(?<=[\\[\\](){}<>=,;+-/*%|&!])");
        for (int i = 0; i < tokens.length; i++) {
            String token = tokens[i].replaceAll("\\s+", "");
            if (token.matches("procedure")) {
                tokens[i] = "PROC";
            } else if (token.matches("int")) {
                tokens[i] = "INT";
            } else if (token.matches("[0-9]+")) {
                // Integer constant
                tokens[i] = "INT_CONST";
            } else if (token.matches("[(]")) {
                tokens[i] = "LP";
            } else if (token.matches("[)]")) {
                tokens[i] = "RP";
            } else if (token.matches("\".*\"")) {
                // String constant
                tokens[i] = "STR_CONST";
            } else if (token.matches("String") || token.matches("string")) {
                // String keyword
                tokens[i] = "STR";
            } else if (token.matches("if")) {
                tokens[i] = "IF";
            } else if (token.matches("for")) {
                tokens[i] = "FOR";
            } else if (token.matches("while")) {
                tokens[i] = "WHILE";
            } else if (token.matches("return")) {
                tokens[i] = "RETURN";
            } else if (token.matches("[;]")) {
                tokens[i] = "SEMI";
            } else if (token.matches("do")) {
                tokens[i] = "DO";
            } else if (token.matches("break")) {
                tokens[i] = "BREAK";
            } else if (token.matches("end")) {
                tokens[i] = "END";
            } else if (token.matches("[a-zA-Z][a-zA-Z0-9]*")) {
                // Identifier
                tokens[i] = "IDENT";
            } else if (token.matches("[=]")) {
                tokens[i] = "ASSIGN";
            } else if (token.matches("[<]")) {
                tokens[i] = "LT";
            } else if (token.matches("[>]")) {
                tokens[i] = "RT";
            } else if (token.matches("[++]")) {
                tokens[i] = "INC";
            } else if (token.matches("[{]")) {
                tokens[i] = "RB";
            } else if (token.matches("[}]")) {
                tokens[i] = "LB";
            } else if (token.matches("[*]")) {
                tokens[i] = "MUL_OP";
            } else if (token.matches("[/]")) {
                tokens[i] = "DIV_OP";
            } else if (token.matches("[>=]")) {
                tokens[i] = "GE";
            }
        }
        return tokens;
    }
}

当前输出结果

LP
IDENT
RP
 

FOR
LP
IDENT
ASSIGN
INT_CONST
SEMI
IDENT
LT
IDENT
SEMI
IDENT
ASSIGN
IDENT
INC
INC
RP
 
RB
IDENT
ASSIGN
IDENT
MUL_OP
 
LP
IDENT
DIV_OP
INT_CONST
RP
SEMI

IF
LP
IDENT
RT
ASSIGN
INT_CONST
RP
BREAK
SEMI
        
LB

IDENT
SEMI

IDENT


IDENT
ASSIGN
STR_CONST
SEMI
SYNTAX ERROR: INVALID IDENTIFIER NAME

解决方案

1. 过滤空token

遍历token数组时,先去除token首尾空白,仅输出非空的token。修改Tokenize方法中的循环部分:

for (String token : tokens) {
    String trimmedToken = token.trim();
    if (!trimmedToken.isEmpty()) {
        System.out.println(trimmedToken);
    }
}

2. 优化分割正则表达式

原split表达式里的\\s+ (多余空格)会导致分割出空字符串,同时添加-1参数保留所有分割结果(避免末尾空元素被自动丢弃),修改tokenizeLine方法中的split语句:

String[] tokens = line.split("\\s+|(?=[\\[\\](){}<>=,;+-/*%|&!])|(?<=[\\[\\](){}<>=,;+-/*%|&!])", -1);

3. 跳过空行

在处理注释后,若整行内容为空,直接跳过token分割步骤,避免生成空token数组:

line = removeComments(line).trim();
if (line.isEmpty()) {
    continue;
}
String[] tokens = tokenizeLine(line);

内容的提问来源于stack exchange,提问作者hamsoace

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 11:45:26