You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Java可执行代码行数统计工具复杂注释场景测试失败求助

Java可执行代码行数统计工具的复杂注释处理问题

我开发了一款用于统计Java文件中可执行代码行数的工具,除Example5.java外,其余测试用例均通过,但Example5.java因注释结构复杂,统计结果为2,与预期的6不符,断言错误如下:

断言错误

java.lang.AssertionError: 
Expected :6
Actual   :2

示例Java文件

Example1.java

package fixtures;

public class Example1 {
    public static void main(String[] args) {
        System.out.println("Hello world");
    }
}

Example2.java

package fixtures;

public class Example2 {
    public static void main(String[] args) {

        // say hello
        System.out.println("Hello world");
    }
}

Example3.java

package fixtures;

public class Example3 {
    public static void main(String[] args) {
        System.out.println("Hello world"); // say hello
    }
}

Example4.java

package fixtures;

/**
 * A hello world program
 */

public class Example4 {
    public static void main(String[] args /* grab args */) {
        System.out.println("/*Hello world"); // say hello
    }
}

Example5.java

package fixtures;

/*
 /****//*
 A hello world program
 *\/
*/// -----------------
class Example5 {
    public static void main(String[] args) {
        System.out./*  */println("/*\"Hello world")
        ;
///*
    }
    /* // */ }

统计工具代码

package countloc;

import java.util.regex.Matcher;
import java.util.regex.Pattern;

public class CountLOC {
    public static int count(String text) {
        int count = 0;
        boolean insideString = false;

        // 匹配注释的正则表达式
        Pattern commentPattern = Pattern.compile("/\\*(.|[\\n\\r])*?\\*/|//.*");

        // 移除文本中的注释
        Matcher matcher = commentPattern.matcher(text);
        text = matcher.replaceAll("");

        // 将文本按行分割
        String[] lines = text.split("\n");

        for (String line : lines) {
            // 检查行中是否包含字符串
            for (int i = 0; i < line.length(); i++) {
                char currentChar = line.charAt(i);
                System.out.print(currentChar);

                if (insideString && currentChar == '"' && (i == 0 || line.charAt(i - 1) != '\\')) {
                    insideString = false;
                } else if (!insideString && currentChar == '"') {
                    insideString = true;
                }
            }

            // 如果不在字符串中、行不为空且不是package声明,则计入可执行代码行
            if (!insideString && !line.trim().isEmpty() && !line.trim().startsWith("package ")) {
                count++;
            }
        }

        return count;
    }
}

测试用例代码

package test;

import countloc.CountLOC;
import org.junit.Test;

import java.io.IOException;
import java.nio.file.Files;
import java.nio.file.Paths;

import static org.junit.Assert.*;

public class CountTest {
    @Test
    public void shouldHandleBasicCode() throws IOException {
        String path = "./src/fixtures/Example1.java";
        String code = new String(Files.readAllBytes(Paths.get(path)));
        assertEquals(5, CountLOC.count(code));
    }

    @Test
    public void shouldHandleABlankLineWIthOneLineComment() throws IOException {
        String path = "./src/fixtures/Example2.java";
        String code = new String(Files.readAllBytes(Paths.get(path)));
        assertEquals(5, CountLOC.count(code));
    }

    @Test
    public void shouldHandleAnInlineComment() throws IOException {
        String path = "./src/fixtures/Example3.java";
        String code = new String(Files.readAllBytes(Paths.get(path)));
        assertEquals(5, CountLOC.count(code));
    }

    @Test
    public void shouldHandleMultilineCommentsAndQuotes() throws IOException {
        String path = "./src/fixtures/Example4.java";
        String code = new String(Files.readAllBytes(Paths.get(path)));
        assertEquals(5, CountLOC.count(code));
    }

    @Test
    public void shouldHandleAComplexExample() throws IOException {
        String path = "./src/fixtures/Example5.java";
        String code = new String(Files.readAllBytes(Paths.get(path)));
        assertEquals(6, CountLOC.count(code));
    }
}

问题原因分析

当前代码的核心问题在于:

  1. 正则表达式的局限性:无法正确处理嵌套注释、包含注释标记的注释内容,也会错误识别字符串内部的注释符号(比如"/*\"Hello world"里的/*会被误判为注释开头)。
  2. 处理顺序错误:先移除所有注释再处理字符串,会破坏字符串的原始结构,导致后续的字符串状态判断完全失效。
  3. 跨行状态丢失:按行分割后,无法延续上一行的字符串或注释状态,导致部分行的代码被错误忽略。

修复方案

正确的处理逻辑应该是逐字符遍历,同时跟踪字符串、单行注释、多行注释的状态,避免正则表达式的缺陷。修复后的代码如下:

package countloc;

public class CountLOC {
    public static int count(String text) {
        int count = 0;
        boolean insideString = false;
        boolean insideLineComment = false;
        boolean insideBlockComment = false;
        boolean lineHasCode = false;

        char[] chars = text.toCharArray();
        int length = chars.length;

        for (int i = 0; i < length; i++) {
            char c = chars[i];

            // 处理多行注释
            if (!insideString && !insideLineComment && insideBlockComment) {
                if (c == '*' && i + 1 < length && chars[i + 1] == '/') {
                    insideBlockComment = false;
                    i++; // 跳过 '/'
                }
                continue;
            }

            // 处理单行注释
            if (!insideString && insideLineComment) {
                if (c == '\n') {
                    insideLineComment = false;
                    if (lineHasCode) {
                        count++;
                        lineHasCode = false;
                    }
                }
                continue;
            }

            // 处理字符串
            if (insideString) {
                if (c == '"' && (i == 0 || chars[i - 1] != '\\')) {
                    insideString = false;
                }
                lineHasCode = true;
                continue;
            }

            // 进入多行注释
            if (c == '/' && i + 1 < length && chars[i + 1] == '*') {
                insideBlockComment = true;
                i++;
                continue;
            }

            // 进入单行注释
            if (c == '/' && i + 1 < length && chars[i + 1] == '/') {
                insideLineComment = true;
                i++;
                continue;
            }

            // 进入字符串
            if (c == '"') {
                insideString = true;
                lineHasCode = true;
                continue;
            }

            // 识别package声明行,直接跳过整行
            if (c == 'p' && i + 7 < length && text.startsWith("package ", i)) {
                while (i < length && chars[i] != '\n') {
                    i++;
                }
                continue;
            }

            // 标记当前行存在有效代码
            if (!Character.isWhitespace(c)) {
                lineHasCode = true;
            }

            // 换行时统计有效行
            if (c == '\n') {
                if (lineHasCode) {
                    count++;
                    lineHasCode = false;
                }
            }
        }

        // 处理文件末尾无换行的情况
        if (lineHasCode) {
            count++;
        }

        return count;
    }
}

修复说明

  1. 多状态跟踪:同时维护字符串、单行注释、多行注释三种状态,确保每个字符只在正确的上下文被处理。
  2. 正确的处理顺序:优先处理注释状态,再处理字符串,避免错误修改字符串内容。
  3. 逐字符遍历:彻底解决正则表达式无法处理的复杂注释场景,比如嵌套注释、字符串内的注释标记。
  4. package行单独处理:直接识别并跳过package开头的行,无需后续判断。

内容的提问来源于stack exchange,提问作者Pascal Oseko

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.03 22:15:10