Java可执行代码行数统计工具复杂注释场景测试失败求助
Java可执行代码行数统计工具的复杂注释处理问题
我开发了一款用于统计Java文件中可执行代码行数的工具,除Example5.java外,其余测试用例均通过,但Example5.java因注释结构复杂,统计结果为2,与预期的6不符,断言错误如下:
断言错误
java.lang.AssertionError: Expected :6 Actual :2
示例Java文件
Example1.java
package fixtures; public class Example1 { public static void main(String[] args) { System.out.println("Hello world"); } }
Example2.java
package fixtures; public class Example2 { public static void main(String[] args) { // say hello System.out.println("Hello world"); } }
Example3.java
package fixtures; public class Example3 { public static void main(String[] args) { System.out.println("Hello world"); // say hello } }
Example4.java
package fixtures; /** * A hello world program */ public class Example4 { public static void main(String[] args /* grab args */) { System.out.println("/*Hello world"); // say hello } }
Example5.java
package fixtures; /* /****//* A hello world program *\/ */// ----------------- class Example5 { public static void main(String[] args) { System.out./* */println("/*\"Hello world") ; ///* } /* // */ }
统计工具代码
package countloc; import java.util.regex.Matcher; import java.util.regex.Pattern; public class CountLOC { public static int count(String text) { int count = 0; boolean insideString = false; // 匹配注释的正则表达式 Pattern commentPattern = Pattern.compile("/\\*(.|[\\n\\r])*?\\*/|//.*"); // 移除文本中的注释 Matcher matcher = commentPattern.matcher(text); text = matcher.replaceAll(""); // 将文本按行分割 String[] lines = text.split("\n"); for (String line : lines) { // 检查行中是否包含字符串 for (int i = 0; i < line.length(); i++) { char currentChar = line.charAt(i); System.out.print(currentChar); if (insideString && currentChar == '"' && (i == 0 || line.charAt(i - 1) != '\\')) { insideString = false; } else if (!insideString && currentChar == '"') { insideString = true; } } // 如果不在字符串中、行不为空且不是package声明,则计入可执行代码行 if (!insideString && !line.trim().isEmpty() && !line.trim().startsWith("package ")) { count++; } } return count; } }
测试用例代码
package test; import countloc.CountLOC; import org.junit.Test; import java.io.IOException; import java.nio.file.Files; import java.nio.file.Paths; import static org.junit.Assert.*; public class CountTest { @Test public void shouldHandleBasicCode() throws IOException { String path = "./src/fixtures/Example1.java"; String code = new String(Files.readAllBytes(Paths.get(path))); assertEquals(5, CountLOC.count(code)); } @Test public void shouldHandleABlankLineWIthOneLineComment() throws IOException { String path = "./src/fixtures/Example2.java"; String code = new String(Files.readAllBytes(Paths.get(path))); assertEquals(5, CountLOC.count(code)); } @Test public void shouldHandleAnInlineComment() throws IOException { String path = "./src/fixtures/Example3.java"; String code = new String(Files.readAllBytes(Paths.get(path))); assertEquals(5, CountLOC.count(code)); } @Test public void shouldHandleMultilineCommentsAndQuotes() throws IOException { String path = "./src/fixtures/Example4.java"; String code = new String(Files.readAllBytes(Paths.get(path))); assertEquals(5, CountLOC.count(code)); } @Test public void shouldHandleAComplexExample() throws IOException { String path = "./src/fixtures/Example5.java"; String code = new String(Files.readAllBytes(Paths.get(path))); assertEquals(6, CountLOC.count(code)); } }
问题原因分析
当前代码的核心问题在于:
- 正则表达式的局限性:无法正确处理嵌套注释、包含注释标记的注释内容,也会错误识别字符串内部的注释符号(比如
"/*\"Hello world"里的/*会被误判为注释开头)。 - 处理顺序错误:先移除所有注释再处理字符串,会破坏字符串的原始结构,导致后续的字符串状态判断完全失效。
- 跨行状态丢失:按行分割后,无法延续上一行的字符串或注释状态,导致部分行的代码被错误忽略。
修复方案
正确的处理逻辑应该是逐字符遍历,同时跟踪字符串、单行注释、多行注释的状态,避免正则表达式的缺陷。修复后的代码如下:
package countloc; public class CountLOC { public static int count(String text) { int count = 0; boolean insideString = false; boolean insideLineComment = false; boolean insideBlockComment = false; boolean lineHasCode = false; char[] chars = text.toCharArray(); int length = chars.length; for (int i = 0; i < length; i++) { char c = chars[i]; // 处理多行注释 if (!insideString && !insideLineComment && insideBlockComment) { if (c == '*' && i + 1 < length && chars[i + 1] == '/') { insideBlockComment = false; i++; // 跳过 '/' } continue; } // 处理单行注释 if (!insideString && insideLineComment) { if (c == '\n') { insideLineComment = false; if (lineHasCode) { count++; lineHasCode = false; } } continue; } // 处理字符串 if (insideString) { if (c == '"' && (i == 0 || chars[i - 1] != '\\')) { insideString = false; } lineHasCode = true; continue; } // 进入多行注释 if (c == '/' && i + 1 < length && chars[i + 1] == '*') { insideBlockComment = true; i++; continue; } // 进入单行注释 if (c == '/' && i + 1 < length && chars[i + 1] == '/') { insideLineComment = true; i++; continue; } // 进入字符串 if (c == '"') { insideString = true; lineHasCode = true; continue; } // 识别package声明行,直接跳过整行 if (c == 'p' && i + 7 < length && text.startsWith("package ", i)) { while (i < length && chars[i] != '\n') { i++; } continue; } // 标记当前行存在有效代码 if (!Character.isWhitespace(c)) { lineHasCode = true; } // 换行时统计有效行 if (c == '\n') { if (lineHasCode) { count++; lineHasCode = false; } } } // 处理文件末尾无换行的情况 if (lineHasCode) { count++; } return count; } }
修复说明
- 多状态跟踪:同时维护字符串、单行注释、多行注释三种状态,确保每个字符只在正确的上下文被处理。
- 正确的处理顺序:优先处理注释状态,再处理字符串,避免错误修改字符串内容。
- 逐字符遍历:彻底解决正则表达式无法处理的复杂注释场景,比如嵌套注释、字符串内的注释标记。
- package行单独处理:直接识别并跳过
package开头的行,无需后续判断。
内容的提问来源于stack exchange,提问作者Pascal Oseko
相关产品推荐
相关产品推荐

