Java读取文本文件时如何忽略括号、逗号、句号等标点符号
Java 读取文本文件时过滤单词旁标点的实现方法
问题背景
现有文本文件中部分单词末尾附带句号、逗号,部分单词两侧有括号,使用现有代码读取文件时,这类标点会被计入单词列表。尝试使用replaceAll方法处理后未达到预期效果。
原有代码
文件读取方法
public static Scanner openTextFile(String fileName) { Scanner data; try{ data = new Scanner(new File(fileName)); return data; } catch(FileNotFoundException e){ System.out.println(fileName + " did not read correctly"); } data = null; return data; }
内容处理方法
public static void readOtherFile(Scanner data, int g[][], Key[] hashTable, int[] keyWordCounter, int modValue) { int lineCounter = 0, wordCounter = 0; String x; String []y; while(data.hasNextLine()){ lineCounter += 1; x = data.nextLine(); /*the following conditional statement takes care of the issue of their being an * entirely blank line encountered before reaching the end of the text file. */ if(x.length() == 0) { x = data.nextLine(); } x = x.toLowerCase(); x = x.replaceAll("\\p{Punct}", ""); y = x.split(" "); wordCounter += y.length; //method compares a token to a key word to see if they are identical. checkForKeyWord(y, g, hashTable, keyWordCounter, modValue); } //method prints statistical results printResults(lineCounter, wordCounter, hashTable, keyWordCounter); }
失效原因
原有代码无法生效主要有三个原因:
- 空行处理逻辑存在漏洞:如果连续出现多行空行,会重复调用
nextLine()跳过有效内容,导致部分文本没有被处理 - 按单空格分割字符串会生成空字符串元素,被误统计为单词
\\p{Punct}仅匹配标准ASCII标点,部分特殊格式的标点、括号无法被完全清除
修复方案
方案1:修改原有处理逻辑(保留行统计功能)
直接修改readOtherFile方法即可,不需要改动文件读取逻辑:
public static void readOtherFile(Scanner data, int g[][], Key[] hashTable, int[] keyWordCounter, int modValue) { int lineCounter = 0, wordCounter = 0; String x; String []y; while(data.hasNextLine()){ lineCounter += 1; x = data.nextLine().trim(); // 空行直接跳过,不处理也不覆盖内容 if(x.isEmpty()) { continue; } x = x.toLowerCase(); // 替换所有非字母、数字、空白的字符,覆盖所有标点和括号 x = x.replaceAll("[^a-z0-9\\s]", ""); // 按任意长度的空白字符分割,避免生成空字符串 y = x.split("\\s+"); wordCounter += y.length; checkForKeyWord(y, g, hashTable, keyWordCounter, modValue); } printResults(lineCounter, wordCounter, hashTable, keyWordCounter); }
方案2:自定义Scanner分隔符(更简洁,适合仅需读取单词的场景)
直接在创建Scanner时指定分隔符为所有非单词字符,不需要额外做字符串处理:
public static Scanner openTextFile(String fileName) { try{ Scanner data = new Scanner(new File(fileName)); // 所有非字母的字符都作为分隔符,直接读取纯单词 data.useDelimiter("[^a-zA-Z]+"); return data; } catch(FileNotFoundException e){ System.out.println(fileName + " 读取失败"); } return null; }
内容的提问来源于stack exchange,提问作者A Mohamed Sakeel
相关产品推荐
相关产品推荐

