You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Java读取文本文件时如何忽略括号、逗号、句号等标点符号

Java 读取文本文件时过滤单词旁标点的实现方法

问题背景

现有文本文件中部分单词末尾附带句号、逗号,部分单词两侧有括号,使用现有代码读取文件时,这类标点会被计入单词列表。尝试使用replaceAll方法处理后未达到预期效果。

原有代码

文件读取方法

public static Scanner openTextFile(String fileName) {
    Scanner data;

    try{
        data = new Scanner(new File(fileName));
        return data;
    }
    catch(FileNotFoundException e){
        System.out.println(fileName  + " did not read correctly");
    }  
    data = null;
    return data;
}

内容处理方法

public static void readOtherFile(Scanner data, int g[][], Key[] hashTable, int[] keyWordCounter, int modValue)  {
    int lineCounter = 0, wordCounter = 0;
    
    String x;
    String []y;
    while(data.hasNextLine()){
        lineCounter += 1;  
        x = data.nextLine();
        
        /*the following conditional statement takes care of the issue of their being an
         * entirely blank line encountered before reaching the end of the text file.
         */
        if(x.length() == 0) {
            x = data.nextLine();
        }
        
        x = x.toLowerCase();
        x = x.replaceAll("\\p{Punct}", "");
      
        
        y = x.split(" ");
        wordCounter +=  y.length;
        
        //method compares a token to a key word to see if they are identical.
        checkForKeyWord(y, g, hashTable, keyWordCounter, modValue);
    }
    //method prints statistical results
    printResults(lineCounter, wordCounter, hashTable, keyWordCounter);
}

失效原因

原有代码无法生效主要有三个原因:

  • 空行处理逻辑存在漏洞:如果连续出现多行空行,会重复调用nextLine()跳过有效内容,导致部分文本没有被处理
  • 按单空格分割字符串会生成空字符串元素,被误统计为单词
  • \\p{Punct}仅匹配标准ASCII标点,部分特殊格式的标点、括号无法被完全清除

修复方案

方案1:修改原有处理逻辑(保留行统计功能)

直接修改readOtherFile方法即可,不需要改动文件读取逻辑:

public static void readOtherFile(Scanner data, int g[][], Key[] hashTable, int[] keyWordCounter, int modValue)  {
    int lineCounter = 0, wordCounter = 0;
    String x;
    String []y;
    while(data.hasNextLine()){
        lineCounter += 1;  
        x = data.nextLine().trim();
        // 空行直接跳过,不处理也不覆盖内容
        if(x.isEmpty()) {
            continue;
        }
        x = x.toLowerCase();
        // 替换所有非字母、数字、空白的字符,覆盖所有标点和括号
        x = x.replaceAll("[^a-z0-9\\s]", "");
        // 按任意长度的空白字符分割,避免生成空字符串
        y = x.split("\\s+");
        wordCounter +=  y.length;
        checkForKeyWord(y, g, hashTable, keyWordCounter, modValue);
    }
    printResults(lineCounter, wordCounter, hashTable, keyWordCounter);
}

方案2:自定义Scanner分隔符(更简洁,适合仅需读取单词的场景)

直接在创建Scanner时指定分隔符为所有非单词字符,不需要额外做字符串处理:

public static Scanner openTextFile(String fileName) {
    try{
        Scanner data = new Scanner(new File(fileName));
        // 所有非字母的字符都作为分隔符,直接读取纯单词
        data.useDelimiter("[^a-zA-Z]+");
        return data;
    }
    catch(FileNotFoundException e){
        System.out.println(fileName  + " 读取失败");
    }  
    return null;
}

内容的提问来源于stack exchange,提问作者A Mohamed Sakeel

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.04 05:45:03