You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

拼写检查场景下如何写入新文本文件并保留空格?

问题解决:拼写检查程序保留原文件空格

你的代码无法保留原文件空格的核心原因是使用了Scanner.next()方法——这个方法会自动跳过所有空白字符(空格、制表符、换行等),只返回连续的非空白"token",导致输出时无法还原原文件的空格布局。此外你的字典读取逻辑还存在一个小bug:会漏掉最后一行的字典单词。

修复方案

  1. 改用Scanner.nextLine()逐行读取目标文件,完整保留每行的空格和换行结构;
  2. 使用正则表达式匹配每行中的「单词片段」和「非单词片段(包括空格、特殊字符)」,逐个处理后拼接回原格式;
  3. 修正字典读取的循环逻辑,确保所有字典单词都被加入HashSet。

修改后的完整代码

import java.io.File;
import java.io.FileReader;
import java.io.PrintWriter;
import java.util.HashSet;
import java.util.Scanner;
import java.util.regex.Matcher;
import java.util.regex.Pattern;

public class SpellChecker {
    public static void spellCheck(File processingFile, File dictionaryFile) {
        try {
            HashSet<String> dictionary = new HashSet<>();
            PrintWriter outputFile = new PrintWriter("test_spellChecked.txt");
            
            // 修复字典读取逻辑:确保所有行都被读取
            Scanner scnr = new Scanner(dictionaryFile);
            while (scnr.hasNextLine()) {
                String word = scnr.nextLine().trim(); // 去掉可能的首尾空白
                if (!word.isEmpty()) {
                    dictionary.add(word.toLowerCase()); // 统一转小写,避免大小写问题
                }
            }
            scnr.close();

            // 逐行读取目标文件,保留原格式
            scnr = new Scanner(processingFile);
            // 匹配所有单词(a-z/A-Z)和非单词片段(包括空格、特殊字符)的正则
            Pattern pattern = Pattern.compile("([a-zA-Z]+)|([^a-zA-Z]+)");
            while (scnr.hasNextLine()) {
                String line = scnr.nextLine();
                Matcher matcher = pattern.matcher(line);
                while (matcher.find()) {
                    String segment = matcher.group();
                    // 判断当前片段是否是纯字母单词
                    if (segment.matches("[a-zA-Z]+")) {
                        if (!dictionary.contains(segment.toLowerCase())) {
                            outputFile.print("<" + segment + ">");
                        } else {
                            outputFile.print(segment);
                        }
                    } else {
                        // 非单词片段(空格、特殊字符等)直接输出,保留原格式
                        outputFile.print(segment);
                    }
                }
                // 输出换行符,保留原文件的换行结构
                outputFile.println();
            }

            outputFile.close();
            System.out.println("文件检查完成。");
        } catch (Exception e) {
            e.printStackTrace();
        }
    }

    public static void main(String[] args) {
        // 示例调用,根据实际路径修改
        spellCheck(new File("input.txt"), new File("dictionary.txt"));
    }
}

关键说明

  • 正则表达式([a-zA-Z]+)|([^a-zA-Z]+):把每行内容拆分成「纯字母单词」和「其他所有内容(包括空格、标点、换行)」两个类型的片段,确保不会丢失任何原格式字符;
  • 逐行处理:通过nextLine()读取完整行,处理后用println()输出换行,完美保留原文件的行结构;
  • 字典统一转小写:避免因为大小写差异导致的误判(比如字典里是"apple",文件里是"Apple"也能匹配);
  • 修复字典读取:用hasNextLine()循环,确保最后一行的单词也被加入HashSet。

内容的提问来源于stack exchange,提问作者h0neybugz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 22:42:45