You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Java实现文本替换:基于HashMap替换过度用词并保留格式

Java文本优化工具:替换过度用词并保留格式

需求概述

  • 已通过构造方法从文件读取过度用词映射,存入HashMap<String, String>(键为全小写的过度用词,值为对应替换词),示例映射包括:
    • amazing:astonishing
    • interesting:intriguing
    • literally:frankly
    • nice:pleasant
    • hard:taxing
    • change:transform
  • 需要完善improveText方法,实现:
    1. 读取目标文本文件,将其中的过度用词替换为对应替换词
    2. 严格保留原文件的标点符号位置和样式
    3. 输入文件中的单词仅为三种格式:全小写、首字母大写、全大写

已实现代码(构造方法)

import java.io.BufferedReader;
import java.io.FileReader;
import java.io.IOException;
import java.util.HashMap;

public class TextImprover {

    private HashMap<String, String> wordMap ;

    /**
     * 构造方法:从文件读取过度用词与替换词的映射
     * 
     * @param wordMapFileName   存储映射关系的文件名
     */
    public TextImprover(String wordMapFileName) { 
        this.wordMap = new HashMap<String,String>();
        try {
            BufferedReader br = new BufferedReader(new FileReader(wordMapFileName));
            String line ;
            while((line = br.readLine())!= null) {
                String[] wordLine = line.split("\t");
                String overUsedWord = wordLine[0].trim();
                String replaceWord = wordLine[1].trim();
                
                wordMap.put(overUsedWord, replaceWord);
            }
            br.close();
                
        }catch(IOException e){
            System.out.println("文件操作异常: " + e.getMessage());   
        }
    }

待完善方法(improveText)

原框架代码:

/**
     * 将指定文件中的过度用词替换为更合适的词汇,基于构造时传入的映射表
     * 
     * @param fileName  需要优化的文本文件名
     */
    public void improveText(String fileName) {
        try {
            BufferedReader br = new BufferedReader(new FileReader(fileName));
            String line ;
            while((line = br.readLine())!= null) {
                String[] lineWords = line.split(" ");
                // 此处需要补充核心替换逻辑
            }
            br.close();
                
        }catch(IOException e){
            System.out.println("文件操作异常: " + e.getMessage());  
        }
    }
}

示例输入输出

  • 原文本内容:

"Amazing" is really the best way I can think of to describe it. Literally, it is hard to express how much I liked it. It was amazingly NICE!!!!! Good, not bad. I wouldn't change a bit of it. Please, be nice and help me fix my writing!! b BB bbb Bb B.

  • 优化后文本内容:

"Astonishing" is really the best way I can think of to describe it. Frankly, it is taxing to express how much I liked it. It was amazingly PLEASANT!!!!! Superior, not inferior. I wouldn't transform a bit of it. Please, be pleasant and help me fix my writing!! cat BB bbb Bb CAT

核心替换逻辑实现

直接按空格分割会导致单词和标点粘连(如Amazing"、Literally,),需用正则表达式拆分单词与标点,再按原格式调整替换词:

完善后的完整improveText方法

import java.io.BufferedWriter;
import java.io.FileWriter;
import java.util.regex.Matcher;
import java.util.regex.Pattern;

// 补充到TextImprover类中
public void improveText(String fileName) {
    // 定义正则:匹配开头的标点+单词+结尾的标点,分组捕获
    Pattern wordPattern = Pattern.compile("(^[^a-zA-Z]*)([a-zA-Z]+)([^a-zA-Z]*$)");
    
    try (BufferedReader br = new BufferedReader(new FileReader(fileName));
         BufferedWriter bw = new BufferedWriter(new FileWriter(fileName.replace(".txt", "_improved.txt")))) { // 输出到新文件避免覆盖原文件
        
        String line;
        while ((line = br.readLine()) != null) {
            StringBuilder newLine = new StringBuilder();
            // 按空格分割每个单元(可能包含单词和标点)
            String[] units = line.split(" ");
            
            for (String unit : units) {
                Matcher matcher = wordPattern.matcher(unit);
                if (matcher.matches()) {
                    String prefix = matcher.group(1); // 单词前的标点
                    String word = matcher.group(2);   // 纯单词部分
                    String suffix = matcher.group(3); // 单词后的标点
                    
                    // 转全小写查映射表
                    String lowerWord = word.toLowerCase();
                    if (wordMap.containsKey(lowerWord)) {
                        String replaceWord = wordMap.get(lowerWord);
                        // 根据原单词格式调整替换词格式
                        if (word.equals(word.toUpperCase())) {
                            // 原单词全大写,替换词也转全大写
                            replaceWord = replaceWord.toUpperCase();
                        } else if (Character.isUpperCase(word.charAt(0)) && word.substring(1).equals(word.substring(1).toLowerCase())) {
                            // 原单词首字母大写,替换词也首字母大写
                            replaceWord = Character.toUpperCase(replaceWord.charAt(0)) + replaceWord.substring(1).toLowerCase();
                        }
                        // 原单词全小写,替换词保持原格式(已全小写)
                        
                        // 拼接前缀+替换词+后缀
                        newLine.append(prefix).append(replaceWord).append(suffix);
                    } else {
                        // 无替换词,直接保留原单元
                        newLine.append(unit);
                    }
                } else {
                    // 不匹配正则(无有效单词),直接保留
                    newLine.append(unit);
                }
                // 添加空格(最后一个单元后不需要,但不影响可读性)
                newLine.append(" ");
            }
            // 移除最后多余的空格,写入新行
            bw.write(newLine.toString().trim());
            bw.newLine();
        }
    } catch (IOException e) {
        System.out.println("文件操作异常: " + e.getMessage());
    }
}

关键逻辑说明

  1. 正则拆分:用(^[^a-zA-Z]*)([a-zA-Z]+)([^a-zA-Z]*$)匹配每个单元,拆分出单词前后的标点和纯单词部分
  2. 格式匹配:判断原单词是全大写、首字母大写还是全小写,对应调整替换词的格式
  3. 文件写入:使用BufferedWriter将优化后的内容写入新文件(避免覆盖原文件)
  4. 异常处理:合并IO异常捕获,简化代码

内容的提问来源于stack exchange,提问作者IO23

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 16:50:22