Java实现文本文件去重:按--分隔块、过滤3行及以上重复内容块
Java文本块去重实现问题
功能需求
- 逐行读取文本文件,以
--为分隔符将内容拆分为多个块 - 首个块直接写入新文件
- 后续每个块需要与已写入新文件的所有块逐行对比,若该块与已有块存在3行及以上的重复行,则不写入该块,否则写入新文件并补充
--分隔符 - 循环处理直至文件结束
现有实现代码
public void removeDuplicateErr(String data) throws IOException { String contents = new String(Files.readAllBytes(Paths.get(data))); String[] blocks = contents.split("--"); String fileName = "output.txt"; PrintWriter pw = new PrintWriter(fileName); int count = 0; int count1 = 0; for (String block : blocks) { boolean flag = false; if(count > 0) { String contents1 = new String(Files.readAllBytes(Paths.get(fileName))); String[] blocks1 = contents1.split("--"); for(String block1 : blocks1) { BufferedReader br1 = new BufferedReader(new StringReader(block1)); String line1 = br1.readLine(); while (line1 != null) { BufferedReader br2 = new BufferedReader(new StringReader(block)); String line2 = br2.readLine(); while (line2 != null) { if(line1.equals(line2)) { count1++; if(count1 >= 3) { flag = true; break; } } line2 = br2.readLine(); } line1 = br1.readLine(); } if (!flag) { pw.print(block); pw.print("--"); pw.flush(); } } } if(count < 1) { pw.print(block); pw.print("--"); pw.flush(); } count++; } pw.close(); }
输入样例
test 1 test 2 test 3 test 4 test 5 -- test 6 test 2 test 3 test 4 test 12 -- test 8 test 9 test 10 test 11 test 12 -- test 1 test 3 test 4 test 21 test 22 -- test 1 test 2 test 3 test 4 test 5 -- test 50 test 51 test 52 test 53 test 54 test 55 -- test 53 test 54 test 55 test 56 test 57
期望输出结果
test 1 test 2 test 3 test 4 test 5 -- test 8 test 9 test 10 test 11 test 12 -- test 50 test 51 test 52 test 53 test 54 test 55
现有代码问题分析
- 重复行数计数器
count1没有在每个新块对比前重置,会累计所有块的重复行数,导致判断逻辑完全错误 - 每次对比都重新读取输出文件,IO效率极低,完全可以用内存集合存储已写入块的行信息
- 未等和所有已有块对比完成就提前写入,只要和其中一个块重复行数不足3就会直接写入,没有判断和其他块的重复情况
- 没有处理块拆分后自带的首尾空白、空换行问题,会导致行对比结果错误
- 每个块写入都加
--分隔符,会导致最终输出末尾多一个多余的分隔符
修复后代码
import java.io.*; import java.nio.file.Files; import java.nio.file.Paths; import java.util.ArrayList; import java.util.HashSet; import java.util.List; import java.util.Set; public class BlockDedup { public void removeDuplicateErr(String data) throws IOException { String contents = new String(Files.readAllBytes(Paths.get(data))); String[] blocks = contents.split("--"); // 内存存储已写入块的去重行集合,避免重复读文件 List<Set<String>> writtenBlockLines = new ArrayList<>(); String fileName = "output.txt"; PrintWriter pw = new PrintWriter(new OutputStreamWriter(new FileOutputStream(fileName), "UTF-8")); boolean isFirstBlock = true; for (String rawBlock : blocks) { // 处理块首尾空白,过滤空块 String block = rawBlock.trim(); if (block.isEmpty()) continue; // 提取当前块的所有非空行 Set<String> currentLines = new HashSet<>(); try (BufferedReader br = new BufferedReader(new StringReader(block))) { String line; while ((line = br.readLine()) != null) { String trimLine = line.trim(); if (!trimLine.isEmpty()) { currentLines.add(trimLine); } } } // 首个块直接写入 if (isFirstBlock) { pw.print(block); writtenBlockLines.add(currentLines); isFirstBlock = false; continue; } // 和所有已写入块对比,只要有一个块重复行>=3就跳过 boolean shouldSkip = false; for (Set<String> writtenLines : writtenBlockLines) { int duplicateCount = 0; for (String line : currentLines) { if (writtenLines.contains(line)) { duplicateCount++; if (duplicateCount >= 3) { shouldSkip = true; break; } } } if (shouldSkip) break; } // 符合写入条件则先加分隔符再写块 if (!shouldSkip) { pw.print("\n--\n"); pw.print(block); pw.flush(); writtenBlockLines.add(currentLines); } } pw.close(); } public static void main(String[] args) throws IOException { new BlockDedup().removeDuplicateErr("input.txt"); } }
内容的提问来源于stack exchange,提问作者Georgi Atanasov
相关产品推荐
相关产品推荐

