Java实现PDF关键词搜索与文件迁移后PDF文件损坏问题排查
问题分析与修复方案
核心问题
当前代码存在两个致命问题:
- 错误使用字符流读取PDF:PDF是二进制格式文件,
FileReader和BufferedReader属于字符流,会把二进制PDF数据当作纯文本解析,既无法正确提取内容做关键词搜索,还可能因文件流未正确释放,导致复制出损坏的PDF文件。 - 重复复制逻辑:代码中只要读到包含关键词的行就触发一次文件复制,多次复制会覆盖目标文件,也可能因文件被占用导致损坏。
修复步骤
1. 使用PDF专用库提取文本
推荐用Apache PDFBox(Java生态常用的PDF处理库)来正确提取PDF中的文本内容。先引入依赖:
如果用Maven,添加到pom.xml:
<dependency> <groupId>org.apache.pdfbox</groupId> <artifactId>pdfbox</artifactId> <version>2.0.32</version> <!-- 可替换为最新稳定版 --> </dependency>
2. 修正代码逻辑
- 用PDFBox提取PDF文本后再做关键词搜索
- 确保文件复制只执行一次(找到关键词后停止搜索并复制)
- 正确管理文件流,避免资源泄漏
修正后的代码示例:
import org.apache.pdfbox.pdmodel.PDDocument; import org.apache.pdfbox.text.PDFTextStripper; import java.io.File; import java.io.IOException; import java.nio.file.Files; import java.nio.file.Paths; import java.util.Scanner; import java.util.HashSet; import java.util.Set; public class PDFSearchCopy { public static void main(String[] args) { Scanner scanner1 = new Scanner(System.in); Set<String> searchedWords = new HashSet<>(); String sourcePdfPath = "C:\\Users\\user012\\Desktop\\Evalution.pdf"; String targetPdfPath = "C:\\Users\\user012\\Desktop\\Search\\Evalution2.pdf"; System.out.println("Enter word you want to search:"); while (scanner1.hasNext()) { String searchWord = scanner1.next(); if (searchWord.equalsIgnoreCase("Exit")) { break; } // 避免重复搜索同一关键词 if (searchedWords.contains(searchWord)) { System.out.println("已搜索过关键词:" + searchWord); continue; } searchedWords.add(searchWord); boolean isFound = false; PDDocument document = null; try { // 用PDFBox加载并提取PDF文本 document = PDDocument.load(new File(sourcePdfPath)); PDFTextStripper stripper = new PDFTextStripper(); String pdfContent = stripper.getText(document); // 检查关键词是否存在 if (pdfContent.contains(searchWord)) { isFound = true; System.out.println("Yes, " + searchWord + " is in the file"); // 二进制复制PDF,保证文件完整 Files.copy(Paths.get(sourcePdfPath), Paths.get(targetPdfPath)); System.out.println("File copied successfully."); } else { System.out.println("No, " + searchWord + " is Not in the file"); } } catch (IOException e) { e.printStackTrace(); } finally { // 关闭PDF文档流 if (document != null) { try { document.close(); } catch (IOException e) { e.printStackTrace(); } } } } scanner1.close(); } }
3. 额外注意事项
- 确保目标目录
C:\\Users\\user012\\Desktop\\Search已存在,否则Files.copy会抛出异常,可提前用Files.createDirectories(Paths.get("C:\\Users\\user012\\Desktop\\Search"))创建目录。 - 若需批量处理多个PDF,可遍历目录下的所有PDF文件逐个处理。
内容的提问来源于stack exchange,提问作者Razan alamri
相关产品推荐
相关产品推荐

