You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Java实现PDF关键词搜索与文件迁移后PDF文件损坏问题排查

问题分析与修复方案

核心问题

当前代码存在两个致命问题:

  1. 错误使用字符流读取PDF:PDF是二进制格式文件,FileReader和BufferedReader属于字符流,会把二进制PDF数据当作纯文本解析,既无法正确提取内容做关键词搜索,还可能因文件流未正确释放,导致复制出损坏的PDF文件。
  2. 重复复制逻辑:代码中只要读到包含关键词的行就触发一次文件复制,多次复制会覆盖目标文件,也可能因文件被占用导致损坏。

修复步骤

1. 使用PDF专用库提取文本

推荐用Apache PDFBox(Java生态常用的PDF处理库)来正确提取PDF中的文本内容。先引入依赖:
如果用Maven,添加到pom.xml:

<dependency>
    <groupId>org.apache.pdfbox</groupId>
    <artifactId>pdfbox</artifactId>
    <version>2.0.32</version> <!-- 可替换为最新稳定版 -->
</dependency>

2. 修正代码逻辑

  • 用PDFBox提取PDF文本后再做关键词搜索
  • 确保文件复制只执行一次(找到关键词后停止搜索并复制)
  • 正确管理文件流,避免资源泄漏

修正后的代码示例:

import org.apache.pdfbox.pdmodel.PDDocument;
import org.apache.pdfbox.text.PDFTextStripper;

import java.io.File;
import java.io.IOException;
import java.nio.file.Files;
import java.nio.file.Paths;
import java.util.Scanner;
import java.util.HashSet;
import java.util.Set;

public class PDFSearchCopy {
    public static void main(String[] args) {
        Scanner scanner1 = new Scanner(System.in);
        Set<String> searchedWords = new HashSet<>();
        String sourcePdfPath = "C:\\Users\\user012\\Desktop\\Evalution.pdf";
        String targetPdfPath = "C:\\Users\\user012\\Desktop\\Search\\Evalution2.pdf";

        System.out.println("Enter word you want to search:");
        while (scanner1.hasNext()) {
            String searchWord = scanner1.next();
            if (searchWord.equalsIgnoreCase("Exit")) {
                break;
            }

            // 避免重复搜索同一关键词
            if (searchedWords.contains(searchWord)) {
                System.out.println("已搜索过关键词:" + searchWord);
                continue;
            }
            searchedWords.add(searchWord);

            boolean isFound = false;
            PDDocument document = null;
            try {
                // 用PDFBox加载并提取PDF文本
                document = PDDocument.load(new File(sourcePdfPath));
                PDFTextStripper stripper = new PDFTextStripper();
                String pdfContent = stripper.getText(document);

                // 检查关键词是否存在
                if (pdfContent.contains(searchWord)) {
                    isFound = true;
                    System.out.println("Yes, " + searchWord + " is in the file");
                    // 二进制复制PDF,保证文件完整
                    Files.copy(Paths.get(sourcePdfPath), Paths.get(targetPdfPath));
                    System.out.println("File copied successfully.");
                } else {
                    System.out.println("No, " + searchWord + " is Not in the file");
                }
            } catch (IOException e) {
                e.printStackTrace();
            } finally {
                // 关闭PDF文档流
                if (document != null) {
                    try {
                        document.close();
                    } catch (IOException e) {
                        e.printStackTrace();
                    }
                }
            }
        }
        scanner1.close();
    }
}

3. 额外注意事项

  • 确保目标目录C:\\Users\\user012\\Desktop\\Search已存在,否则Files.copy会抛出异常,可提前用Files.createDirectories(Paths.get("C:\\Users\\user012\\Desktop\\Search"))创建目录。
  • 若需批量处理多个PDF,可遍历目录下的所有PDF文件逐个处理。

内容的提问来源于stack exchange,提问作者Razan alamri

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 08:01:20