You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过Java代码提取PDF文件中的嵌入式附件?

完整Java实现提取PDF嵌入式附件

我来给你补全并整理好完整的可用代码,基于你提供的iText代码片段扩展,用来提取PDF中的嵌入式附件:

首先确认依赖

如果你用Maven管理项目,需要在pom.xml中添加iText 5的依赖(注意这是iText 5版本,和iText 7的包结构不同):

<dependency>
    <groupId>com.itextpdf</groupId>
    <artifactId>itextpdf</artifactId>
    <version>5.5.13.3</version>
</dependency>

完整实现代码

import java.io.File;
import java.io.FileOutputStream;
import java.io.IOException;
import com.itextpdf.text.pdf.PRStream;
import com.itextpdf.text.pdf.PdfArray;
import com.itextpdf.text.pdf.PdfDictionary;
import com.itextpdf.text.pdf.PdfName;
import com.itextpdf.text.pdf.PdfReader;
import com.itextpdf.text.pdf.PdfString;

public class ExtractAttachments {

    public ExtractAttachments(String srcPdfPath, String outputDir) throws IOException {
        // 创建输出目录(不存在则自动创建)
        File outputFolder = new File(outputDir);
        if (!outputFolder.exists()) {
            outputFolder.mkdirs();
        }

        // 初始化PDF读取器
        PdfReader reader = new PdfReader(srcPdfPath);
        try {
            // 获取PDF的Names字典,里面存储了嵌入式附件的信息
            PdfDictionary catalog = reader.getCatalog();
            PdfDictionary names = catalog.getAsDict(PdfName.NAMES);
            if (names == null) {
                System.out.println("该PDF没有嵌入式附件");
                return;
            }

            PdfDictionary embeddedFiles = names.getAsDict(PdfName.EMBEDDEDFILES);
            if (embeddedFiles == null) {
                System.out.println("该PDF没有嵌入式附件");
                return;
            }

            PdfArray filesArray = embeddedFiles.getAsArray(PdfName.NAMES);
            if (filesArray == null || filesArray.size() == 0) {
                System.out.println("该PDF没有嵌入式附件");
                return;
            }

            // 遍历附件列表:filesArray的结构是 [文件名1, 文件字典1, 文件名2, 文件字典2,...]
            for (int i = 0; i < filesArray.size(); i += 2) {
                // 获取附件文件名
                PdfString fileName = filesArray.getAsString(i);
                if (fileName == null) {
                    continue;
                }
                String attachmentName = fileName.toUnicodeString();

                // 获取附件对应的文件字典
                PdfDictionary fileDict = filesArray.getAsDict(i + 1);
                if (fileDict == null) {
                    continue;
                }

                // 获取文件流字典
                PdfDictionary fileStreamDict = fileDict.getAsDict(PdfName.EF);
                if (fileStreamDict == null) {
                    continue;
                }

                // 提取文件流(这里默认取F条目,大多数附件存在这里)
                PRStream fileStream = (PRStream) PdfReader.getPdfObject(fileStreamDict.get(PdfName.F));
                if (fileStream == null) {
                    continue;
                }

                // 读取流中的字节内容
                byte[] fileContent = PdfReader.getStreamBytes(fileStream);

                // 将附件写入输出目录
                File outputFile = new File(outputFolder, attachmentName);
                try (FileOutputStream fos = new FileOutputStream(outputFile)) {
                    fos.write(fileContent);
                    System.out.println("成功提取附件:" + attachmentName);
                } catch (IOException e) {
                    System.err.println("写入附件失败:" + attachmentName + ",错误信息:" + e.getMessage());
                }
            }
        } finally {
            // 关闭PDF读取器,释放资源
            reader.close();
        }
    }

    // 主方法用于测试
    public static void main(String[] args) {
        try {
            // 替换为你的源PDF路径和输出目录路径
            String srcPdf = "path/to/your/source.pdf";
            String outputDir = "path/to/your/output/folder";
            new ExtractAttachments(srcPdf, outputDir);
        } catch (IOException e) {
            e.printStackTrace();
        }
    }
}

关键代码说明

  • 目录初始化:自动创建输出目录,避免因目录不存在导致写入失败
  • PDF结构解析:通过PDF的Catalog -> Names -> EmbeddedFiles路径定位附件信息,这是PDF存储嵌入式附件的标准结构
  • 流处理:从PRStream中读取附件的字节内容,确保完整获取附件数据
  • 资源释放:在finally块中关闭PdfReader,防止资源泄漏
  • 异常处理:对文件写入过程的异常进行捕获和提示,避免程序崩溃

注意事项

  1. 这段代码基于iText 5.x版本开发,如果你使用iText 7,包结构和API会有较大差异,需要调整代码
  2. 部分PDF可能会用不同的键存储附件流(比如UF),如果遇到提取失败的情况,可以尝试替换fileStreamDict.get(PdfName.F)为fileStreamDict.get(PdfName.UF)
  3. 确保程序有读取源PDF和写入输出目录的权限

内容的提问来源于stack exchange,提问作者Sonia

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 07:44:37