如何通过Java代码提取PDF文件中的嵌入式附件?
完整Java实现提取PDF嵌入式附件
我来给你补全并整理好完整的可用代码,基于你提供的iText代码片段扩展,用来提取PDF中的嵌入式附件:
首先确认依赖
如果你用Maven管理项目,需要在pom.xml中添加iText 5的依赖(注意这是iText 5版本,和iText 7的包结构不同):
<dependency> <groupId>com.itextpdf</groupId> <artifactId>itextpdf</artifactId> <version>5.5.13.3</version> </dependency>
完整实现代码
import java.io.File; import java.io.FileOutputStream; import java.io.IOException; import com.itextpdf.text.pdf.PRStream; import com.itextpdf.text.pdf.PdfArray; import com.itextpdf.text.pdf.PdfDictionary; import com.itextpdf.text.pdf.PdfName; import com.itextpdf.text.pdf.PdfReader; import com.itextpdf.text.pdf.PdfString; public class ExtractAttachments { public ExtractAttachments(String srcPdfPath, String outputDir) throws IOException { // 创建输出目录(不存在则自动创建) File outputFolder = new File(outputDir); if (!outputFolder.exists()) { outputFolder.mkdirs(); } // 初始化PDF读取器 PdfReader reader = new PdfReader(srcPdfPath); try { // 获取PDF的Names字典,里面存储了嵌入式附件的信息 PdfDictionary catalog = reader.getCatalog(); PdfDictionary names = catalog.getAsDict(PdfName.NAMES); if (names == null) { System.out.println("该PDF没有嵌入式附件"); return; } PdfDictionary embeddedFiles = names.getAsDict(PdfName.EMBEDDEDFILES); if (embeddedFiles == null) { System.out.println("该PDF没有嵌入式附件"); return; } PdfArray filesArray = embeddedFiles.getAsArray(PdfName.NAMES); if (filesArray == null || filesArray.size() == 0) { System.out.println("该PDF没有嵌入式附件"); return; } // 遍历附件列表:filesArray的结构是 [文件名1, 文件字典1, 文件名2, 文件字典2,...] for (int i = 0; i < filesArray.size(); i += 2) { // 获取附件文件名 PdfString fileName = filesArray.getAsString(i); if (fileName == null) { continue; } String attachmentName = fileName.toUnicodeString(); // 获取附件对应的文件字典 PdfDictionary fileDict = filesArray.getAsDict(i + 1); if (fileDict == null) { continue; } // 获取文件流字典 PdfDictionary fileStreamDict = fileDict.getAsDict(PdfName.EF); if (fileStreamDict == null) { continue; } // 提取文件流(这里默认取F条目,大多数附件存在这里) PRStream fileStream = (PRStream) PdfReader.getPdfObject(fileStreamDict.get(PdfName.F)); if (fileStream == null) { continue; } // 读取流中的字节内容 byte[] fileContent = PdfReader.getStreamBytes(fileStream); // 将附件写入输出目录 File outputFile = new File(outputFolder, attachmentName); try (FileOutputStream fos = new FileOutputStream(outputFile)) { fos.write(fileContent); System.out.println("成功提取附件:" + attachmentName); } catch (IOException e) { System.err.println("写入附件失败:" + attachmentName + ",错误信息:" + e.getMessage()); } } } finally { // 关闭PDF读取器,释放资源 reader.close(); } } // 主方法用于测试 public static void main(String[] args) { try { // 替换为你的源PDF路径和输出目录路径 String srcPdf = "path/to/your/source.pdf"; String outputDir = "path/to/your/output/folder"; new ExtractAttachments(srcPdf, outputDir); } catch (IOException e) { e.printStackTrace(); } } }
关键代码说明
- 目录初始化:自动创建输出目录,避免因目录不存在导致写入失败
- PDF结构解析:通过PDF的Catalog -> Names -> EmbeddedFiles路径定位附件信息,这是PDF存储嵌入式附件的标准结构
- 流处理:从PRStream中读取附件的字节内容,确保完整获取附件数据
- 资源释放:在finally块中关闭PdfReader,防止资源泄漏
- 异常处理:对文件写入过程的异常进行捕获和提示,避免程序崩溃
注意事项
- 这段代码基于iText 5.x版本开发,如果你使用iText 7,包结构和API会有较大差异,需要调整代码
- 部分PDF可能会用不同的键存储附件流(比如
UF),如果遇到提取失败的情况,可以尝试替换fileStreamDict.get(PdfName.F)为fileStreamDict.get(PdfName.UF) - 确保程序有读取源PDF和写入输出目录的权限
内容的提问来源于stack exchange,提问作者Sonia
相关产品推荐
相关产品推荐

