使用PDFBox修改已有PDF的作者/标题元数据不生效问题
问题描述
使用PDFBox v2.0.26修改已有PDF文件(包括Adobe Acrobat Pro DC创建的、PrinceXML生成后经PDFBox CLI合并的文件)的作者、标题等元数据时出现异常:通过代码读取处理后的PDF能获取到新的元数据值,但在Adobe Acrobat Pro DC(2022.002版本)中查看元数据仍显示旧值。该代码在创建新PDF时可正常生效,仅修改已有PDF时失效。
复现步骤
- 打开Adobe Acrobat Pro DC(2022.002版本)
- 文件>创建>空白页面
- 按Ctrl+D,输入作者(如“author here”)和标题(如“title here”)
- 关闭对话框并保存文档
- 运行下方Java代码
首次运行代码输出:
Existing author: author here
Existing title: title here
第二次运行代码输出:
Existing author: My new author
Existing title: My new title
但在Acrobat中按Ctrl+D查看元数据时,仍显示:
Title: title here
Author: author here
原代码如下:
import java.io.File; import java.io.IOException; import org.apache.pdfbox.pdmodel.PDDocument; import org.apache.pdfbox.pdmodel.PDDocumentInformation; public class setTitle { public static void main(String args[]) throws IOException { String filepath = "C:\\Temp\\temp.pdf"; //Loading an existing document File file = new File(filepath); PDDocument document = PDDocument.load(file); //Creating the PDDocumentInformation object PDDocumentInformation pdd = document.getDocumentInformation(); System.out.println("Existing author: " + pdd.getAuthor()); System.out.println("Existing title: " + pdd.getTitle()); //Setting the author of the document pdd.setAuthor("My new author"); // Setting the title of the document pdd.setTitle("My new title"); //Saving the document document.save("C:/Temp/temp.pdf"); //Closing the document document.close(); } }
问题原因
Adobe Acrobat在创建或保存PDF时,会将元数据同时存储在两个位置:
- PDF文档根目录的Info字典(这是PDFBox默认修改的位置)
- XMP元数据(嵌入在PDF中的XML格式元数据,Acrobat优先读取这个位置的内容)
原代码仅修改了Info字典,但Acrobat显示元数据时优先读取XMP中的值,所以导致代码读取正常,但Acrobat显示旧值。
解决方案
需要同时更新Info字典和XMP元数据,修改后的代码如下:
import java.io.File; import java.io.IOException; import org.apache.pdfbox.pdmodel.PDDocument; import org.apache.pdfbox.pdmodel.PDDocumentCatalog; import org.apache.pdfbox.pdmodel.PDDocumentInformation; import org.apache.pdfbox.pdmodel.xmp.PDXMPMetadata; import org.apache.xmpbox.XMPMetadata; import org.apache.xmpbox.schema.DublinCoreSchema; import org.apache.xmpbox.schema.PDFSchema; import org.apache.xmpbox.xml.XmpSerializer; public class UpdatePdfMetadata { public static void main(String args[]) throws IOException { String filepath = "C:\\Temp\\temp.pdf"; // 加载现有文档 File file = new File(filepath); try (PDDocument document = PDDocument.load(file)) { // 更新Info字典元数据 PDDocumentInformation info = document.getDocumentInformation(); System.out.println("Existing author: " + info.getAuthor()); System.out.println("Existing title: " + info.getTitle()); info.setAuthor("My new author"); info.setTitle("My new title"); // 更新XMP元数据 PDDocumentCatalog catalog = document.getDocumentCatalog(); PDXMPMetadata xmp = catalog.getMetadata(); if (xmp == null) { // 如果文档没有XMP元数据,创建新的 xmp = new PDXMPMetadata(document); catalog.setMetadata(xmp); } XMPMetadata metadata = xmp.getXMPMetadata(); // 更新Dublin Core schema中的作者和标题 DublinCoreSchema dcSchema = metadata.getDublinCoreSchema(); if (dcSchema == null) { dcSchema = metadata.createAndAddDublinCoreSchema(); } dcSchema.setCreator("My new author"); dcSchema.setTitle("My new title"); // 更新PDF schema中的元数据(确保一致性) PDFSchema pdfSchema = metadata.getPDFSchema(); if (pdfSchema == null) { pdfSchema = metadata.createAndAddPDFSchema(); } pdfSchema.setAuthor("My new author"); pdfSchema.setTitle("My new title"); // 序列化XMP元数据并保存回文档 xmp.embedXMPMetadata(new XmpSerializer().serializeToString(metadata)); // 保存文档 document.save(filepath); } } }
说明
- PDFBox v2.0.26已包含XMPBox相关依赖,无需额外引入,但需确保项目依赖完整
- 代码同时更新了Info字典和XMP元数据的DublinCore、PDF两个schema,确保Acrobat读取时能获取到新值
- 使用try-with-resources语法自动关闭文档,避免资源泄漏
内容的提问来源于stack exchange,提问作者Crafty Cat
相关产品推荐
相关产品推荐

