You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用PDFBox修改已有PDF的作者/标题元数据不生效问题

问题描述

使用PDFBox v2.0.26修改已有PDF文件(包括Adobe Acrobat Pro DC创建的、PrinceXML生成后经PDFBox CLI合并的文件)的作者、标题等元数据时出现异常:通过代码读取处理后的PDF能获取到新的元数据值,但在Adobe Acrobat Pro DC(2022.002版本)中查看元数据仍显示旧值。该代码在创建新PDF时可正常生效,仅修改已有PDF时失效。

复现步骤

  • 打开Adobe Acrobat Pro DC(2022.002版本)
  • 文件>创建>空白页面
  • 按Ctrl+D,输入作者(如“author here”)和标题(如“title here”)
  • 关闭对话框并保存文档
  • 运行下方Java代码

首次运行代码输出:

Existing author: author here
Existing title: title here

第二次运行代码输出:

Existing author: My new author
Existing title: My new title

但在Acrobat中按Ctrl+D查看元数据时,仍显示:

Title: title here
Author: author here

原代码如下:

import java.io.File;
import java.io.IOException; 
import org.apache.pdfbox.pdmodel.PDDocument;
import org.apache.pdfbox.pdmodel.PDDocumentInformation;

public class setTitle {
   public static void main(String args[]) throws IOException {

       String filepath = "C:\\Temp\\temp.pdf";
    
       //Loading an existing document 
       File file = new File(filepath); 
       PDDocument document = PDDocument.load(file); 
     
       //Creating the PDDocumentInformation object 
       PDDocumentInformation pdd = document.getDocumentInformation();
    
       System.out.println("Existing author: " + pdd.getAuthor());
       System.out.println("Existing title: " + pdd.getTitle());
       
       //Setting the author of the document
       pdd.setAuthor("My new author");
       
       // Setting the title of the document
       pdd.setTitle("My new title"); 
       
       //Saving the document 
       document.save("C:/Temp/temp.pdf");
    
      //Closing the document
      document.close();
    
   }
}
问题原因

Adobe Acrobat在创建或保存PDF时,会将元数据同时存储在两个位置:

  1. PDF文档根目录的Info字典(这是PDFBox默认修改的位置)
  2. XMP元数据(嵌入在PDF中的XML格式元数据,Acrobat优先读取这个位置的内容)

原代码仅修改了Info字典,但Acrobat显示元数据时优先读取XMP中的值,所以导致代码读取正常,但Acrobat显示旧值。

解决方案

需要同时更新Info字典和XMP元数据,修改后的代码如下:

import java.io.File;
import java.io.IOException;
import org.apache.pdfbox.pdmodel.PDDocument;
import org.apache.pdfbox.pdmodel.PDDocumentCatalog;
import org.apache.pdfbox.pdmodel.PDDocumentInformation;
import org.apache.pdfbox.pdmodel.xmp.PDXMPMetadata;
import org.apache.xmpbox.XMPMetadata;
import org.apache.xmpbox.schema.DublinCoreSchema;
import org.apache.xmpbox.schema.PDFSchema;
import org.apache.xmpbox.xml.XmpSerializer;

public class UpdatePdfMetadata {
    public static void main(String args[]) throws IOException {
        String filepath = "C:\\Temp\\temp.pdf";
        
        // 加载现有文档
        File file = new File(filepath);
        try (PDDocument document = PDDocument.load(file)) {
            // 更新Info字典元数据
            PDDocumentInformation info = document.getDocumentInformation();
            System.out.println("Existing author: " + info.getAuthor());
            System.out.println("Existing title: " + info.getTitle());
            
            info.setAuthor("My new author");
            info.setTitle("My new title");
            
            // 更新XMP元数据
            PDDocumentCatalog catalog = document.getDocumentCatalog();
            PDXMPMetadata xmp = catalog.getMetadata();
            if (xmp == null) {
                // 如果文档没有XMP元数据,创建新的
                xmp = new PDXMPMetadata(document);
                catalog.setMetadata(xmp);
            }
            
            XMPMetadata metadata = xmp.getXMPMetadata();
            // 更新Dublin Core schema中的作者和标题
            DublinCoreSchema dcSchema = metadata.getDublinCoreSchema();
            if (dcSchema == null) {
                dcSchema = metadata.createAndAddDublinCoreSchema();
            }
            dcSchema.setCreator("My new author");
            dcSchema.setTitle("My new title");
            
            // 更新PDF schema中的元数据(确保一致性)
            PDFSchema pdfSchema = metadata.getPDFSchema();
            if (pdfSchema == null) {
                pdfSchema = metadata.createAndAddPDFSchema();
            }
            pdfSchema.setAuthor("My new author");
            pdfSchema.setTitle("My new title");
            
            // 序列化XMP元数据并保存回文档
            xmp.embedXMPMetadata(new XmpSerializer().serializeToString(metadata));
            
            // 保存文档
            document.save(filepath);
        }
    }
}

说明

  1. PDFBox v2.0.26已包含XMPBox相关依赖,无需额外引入,但需确保项目依赖完整
  2. 代码同时更新了Info字典和XMP元数据的DublinCore、PDF两个schema,确保Acrobat读取时能获取到新值
  3. 使用try-with-resources语法自动关闭文档,避免资源泄漏

内容的提问来源于stack exchange,提问作者Crafty Cat

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 16:09:25