You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在iText中识别“回复”注释并关联被回复的原注释

解决iText中PDF回复注释与目标注释的关联问题

我明白你现在的困扰——你能通过PdfName.IRT拿到一段字符串,但不知道怎么用它找到被回复的目标注释对吧?其实问题出在你对IRT字段的理解上:在PDF规范里,IRT(In Reply To)存储的是被回复注释的间接引用(PdfIndirectReference),而不是普通字符串,所以直接用getAsString()是拿不到有效关联的,得用正确的方法获取引用并匹配目标注释。

给你一个修改后的实现思路和代码,核心是先把所有注释的引用和字典做一个映射,这样就能快速通过IRT找到对应的目标注释:

import com.itextpdf.kernel.pdf.*;
import java.io.IOException;
import java.util.HashMap;
import java.util.Map;

public class PdfCommentReplyHandler {
    public static void main(String[] args) throws IOException {
        PdfReader reader = new PdfReader("C:\\PDF\\test.pdf");
        
        // 第一步:先收集所有注释的间接引用与对应字典的映射
        Map<PdfIndirectReference, PdfDictionary> allAnnotations = new HashMap<>();
        
        // 遍历所有页面收集注释
        for (int i = 1; i <= reader.getNumberOfPages(); i++) {
            PdfArray annotsArray = reader.getPageN(i).getAsArray(PdfName.ANNOTS);
            if (annotsArray == null) continue;
            
            for (int j = 0; j < annotsArray.size(); j++) {
                PdfObject annotObj = annotsArray.get(j);
                // 确认是间接引用的注释字典
                if (annotObj instanceof PdfIndirectReference) {
                    PdfDictionary annotDict = reader.getPdfObject((PdfIndirectReference) annotObj).getAsDict();
                    if (annotDict != null) {
                        allAnnotations.put((PdfIndirectReference) annotObj, annotDict);
                    }
                }
            }
        }
        
        // 第二步:遍历注释并关联回复目标
        for (int i = 1; i <= reader.getNumberOfPages(); i++) {
            PdfArray annotsArray = reader.getPageN(i).getAsArray(PdfName.ANNOTS);
            if (annotsArray == null) continue;
            
            for (int j = 0; j < annotsArray.size(); j++) {
                PdfObject annotObj = annotsArray.get(j);
                if (!(annotObj instanceof PdfIndirectReference)) continue;
                
                PdfDictionary currentAnnot = allAnnotations.get(annotObj);
                if (currentAnnot == null || currentAnnot.getAsString(PdfName.CONTENTS) == null) continue;
                
                // 打印当前注释信息
                System.out.println("当前注释内容: " + currentAnnot.getAsString(PdfName.CONTENTS));
                System.out.println("当前注释作者: " + currentAnnot.getAsString(PdfName.T));
                System.out.println("当前注释创建时间: " + currentAnnot.get(PdfName.CREATIONDATE));
                
                // 处理IRT关联
                PdfObject irtObj = currentAnnot.get(PdfName.IRT);
                if (irtObj instanceof PdfIndirectReference) {
                    PdfDictionary targetAnnot = allAnnotations.get(irtObj);
                    if (targetAnnot != null) {
                        System.out.println("回复的目标注释内容: " + targetAnnot.getAsString(PdfName.CONTENTS));
                        System.out.println("回复的目标注释作者: " + targetAnnot.getAsString(PdfName.T));
                    } else {
                        System.out.println("未找到关联的目标注释");
                    }
                } else if (irtObj != null) {
                    // 极少数情况IRT是数组(回复多个注释),这里做简单处理
                    System.out.println("IRT是数组类型,需额外处理");
                } else {
                    System.out.println("当前注释不是回复类型");
                }
                
                System.out.println("--------");
            }
        }
        
        reader.close();
    }
}

关键说明:

  • 注释映射表:先遍历所有页面,把每个注释的PdfIndirectReference作为键,对应的PdfDictionary作为值存入Map,这样后续查找效率很高。
  • 正确获取IRT:用currentAnnot.get(PdfName.IRT)获取原始对象,判断是否为PdfIndirectReference,再用这个引用去映射表里找目标注释。
  • 特殊情况处理:虽然大部分场景下IRT是单个引用,但少数PDF可能用数组存储多个回复目标,你可以根据需求扩展数组的处理逻辑。

这样就能完美区分回复类型的注释,并关联到它的目标注释啦!

内容的提问来源于stack exchange,提问作者atul bharadwaj

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 09:59:38