Apache PDFBox提取无表单带注释PDF的XFDF失败求助
问题原因
你的代码无效主要有两个核心问题:一是手动构建的FDFDocument缺少关联源PDF的必要信息,二是extractAnnotations()生成的FDFAnnotation可能未绑定正确的页面索引,导致序列化时被忽略。
修改方案1:直接使用PDDocument自带的exportXFDF方法(推荐)
PDFBox的PDDocument内置的exportXFDF方法无需依赖表单(AcroForm为null也不影响),可直接导出所有注释:
private String extractXFDF(Path path) throws ProcessingException { try (PDDocument pdfDoc = Loader.loadPDF(new RandomAccessReadBufferedFile(path.toFile().getAbsolutePath()))) { StringWriter writer = new StringWriter(); // 直接调用方法完成注释导出 pdfDoc.exportXFDF(writer); return writer.toString(); } catch (IOException e) { throw new CustomException(e.getMessage()); } }
修改方案2:修复手动构建FDFDocument的代码
如果需要保留手动构建逻辑,需补充两个关键步骤:
步骤1:修复extractAnnotations()方法,绑定页面索引
XFDF的页面编号从1开始,必须为每个FDFAnnotation设置对应页面:
private List<FDFAnnotation> extractAnnotations(PDDocument pdfDoc) { List<FDFAnnotation> fdfAnnotations = new ArrayList<>(); for (int pageIndex = 0; pageIndex < pdfDoc.getNumberOfPages(); pageIndex++) { PDPage page = pdfDoc.getPage(pageIndex); for (PDAnnotation pdfAnnotation : page.getAnnotations()) { FDFAnnotation fdfAnnotation = FDFAnnotation.createFromPDFAnnotation(pdfAnnotation); // 绑定页面索引(XFDF规范从1开始计数) fdfAnnotation.setPage(pageIndex + 1); fdfAnnotations.add(fdfAnnotation); } } return fdfAnnotations; }
步骤2:给FDF对象设置关联的PDF文件名
XFDF格式要求必须关联源PDF文件,否则注释不会被序列化:
private String extractXFDF(Path path) throws ProcessingException { try (PDDocument pdfDoc = Loader.loadPDF(new RandomAccessReadBufferedFile(path.toFile().getAbsolutePath())); FDFDocument fdfDoc = new FDFDocument(); ) { List<FDFAnnotation> fdfAnnotations = extractAnnotations(pdfDoc); FDF fdf = fdfDoc.getCatalog().getFDF(); // 设置关联的PDF文件名(XFDF规范必要项) fdf.setFile(path.getFileName().toString()); fdf.setAnnotations(fdfAnnotations); var writer = new StringWriter(); fdfDoc.saveXFDF(writer); return writer.toString(); } catch (IOException e) { throw new CustomException(e.getMessage()); } }
内容的提问来源于stack exchange,提问作者dinoooze
相关产品推荐
相关产品推荐

