Spring Boot中使用Apache PDFBox修改PDF指定文本无效果求助
问题排查:Apache PDFBox 修改PDF文本无效的原因及修复方案
你的代码能正常运行但无法修改PDF文本,核心问题是你只修改了从PDF中提取出的字符串,并没有把修改后的内容写回原PDF文档。PDFTextStripper只是用来读取PDF文本的工具,和PDF的实际内容存储完全无关,修改它提取出的字符串不会对PDF文件有任何影响。
具体错误点:
- 你通过
stripper.getText(document)提取文本后,修改了texte变量,但这个变量只是内存中的字符串,和加载的PDDocument对象没有关联。 - 后续调用
stripper.getText(document)只是再次读取文本,没有任何修改PDF的作用,直接保存原文档自然看不到变化。
修复方案:
要修改PDF中的文本,需要遍历PDF页面的内容流,找到文本绘制操作(如Tj、TJ操作符),替换其中的目标文本。以下是修正后的PdfProcessingService代码:
import org.apache.pdfbox.cos.*; import org.apache.pdfbox.io.IOUtils; import org.apache.pdfbox.pdfwriter.ContentStreamWriter; import org.apache.pdfbox.pdmodel.*; import org.apache.pdfbox.pdmodel.font.PDFont; import org.apache.pdfbox.util.PDFStreamParser; import java.io.*; import java.nio.charset.StandardCharsets; import java.util.List; import java.util.Map; public class PdfProcessingService { public boolean rechercherEtRemplacerPDF(String filePath, String recherche, String remplacement) { try (PDDocument document = PDDocument.load(new File(filePath))) { // 遍历所有页面处理文本 for (PDPage page : document.getPages()) { processPageContent(page, recherche, remplacement); } document.save("/home/mohamed/Downloads/nouveau_fichier.pdf"); System.out.println("文本替换完成"); return true; } catch (IOException e) { e.printStackTrace(); System.out.println("处理出错"); return false; } } private void processPageContent(PDPage page, String recherche, String remplacement) throws IOException { PDStream contents = page.getContents(); if (contents == null) return; PDFStreamParser parser = new PDFStreamParser(contents); parser.parse(); List<Object> tokens = parser.getTokens(); for (int i = 0; i < tokens.size(); i++) { Object token = tokens.get(i); // 处理单个文本操作符Tj if (token instanceof Operator && ((Operator) token).getName().equals("Tj")) { COSString textString = (COSString) tokens.get(i - 1); String originalText = textString.getString(); if (originalText.contains(recherche)) { String newText = originalText.replace(recherche, remplacement); textString.setValue(newText.getBytes(StandardCharsets.UTF_8)); } } // 处理文本数组操作符TJ else if (token instanceof Operator && ((Operator) token).getName().equals("TJ")) { COSArray textArray = (COSArray) tokens.get(i - 1); for (int j = 0; j < textArray.size(); j++) { Object arrayItem = textArray.getObject(j); if (arrayItem instanceof COSString) { COSString textString = (COSString) arrayItem; String originalText = textString.getString(); if (originalText.contains(recherche)) { String newText = originalText.replace(recherche, remplacement); textString.setValue(newText.getBytes(StandardCharsets.UTF_8)); } } } } } // 将修改后的内容写回页面 PDStream newStream = new PDStream(page.getDocument()); try (OutputStream out = newStream.createOutputStream(COSName.FLATE_DECODE)) { ContentStreamWriter writer = new ContentStreamWriter(out); writer.writeTokens(tokens); } page.setContents(newStream); } }
额外说明:
- 上述代码处理了PDF中两种最常见的文本绘制操作符
Tj(单个文本块)和TJ(文本数组),覆盖了大多数普通PDF的文本替换场景。 - 如果你的PDF存在文本被拆分、使用加密字体或复杂布局的情况,可能需要额外适配,但当前代码足以解决你的基础替换需求。
- 使用
try-with-resources语法自动管理PDDocument资源,避免内存泄漏。
内容的提问来源于stack exchange,提问作者mido med
相关产品推荐
相关产品推荐

