如何删除或替换PDF中由多条连续绘制指令生成的指定文本
多指令拼接文本的替换解决方案
你现有代码的核心问题是每次处理完文本绘制指令后就立刻清空字符缓存,仅匹配单条指令内绘制的文本,无法覆盖目标文本被拆分为多条Tj/TJ等指令输出的场景。
核心实现思路
- 调整缓存逻辑:不再单指令清空缓存,持续累计绘制的字符,同时为每个字符关联对应的绘制上下文(所属指令、坐标位置、字体等信息)
- 增加连续文本判定:遇到非文本绘制指令(如位置移动指令
Td/Tm、文本块结束指令ET)时,判断文本是否出现断裂,若前后字符位置不连续则清空缓存 - 跨指令匹配:当累计的字符序列命中目标替换文本时,跳过对应字符关联的所有原始绘制指令,再插入替换后的新文本绘制指令
修改后的核心代码示例
// 用于存储字符关联的元信息 static class CharMeta { String glyphStr; float x; float y; Operator relatedOperator; List<COSBase> relatedOperands; } for (PDPage page : document.getDocumentCatalog().getPages()) { PdfContentStreamEditor editor = new PdfContentStreamEditor(document, page) { final List<CharMeta> charMetaCache = new ArrayList<>(); float lastCharX = -1; final String TARGET_TEXT = "Text which is to be replace"; @Override protected void showGlyph(Matrix textRenderingMatrix, PDFont font, int code, Vector displacement) throws IOException { String string = font.toUnicode(code); if (string != null) { CharMeta meta = new CharMeta(); meta.glyphStr = string; // 取当前字符的绘制坐标 meta.x = textRenderingMatrix.getTranslateX(); meta.y = textRenderingMatrix.getTranslateY(); charMetaCache.add(meta); } super.showGlyph(textRenderingMatrix, font, code, displacement); } @Override protected void write(ContentStreamWriter contentStreamWriter, Operator operator, List<COSBase> operands) throws IOException { String operatorString = operator.getName(); // 非文本绘制指令:先校验缓存中是否有匹配文本,再执行输出 if (!TEXT_SHOWING_OPERATORS.contains(operatorString)) { // 先检查缓存中是否命中目标文本 String cachedStr = charMetaCache.stream().map(m -> m.glyphStr).collect(Collectors.joining()); if (cachedStr.contains(TARGET_TEXT)) { // 命中目标则跳过所有对应文本的绘制指令,此处可插入你替换的新文本绘制逻辑 charMetaCache.clear(); lastCharX = -1; return; } // 遇到文本位置移动/文本块结束指令,清空缓存 if (Arrays.asList("ET", "Td", "Tm", "TD").contains(operatorString)) { charMetaCache.clear(); lastCharX = -1; } super.write(contentStreamWriter, operator, operands); return; } // 文本绘制指令:关联到缓存中对应字符 float currentX = charMetaCache.isEmpty() ? -1 : charMetaCache.get(charMetaCache.size()-1).x; // 校验字符是否连续,间距过大则判定为不同文本片段,清空缓存 if (lastCharX > 0 && currentX - lastCharX > 2 * getGraphicsState().getTextState().getFontSize()) { charMetaCache.clear(); } lastCharX = currentX; // 检查当前累计的字符串是否已命中目标 String cachedStr = charMetaCache.stream().map(m -> m.glyphStr).collect(Collectors.joining()); if (cachedStr.equals(TARGET_TEXT)) { // 命中则跳过当前指令,清空缓存 charMetaCache.clear(); lastCharX = -1; return; } super.write(contentStreamWriter, operator, operands); } final List<String> TEXT_SHOWING_OPERATORS = Arrays.asList("Tj", "'", "\"", "TJ"); }; editor.processPage(page); } document.save("watermark-RemoveByText.pdf");
注意事项
- 可根据实际PDF的排版调整连续字符的间距判定阈值,避免把不连续的文本误判为目标文本
- 替换新文本时需复用原始文本的字体、字号、填充色等属性,避免出现排版错位
- 若处理的PDF使用自定义字体编码,需提前确认字体的
toUnicode映射表完整,否则会出现字符识别错误
内容的提问来源于stack exchange,提问作者Akhil NagaSai
相关产品推荐
相关产品推荐

