基于标记顺序调整PDF内容流文本顺序的Itext/PDFBox实现问询
问题背景与需求
- PDF内容流的文本绘制顺序随机,文本可放置在页面任意位置,导致内容顺序不符合人类阅读逻辑
- 部分屏幕阅读器(如Orbit Note、Read & Write)读取PDF时,会按照内容面板顺序而非标记定义的逻辑顺序读取
- Adobe标记内容时,内容面板顺序会依据标记顺序(如MCID顺序)排序,需实现相同效果:
- 明确知晓文本应遵循的阅读逻辑顺序
- 已获取每个
Tj操作的BBOX(矩形区域)及图形状态信息
- 疑问:是否可通过iText或PDFBox调整PDF内容流顺序并保存?若无法直接实现,重写这些库的部分文件能否达成目标?
附尝试但未生效的PDFBox代码:
@Component public class Contentorder { @Autowired public static List<Object> sortContentByMCID(List<Object> tokens) { int min = 2147483647;// max integer value int max = 0; Map<Integer, ArrayList> indexes = new HashMap<>(); // ArrayList<ArrayList> indexes = new ArrayList<ArrayList>(); ArrayList indices = new ArrayList(); for (int ind=0; ind < tokens.size(); ind++) { if (tokens.get(ind) instanceof Operator) { Operator op = (Operator) tokens.get(ind); if(op.getName().equals("EMC") && indices.size() == 2){ indices.add(ind+1); int key = (int) indices.get(0); indexes.put(key, indices); } }else if (tokens.get(ind) instanceof COSDictionary) { if( ((COSDictionary)tokens.get(ind)).containsKey("MCID") && tokens.get(ind+1) instanceof Operator && ((Operator) tokens.get(ind+1)).getName().equals("BDC") && !((COSName) tokens.get(ind-1)).getName().equals("Figure") ) { if(min > ((COSInteger)((COSDictionary)tokens.get(ind)).getItem("MCID")).intValue()) min = ((COSInteger)((COSDictionary)tokens.get(ind)).getItem("MCID")).intValue(); if(max < ((COSInteger)((COSDictionary)tokens.get(ind)).getItem("MCID")).intValue()) max = ((COSInteger)((COSDictionary)tokens.get(ind)).getItem("MCID")).intValue(); indices = new ArrayList(); indices.add(((COSInteger)((COSDictionary)tokens.get(ind)).getItem("MCID")).intValue()); indices.add(getLastTf(tokens, ind)); } } } System.out.println("print mcid min, max"); System.out.println(indexes); for (Integer key : indexes.keySet()) { ArrayList lst = indexes.get(key); for (int ind=0; ind < tokens.size(); ind++) { if (tokens.get(ind) instanceof COSDictionary) { if( ((COSDictionary)tokens.get(ind)).containsKey("MCID") && tokens.get(ind+1) instanceof Operator && ((Operator) tokens.get(ind+1)).getName().equals("BDC") && !((COSName) tokens.get(ind-1)).getName().equals("Figure") ) { //some code if(((COSInteger)((COSDictionary)tokens.get(ind)).getItem("MCID")).intValue() < key){ continue; }else if(((COSInteger)((COSDictionary)tokens.get(ind)).getItem("MCID")).intValue() == key){ break; }else{ List<Object> toks = tokens.subList((Integer) lst.get(1), (Integer) lst.get(2)); Integer size = toks.size(); Integer index = getLastTf(tokens, ind); if(index < (Integer) lst.get(1)) { tokens.addAll(index, toks); tokens.subList((Integer) lst.get(1)+size, (Integer) lst.get(2)+size).clear(); }else{ tokens.addAll(index, toks); tokens.subList((Integer) lst.get(1), (Integer) lst.get(2)).clear(); } break; } } } } } return tokens; } private static int getLastTf(List<Object> tokens, int ind) { for(int i = ind; i > 0; i--){ if(tokens.get(i) instanceof Operator){ if( ((Operator) tokens.get(i)).getName().equals("Tf")){ return i-2; } } } return ind; } }
可行性分析与解决方案
1. iText/PDFBox 能否实现内容流重排?
完全可以实现,无需重写库核心文件,仅需基于现有API做上层逻辑开发:
- PDFBox:通过
ContentStreamParser解析内容流为token列表,修改顺序后用ContentStreamWriter重新生成内容流 - iText 7:通过
PdfCanvasProcessor解析内容流,使用PdfCanvas按指定顺序重新绘制内容操作
2. 现有代码未生效的问题分析
你的PDFBox代码存在几个核心问题:
- token操作逻辑混乱:移动子列表时边界计算错误,修改原列表未考虑索引偏移,易导致内容丢失或重复
- MCID排序逻辑缺失:仅遍历
indexes的keySet,未按MCID从小到大顺序处理,无法保证最终顺序符合要求 getLastTf逻辑错误:Tf操作包含字体、字号两个参数,i-2不一定是正确起始位置,且未处理嵌套标记内容- 非标记内容未处理:内容流中的非MCID标记文本/图形未做合理安排,会破坏页面布局
3. 可行的实现思路(基于PDFBox)
- 解析内容流:用
ContentStreamParser将页面内容流解析为List<Object>格式的token列表 - 提取标记内容块:遍历token列表,识别
BDC/EMC包裹的标记内容块,记录每个块的MCID、起始/结束索引、BBOX及图形状态 - 排序内容块:根据已知阅读顺序(或MCID顺序)对标记内容块排序
- 重组内容流:
- 保留非标记内容的相对顺序,按排序后的顺序插入标记内容块
- 维护图形状态一致性(如字体、颜色、变换矩阵),避免重排后页面样式错乱
- 重写内容流:将重组后的token列表通过
ContentStreamWriter写入PDF页面,替换原内容流
4. 简化示例代码片段
public void reorderContentByMCID(PDPage page, Map<Integer, ContentBlock> sortedBlocks) throws IOException { // 解析原内容流 ContentStreamParser parser = new ContentStreamParser(page); List<Object> tokens = parser.parse(); List<Object> newTokens = new ArrayList<>(); Set<Integer> processedMCIDs = new HashSet<>(); // 遍历原tokens,插入排序后的内容块 int i = 0; while (i < tokens.size()) { Object token = tokens.get(i); if (token instanceof COSDictionary && ((COSDictionary) token).containsKey("MCID")) { COSDictionary mcidDict = (COSDictionary) token; int mcid = ((COSInteger) mcidDict.getItem("MCID")).intValue(); if (sortedBlocks.containsKey(mcid) && !processedMCIDs.contains(mcid)) { // 添加排序后的内容块 newTokens.addAll(sortedBlocks.get(mcid).getTokens()); processedMCIDs.add(mcid); // 跳过原内容块的token i = findEMCIndex(tokens, i) + 1; } else { newTokens.add(token); i++; } } else { newTokens.add(token); i++; } } // 重写页面内容流 PDPageContentStream contentStream = new PDPageContentStream(page.getDocument(), page, PDPageContentStream.AppendMode.OVERWRITE, true, true); ContentStreamWriter writer = new ContentStreamWriter(contentStream); writer.writeTokens(newTokens); contentStream.close(); } // 辅助方法:找到EMC操作的索引 private int findEMCIndex(List<Object> tokens, int startIndex) { for (int i = startIndex; i < tokens.size(); i++) { if (tokens.get(i) instanceof Operator && ((Operator) tokens.get(i)).getName().equals("EMC")) { return i; } } return tokens.size(); } // 自定义ContentBlock类,存储内容块的token、MCID、BBOX等信息 static class ContentBlock { private int mcid; private List<Object> tokens; private PDRectangle bbox; // 构造方法、getter/setter }
内容的提问来源于stack exchange,提问作者fascinating coder
相关产品推荐
相关产品推荐

