You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于标记顺序调整PDF内容流文本顺序的Itext/PDFBox实现问询

问题背景与需求
  • PDF内容流的文本绘制顺序随机,文本可放置在页面任意位置,导致内容顺序不符合人类阅读逻辑
  • 部分屏幕阅读器(如Orbit Note、Read & Write)读取PDF时,会按照内容面板顺序而非标记定义的逻辑顺序读取
  • Adobe标记内容时,内容面板顺序会依据标记顺序(如MCID顺序)排序,需实现相同效果:
    • 明确知晓文本应遵循的阅读逻辑顺序
    • 已获取每个Tj操作的BBOX(矩形区域)及图形状态信息
  • 疑问:是否可通过iText或PDFBox调整PDF内容流顺序并保存?若无法直接实现,重写这些库的部分文件能否达成目标?

附尝试但未生效的PDFBox代码:

@Component
public class Contentorder {
@Autowired
public static List<Object> sortContentByMCID(List<Object> tokens)         
{
    int min = 2147483647;// max integer value
    int max = 0;
    Map<Integer, ArrayList> indexes = new HashMap<>();
//        ArrayList<ArrayList> indexes = new ArrayList<ArrayList>();
    ArrayList indices = new ArrayList();
    for (int ind=0; ind < tokens.size(); ind++) {
        if (tokens.get(ind) instanceof Operator) {
            Operator op = (Operator) tokens.get(ind);
            if(op.getName().equals("EMC") && indices.size() == 2){
                indices.add(ind+1);
                int key = (int) indices.get(0);
                indexes.put(key, indices);
            }
        }else if (tokens.get(ind) instanceof COSDictionary) {
            if( ((COSDictionary)tokens.get(ind)).containsKey("MCID")
                    &&  tokens.get(ind+1) instanceof Operator
                    && ((Operator) tokens.get(ind+1)).getName().equals("BDC")
                    && !((COSName) tokens.get(ind-1)).getName().equals("Figure") ) {
                if(min > ((COSInteger)((COSDictionary)tokens.get(ind)).getItem("MCID")).intValue())
                    min = ((COSInteger)((COSDictionary)tokens.get(ind)).getItem("MCID")).intValue();
                if(max < ((COSInteger)((COSDictionary)tokens.get(ind)).getItem("MCID")).intValue())
                    max = ((COSInteger)((COSDictionary)tokens.get(ind)).getItem("MCID")).intValue();
                indices = new ArrayList();
                indices.add(((COSInteger)((COSDictionary)tokens.get(ind)).getItem("MCID")).intValue());
                indices.add(getLastTf(tokens, ind));
            }
        }
    }
    System.out.println("print mcid min, max");
    System.out.println(indexes);
    for (Integer key : indexes.keySet()) {
        ArrayList lst = indexes.get(key);
        for (int ind=0; ind < tokens.size(); ind++) {
            if (tokens.get(ind) instanceof COSDictionary) {
                if( ((COSDictionary)tokens.get(ind)).containsKey("MCID")
                        &&  tokens.get(ind+1) instanceof Operator
                        && ((Operator) tokens.get(ind+1)).getName().equals("BDC")
                        && !((COSName) tokens.get(ind-1)).getName().equals("Figure") ) {
                    //some code
                    if(((COSInteger)((COSDictionary)tokens.get(ind)).getItem("MCID")).intValue() < key){
                        continue;
                    }else if(((COSInteger)((COSDictionary)tokens.get(ind)).getItem("MCID")).intValue() == key){
                        break;
                    }else{
                        List<Object> toks = tokens.subList((Integer) lst.get(1), (Integer) lst.get(2));
                        Integer size = toks.size();
                        Integer index = getLastTf(tokens, ind);
                        if(index < (Integer) lst.get(1)) {
                            tokens.addAll(index, toks);
                            tokens.subList((Integer) lst.get(1)+size, (Integer) lst.get(2)+size).clear();
                        }else{
                            tokens.addAll(index, toks);
                            tokens.subList((Integer) lst.get(1), (Integer) lst.get(2)).clear();
                        }
                        break;
                    }
                }
            }
        }
    }

    return tokens;
}

private static int getLastTf(List<Object> tokens, int ind) {
    for(int i = ind; i > 0; i--){
        if(tokens.get(i) instanceof Operator){
            if( ((Operator) tokens.get(i)).getName().equals("Tf")){
                return i-2;
            }
        }
    }
    return ind;
}
}
可行性分析与解决方案

1. iText/PDFBox 能否实现内容流重排?

完全可以实现,无需重写库核心文件,仅需基于现有API做上层逻辑开发:

  • PDFBox:通过ContentStreamParser解析内容流为token列表,修改顺序后用ContentStreamWriter重新生成内容流
  • iText 7:通过PdfCanvasProcessor解析内容流,使用PdfCanvas按指定顺序重新绘制内容操作

2. 现有代码未生效的问题分析

你的PDFBox代码存在几个核心问题:

  • token操作逻辑混乱:移动子列表时边界计算错误,修改原列表未考虑索引偏移,易导致内容丢失或重复
  • MCID排序逻辑缺失:仅遍历indexes的keySet,未按MCID从小到大顺序处理,无法保证最终顺序符合要求
  • getLastTf逻辑错误:Tf操作包含字体、字号两个参数,i-2不一定是正确起始位置,且未处理嵌套标记内容
  • 非标记内容未处理:内容流中的非MCID标记文本/图形未做合理安排,会破坏页面布局

3. 可行的实现思路(基于PDFBox)

  1. 解析内容流:用ContentStreamParser将页面内容流解析为List<Object>格式的token列表
  2. 提取标记内容块:遍历token列表,识别BDC/EMC包裹的标记内容块,记录每个块的MCID、起始/结束索引、BBOX及图形状态
  3. 排序内容块:根据已知阅读顺序(或MCID顺序)对标记内容块排序
  4. 重组内容流:
    • 保留非标记内容的相对顺序,按排序后的顺序插入标记内容块
    • 维护图形状态一致性(如字体、颜色、变换矩阵),避免重排后页面样式错乱
  5. 重写内容流:将重组后的token列表通过ContentStreamWriter写入PDF页面,替换原内容流

4. 简化示例代码片段

public void reorderContentByMCID(PDPage page, Map<Integer, ContentBlock> sortedBlocks) throws IOException {
    // 解析原内容流
    ContentStreamParser parser = new ContentStreamParser(page);
    List<Object> tokens = parser.parse();
    
    List<Object> newTokens = new ArrayList<>();
    Set<Integer> processedMCIDs = new HashSet<>();
    
    // 遍历原tokens,插入排序后的内容块
    int i = 0;
    while (i < tokens.size()) {
        Object token = tokens.get(i);
        if (token instanceof COSDictionary && ((COSDictionary) token).containsKey("MCID")) {
            COSDictionary mcidDict = (COSDictionary) token;
            int mcid = ((COSInteger) mcidDict.getItem("MCID")).intValue();
            if (sortedBlocks.containsKey(mcid) && !processedMCIDs.contains(mcid)) {
                // 添加排序后的内容块
                newTokens.addAll(sortedBlocks.get(mcid).getTokens());
                processedMCIDs.add(mcid);
                // 跳过原内容块的token
                i = findEMCIndex(tokens, i) + 1;
            } else {
                newTokens.add(token);
                i++;
            }
        } else {
            newTokens.add(token);
            i++;
        }
    }
    
    // 重写页面内容流
    PDPageContentStream contentStream = new PDPageContentStream(page.getDocument(), page, PDPageContentStream.AppendMode.OVERWRITE, true, true);
    ContentStreamWriter writer = new ContentStreamWriter(contentStream);
    writer.writeTokens(newTokens);
    contentStream.close();
}

// 辅助方法:找到EMC操作的索引
private int findEMCIndex(List<Object> tokens, int startIndex) {
    for (int i = startIndex; i < tokens.size(); i++) {
        if (tokens.get(i) instanceof Operator && ((Operator) tokens.get(i)).getName().equals("EMC")) {
            return i;
        }
    }
    return tokens.size();
}

// 自定义ContentBlock类,存储内容块的token、MCID、BBOX等信息
static class ContentBlock {
    private int mcid;
    private List<Object> tokens;
    private PDRectangle bbox;
    // 构造方法、getter/setter
}

内容的提问来源于stack exchange,提问作者fascinating coder

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 14:45:00