You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Java的iText 5为合并后的PDF生成带页码的动态目录

使用iText 5为合并后的PDF生成动态目录方案

核心思路

要实现需求,关键分为两步:

  1. 捕获标题与对应页码:要么在合并子PDF时直接记录(效率更高),要么合并后解析PDF提取标题及页码
  2. 生成可跳转目录:将目录插入到合并后PDF的开头,每个条目关联对应页码的跳转动作

具体实现方案

方案一:合并阶段直接记录标题(推荐)

如果Crystal Reports生成的每个子PDF第一页就是该部分标题,可在合并时直接提取标题并记录其在最终PDF中的起始页码,避免后续解析开销。

修改后的完整代码

import com.itextpdf.text.*;
import com.itextpdf.text.pdf.*;
import com.itextpdf.text.pdf.parser.LocationTextExtractionStrategy;
import com.itextpdf.text.pdf.parser.PdfTextExtractor;
import com.itextpdf.text.pdf.parser.TextExtractionStrategy;

import java.io.File;
import java.io.FileOutputStream;
import java.io.IOException;
import java.util.ArrayList;
import java.util.List;

public class PDFMerge {
    private static final String fileName = "_My Report.pdf";
    private static final String tempDest = "temp_MergePDFs.pdf";
    private static final String finalDest = "3_MergePDFs.pdf";
    // 存储目录条目:[标题文本, 对应页码]
    private List<String[]> tocEntries = new ArrayList<>();

    public void generateMergedPDF(int fileCount) throws IOException, DocumentException {
        Document document = new Document();
        PdfCopy copy = new PdfCopy(document, new FileOutputStream(tempDest));
        document.open();

        int currentFinalPage = 1; // 记录合并后PDF的当前页码
        for (int inc = 1; inc <= fileCount; inc++) {
            PdfReader reader = new PdfReader(inc + fileName);
            int subPageCount = reader.getNumberOfPages();
            
            // 提取子PDF第一页的标题文本(可根据实际调整标题提取规则)
            TextExtractionStrategy strategy = new LocationTextExtractionStrategy();
            String rawTitle = PdfTextExtractor.getTextFromPage(reader, 1, strategy).trim();
            // 去除换行,取首行作为标题(示例处理,需匹配你的标题格式)
            String cleanTitle = rawTitle.split("\\r?\\n")[0];
            tocEntries.add(new String[]{cleanTitle, String.valueOf(currentFinalPage)});
            
            // 将子PDF所有页面加入临时文件
            for (int i = 1; i <= subPageCount; i++) {
                PdfImportedPage page = copy.getImportedPage(reader, i);
                copy.addPage(page);
                currentFinalPage++;
            }
            reader.close();
        }

        document.close();
        // 生成带目录的最终PDF
        addTocToFinalPdf();
    }

    // 生成目录并插入到合并后PDF开头
    private void addTocToFinalPdf() throws IOException, DocumentException {
        PdfReader tempReader = new PdfReader(tempDest);
        Document finalDoc = new Document();
        // 使用PdfSmartCopy避免冗余资源,优化文件大小
        PdfSmartCopy copyWith = new PdfSmartCopy(finalDoc, new FileOutputStream(finalDest));
        finalDoc.open();

        // 生成目录页内容
        Paragraph tocHeader = new Paragraph("目录", FontFactory.getFont(FontFactory.HELVETICA_BOLD, 16));
        tocHeader.setAlignment(Element.ALIGN_CENTER);
        finalDoc.add(tocHeader);
        finalDoc.add(new Paragraph("\n"));

        // 逐个添加目录条目,实现点击跳转
        for (String[] entry : tocEntries) {
            String title = entry[0];
            int targetPage = Integer.parseInt(entry[1]) + 1; // 因新增目录页,原页码+1
            
            Paragraph tocItem = new Paragraph();
            // 标题部分添加跳转动作
            Chunk titleChunk = new Chunk(title);
            titleChunk.setAction(PdfAction.gotoLocalPage(targetPage, new PdfDestination(PdfDestination.FIT), copyWith));
            tocItem.add(titleChunk);
            
            // 添加标题与页码之间的虚线分隔
            float pageWidth = finalDoc.getPageSize().getWidth() - finalDoc.leftMargin() - finalDoc.rightMargin();
            float titleWidth = titleChunk.getWidthPoint();
            float dotsWidth = pageWidth - titleWidth - 25; // 预留页码宽度
            Chunk dottedLine = new Chunk(new DottedLineSeparator());
            dottedLine.setWidth(dotsWidth);
            tocItem.add(dottedLine);
            
            // 添加页码
            tocItem.add(new Chunk(String.valueOf(targetPage)));
            
            finalDoc.add(tocItem);
            finalDoc.add(new Paragraph("\n"));
        }

        // 添加原合并PDF的所有页面
        int totalTempPages = tempReader.getNumberOfPages();
        for (int i = 1; i <= totalTempPages; i++) {
            PdfImportedPage page = copyWith.getImportedPage(tempReader, i);
            copyWith.addPage(page);
        }

        finalDoc.close();
        tempReader.close();
        // 删除临时文件
        new File(tempDest).delete();
    }
}

方案二:合并后解析识别标题

如果标题不固定在子PDF第一页,需通过字体特征(如加粗、字号)识别,可自定义文本提取策略:

自定义标题识别策略

class TitleExtractionStrategy extends LocationTextExtractionStrategy {
    private List<String[]> titleList = new ArrayList<>();
    // 定义标题的字体阈值
    private final float MIN_TITLE_SIZE = 14f;

    @Override
    public void renderText(TextRenderInfo renderInfo) {
        super.renderText(renderInfo);
        // 判断当前文本是否为标题:加粗+字号达标
        if (renderInfo.getFont().isBold() && renderInfo.getFontSize() >= MIN_TITLE_SIZE) {
            String title = renderInfo.getText().trim();
            int pageNum = renderInfo.getPageNumber();
            titleList.add(new String[]{title, String.valueOf(pageNum)});
        }
    }

    public List<String[]> getTitleList() {
        return titleList;
    }
}

使用方式

在合并完成后,遍历临时PDF提取标题:

PdfReader tempReader = new PdfReader(tempDest);
TitleExtractionStrategy strategy = new TitleExtractionStrategy();
for (int i = 1; i <= tempReader.getNumberOfPages(); i++) {
    PdfTextExtractor.getTextFromPage(tempReader, i, strategy);
}
tocEntries = strategy.getTitleList();

关键注意事项

  1. 标题提取规则:需根据实际PDF的标题格式调整提取逻辑,比如多行标题、特定字体名称等
  2. 页码偏移:插入目录页后,原合并PDF的所有页码需+1,确保跳转准确
  3. 性能优化:优先选择合并时记录标题的方案,避免合并后全量解析PDF的额外开销

内容的提问来源于stack exchange,提问作者Usama

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.19 19:13:20