如何使用Java的iText 5为合并后的PDF生成带页码的动态目录
使用iText 5为合并后的PDF生成动态目录方案
核心思路
要实现需求,关键分为两步:
- 捕获标题与对应页码:要么在合并子PDF时直接记录(效率更高),要么合并后解析PDF提取标题及页码
- 生成可跳转目录:将目录插入到合并后PDF的开头,每个条目关联对应页码的跳转动作
具体实现方案
方案一:合并阶段直接记录标题(推荐)
如果Crystal Reports生成的每个子PDF第一页就是该部分标题,可在合并时直接提取标题并记录其在最终PDF中的起始页码,避免后续解析开销。
修改后的完整代码
import com.itextpdf.text.*; import com.itextpdf.text.pdf.*; import com.itextpdf.text.pdf.parser.LocationTextExtractionStrategy; import com.itextpdf.text.pdf.parser.PdfTextExtractor; import com.itextpdf.text.pdf.parser.TextExtractionStrategy; import java.io.File; import java.io.FileOutputStream; import java.io.IOException; import java.util.ArrayList; import java.util.List; public class PDFMerge { private static final String fileName = "_My Report.pdf"; private static final String tempDest = "temp_MergePDFs.pdf"; private static final String finalDest = "3_MergePDFs.pdf"; // 存储目录条目:[标题文本, 对应页码] private List<String[]> tocEntries = new ArrayList<>(); public void generateMergedPDF(int fileCount) throws IOException, DocumentException { Document document = new Document(); PdfCopy copy = new PdfCopy(document, new FileOutputStream(tempDest)); document.open(); int currentFinalPage = 1; // 记录合并后PDF的当前页码 for (int inc = 1; inc <= fileCount; inc++) { PdfReader reader = new PdfReader(inc + fileName); int subPageCount = reader.getNumberOfPages(); // 提取子PDF第一页的标题文本(可根据实际调整标题提取规则) TextExtractionStrategy strategy = new LocationTextExtractionStrategy(); String rawTitle = PdfTextExtractor.getTextFromPage(reader, 1, strategy).trim(); // 去除换行,取首行作为标题(示例处理,需匹配你的标题格式) String cleanTitle = rawTitle.split("\\r?\\n")[0]; tocEntries.add(new String[]{cleanTitle, String.valueOf(currentFinalPage)}); // 将子PDF所有页面加入临时文件 for (int i = 1; i <= subPageCount; i++) { PdfImportedPage page = copy.getImportedPage(reader, i); copy.addPage(page); currentFinalPage++; } reader.close(); } document.close(); // 生成带目录的最终PDF addTocToFinalPdf(); } // 生成目录并插入到合并后PDF开头 private void addTocToFinalPdf() throws IOException, DocumentException { PdfReader tempReader = new PdfReader(tempDest); Document finalDoc = new Document(); // 使用PdfSmartCopy避免冗余资源,优化文件大小 PdfSmartCopy copyWith = new PdfSmartCopy(finalDoc, new FileOutputStream(finalDest)); finalDoc.open(); // 生成目录页内容 Paragraph tocHeader = new Paragraph("目录", FontFactory.getFont(FontFactory.HELVETICA_BOLD, 16)); tocHeader.setAlignment(Element.ALIGN_CENTER); finalDoc.add(tocHeader); finalDoc.add(new Paragraph("\n")); // 逐个添加目录条目,实现点击跳转 for (String[] entry : tocEntries) { String title = entry[0]; int targetPage = Integer.parseInt(entry[1]) + 1; // 因新增目录页,原页码+1 Paragraph tocItem = new Paragraph(); // 标题部分添加跳转动作 Chunk titleChunk = new Chunk(title); titleChunk.setAction(PdfAction.gotoLocalPage(targetPage, new PdfDestination(PdfDestination.FIT), copyWith)); tocItem.add(titleChunk); // 添加标题与页码之间的虚线分隔 float pageWidth = finalDoc.getPageSize().getWidth() - finalDoc.leftMargin() - finalDoc.rightMargin(); float titleWidth = titleChunk.getWidthPoint(); float dotsWidth = pageWidth - titleWidth - 25; // 预留页码宽度 Chunk dottedLine = new Chunk(new DottedLineSeparator()); dottedLine.setWidth(dotsWidth); tocItem.add(dottedLine); // 添加页码 tocItem.add(new Chunk(String.valueOf(targetPage))); finalDoc.add(tocItem); finalDoc.add(new Paragraph("\n")); } // 添加原合并PDF的所有页面 int totalTempPages = tempReader.getNumberOfPages(); for (int i = 1; i <= totalTempPages; i++) { PdfImportedPage page = copyWith.getImportedPage(tempReader, i); copyWith.addPage(page); } finalDoc.close(); tempReader.close(); // 删除临时文件 new File(tempDest).delete(); } }
方案二:合并后解析识别标题
如果标题不固定在子PDF第一页,需通过字体特征(如加粗、字号)识别,可自定义文本提取策略:
自定义标题识别策略
class TitleExtractionStrategy extends LocationTextExtractionStrategy { private List<String[]> titleList = new ArrayList<>(); // 定义标题的字体阈值 private final float MIN_TITLE_SIZE = 14f; @Override public void renderText(TextRenderInfo renderInfo) { super.renderText(renderInfo); // 判断当前文本是否为标题:加粗+字号达标 if (renderInfo.getFont().isBold() && renderInfo.getFontSize() >= MIN_TITLE_SIZE) { String title = renderInfo.getText().trim(); int pageNum = renderInfo.getPageNumber(); titleList.add(new String[]{title, String.valueOf(pageNum)}); } } public List<String[]> getTitleList() { return titleList; } }
使用方式
在合并完成后,遍历临时PDF提取标题:
PdfReader tempReader = new PdfReader(tempDest); TitleExtractionStrategy strategy = new TitleExtractionStrategy(); for (int i = 1; i <= tempReader.getNumberOfPages(); i++) { PdfTextExtractor.getTextFromPage(tempReader, i, strategy); } tocEntries = strategy.getTitleList();
关键注意事项
- 标题提取规则:需根据实际PDF的标题格式调整提取逻辑,比如多行标题、特定字体名称等
- 页码偏移:插入目录页后,原合并PDF的所有页码需+1,确保跳转准确
- 性能优化:优先选择合并时记录标题的方案,避免合并后全量解析PDF的额外开销
内容的提问来源于stack exchange,提问作者Usama
相关产品推荐
相关产品推荐

