如何用Java提取PDF书籍各章节的起始页码?
基于PDFBox处理PDF章节索引页的实用方案
针对你遇到的三个索引页处理问题,直接给可落地的解决方法:
1. 识别PDF的索引页
索引页通常有明显特征,通过「关键词检测+结构占比判断」组合识别:
- 优先扫描PDF的前5页和后10页(多数文档的索引要么靠前要么在末尾);
- 检查页面文本是否包含“Index”“Contents”“目录”这类索引关键词(注意大小写和多语言场景);
- 统计页面中符合「文本+数字/罗马数字」格式的行占比,占比超过60%的页面大概率是索引页。
示例代码:
PDFTextStripper stripper = new PDFTextStripper(); PDDocument document = PDDocument.load(new File("target.pdf")); for (int pageNum = 1; pageNum <= document.getNumberOfPages(); pageNum++) { stripper.setStartPage(pageNum); stripper.setEndPage(pageNum); String pageText = stripper.getText(document); // 检查索引关键词 boolean hasIndexKeyword = pageText.contains("Index") || pageText.contains("Contents") || pageText.contains("目录"); if (!hasIndexKeyword) continue; // 统计符合索引结构的行占比 String[] lines = pageText.split("\n"); int indexLineCount = 0; Pattern pagePattern = Pattern.compile(".*\\s+([0-9]+|I{1,3}|IV|V|VI{0,3})$"); // 匹配数字/罗马数字结尾的行 for (String line : lines) { if (pagePattern.matcher(line.trim()).find()) { indexLineCount++; } } float ratio = (float) indexLineCount / lines.length; if (ratio > 0.6) { System.out.println("识别到索引页:第" + pageNum + "页"); // 后续处理该页面 } } document.close();
2. 过滤索引页中的非章节条目
通过「章节特征匹配+排除列表」筛选有效章节:
- 自定义章节标题的正则规则(根据你遇到的格式扩展);
- 维护排除列表,过滤“参考文献”“附录”这类非章节条目。
示例代码:
// 自定义章节标题匹配规则 List<Pattern> chapterPatterns = Arrays.asList( Pattern.compile("(Chapter|Section|Part)\\s+[0-9IVXLCDM]+\\s+.+"), // 英文章节格式 Pattern.compile("第[0-9一二三四五六七八九十]+章\\s+.+"), // 中文章节格式 Pattern.compile("0x[0-9A-Fa-f]+\\s+.+"), // 十六进制前缀格式 Pattern.compile("^[0-9]+\\.\\s+.+") // 层级标题格式(如1. Introduction) ); // 非章节条目排除列表 List<String> excludeList = Arrays.asList("参考文献", "附录", "致谢", "References", "Appendix", "Acknowledgements"); String indexPageText = stripper.getText(document); // 已提取的索引页文本 for (String line : indexPageText.split("\n")) { String trimmedLine = line.trim(); if (trimmedLine.isEmpty()) continue; // 跳过排除条目 boolean isExcluded = excludeList.stream().anyMatch(trimmedLine::contains); if (isExcluded) continue; // 匹配章节格式 boolean isChapter = chapterPatterns.stream().anyMatch(p -> p.matcher(trimmedLine).find()); if (isChapter) { // 提取章节标题和页码 String[] parts = trimmedLine.split("\\s+(?=[0-9IVXLCDM]+$)"); // 从末尾的页码处分割 String chapterTitle = parts[0].trim(); String pageStr = parts[1].trim(); // 转换页码为数字(罗马数字转int的方法自行补充) int pageNumber = convertRomanToInt(pageStr); if (pageNumber == -1) { try { pageNumber = Integer.parseInt(pageStr); } catch (NumberFormatException e) { continue; // 无效页码跳过 } } // 执行入库操作 // saveChapter(chapterTitle, pageNumber); } }
3. 处理PDF转文本时的开头无关内容
通过「文本清理+起始行定位」跳过无关内容:
- 先清理文本中的非打印字符,过滤空白行;
- 找到第一个符合索引结构的行,从该行开始处理。
示例代码:
String rawPageText = stripper.getText(document); // 清理非打印字符(保留换行、制表符) String cleanedText = rawPageText.replaceAll("[\\p{Cntrl}&&[^\r\n\t]]", ""); // 过滤空白行,转为有序列表 List<String> validLines = Arrays.stream(cleanedText.split("\n")) .map(String::trim) .filter(line -> !line.isEmpty()) .collect(Collectors.toList()); // 定位第一个索引行的起始位置 int startLine = 0; Pattern indexLinePattern = Pattern.compile(".*\\s+([0-9]+|I{1,3}|IV|V|VI{0,3})$"); while (startLine < validLines.size()) { if (indexLinePattern.matcher(validLines.get(startLine)).find()) { break; } startLine++; } // 从startLine开始处理索引内容 for (int i = startLine; i < validLines.size(); i++) { String line = validLines.get(i); // 执行章节提取逻辑 }
内容的提问来源于stack exchange,提问作者Aakash
相关产品推荐
相关产品推荐

