You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Java提取PDF书籍各章节的起始页码?

基于PDFBox处理PDF章节索引页的实用方案

针对你遇到的三个索引页处理问题,直接给可落地的解决方法:

1. 识别PDF的索引页

索引页通常有明显特征,通过「关键词检测+结构占比判断」组合识别:

  • 优先扫描PDF的前5页和后10页(多数文档的索引要么靠前要么在末尾);
  • 检查页面文本是否包含“Index”“Contents”“目录”这类索引关键词(注意大小写和多语言场景);
  • 统计页面中符合「文本+数字/罗马数字」格式的行占比,占比超过60%的页面大概率是索引页。

示例代码:

PDFTextStripper stripper = new PDFTextStripper();
PDDocument document = PDDocument.load(new File("target.pdf"));

for (int pageNum = 1; pageNum <= document.getNumberOfPages(); pageNum++) {
    stripper.setStartPage(pageNum);
    stripper.setEndPage(pageNum);
    String pageText = stripper.getText(document);
    
    // 检查索引关键词
    boolean hasIndexKeyword = pageText.contains("Index") || pageText.contains("Contents") || pageText.contains("目录");
    if (!hasIndexKeyword) continue;
    
    // 统计符合索引结构的行占比
    String[] lines = pageText.split("\n");
    int indexLineCount = 0;
    Pattern pagePattern = Pattern.compile(".*\\s+([0-9]+|I{1,3}|IV|V|VI{0,3})$"); // 匹配数字/罗马数字结尾的行
    for (String line : lines) {
        if (pagePattern.matcher(line.trim()).find()) {
            indexLineCount++;
        }
    }
    float ratio = (float) indexLineCount / lines.length;
    if (ratio > 0.6) {
        System.out.println("识别到索引页:第" + pageNum + "页");
        // 后续处理该页面
    }
}
document.close();

2. 过滤索引页中的非章节条目

通过「章节特征匹配+排除列表」筛选有效章节:

  • 自定义章节标题的正则规则(根据你遇到的格式扩展);
  • 维护排除列表,过滤“参考文献”“附录”这类非章节条目。

示例代码:

// 自定义章节标题匹配规则
List<Pattern> chapterPatterns = Arrays.asList(
    Pattern.compile("(Chapter|Section|Part)\\s+[0-9IVXLCDM]+\\s+.+"), // 英文章节格式
    Pattern.compile("第[0-9一二三四五六七八九十]+章\\s+.+"), // 中文章节格式
    Pattern.compile("0x[0-9A-Fa-f]+\\s+.+"), // 十六进制前缀格式
    Pattern.compile("^[0-9]+\\.\\s+.+") // 层级标题格式(如1. Introduction)
);

// 非章节条目排除列表
List<String> excludeList = Arrays.asList("参考文献", "附录", "致谢", "References", "Appendix", "Acknowledgements");

String indexPageText = stripper.getText(document); // 已提取的索引页文本
for (String line : indexPageText.split("\n")) {
    String trimmedLine = line.trim();
    if (trimmedLine.isEmpty()) continue;
    
    // 跳过排除条目
    boolean isExcluded = excludeList.stream().anyMatch(trimmedLine::contains);
    if (isExcluded) continue;
    
    // 匹配章节格式
    boolean isChapter = chapterPatterns.stream().anyMatch(p -> p.matcher(trimmedLine).find());
    if (isChapter) {
        // 提取章节标题和页码
        String[] parts = trimmedLine.split("\\s+(?=[0-9IVXLCDM]+$)"); // 从末尾的页码处分割
        String chapterTitle = parts[0].trim();
        String pageStr = parts[1].trim();
        
        // 转换页码为数字(罗马数字转int的方法自行补充)
        int pageNumber = convertRomanToInt(pageStr);
        if (pageNumber == -1) {
            try {
                pageNumber = Integer.parseInt(pageStr);
            } catch (NumberFormatException e) {
                continue; // 无效页码跳过
            }
        }
        
        // 执行入库操作
        // saveChapter(chapterTitle, pageNumber);
    }
}

3. 处理PDF转文本时的开头无关内容

通过「文本清理+起始行定位」跳过无关内容:

  • 先清理文本中的非打印字符,过滤空白行;
  • 找到第一个符合索引结构的行,从该行开始处理。

示例代码:

String rawPageText = stripper.getText(document);
// 清理非打印字符(保留换行、制表符)
String cleanedText = rawPageText.replaceAll("[\\p{Cntrl}&&[^\r\n\t]]", "");

// 过滤空白行,转为有序列表
List<String> validLines = Arrays.stream(cleanedText.split("\n"))
    .map(String::trim)
    .filter(line -> !line.isEmpty())
    .collect(Collectors.toList());

// 定位第一个索引行的起始位置
int startLine = 0;
Pattern indexLinePattern = Pattern.compile(".*\\s+([0-9]+|I{1,3}|IV|V|VI{0,3})$");
while (startLine < validLines.size()) {
    if (indexLinePattern.matcher(validLines.get(startLine)).find()) {
        break;
    }
    startLine++;
}

// 从startLine开始处理索引内容
for (int i = startLine; i < validLines.size(); i++) {
    String line = validLines.get(i);
    // 执行章节提取逻辑
}

内容的提问来源于stack exchange,提问作者Aakash

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 14:05:19