You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何实现多格式文档(Docx、HTML等)页数统计?附现有PDF实现代码

Hey there! Let's figure out how to extend your existing PDF page-counting code to handle Docx, HTML, and other formats. I'll walk you through a clean, maintainable approach that follows good design principles:

Extending Multi-Format Document Page Counting

1. Use the Strategy Pattern to Decouple Logic

First, let's separate the page-counting logic for each format so we can add new formats without rewriting existing code. Start with a unified interface for all counters:

public interface DocumentPageCounter {
    int countPages(MultipartFile file) throws Exception;
    boolean supports(String fileExtension); // Check if this counter handles the format
}

2. Implement Counters for Each Format

PDF Counter (Reuse Your Existing Code)

We'll wrap your current PDFBox logic into a dedicated class:

public class PdfPageCounter implements DocumentPageCounter {
    @Override
    public int countPages(MultipartFile file) throws Exception {
        // Use try-with-resources to auto-close streams/documents and avoid leaks
        try (InputStream is = file.getInputStream();
             PDDocument document = PDDocument.load(is)) {
            return document != null ? document.getNumberOfPages() : 0;
        }
    }

    @Override
    public boolean supports(String fileExtension) {
        return "pdf".equalsIgnoreCase(fileExtension);
    }
}

Docx Counter (Using Apache POI)

Docx doesn't store a reliable native page count (it depends on rendering settings), but we have two solid options:

public class DocxPageCounter implements DocumentPageCounter {
    @Override
    public int countPages(MultipartFile file) throws Exception {
        try (XWPFDocument docx = new XWPFDocument(file.getInputStream())) {
            // Option 1: Use the document's saved page count (fast but may be outdated)
            Integer savedPageCount = docx.getProperties().getExtendedProperties()
                .getUnderlyingProperties().getPages();
            if (savedPageCount != null) {
                return savedPageCount;
            }
            
            // Option 2: Manually count page breaks + 1 (more reliable for edited docs)
            int pageBreakCount = 0;
            for (XWPFParagraph para : docx.getParagraphs()) {
                for (CTPPr paraProps : para.getCTP().getPPrList()) {
                    if (paraProps.getPgBr() != null) {
                        pageBreakCount++;
                    }
                }
            }
            return pageBreakCount + 1; // At least 1 page exists
        }
    }

    @Override
    public boolean supports(String fileExtension) {
        return "docx".equalsIgnoreCase(fileExtension);
    }
}

HTML Counter (Using Flying Saucer + PDFBox)

HTML is responsive, so page counts depend on rendering. The easiest reliable method is to convert HTML to a temporary PDF, then reuse your PDF counting logic:

public class HtmlPageCounter implements DocumentPageCounter {
    @Override
    public int countPages(MultipartFile file) throws Exception {
        String htmlContent = new String(file.getBytes(), StandardCharsets.UTF_8);
        try (ByteArrayOutputStream tempPdfStream = new ByteArrayOutputStream()) {
            // Convert HTML to PDF
            ITextRenderer renderer = new ITextRenderer();
            renderer.setDocumentFromString(htmlContent);
            renderer.layout();
            renderer.createPDF(tempPdfStream);
            
            // Count pages from the temporary PDF
            try (InputStream pdfIs = new ByteArrayInputStream(tempPdfStream.toByteArray());
                 PDDocument pdfDoc = PDDocument.load(pdfIs)) {
                return pdfDoc.getNumberOfPages();
            }
        }
    }

    @Override
    public boolean supports(String fileExtension) {
        return "html".equalsIgnoreCase(fileExtension) || "htm".equalsIgnoreCase(fileExtension);
    }
}

3. Combine Counters into a Unified Service

Create a service that routes files to the correct counter based on their extension:

@Service
public class DocumentPageService {
    private final List<DocumentPageCounter> counters;

    // For non-Spring apps, initialize counters manually
    public DocumentPageService() {
        this.counters = Arrays.asList(
                new PdfPageCounter(),
                new DocxPageCounter(),
                new HtmlPageCounter()
        );
    }

    // For Spring apps, autowire all implementations:
    // public DocumentPageService(List<DocumentPageCounter> counters) {
    //     this.counters = counters;
    // }

    public int numberOfPages(@RequestBody MultipartFile inputFile) throws Exception {
        String fileName = inputFile.getOriginalFilename();
        if (fileName == null) {
            throw new IllegalArgumentException("File name cannot be null");
        }
        String extension = fileName.substring(fileName.lastIndexOf(".") + 1);

        // Find the right counter for the file format
        DocumentPageCounter counter = counters.stream()
                .filter(c -> c.supports(extension))
                .findFirst()
                .orElseThrow(() -> new UnsupportedOperationException("Unsupported file format: " + extension));

        return counter.countPages(inputFile);
    }
}

4. Add Required Dependencies (Maven Example)

Add these to your pom.xml to pull in the necessary libraries:

<!-- PDFBox for PDF handling -->
<dependency>
    <groupId>org.apache.pdfbox</groupId>
    <artifactId>pdfbox</artifactId>
    <version>2.0.32</version>
</dependency>
<!-- Apache POI for Docx handling -->
<dependency>
    <groupId>org.apache.poi</groupId>
    <artifactId>poi-ooxml</artifactId>
    <version>5.2.3</version>
</dependency>
<!-- Flying Saucer for HTML to PDF conversion -->
<dependency>
    <groupId>org.xhtmlrenderer</groupId>
    <artifactId>flying-saucer-pdf</artifactId>
    <version>9.1.22</version>
</dependency>

Key Notes

  • Docx Accuracy: The saved page count is fast but may be outdated if the document was edited and not re-saved. Manual page break counting is more reliable for dynamic docs.
  • HTML Accuracy: Different renderers (Flying Saucer vs. browsers) may produce slightly different page counts. For browser-matching counts, you could use Headless Chrome, but that adds complexity.
  • Resource Management: Always use try-with-resources to close streams and document objects—this prevents memory leaks.

内容的提问来源于stack exchange,提问作者Gerard Lopez

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 06:29:42