如何实现多格式文档(Docx、HTML等)页数统计?附现有PDF实现代码
Hey there! Let's figure out how to extend your existing PDF page-counting code to handle Docx, HTML, and other formats. I'll walk you through a clean, maintainable approach that follows good design principles:
1. Use the Strategy Pattern to Decouple Logic
First, let's separate the page-counting logic for each format so we can add new formats without rewriting existing code. Start with a unified interface for all counters:
public interface DocumentPageCounter { int countPages(MultipartFile file) throws Exception; boolean supports(String fileExtension); // Check if this counter handles the format }
2. Implement Counters for Each Format
PDF Counter (Reuse Your Existing Code)
We'll wrap your current PDFBox logic into a dedicated class:
public class PdfPageCounter implements DocumentPageCounter { @Override public int countPages(MultipartFile file) throws Exception { // Use try-with-resources to auto-close streams/documents and avoid leaks try (InputStream is = file.getInputStream(); PDDocument document = PDDocument.load(is)) { return document != null ? document.getNumberOfPages() : 0; } } @Override public boolean supports(String fileExtension) { return "pdf".equalsIgnoreCase(fileExtension); } }
Docx Counter (Using Apache POI)
Docx doesn't store a reliable native page count (it depends on rendering settings), but we have two solid options:
public class DocxPageCounter implements DocumentPageCounter { @Override public int countPages(MultipartFile file) throws Exception { try (XWPFDocument docx = new XWPFDocument(file.getInputStream())) { // Option 1: Use the document's saved page count (fast but may be outdated) Integer savedPageCount = docx.getProperties().getExtendedProperties() .getUnderlyingProperties().getPages(); if (savedPageCount != null) { return savedPageCount; } // Option 2: Manually count page breaks + 1 (more reliable for edited docs) int pageBreakCount = 0; for (XWPFParagraph para : docx.getParagraphs()) { for (CTPPr paraProps : para.getCTP().getPPrList()) { if (paraProps.getPgBr() != null) { pageBreakCount++; } } } return pageBreakCount + 1; // At least 1 page exists } } @Override public boolean supports(String fileExtension) { return "docx".equalsIgnoreCase(fileExtension); } }
HTML Counter (Using Flying Saucer + PDFBox)
HTML is responsive, so page counts depend on rendering. The easiest reliable method is to convert HTML to a temporary PDF, then reuse your PDF counting logic:
public class HtmlPageCounter implements DocumentPageCounter { @Override public int countPages(MultipartFile file) throws Exception { String htmlContent = new String(file.getBytes(), StandardCharsets.UTF_8); try (ByteArrayOutputStream tempPdfStream = new ByteArrayOutputStream()) { // Convert HTML to PDF ITextRenderer renderer = new ITextRenderer(); renderer.setDocumentFromString(htmlContent); renderer.layout(); renderer.createPDF(tempPdfStream); // Count pages from the temporary PDF try (InputStream pdfIs = new ByteArrayInputStream(tempPdfStream.toByteArray()); PDDocument pdfDoc = PDDocument.load(pdfIs)) { return pdfDoc.getNumberOfPages(); } } } @Override public boolean supports(String fileExtension) { return "html".equalsIgnoreCase(fileExtension) || "htm".equalsIgnoreCase(fileExtension); } }
3. Combine Counters into a Unified Service
Create a service that routes files to the correct counter based on their extension:
@Service public class DocumentPageService { private final List<DocumentPageCounter> counters; // For non-Spring apps, initialize counters manually public DocumentPageService() { this.counters = Arrays.asList( new PdfPageCounter(), new DocxPageCounter(), new HtmlPageCounter() ); } // For Spring apps, autowire all implementations: // public DocumentPageService(List<DocumentPageCounter> counters) { // this.counters = counters; // } public int numberOfPages(@RequestBody MultipartFile inputFile) throws Exception { String fileName = inputFile.getOriginalFilename(); if (fileName == null) { throw new IllegalArgumentException("File name cannot be null"); } String extension = fileName.substring(fileName.lastIndexOf(".") + 1); // Find the right counter for the file format DocumentPageCounter counter = counters.stream() .filter(c -> c.supports(extension)) .findFirst() .orElseThrow(() -> new UnsupportedOperationException("Unsupported file format: " + extension)); return counter.countPages(inputFile); } }
4. Add Required Dependencies (Maven Example)
Add these to your pom.xml to pull in the necessary libraries:
<!-- PDFBox for PDF handling --> <dependency> <groupId>org.apache.pdfbox</groupId> <artifactId>pdfbox</artifactId> <version>2.0.32</version> </dependency> <!-- Apache POI for Docx handling --> <dependency> <groupId>org.apache.poi</groupId> <artifactId>poi-ooxml</artifactId> <version>5.2.3</version> </dependency> <!-- Flying Saucer for HTML to PDF conversion --> <dependency> <groupId>org.xhtmlrenderer</groupId> <artifactId>flying-saucer-pdf</artifactId> <version>9.1.22</version> </dependency>
Key Notes
- Docx Accuracy: The saved page count is fast but may be outdated if the document was edited and not re-saved. Manual page break counting is more reliable for dynamic docs.
- HTML Accuracy: Different renderers (Flying Saucer vs. browsers) may produce slightly different page counts. For browser-matching counts, you could use Headless Chrome, but that adds complexity.
- Resource Management: Always use
try-with-resourcesto close streams and document objects—this prevents memory leaks.
内容的提问来源于stack exchange,提问作者Gerard Lopez

