You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Java实现PDF转Excel:求示例代码及API推荐(iText生成PDF)

Java PDF to Excel Conversion (for iText-generated PDFs)

Hey there! Since you're already working with iText for PDF generation, I've got two reliable approaches to convert those PDFs to Excel in Java—one leveraging the iText ecosystem, and another using free open-source tools. Let's dive in:

Option 1: iText 7 + pdf2Data (iText Ecosystem)

Since you're already using iText, this is a natural fit. pdf2Data is an iText library designed specifically for extracting structured data (like tables) from PDFs and outputting it into formats like CSV or Excel. It works great with iText-generated PDFs because the structure is well-defined.

Step 1: Add Dependencies

First, include the necessary Maven dependencies (adjust versions as needed):

<dependency>
    <groupId>com.itextpdf</groupId>
    <artifactId>itext7-core</artifactId>
    <version>7.2.5</version>
</dependency>
<dependency>
    <groupId>com.itextpdf</groupId>
    <artifactId>pdf2data</artifactId>
    <version>3.0.3</version>
</dependency>
<dependency>
    <groupId>org.apache.poi</groupId>
    <artifactId>poi</artifactId>
    <version>5.2.3</version>
</dependency>
<dependency>
    <groupId>org.apache.poi</groupId>
    <artifactId>poi-ooxml</artifactId>
    <version>5.2.3</version>
</dependency>

Step 2: Example Code

This example extracts table data from a PDF and writes it to an Excel file:

import com.itextpdf.pdf2data.Pdf2DataExtractor;
import com.itextpdf.pdf2data.ResultElement;
import com.itextpdf.pdf2data.extractionstrategy.ExtractionResult;
import org.apache.poi.ss.usermodel.*;
import org.apache.poi.xssf.usermodel.XSSFWorkbook;

import java.io.FileInputStream;
import java.io.FileOutputStream;
import java.io.IOException;
import java.util.List;

public class PdfToExcelItext {
    public static void main(String[] args) throws IOException {
        // Path to your iText-generated PDF
        String pdfPath = "input.pdf";
        // Path for output Excel file
        String excelPath = "output.xlsx";

        // Initialize pdf2Data extractor
        Pdf2DataExtractor extractor = new Pdf2DataExtractor();
        ExtractionResult result = extractor.extract(new FileInputStream(pdfPath));

        // Create Excel workbook and sheet
        Workbook workbook = new XSSFWorkbook();
        Sheet sheet = workbook.createSheet("PDF Data");

        // Iterate over extracted data rows
        int rowNum = 0;
        for (List<ResultElement> rowElements : result.getRows()) {
            Row row = sheet.createRow(rowNum++);
            int colNum = 0;
            for (ResultElement element : rowElements) {
                Cell cell = row.createCell(colNum++);
                cell.setCellValue(element.getText());
            }
        }

        // Auto-size columns for readability
        for (int i = 0; i < result.getRows().get(0).size(); i++) {
            sheet.autoSizeColumn(i);
        }

        // Write workbook to file
        try (FileOutputStream fos = new FileOutputStream(excelPath)) {
            workbook.write(fos);
        }
        workbook.close();

        System.out.println("PDF converted to Excel successfully!");
    }
}

Option 2: Apache PDFBox + Apache POI (Free Open-Source)

If you prefer a completely free, open-source stack without commercial libraries, Apache PDFBox (for PDF parsing) paired with Apache POI (for Excel generation) is a solid choice. It works well with iText-generated PDFs since they're text-based (not scanned).

Step 1: Add Dependencies

Maven dependencies for PDFBox and POI:

<dependency>
    <groupId>org.apache.pdfbox</groupId>
    <artifactId>pdfbox</artifactId>
    <version>2.0.29</version>
</dependency>
<dependency>
    <groupId>org.apache.poi</groupId>
    <artifactId>poi</artifactId>
    <version>5.2.3</version>
</dependency>
<dependency>
    <groupId>org.apache.poi</groupId>
    <artifactId>poi-ooxml</artifactId>
    <version>5.2.3</version>
</dependency>

Step 2: Example Code

This example extracts text content (assuming table structure is separated by spaces/tabs) and writes it to Excel. For more complex tables, you might need to add logic to detect table boundaries:

import org.apache.pdfbox.pdmodel.PDDocument;
import org.apache.pdfbox.text.PDFTextStripper;
import org.apache.poi.ss.usermodel.*;
import org.apache.poi.xssf.usermodel.XSSFWorkbook;

import java.io.FileInputStream;
import java.io.FileOutputStream;
import java.io.IOException;
import java.util.Arrays;

public class PdfToExcelPdfBox {
    public static void main(String[] args) throws IOException {
        String pdfPath = "input.pdf";
        String excelPath = "output.xlsx";

        // Load PDF document
        try (PDDocument document = PDDocument.load(new FileInputStream(pdfPath))) {
            PDFTextStripper stripper = new PDFTextStripper();
            String pdfText = stripper.getText(document);

            // Split text into rows (adjust delimiter if needed)
            String[] rows = pdfText.split("\\r?\\n");

            // Create Excel workbook
            Workbook workbook = new XSSFWorkbook();
            Sheet sheet = workbook.createSheet("PDF Data");

            // Write rows to Excel
            int rowNum = 0;
            for (String rowText : rows) {
                // Split row into columns (assuming space-separated; adjust for your table's delimiter)
                String[] columns = rowText.split("\\s{2,}"); // Split on multiple spaces
                Row row = sheet.createRow(rowNum++);
                int colNum = 0;
                for (String colText : columns) {
                    Cell cell = row.createCell(colNum++);
                    cell.setCellValue(colText.trim());
                }
            }

            // Auto-size columns
            for (int i = 0; i < rows[0].split("\\s{2,").length; i++) {
                sheet.autoSizeColumn(i);
            }

            // Save Excel file
            try (FileOutputStream fos = new FileOutputStream(excelPath)) {
                workbook.write(fos);
            }
            workbook.close();

            System.out.println("Conversion complete!");
        }
    }
}

Key Notes

  • Both approaches work best with text-based PDFs (like those generated by iText). If your PDFs were scanned (image-based), you'd need to add an OCR step (e.g., with Tesseract) first.
  • For complex tables, pdf2Data offers better structured extraction since it can recognize table boundaries and columns more reliably than basic text splitting with PDFBox.

内容的提问来源于stack exchange,提问作者T.Samy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 10:18:20