You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python 3.6提取PDF中的表格(CSV)与图表(图片)?

Solution for Extracting Tables (CSV) and Charts (Images) from PDFs

Got it, let's solve this problem properly. Your byte-matching trick for JPGs only works for embedded raster images, but most professional PDFs (like annual reports) use vector-based tables and charts. Here's a robust, Python 3.6-compatible solution using two powerful libraries that handle both structured tables and vector/raster charts.

Required Libraries

We'll use:

  • pdfplumber: Great for extracting structured tables from PDFs (even if they're drawn with vectors, not just image tables)
  • PyMuPDF (fitz): Efficiently renders PDF pages (or elements) into high-quality images, perfect for capturing vector charts that can't be pulled via byte markers

Install them with this command (specifying a PyMuPDF version compatible with Python 3.6):

pip install pdfplumber pymupdf==1.20.2

Full Code Implementation

This code downloads the target PDF, extracts the table from page 11 as CSV, and the chart from page 12 as a high-resolution PNG:

import pdfplumber
import fitz  # PyMuPDF
import urllib.request

def extract_table_to_csv(pdf_path, page_number, output_file):
    """Extract structured table from a specific PDF page to CSV"""
    with pdfplumber.open(pdf_path) as pdf:
        # Convert 1-based page number to 0-based index
        target_page = pdf.pages[page_number - 1]
        # Extract table (handles vector-drawn and properly formatted tables)
        table_data = target_page.extract_table()
        
        if not table_data:
            print(f"No structured table found on page {page_number}")
            return
        
        # Write cleaned table data to CSV
        with open(output_file, 'w', encoding='utf-8') as csv_file:
            for row in table_data:
                # Clean empty cells and format as CSV
                cleaned_row = [str(cell).strip() if cell else '' for cell in row]
                csv_file.write(','.join(cleaned_row) + '\n')
        print(f"Table saved to {output_file} successfully!")

def extract_chart_to_image(pdf_path, page_number, output_file, dpi=300):
    """Render a PDF page to a high-quality image (captures vector/raster charts)"""
    doc = fitz.open(pdf_path)
    target_page = doc[page_number - 1]
    
    # Render page with specified DPI for sharpness
    pixmap = target_page.get_pixmap(dpi=dpi)
    pixmap.save(output_file)
    doc.close()
    print(f"Chart image saved to {output_file} successfully!")

# --- Main Execution ---
# Download the target PDF first (libraries need local file access)
pdf_url = "https://hartmannazurecdn.azureedge.net/media/2369/annual-report-2017.pdf"
local_pdf = "annual-report-2017.pdf"
urllib.request.urlretrieve(pdf_url, local_pdf)

# Extract table from page 11
extract_table_to_csv(local_pdf, 11, "page11_financial_table.csv")

# Extract chart from page 12
extract_chart_to_image(local_pdf, 12, "page12_performance_chart.png")

Key Advantages Over Your Previous Approach

  • Table Extraction: pdfplumber analyzes the PDF's internal structure to pull tables as structured text, not just pixel data. This means you get editable CSV instead of needing OCR for image-based tables (though if your table is a scanned image, you can add a pytesseract OCR step easily).
  • Chart Extraction: PyMuPDF renders the entire page (or specific elements) into a raster image, so it captures both vector-based charts and embedded raster images. The high DPI setting ensures the output is crisp and usable.

Edge Case Handling

If you ever need to extract only the chart (not the whole page), you can modify the extract_chart_to_image function to crop the pixmap using page coordinates. You can get coordinates using PyMuPDF's page inspection tools, but for most cases, rendering the full page is efficient enough.

内容的提问来源于stack exchange,提问作者Aakash Basu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 05:09:49