如何用Python 3.6提取PDF中的表格(CSV)与图表(图片)?
Got it, let's solve this problem properly. Your byte-matching trick for JPGs only works for embedded raster images, but most professional PDFs (like annual reports) use vector-based tables and charts. Here's a robust, Python 3.6-compatible solution using two powerful libraries that handle both structured tables and vector/raster charts.
Required Libraries
We'll use:
- pdfplumber: Great for extracting structured tables from PDFs (even if they're drawn with vectors, not just image tables)
- PyMuPDF (fitz): Efficiently renders PDF pages (or elements) into high-quality images, perfect for capturing vector charts that can't be pulled via byte markers
Install them with this command (specifying a PyMuPDF version compatible with Python 3.6):
pip install pdfplumber pymupdf==1.20.2
Full Code Implementation
This code downloads the target PDF, extracts the table from page 11 as CSV, and the chart from page 12 as a high-resolution PNG:
import pdfplumber import fitz # PyMuPDF import urllib.request def extract_table_to_csv(pdf_path, page_number, output_file): """Extract structured table from a specific PDF page to CSV""" with pdfplumber.open(pdf_path) as pdf: # Convert 1-based page number to 0-based index target_page = pdf.pages[page_number - 1] # Extract table (handles vector-drawn and properly formatted tables) table_data = target_page.extract_table() if not table_data: print(f"No structured table found on page {page_number}") return # Write cleaned table data to CSV with open(output_file, 'w', encoding='utf-8') as csv_file: for row in table_data: # Clean empty cells and format as CSV cleaned_row = [str(cell).strip() if cell else '' for cell in row] csv_file.write(','.join(cleaned_row) + '\n') print(f"Table saved to {output_file} successfully!") def extract_chart_to_image(pdf_path, page_number, output_file, dpi=300): """Render a PDF page to a high-quality image (captures vector/raster charts)""" doc = fitz.open(pdf_path) target_page = doc[page_number - 1] # Render page with specified DPI for sharpness pixmap = target_page.get_pixmap(dpi=dpi) pixmap.save(output_file) doc.close() print(f"Chart image saved to {output_file} successfully!") # --- Main Execution --- # Download the target PDF first (libraries need local file access) pdf_url = "https://hartmannazurecdn.azureedge.net/media/2369/annual-report-2017.pdf" local_pdf = "annual-report-2017.pdf" urllib.request.urlretrieve(pdf_url, local_pdf) # Extract table from page 11 extract_table_to_csv(local_pdf, 11, "page11_financial_table.csv") # Extract chart from page 12 extract_chart_to_image(local_pdf, 12, "page12_performance_chart.png")
Key Advantages Over Your Previous Approach
- Table Extraction: pdfplumber analyzes the PDF's internal structure to pull tables as structured text, not just pixel data. This means you get editable CSV instead of needing OCR for image-based tables (though if your table is a scanned image, you can add a pytesseract OCR step easily).
- Chart Extraction: PyMuPDF renders the entire page (or specific elements) into a raster image, so it captures both vector-based charts and embedded raster images. The high DPI setting ensures the output is crisp and usable.
Edge Case Handling
If you ever need to extract only the chart (not the whole page), you can modify the extract_chart_to_image function to crop the pixmap using page coordinates. You can get coordinates using PyMuPDF's page inspection tools, but for most cases, rendering the full page is efficient enough.
内容的提问来源于stack exchange,提问作者Aakash Basu

