You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python批量提取多份PDF指定位置数据并导出至Excel可行吗?

Absolutely! Since your PDFs have a consistent layout, extracting position-specific values and exporting them to a single Excel file in bulk is totally doable with Python. Here's a practical, step-by-step guide tailored to your lab test result use case:

Step 1: Install Required Libraries

First, install the tools we'll need for PDF parsing and Excel handling:

pip install pdfplumber pandas openpyxl
  • pdfplumber: Lets us precisely target text by its on-page coordinates (perfect for consistent layouts)
  • pandas: Makes organizing and exporting data to Excel straightforward
  • openpyxl: Supports writing Excel .xlsx files
Step 2: Identify Target Coordinates

Since your PDFs are identical in structure, you only need to map the coordinates of your target fields once:

  1. Open a sample PDF in a tool like Adobe Acrobat, use the "Measure" tool to note the x0, y0 (top-left) and x1, y1 (bottom-right) coordinates for:
    • Customer name
    • Each lab test value you want to extract
  2. Alternatively, use pdfplumber's debug mode to print text positions:
    import pdfplumber
    with pdfplumber.open("sample_lab.pdf") as pdf:
        page = pdf.pages[0]
        print(page.chars)  # Prints all text chunks with their coordinates
    
Step 3: Write the Bulk Extraction Script

This script will loop through all PDFs in a folder, extract your target values, and compile everything into an Excel file:

import pdfplumber
import pandas as pd
import os

def extract_single_pdf(pdf_path):
    # Replace these coordinates with your actual mapped values
    target_regions = {
        "Customer Name": {"x0": 60, "y0": 110, "x1": 320, "y1": 135},
        "Hemoglobin": {"x0": 450, "y0": 220, "x1": 520, "y1": 240},
        "White Blood Cells": {"x0": 450, "y0": 260, "x1": 520, "y1": 280},
        "Platelets": {"x0": 450, "y0": 300, "x1": 520, "y1": 320}
    }

    extracted_data = {}
    with pdfplumber.open(pdf_path) as pdf:
        # Assume target data is on the first page; adjust index if needed
        page = pdf.pages[0]
        
        for field, coords in target_regions.items():
            # Extract text from the defined bounding box
            text = page.within_bbox((coords["x0"], coords["y0"], coords["x1"], coords["y1"])).extract_text()
            # Clean up extra whitespace/newlines
            extracted_data[field] = text.strip() if text else "No Data"
    
    # Add the original filename for reference
    extracted_data["Source PDF"] = os.path.basename(pdf_path)
    return extracted_data

# Process all PDFs in your target folder
pdf_directory = "./lab_test_pdfs"  # Replace with your folder path
all_results = []

for filename in os.listdir(pdf_directory):
    if filename.lower().endswith(".pdf"):
        full_path = os.path.join(pdf_directory, filename)
        print(f"Processing {filename}...")
        all_results.append(extract_single_pdf(full_path))

# Export compiled data to Excel
results_df = pd.DataFrame(all_results)
results_df.to_excel("consolidated_lab_results.xlsx", index=False, engine="openpyxl")
print("Done! All results saved to consolidated_lab_results.xlsx")
Step 4: Adjust and Test
  • Run the script with one sample PDF first to verify coordinates are correct
  • If your PDFs are scanned images (not editable text), add OCR support with pytesseract:
    pip install pytesseract pillow
    
    Then modify the extraction part to convert the page to an image and run OCR:
    from PIL import Image
    import pytesseract
    
    # Inside the extract_single_pdf function, replace the text extraction line with:
    page_image = page.to_image().original
    text = pytesseract.image_to_string(page_image).strip()
    

内容的提问来源于stack exchange,提问作者Alex Weiner

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 07:07:38