使用Python批量提取多份PDF指定位置数据并导出至Excel可行吗?
Absolutely! Since your PDFs have a consistent layout, extracting position-specific values and exporting them to a single Excel file in bulk is totally doable with Python. Here's a practical, step-by-step guide tailored to your lab test result use case:
Step 1: Install Required Libraries
First, install the tools we'll need for PDF parsing and Excel handling:
pip install pdfplumber pandas openpyxl
pdfplumber: Lets us precisely target text by its on-page coordinates (perfect for consistent layouts)pandas: Makes organizing and exporting data to Excel straightforwardopenpyxl: Supports writing Excel.xlsxfiles
Step 2: Identify Target Coordinates
Since your PDFs are identical in structure, you only need to map the coordinates of your target fields once:
- Open a sample PDF in a tool like Adobe Acrobat, use the "Measure" tool to note the
x0, y0(top-left) andx1, y1(bottom-right) coordinates for:- Customer name
- Each lab test value you want to extract
- Alternatively, use
pdfplumber's debug mode to print text positions:import pdfplumber with pdfplumber.open("sample_lab.pdf") as pdf: page = pdf.pages[0] print(page.chars) # Prints all text chunks with their coordinates
Step 3: Write the Bulk Extraction Script
This script will loop through all PDFs in a folder, extract your target values, and compile everything into an Excel file:
import pdfplumber import pandas as pd import os def extract_single_pdf(pdf_path): # Replace these coordinates with your actual mapped values target_regions = { "Customer Name": {"x0": 60, "y0": 110, "x1": 320, "y1": 135}, "Hemoglobin": {"x0": 450, "y0": 220, "x1": 520, "y1": 240}, "White Blood Cells": {"x0": 450, "y0": 260, "x1": 520, "y1": 280}, "Platelets": {"x0": 450, "y0": 300, "x1": 520, "y1": 320} } extracted_data = {} with pdfplumber.open(pdf_path) as pdf: # Assume target data is on the first page; adjust index if needed page = pdf.pages[0] for field, coords in target_regions.items(): # Extract text from the defined bounding box text = page.within_bbox((coords["x0"], coords["y0"], coords["x1"], coords["y1"])).extract_text() # Clean up extra whitespace/newlines extracted_data[field] = text.strip() if text else "No Data" # Add the original filename for reference extracted_data["Source PDF"] = os.path.basename(pdf_path) return extracted_data # Process all PDFs in your target folder pdf_directory = "./lab_test_pdfs" # Replace with your folder path all_results = [] for filename in os.listdir(pdf_directory): if filename.lower().endswith(".pdf"): full_path = os.path.join(pdf_directory, filename) print(f"Processing {filename}...") all_results.append(extract_single_pdf(full_path)) # Export compiled data to Excel results_df = pd.DataFrame(all_results) results_df.to_excel("consolidated_lab_results.xlsx", index=False, engine="openpyxl") print("Done! All results saved to consolidated_lab_results.xlsx")
Step 4: Adjust and Test
- Run the script with one sample PDF first to verify coordinates are correct
- If your PDFs are scanned images (not editable text), add OCR support with
pytesseract:
Then modify the extraction part to convert the page to an image and run OCR:pip install pytesseract pillowfrom PIL import Image import pytesseract # Inside the extract_single_pdf function, replace the text extraction line with: page_image = page.to_image().original text = pytesseract.image_to_string(page_image).strip()
内容的提问来源于stack exchange,提问作者Alex Weiner
相关产品推荐
相关产品推荐

