如何提取PDF中图片内的文本?需保留文本顺序及优化现有代码
Extract Native PDF Text + Image Text in Order
Hey there! Your current code uses PyPDF2 which works well for pulling native text from PDFs, but it can't read text embedded inside images. To capture both types of text while keeping their original order on each page, we'll combine PDF parsing with OCR (Optical Character Recognition). Here's a complete solution:
Step 1: Install Required Libraries
First, install the tools we'll need:
pip install pdfplumber pytesseract pillow
- pdfplumber: Lets us extract both text chunks and images from PDFs, along with their positional data (critical for maintaining order)
- pytesseract: The OCR engine that reads text from images
- Pillow: Handles image processing for OCR
Note: You'll also need to install Tesseract itself:
- On Ubuntu/Debian:
sudo apt install tesseract-ocr- On Windows: Download the official installer, then add its path to your system environment variables (or specify it directly in code as shown below)
Step 2: Updated Code
Replace your existing convertPdfToText method with this version. It will extract native text and image text in the correct order:
import os import sys import pdfplumber import pytesseract from PIL import Image from io import BytesIO def convertPdfToText(self, fileToConvert, outputTextFile): # Configure Tesseract path (uncomment and update if needed for Windows) # pytesseract.pytesseract.tesseract_cmd = r'C:\Program Files\Tesseract-OCR\tesseract.exe' try: with open(outputTextFile, 'w', encoding='utf-8') as text_file: with pdfplumber.open(fileToConvert) as pdf: for page_num, page in enumerate(pdf.pages, start=1): # Collect all page elements (text chunks and images) with position data elements = [] # Add individual text chunks to elements list for text_chunk in page.extract_text_words(): elements.append({ 'type': 'text', 'content': text_chunk['text'], 'position': text_chunk['top'] }) # Add OCR'd image text to elements list for img in page.images: # Extract image data and convert to a Pillow Image object img_data = page.extract_image(img['name'])['image'] img_obj = Image.open(BytesIO(img_data)) # Run OCR to extract text from the image ocr_text = pytesseract.image_to_string(img_obj) elements.append({ 'type': 'image_text', 'content': ocr_text, 'position': img['top'] }) # Sort elements by vertical position (higher top = appears earlier on page) elements.sort(key=lambda x: x['position'], reverse=True) # Write all elements to the output file for elem in elements: text_file.write(elem['content'] + ' ') # Add a page break for readability text_file.write(f"\n--- Page {page_num} ---\n") print(f"Successfully extracted text to {outputTextFile}") except Exception as e: sys.exit(f"Error occurred: {str(e)}")
Key Details Explained
- Positional Sorting: We sort all elements by their
topcoordinate. Since PDF coordinates start at the bottom-left of the page, a highertopvalue means the element is closer to the top—sorting in reverse order keeps content in natural reading order. - Text Chunks: Instead of extracting full page text at once, we pull individual words with position data to ensure precise ordering relative to images.
- Image OCR: For each image, we extract raw data, convert it to a Pillow Image, then use Tesseract to read embedded text.
- Encoding: We use
utf-8for the output file to support special characters across languages.
Optional Enhancements
- Language Support: Add Tesseract language packs and specify languages in the OCR call, e.g.:
ocr_text = pytesseract.image_to_string(img_obj, lang='eng+chi_sim') # English + Simplified Chinese - Image Preprocessing: Improve OCR accuracy by preprocessing images (grayscale conversion, thresholding) before running Tesseract.
- Granular Error Handling: Add specific checks for missing Tesseract installations or corrupted PDF files.
内容的提问来源于stack exchange,提问作者darknight
相关产品推荐
相关产品推荐

