You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何提取PDF中图片内的文本?需保留文本顺序及优化现有代码

Extract Native PDF Text + Image Text in Order

Hey there! Your current code uses PyPDF2 which works well for pulling native text from PDFs, but it can't read text embedded inside images. To capture both types of text while keeping their original order on each page, we'll combine PDF parsing with OCR (Optical Character Recognition). Here's a complete solution:

Step 1: Install Required Libraries

First, install the tools we'll need:

pip install pdfplumber pytesseract pillow
  • pdfplumber: Lets us extract both text chunks and images from PDFs, along with their positional data (critical for maintaining order)
  • pytesseract: The OCR engine that reads text from images
  • Pillow: Handles image processing for OCR

Note: You'll also need to install Tesseract itself:

  • On Ubuntu/Debian: sudo apt install tesseract-ocr
  • On Windows: Download the official installer, then add its path to your system environment variables (or specify it directly in code as shown below)

Step 2: Updated Code

Replace your existing convertPdfToText method with this version. It will extract native text and image text in the correct order:

import os
import sys
import pdfplumber
import pytesseract
from PIL import Image
from io import BytesIO

def convertPdfToText(self, fileToConvert, outputTextFile):
    # Configure Tesseract path (uncomment and update if needed for Windows)
    # pytesseract.pytesseract.tesseract_cmd = r'C:\Program Files\Tesseract-OCR\tesseract.exe'

    try:
        with open(outputTextFile, 'w', encoding='utf-8') as text_file:
            with pdfplumber.open(fileToConvert) as pdf:
                for page_num, page in enumerate(pdf.pages, start=1):
                    # Collect all page elements (text chunks and images) with position data
                    elements = []

                    # Add individual text chunks to elements list
                    for text_chunk in page.extract_text_words():
                        elements.append({
                            'type': 'text',
                            'content': text_chunk['text'],
                            'position': text_chunk['top']
                        })

                    # Add OCR'd image text to elements list
                    for img in page.images:
                        # Extract image data and convert to a Pillow Image object
                        img_data = page.extract_image(img['name'])['image']
                        img_obj = Image.open(BytesIO(img_data))

                        # Run OCR to extract text from the image
                        ocr_text = pytesseract.image_to_string(img_obj)

                        elements.append({
                            'type': 'image_text',
                            'content': ocr_text,
                            'position': img['top']
                        })

                    # Sort elements by vertical position (higher top = appears earlier on page)
                    elements.sort(key=lambda x: x['position'], reverse=True)

                    # Write all elements to the output file
                    for elem in elements:
                        text_file.write(elem['content'] + ' ')
                    # Add a page break for readability
                    text_file.write(f"\n--- Page {page_num} ---\n")

        print(f"Successfully extracted text to {outputTextFile}")

    except Exception as e:
        sys.exit(f"Error occurred: {str(e)}")

Key Details Explained

  • Positional Sorting: We sort all elements by their top coordinate. Since PDF coordinates start at the bottom-left of the page, a higher top value means the element is closer to the top—sorting in reverse order keeps content in natural reading order.
  • Text Chunks: Instead of extracting full page text at once, we pull individual words with position data to ensure precise ordering relative to images.
  • Image OCR: For each image, we extract raw data, convert it to a Pillow Image, then use Tesseract to read embedded text.
  • Encoding: We use utf-8 for the output file to support special characters across languages.

Optional Enhancements

  • Language Support: Add Tesseract language packs and specify languages in the OCR call, e.g.:
    ocr_text = pytesseract.image_to_string(img_obj, lang='eng+chi_sim')  # English + Simplified Chinese
    
  • Image Preprocessing: Improve OCR accuracy by preprocessing images (grayscale conversion, thresholding) before running Tesseract.
  • Granular Error Handling: Add specific checks for missing Tesseract installations or corrupted PDF files.

内容的提问来源于stack exchange,提问作者darknight

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:42:28