You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用pytesseract与PyMuPDF提取PDF图片文本报错:'dict'无width属性

解决PDF图片文本提取中的'dict' object has no attribute 'width'错误

错误原因

doc.extract_image(xref)返回的是字典类型,包含图片的宽、高、二进制数据等元信息(键如width、height、image)。你的代码直接将这个字典传入extract_text_from_image函数,试图访问image.width,但字典没有width属性,因此触发错误。

修正方案

修改extract_text_from_image函数,从传入的字典中提取所需参数,正确创建PIL Image对象;同时优化图片模式的处理(根据PDF中图片的色彩空间自动适配)。

修正后的完整代码

import os
import pytesseract
import fitz  # PyMuPDF
from PIL import Image

def extract_text_from_image(img_dict):
    try:
        # 从字典中提取图片参数
        width = img_dict['width']
        height = img_dict['height']
        img_data = img_dict['image']
        colorspace = img_dict['colorspace']
        
        # 根据色彩空间确定图片模式
        mode = "RGB" if colorspace == 3 else "L" if colorspace == 1 else "RGBA" if colorspace == 4 else "RGB"
        # 从二进制数据创建PIL Image
        image = Image.frombytes(mode, (width, height), img_data)
        
        # 转灰度提升识别准确率
        image = image.convert("L")
        extracted_text = pytesseract.image_to_string(image, config='--oem 3 --psm 6')
        return extracted_text
    except Exception as e:
        print(f"Error processing image: {str(e)}")
        return ""

def process_pdf(pdf_path):
    try:
        doc = fitz.open(pdf_path)
        total_pages = doc.page_count
        for i in range(total_pages):
            page = doc[i]
            image_list = page.get_images(full=True)
            if image_list:
                for img in image_list:
                    xref = img[0]
                    base_image = doc.extract_image(xref)
                    extracted_text = extract_text_from_image(base_image)
                    print(f"Page {i + 1}/{total_pages}: Extracted text: {extracted_text}")
            else:
                print(f"Page {i + 1}/{total_pages}: No images found")
            # 显示处理进度
            progress_percent = (i + 1) / total_pages * 100
            print(f"Processing progress: [{'#' * int(progress_percent / 2):50s}] {progress_percent:.2f}%")
        # 检查文件大小
        file_size_mb = os.path.getsize(pdf_path) / (1024 * 1024)
        if file_size_mb > 10:
            print(f"File size ({file_size_mb:.2f} MB) exceeds threshold.")
    except Exception as e:
        print(f"Error processing PDF {pdf_path}: {str(e)}")

process_pdf(r"your_pdf.pdf")

额外说明

  • 函数参数名从image改为img_dict,更直观表示传入的是字典
  • 增加了色彩空间判断,适配不同类型的PDF图片(灰度、RGB、RGBA等)
  • 保留了原有的错误处理和进度显示逻辑

内容的提问来源于stack exchange,提问作者Neon delphifer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.28 14:15:22