You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python提取影印/扫描版PDF文件中的文本?

提取扫描/影印PDF文本的Python方案

扫描或影印生成的PDF本质是由图像组成的,没有可直接读取的文本层,所以PyPDF2这类仅处理PDF文本层的工具无法提取内容,需要借助**OCR(光学字符识别)**技术。下面是最常用的实现方案:

依赖准备

需要安装两个Python库,以及本地的OCR引擎:

  • pdf2image:用于将PDF页面转换为图像
  • pytesseract:用于调用Tesseract OCR引擎识别图像文本
  • Tesseract OCR引擎:本地安装(Windows/macOS/Linux都有对应安装包)

安装命令:

pip install pdf2image pytesseract

Windows用户需额外安装Tesseract OCR,并确保其路径加入系统环境变量;macOS可通过brew install tesseract安装;Linux用sudo apt install tesseract-ocr。

代码实现

from pdf2image import convert_from_path
import pytesseract

# 如果Tesseract不在系统环境变量中,需指定路径(Windows示例)
# pytesseract.pytesseract.tesseract_cmd = r'C:\Program Files\Tesseract-OCR\tesseract.exe'

def extract_scanned_pdf_text(pdf_path):
    # 将PDF转换为图像列表
    pages = convert_from_path(pdf_path)
    full_text = ""
    
    for page_num, page_img in enumerate(pages, start=1):
        # 识别单页图像的文本,lang指定语言(中文用chi_sim,英文用eng)
        page_text = pytesseract.image_to_string(page_img, lang='chi_sim')
        full_text += f"--- 第{page_num}页 ---\n{page_text}\n"
    
    return full_text

# 使用示例
pdf_path = r'C:\filepath\file.pdf'
extracted_text = extract_scanned_pdf_text(pdf_path)
print(extracted_text)

关键说明

  • lang参数:根据PDF语言选择,支持多语言组合(如eng+chi_sim同时识别中英)
  • 图像预处理:如果PDF图像模糊、倾斜,可先使用OpenCV进行降噪、矫正操作,提升识别准确率
  • 替代方案:也可以使用easyocr库(无需额外安装Tesseract),但识别速度和准确率略低于Tesseract

内容的提问来源于stack exchange,提问作者syntax_of_vectors

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 03:55:15