You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python如何提取AutoCAD生成的图片型PDF中的文本?

提取AutoCAD生成的图片类PDF文本

这类PDF本质是图像集合,无法通过PyPDF2、PDFMiner这类针对可编辑文本的工具提取内容,必须使用OCR(光学字符识别)工具来提取图像中的文本。以下是几个实用的Python库方案:

1. PyTesseract + pdf2image

PyTesseract是Google Tesseract OCR引擎的Python封装,需要配合pdf2image将PDF页面转为图像,再进行识别。

依赖安装

pip install pytesseract pillow pdf2image

注意:需额外安装Tesseract OCR引擎,以及Poppler工具用于PDF转图像(需将工具路径配置到系统环境变量,或在代码中指定)。

示例代码

import pytesseract
from pdf2image import convert_from_path
from PIL import Image

# 若Tesseract不在系统PATH中,指定其安装路径
pytesseract.pytesseract.tesseract_cmd = r'C:\Program Files\Tesseract-OCR\tesseract.exe'

# 将PDF转换为图像列表
pdf_pages = convert_from_path('autocad_output.pdf')

# 逐页识别文本
for idx, page_img in enumerate(pdf_pages):
    # 可选:图像预处理提升识别率(转灰度、二值化)
    gray_img = page_img.convert('L')
    binary_img = gray_img.point(lambda x: 0 if x < 127 else 255, '1')
    
    # 识别文本,lang参数指定语言(中文用chi_sim,英文用eng)
    page_text = pytesseract.image_to_string(binary_img, lang='chi_sim')
    print(f"第{idx+1}页识别结果:\n{page_text}\n")

2. EasyOCR

EasyOCR无需额外安装Tesseract引擎,自带预训练模型,使用更简便,支持多语言识别。

依赖安装

pip install easyocr pdf2image

示例代码

import easyocr
from pdf2image import convert_from_path

# 初始化OCR阅读器,指定需识别的语言(中文+英文)
reader = easyocr.Reader(['ch_sim', 'en'])

pdf_pages = convert_from_path('autocad_output.pdf')

for idx, page_img in enumerate(pdf_pages):
    # 识别图像中的文本
    recognition_result = reader.readtext(page_img)
    # 提取并整理识别出的文本
    page_text = '\n'.join([item[1] for item in recognition_result])
    print(f"第{idx+1}页识别结果:\n{page_text}\n")

优化建议

AutoCAD生成的PDF常包含工程类字体或复杂排版,可通过以下方式提升识别准确率:

  • 确保使用对应语言的OCR模型(比如中文需加载chi_sim模型)
  • 对图像进行预处理(灰度化、二值化、降噪)
  • 调整OCR的参数(如指定识别区域、调整置信度阈值)

内容的提问来源于stack exchange,提问作者Saeed isa

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.10 12:42:48