如何使用Tesseract的--psm 2模式仅实现页面分割(无需OCR)?
解决方案
一、正确用Tesseract仅执行页面分割(无OCR)
Tesseract的--psm 2确实对应「自动页面分割,无OSD或OCR」,但默认文本输出格式因无OCR结果不会生成有效内容,你需要选择能保存布局边界框(BBox)的输出格式,比如hocr或调整PSM模式获取更细粒度的分割信息:
命令行方式
- 生成HOCR格式(包含所有分割区域的坐标):
tesseract img.png outfile --psm 2 hocr
生成的outfile.hocr是HTML文件,其中<div class="ocr_page">、<div class="ocr_carea">等标签的title属性包含边界框数据(格式如bbox 0 0 100 200),解析这些属性即可提取布局坐标。
- 获取更细粒度分割(块、行):
若需要段落、行级别的布局,可改用--psm 3(默认自动页面分割,无OSD)搭配HOCR输出,实际测试中该模式在不执行OCR的情况下也能输出更详细的布局:
tesseract img.png outfile --psm 3 hocr
二、Python中实现仅页面分割
方法1:调用Tesseract命令行解析HOCR
用subprocess调用Tesseract生成HOCR,再通过BeautifulSoup解析边界框:
import subprocess from bs4 import BeautifulSoup def get_tesseract_layout(img_path): # 调用Tesseract生成HOCR文件 subprocess.run([ "tesseract", img_path, "temp_layout", "--psm", "3", "hocr" ], check=True) # 解析HOCR提取布局坐标 with open("temp_layout.hocr", "r") as f: soup = BeautifulSoup(f.read(), "html.parser") layout_data = [] # 提取块级区域(ocr_carea) for area in soup.find_all("div", class_="ocr_carea"): bbox_str = area["title"].split("bbox ")[1].split(";")[0] x1, y1, x2, y2 = map(int, bbox_str.split()) layout_data.append({ "type": "block", "bbox": (x1, y1, x2, y2) }) # 可按需提取行(ocr_line)或单词级区域 return layout_data # 使用示例 layout = get_tesseract_layout("img.png")
方法2:优化layoutparser调用,强制禁用OCR
layoutparser的TesseractAgent默认会执行OCR,可通过传递Tesseract底层参数强制关闭OCR,仅保留页面分割:
import layoutparser as lp ocr_agent = lp.TesseractAgent(languages='eng') res = ocr_agent.detect( img_path, return_response=True, tessedit_do_ocr=0, # 核心参数:禁用OCR psm=3 ) layout_info = res['data']
三、加快页面分割/OCR速度的技巧
- 缩小图片尺寸:将图片缩至长边不超过1000像素,减少处理像素量:
from PIL import Image img = Image.open("img.png") max_dim = 1000 width, height = img.size if max(width, height) > max_dim: scale = max_dim / max(width, height) img = img.resize((int(width*scale), int(height*scale)), Image.Resampling.LANCZOS) img.save("resized_img.png")
- 启用多线程:Tesseract支持多线程处理,添加
--threads参数提升速度:
# 命令行示例 tesseract img.png outfile --psm 3 hocr --threads 4
# pytesseract示例 import pytesseract pytesseract.image_to_data(img, config='--psm 3 --threads 4')
- 图片预处理:对图片做灰度化、二值化、去噪,减少干扰:
from PIL import Image, ImageOps img = Image.open("img.png").convert("L") # 转灰度图 img = ImageOps.autocontrast(img) # 自动提升对比度 img = img.point(lambda x: 0 if x < 128 else 255, '1') # 二值化
- 切换轻量OCR引擎:若必须执行OCR,可指定
--oem 0使用传统引擎,比LSTM引擎(--oem 1)更快:
tesseract img.png outfile --oem 0 --psm 3
内容的提问来源于stack exchange,提问作者Vera Bernhard
相关产品推荐
相关产品推荐

