如何解决Tesseract无法识别LaTeX生成PDF页码的问题?
问题:LaTeX生成PDF的页码无法被pytesseract识别
我有一份LaTeX生成的PDF文档,想用pytesseract做OCR扫描,但所有页面的页码都识别不出来。PDF是通过Streamlit上传的,以类BytesIO对象传入处理函数,函数返回字符串组成的单词数组。
我试过不同的psm模式,没用;也做了图像预处理(转二值图、放大图像),还是没效果。
相关代码如下:
def get_text_from_ocr(uploaded_file): images = [] config = r"--psm 3" # 3: Fully automatic page segmentation, but no OSD. (Default) # pdf to images uploaded_file.seek(0) pdf_bytes = uploaded_file.read() doc = pymupdf.open(stream=pdf_bytes, filetype="pdf") for page in doc: pix = page.get_pixmap(dpi=300) img = Image.open(BytesIO(pix.tobytes("png"))) images.append(img) # Do OCR text = [word for img in images for word in pytesseract.image_to_string(img, config=config).split()] return text
def get_text_from_ocr(uploaded_file): images = [] config = r"--psm 3" # 3: Fully automatic page segmentation, but no OSD. (Default) # pdf to images uploaded_file.seek(0) pdf_bytes = uploaded_file.read() doc = pymupdf.open(stream=pdf_bytes, filetype="pdf") for page in doc: pix = page.get_pixmap(dpi=300) img = Image.open(BytesIO(pix.tobytes("png"))) # Preprocessing gray = img.convert("L") # "L" = 8-bit grayscale # Tune threshold value as needed (e.g., 180, 200) binary = gray.point(lambda x: 0 if x < 180 else 255, '1') # '1' mode = black & white scale = 2 resized = img.resize((img.width * scale, img.height * scale), Image.LANCZOS) images.append(resized) # Do OCR text = [word for img in images for word in pytesseract.image_to_string(img, config=config).split()] return text
页面截图:

解决思路
1. 单独处理页脚页码区域
LaTeX生成的页码通常固定在页脚,可裁剪该区域单独做OCR,减少其他内容干扰:
def get_text_from_ocr(uploaded_file): page_numbers = [] # 针对短文本块的psm模式 config = r"--oem 3 --psm 6 -l eng" uploaded_file.seek(0) pdf_bytes = uploaded_file.read() doc = pymupdf.open(stream=pdf_bytes, filetype="pdf") for page in doc: pix = page.get_pixmap(dpi=300) img = Image.open(BytesIO(pix.tobytes("png"))) # 裁剪页脚区域(根据实际页码位置调整,示例取底部10%) footer_height = int(img.height * 0.1) footer_img = img.crop((0, img.height - footer_height, img.width, img.height)) # 预处理增强对比度 gray = footer_img.convert("L") # 自动增强对比度 contrast_img = ImageOps.autocontrast(gray, cutoff=2) # 放大提升识别率 resized = contrast_img.resize((contrast_img.width * 2, contrast_img.height * 2), Image.LANCZOS) # 提取页码 number = pytesseract.image_to_string(resized, config=config).strip() page_numbers.append(number) return page_numbers
2. 优化Tesseract配置
- 使用
--psm 6:假设输入为单一文本块,适合页码这类短文本 - 添加
--oem 3:结合LSTM和传统引擎,提升兼容性 - 指定语言参数
-l eng(如果是英文页码)
示例配置:
config = r"--oem 3 --psm 6 -l eng"
3. 改进图像预处理
LaTeX页码颜色可能偏浅,需针对性调整预处理步骤:
# 替换原有预处理代码 gray = img.convert("L") # 中值滤波降噪 gray_filtered = gray.filter(ImageFilter.MedianFilter(size=3)) # 自适应阈值二值化,增强对比度 binary = gray_filtered.point(lambda x: 0 if x < 210 else 255, '1')
4. 提升PDF转图片质量
尝试更高dpi(如600)生成图片,确保页码细节清晰:
pix = page.get_pixmap(dpi=600)
内容的提问来源于stack exchange,提问作者mike3467
相关产品推荐
相关产品推荐

