You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何解决Tesseract无法识别LaTeX生成PDF页码的问题?

问题:LaTeX生成PDF的页码无法被pytesseract识别

我有一份LaTeX生成的PDF文档,想用pytesseract做OCR扫描,但所有页面的页码都识别不出来。PDF是通过Streamlit上传的,以类BytesIO对象传入处理函数,函数返回字符串组成的单词数组。

我试过不同的psm模式,没用;也做了图像预处理(转二值图、放大图像),还是没效果。

相关代码如下:

def get_text_from_ocr(uploaded_file):
    images = []
    config = r"--psm 3"  # 3: Fully automatic page segmentation, but no OSD. (Default)

    # pdf to images
    uploaded_file.seek(0)
    pdf_bytes = uploaded_file.read()
    doc = pymupdf.open(stream=pdf_bytes, filetype="pdf")

    for page in doc:
        pix = page.get_pixmap(dpi=300)
        img = Image.open(BytesIO(pix.tobytes("png")))
        images.append(img)

    # Do OCR
    text = [word for img in images for word in pytesseract.image_to_string(img, config=config).split()]
    return text
def get_text_from_ocr(uploaded_file):
    images = []
    config = r"--psm 3"  # 3: Fully automatic page segmentation, but no OSD. (Default)

    # pdf to images
    uploaded_file.seek(0)
    pdf_bytes = uploaded_file.read()
    doc = pymupdf.open(stream=pdf_bytes, filetype="pdf")

    for page in doc:
        pix = page.get_pixmap(dpi=300)
        img = Image.open(BytesIO(pix.tobytes("png")))
     
        # Preprocessing  
        gray = img.convert("L")  # "L" = 8-bit grayscale
        # Tune threshold value as needed (e.g., 180, 200)
        binary = gray.point(lambda x: 0 if x < 180 else 255, '1')  # '1' mode = black & white
        scale = 2 
        resized = img.resize((img.width * scale, img.height * scale), Image.LANCZOS)

        images.append(resized)

    # Do OCR
    text = [word for img in images for word in pytesseract.image_to_string(img, config=config).split()]
    return text

页面截图:
页面1截图
页面2截图


解决思路

1. 单独处理页脚页码区域

LaTeX生成的页码通常固定在页脚,可裁剪该区域单独做OCR,减少其他内容干扰:

def get_text_from_ocr(uploaded_file):
    page_numbers = []
    # 针对短文本块的psm模式
    config = r"--oem 3 --psm 6 -l eng"

    uploaded_file.seek(0)
    pdf_bytes = uploaded_file.read()
    doc = pymupdf.open(stream=pdf_bytes, filetype="pdf")

    for page in doc:
        pix = page.get_pixmap(dpi=300)
        img = Image.open(BytesIO(pix.tobytes("png")))
        
        # 裁剪页脚区域(根据实际页码位置调整,示例取底部10%)
        footer_height = int(img.height * 0.1)
        footer_img = img.crop((0, img.height - footer_height, img.width, img.height))
        
        # 预处理增强对比度
        gray = footer_img.convert("L")
        # 自动增强对比度
        contrast_img = ImageOps.autocontrast(gray, cutoff=2)
        # 放大提升识别率
        resized = contrast_img.resize((contrast_img.width * 2, contrast_img.height * 2), Image.LANCZOS)
        
        # 提取页码
        number = pytesseract.image_to_string(resized, config=config).strip()
        page_numbers.append(number)

    return page_numbers

2. 优化Tesseract配置

  • 使用--psm 6:假设输入为单一文本块,适合页码这类短文本
  • 添加--oem 3:结合LSTM和传统引擎,提升兼容性
  • 指定语言参数-l eng(如果是英文页码)

示例配置:

config = r"--oem 3 --psm 6 -l eng"

3. 改进图像预处理

LaTeX页码颜色可能偏浅,需针对性调整预处理步骤:

# 替换原有预处理代码
gray = img.convert("L")
# 中值滤波降噪
gray_filtered = gray.filter(ImageFilter.MedianFilter(size=3))
# 自适应阈值二值化,增强对比度
binary = gray_filtered.point(lambda x: 0 if x < 210 else 255, '1')

4. 提升PDF转图片质量

尝试更高dpi(如600)生成图片,确保页码细节清晰:

pix = page.get_pixmap(dpi=600)

内容的提问来源于stack exchange,提问作者mike3467

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.13 01:35:19