You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PyTesseract OCR返回乱码,多种调试操作无效求助

pytesseract OCR识别返回乱码的排查与解决方法

使用以下Python代码进行OCR识别时返回乱码文本,已尝试调整PSM配置、修改/移除图像滤镜、通过pip更新pytesseract等操作,问题仍未解决:

pytesseract.pytesseract.tesseract_cmd = "C:\Program Files (x86)\Tesseract-OCR\tesseract.exe"
def extract_text(image):
    gray = image.convert('L')
    enhancer = ImageEnhance.Contrast(gray)
    enhanced_image = enhancer.enhance(2)
    enhanced_image = enhanced_image.filter(ImageFilter.GaussianBlur(radius=.3))
    enhanced_image.save("processed_image.png")
    text = pytesseract.image_to_string(enhanced_image, config='--psm 12')
    print(f'{text=}')
    return text

@interactions.slash_command("predict", description="Predict the location of an uploaded image")
@interactions.slash_option("image", "Upload an image", opt_type=interactions.OptionType.ATTACHMENT, required=True)
async def predict(ctx: interactions.SlashContext, image: interactions.Attachment):
    await ctx.defer()
    url = image.url
    image = Image.open(requests.get(url, stream=True).raw).convert("RGB")
    enhancer = ImageEnhance.Contrast(image)
    enhanced_image = enhancer.enhance(1.3)

    # Convert the enhanced image to a tensor and ensure values are in the range [0, 1]
    transform = transforms.Compose([
        transforms.ToTensor()
    ])
    image_tensor = transform(enhanced_image).unsqueeze(0).to(device)

    extracted_text = extract_text(enhanced_image)

识别返回的乱码示例:

text = "m ms

nu MvInIA6=Fmcem'x4Ir;u
q Dnvuvauou 53523

3 mm nyzenq795nx3I)1r>cmePmoe:.x
- 531463 my

1; I92I)xl2I)0,IMHz

"

排查与解决步骤

1. 确认Tesseract语言包安装

Tesseract默认仅支持英文,如果目标图像包含其他语言(如中文),必须安装对应语言包,并在识别时指定语言参数:

  • 下载对应语言包(如中文简体chi_sim)放入Tesseract的tessdata目录(通常是C:\Program Files (x86)\Tesseract-OCR\tessdata)
  • 修改image_to_string的config参数,添加语言指定:
    text = pytesseract.image_to_string(enhanced_image, config='--psm 6 -l chi_sim+eng')  # 同时识别中英
    

2. 优化图像预处理流程

当前的预处理可能过度增强对比度或模糊导致文字失真,建议调整:

  • 移除高斯模糊滤镜:模糊会弱化文字边缘,降低识别准确率
  • 调整对比度增强系数:将enhance(2)改为enhance(1.5)或更低,避免噪点被过度放大
  • 尝试二值化处理:将灰度图转为黑白二值,减少干扰
    修改后的extract_text示例:
    def extract_text(image):
        gray = image.convert('L')
        # 二值化处理(阈值可根据实际图像调整)
        threshold = 128
        binary_image = gray.point(lambda x: 0 if x < threshold else 255, '1')
        enhancer = ImageEnhance.Contrast(binary_image)
        enhanced_image = enhancer.enhance(1.5)
        enhanced_image.save("processed_image.png")
        # 切换更适合的PSM模式,比如--psm 6(假设文字为单一区块)
        text = pytesseract.image_to_string(enhanced_image, config='--psm 6 -l chi_sim+eng')
        print(f'{text=}')
        return text
    

3. 尝试匹配场景的PSM模式

不同PSM模式对应不同的文本布局,建议测试以下几种:

  • --psm 3:默认模式,自动检测文本布局
  • --psm 6:假设图像是单一统一的文本块
  • --psm 11:适合稀疏分布的文本
    避免盲目使用--psm 12(极端稀疏文本场景),除非你的图像确实是零散的文字碎片。

4. 验证图像质量与预处理结果

  • 打开保存的processed_image.png,检查预处理后的图像是否清晰:文字边缘是否锐利,有没有过多噪点或模糊
  • 确保上传的原始图像分辨率足够,没有被压缩失真(建议分辨率不低于300DPI)

5. 更新Tesseract-OCR本体

仅更新pytesseract Python包不够,需确保Tesseract-OCR核心程序是最新版本:

  • 前往Tesseract官网下载最新稳定版安装,覆盖旧版本
  • 确认tesseract_cmd指向的是最新安装路径

内容的提问来源于stack exchange,提问作者billy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.23 19:47:20