PyTesseract OCR返回乱码,多种调试操作无效求助
pytesseract OCR识别返回乱码的排查与解决方法
使用以下Python代码进行OCR识别时返回乱码文本,已尝试调整PSM配置、修改/移除图像滤镜、通过pip更新pytesseract等操作,问题仍未解决:
pytesseract.pytesseract.tesseract_cmd = "C:\Program Files (x86)\Tesseract-OCR\tesseract.exe" def extract_text(image): gray = image.convert('L') enhancer = ImageEnhance.Contrast(gray) enhanced_image = enhancer.enhance(2) enhanced_image = enhanced_image.filter(ImageFilter.GaussianBlur(radius=.3)) enhanced_image.save("processed_image.png") text = pytesseract.image_to_string(enhanced_image, config='--psm 12') print(f'{text=}') return text @interactions.slash_command("predict", description="Predict the location of an uploaded image") @interactions.slash_option("image", "Upload an image", opt_type=interactions.OptionType.ATTACHMENT, required=True) async def predict(ctx: interactions.SlashContext, image: interactions.Attachment): await ctx.defer() url = image.url image = Image.open(requests.get(url, stream=True).raw).convert("RGB") enhancer = ImageEnhance.Contrast(image) enhanced_image = enhancer.enhance(1.3) # Convert the enhanced image to a tensor and ensure values are in the range [0, 1] transform = transforms.Compose([ transforms.ToTensor() ]) image_tensor = transform(enhanced_image).unsqueeze(0).to(device) extracted_text = extract_text(enhanced_image)识别返回的乱码示例:
text = "m ms nu MvInIA6=Fmcem'x4Ir;u q Dnvuvauou 53523 3 mm nyzenq795nx3I)1r>cmePmoe:.x - 531463 my 1; I92I)xl2I)0,IMHz "
排查与解决步骤
1. 确认Tesseract语言包安装
Tesseract默认仅支持英文,如果目标图像包含其他语言(如中文),必须安装对应语言包,并在识别时指定语言参数:
- 下载对应语言包(如中文简体
chi_sim)放入Tesseract的tessdata目录(通常是C:\Program Files (x86)\Tesseract-OCR\tessdata) - 修改
image_to_string的config参数,添加语言指定:text = pytesseract.image_to_string(enhanced_image, config='--psm 6 -l chi_sim+eng') # 同时识别中英
2. 优化图像预处理流程
当前的预处理可能过度增强对比度或模糊导致文字失真,建议调整:
- 移除高斯模糊滤镜:模糊会弱化文字边缘,降低识别准确率
- 调整对比度增强系数:将
enhance(2)改为enhance(1.5)或更低,避免噪点被过度放大 - 尝试二值化处理:将灰度图转为黑白二值,减少干扰
修改后的extract_text示例:def extract_text(image): gray = image.convert('L') # 二值化处理(阈值可根据实际图像调整) threshold = 128 binary_image = gray.point(lambda x: 0 if x < threshold else 255, '1') enhancer = ImageEnhance.Contrast(binary_image) enhanced_image = enhancer.enhance(1.5) enhanced_image.save("processed_image.png") # 切换更适合的PSM模式,比如--psm 6(假设文字为单一区块) text = pytesseract.image_to_string(enhanced_image, config='--psm 6 -l chi_sim+eng') print(f'{text=}') return text
3. 尝试匹配场景的PSM模式
不同PSM模式对应不同的文本布局,建议测试以下几种:
--psm 3:默认模式,自动检测文本布局--psm 6:假设图像是单一统一的文本块--psm 11:适合稀疏分布的文本
避免盲目使用--psm 12(极端稀疏文本场景),除非你的图像确实是零散的文字碎片。
4. 验证图像质量与预处理结果
- 打开保存的
processed_image.png,检查预处理后的图像是否清晰:文字边缘是否锐利,有没有过多噪点或模糊 - 确保上传的原始图像分辨率足够,没有被压缩失真(建议分辨率不低于300DPI)
5. 更新Tesseract-OCR本体
仅更新pytesseract Python包不够,需确保Tesseract-OCR核心程序是最新版本:
- 前往Tesseract官网下载最新稳定版安装,覆盖旧版本
- 确认
tesseract_cmd指向的是最新安装路径
内容的提问来源于stack exchange,提问作者billy
相关产品推荐
相关产品推荐

