You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

褪色倾斜收据图像预处理:文本清晰度提升难题

收据图像清晰度提升与OCR优化方案

针对褪色收据预处理后OCR准确率不足的问题,可通过以下流程优化,重点解决对比度低、文本模糊的核心问题:

核心优化步骤

1. 局部对比度增强

褪色图像的文本与背景对比度极低,用CLAHE(限制对比度自适应直方图均衡)替代简单反色,精准提升局部文本清晰度:

import cv2
import pytesseract
import numpy as np

img_path = "rec.jpg"
img = cv2.imread(img_path)

# 转灰度+CLAHE增强
gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)
clahe = cv2.createCLAHE(clipLimit=2.0, tileGridSize=(8,8))
enhanced_gray = clahe.apply(gray)

2. 自适应阈值二值化

针对光照不均的褪色图像,自适应阈值比全局OTSU阈值更能保留文本细节:

# 高斯自适应阈值,反转前景/背景
thresh = cv2.adaptiveThreshold(enhanced_gray, 255, cv2.ADAPTIVE_THRESH_GAUSSIAN_C, 
                               cv2.THRESH_BINARY_INV, 11, 2)

3. 去噪+精细化形态学操作

先去噪再修复文本边缘,避免噪声被放大导致文本粘连:

# 中值滤波去噪
denoised = cv2.medianBlur(thresh, 3)
# 小核腐蚀+膨胀,修复文本断裂
kernel = np.ones((1,1), np.uint8)
eroded = cv2.erode(denoised, kernel, iterations=1)
processed = cv2.dilate(eroded, kernel, iterations=1)

4. 基于增强图的精准纠偏

保留原有纠偏逻辑,但改用增强后的二值图计算角度,提升纠偏准确性:

coords = np.column_stack(np.where(processed > 0))
angle = cv2.minAreaRect(coords)[-1]

if angle < -45:
    angle = -(90 + angle)
else:
    angle = -angle

(h, w) = img.shape[:2]
center = (w // 2, h // 2)
M = cv2.getRotationMatrix2D(center, angle, 1.0)
rotated = cv2.warpAffine(processed, M, (w, h), flags=cv2.INTER_CUBIC, borderMode=cv2.BORDER_REPLICATE)

5. Tesseract参数针对性优化

针对收据类文本,限制识别范围并调整识别模式:

# 配置:启用默认引擎,单块文本识别,只识别字母数字
custom_config = r'--oem 3 --psm 6 -c tessedit_char_whitelist=0123456789ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz'
text = pytesseract.image_to_string(rotated, config=custom_config)
print("识别结果:\n", text)

完整优化代码

import cv2
import pytesseract
import numpy as np

img_path = "rec.jpg"
img = cv2.imread(img_path)

# 1. 灰度转换与对比度增强
gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)
clahe = cv2.createCLAHE(clipLimit=2.0, tileGridSize=(8,8))
enhanced_gray = clahe.apply(gray)

# 2. 自适应阈值二值化
thresh = cv2.adaptiveThreshold(enhanced_gray, 255, cv2.ADAPTIVE_THRESH_GAUSSIAN_C, 
                               cv2.THRESH_BINARY_INV, 11, 2)

# 3. 去噪与形态学修复
denoised = cv2.medianBlur(thresh, 3)
kernel = np.ones((1,1), np.uint8)
eroded = cv2.erode(denoised, kernel, iterations=1)
processed = cv2.dilate(eroded, kernel, iterations=1)

# 4. 图像纠偏
coords = np.column_stack(np.where(processed > 0))
angle = cv2.minAreaRect(coords)[-1]

if angle < -45:
    angle = -(90 + angle)
else:
    angle = -angle

(h, w) = img.shape[:2]
center = (w // 2, h // 2)
M = cv2.getRotationMatrix2D(center, angle, 1.0)
rotated = cv2.warpAffine(processed, M, (w, h), flags=cv2.INTER_CUBIC, borderMode=cv2.BORDER_REPLICATE)

# 5. OCR识别
custom_config = r'--oem 3 --psm 6 -c tessedit_char_whitelist=0123456789ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz'
text = pytesseract.image_to_string(rotated, config=custom_config)

# 显示与输出
cv2.imshow("Enhanced", enhanced_gray)
cv2.imshow("Final Processed", rotated)
print("识别结果:\n", text)
cv2.waitKey(0)
cv2.destroyAllWindows()

关键说明

  • CLAHE增强可针对性提升局部文本对比度,解决褪色导致的文本暗淡问题;
  • 自适应阈值能适配光照不均区域,避免部分文本丢失;
  • 先去噪再做形态学操作,可避免噪声被放大导致的文本粘连;
  • Tesseract的字符白名单和PSM模式设置,能过滤无关干扰,聚焦收据类文本识别。

内容的提问来源于stack exchange,提问作者Furkan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.10 23:05:20