You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何提升发票图片的文本检索与数据提取准确率

如何提升发票图片的文本检索与数据提取准确率

原代码

import cv2
import pytesseract
import numpy as np

image_path = "elecBill.jpg"

img = cv2.imread(image_path)

# Resize (VERY IMPORTANT)
img = cv2.resize(img, None, fx=2, fy=2, interpolation=cv2.INTER_CUBIC)

# Convert to grayscale
gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)

# Increase contrast using CLAHE
clahe = cv2.createCLAHE(clipLimit=3.0, tileGridSize=(8,8))
gray = clahe.apply(gray)

# Denoise
gray = cv2.bilateralFilter(gray, 9, 75, 75)

# Strong adaptive threshold
thresh = cv2.adaptiveThreshold(
    gray,
    255,
    cv2.ADAPTIVE_THRESH_GAUSSIAN_C,
    cv2.THRESH_BINARY,
    31,
    10
)

# OCR with better config for receipts
custom_config = r'--oem 3 --psm 4'

text = pytesseract.image_to_string(thresh, config=custom_config)

import re
def extract_data(text):
    data = {}

    clean_text = text.replace('\n', ' ').replace('\r', ' ')

    # Extract all dates
    dates = re.findall(r'\d{2}[/-]\d{2}[/-]\d{2,4}', clean_text)
    if dates:
        data["dates_found"] = dates

    # Extract all decimal numbers
    numbers = re.findall(r'\d+\.\d+', clean_text)

    if numbers:
        # Convert to float for comparison
        float_numbers = [float(n) for n in numbers]

        # Largest number often = total payable
        largest = max(float_numbers)
        data["largest_amount_guess"] = largest

    # Keyword based search
    keyword_patterns = {
        "bill_amount": r'(Bill|Total|Net).{0,20}?(\d+\.\d+)',
        "units": r'Units.{0,10}?(\d+)'
    }

    for key, pattern in keyword_patterns.items():
        match = re.search(pattern, clean_text, re.IGNORECASE)
        if match:
            data[key] = match.group(match.lastindex)

    return data

extracted = extract_data(text)

print("\n------ EXTRACTED DATA ------")
print(extracted)

提问者问题

Hey guys here is the complete code. I'm using this program to extract the text from the invoices. And I would like to improve the accuracy to get precise text. How can i do so. Most of the times I get somewhat gibberish,etc.

Any help would be beneficial. And this is my complete code for the program i had made.


提升准确率的实用建议

一、图像预处理再优化

图像是OCR的基础,预处理不到位直接导致识别乱码:

  • 调整阈值参数:你当前用的自适应阈值blockSize=31和C=10强度过高,容易把字符的笔画“磨掉”。可以尝试把blockSize改成11、17这类更小的奇数,C值降到5左右,保留更多文本细节。
  • 添加形态学修复:阈值处理后,用形态学运算修复字符的缺口或消除小噪声:
    import cv2
    # 用小核进行闭运算填补字符缺口,开运算消除噪声点
    kernel = np.ones((2,2), np.uint8)
    thresh = cv2.morphologyEx(thresh, cv2.MORPH_CLOSE, kernel, iterations=1)
    thresh = cv2.morphologyEx(thresh, cv2.MORPH_OPEN, kernel, iterations=1)
    
  • 强制倾斜校正:拍摄的发票几乎都有倾斜,Tesseract对倾斜文本识别率极低。可以用轮廓检测自动校正:
    # 检测文本区域的倾斜角度
    coords = np.column_stack(np.where(thresh > 0))
    angle = cv2.minAreaRect(coords)[-1]
    # 调整角度范围
    if angle < -45:
        angle = -(90 + angle)
    else:
        angle = -angle
    # 旋转图像校正
    (h, w) = thresh.shape[:2]
    center = (w // 2, h // 2)
    M = cv2.getRotationMatrix2D(center, angle, 1.0)
    rotated = cv2.warpAffine(thresh, M, (w, h), flags=cv2.INTER_CUBIC, borderMode=cv2.BORDER_REPLICATE)
    thresh = rotated
    

二、Tesseract配置精准调优

Tesseract的参数对识别结果影响极大,别只固定用一套配置:

  • 调整PSM模式:你用的--psm 4适合单列文本,但发票布局复杂,建议测试--psm 6(假设文本是统一块)或--psm 11(保留布局的灵活识别),不同发票适配的模式不一样。
  • 切换OEM模式:如果你的Tesseract是4.0+版本,试试--oem 1(纯LSTM模型),有时候纯LSTM对打印体文本的识别比混合模式更稳定。
  • 训练自定义字库:如果你的发票有固定字体或特殊字符(比如行业专用符号),可以用Tesseract的训练工具生成专属字库,或者找现成的票据类预训练模型,能大幅提升特定场景的识别率。

三、数据提取规则细化

即使OCR有小错误,通过规则优化也能拿到正确数据:

  • 更细致的文本清洗:除了替换换行,还要过滤非打印字符,避免乱码干扰正则:
    import string
    printable = set(string.printable)
    clean_text = ''.join(filter(lambda x: x in printable, clean_text))
    
  • 正则规则贴合发票逻辑:你当前的正则太宽泛,比如金额提取可以结合发票的逻辑(总金额通常和“Total”“合计”“应付”绑定,且带货币符号):
    # 针对人民币发票的总金额提取
    total_pattern = r'(?:Total|合计|应付金额|实付金额)[^\d]*(\d+\.\d+)'
    match = re.search(total_pattern, clean_text, re.IGNORECASE)
    
    日期提取也可以区分开票日期(通常是年-月-日格式)和到期日,用更精准的正则:r'\d{4}[/-]\d{2}[/-]\d{2}'。
  • 多轮逻辑验证:比如提取到的总金额应该是所有数字里最大的,且和小计、税额有逻辑关系(总金额=小计+税额),通过这种验证可以排除OCR识别错误的数值。

四、其他小技巧

  • 裁剪无关区域:如果发票有大面积空白、logo或广告,先自动/手动裁剪出文本区域,减少Tesseract的识别干扰。
  • 优化图像来源:拍摄时保证光线均匀、无阴影,对焦清晰;扫描件用300DPI以上的分辨率,识别效果会好很多。

备注:内容来源于stack exchange,提问作者Sidharth Kumar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.13 18:04:35