如何提升发票图片的文本检索与数据提取准确率
如何提升发票图片的文本检索与数据提取准确率
原代码
import cv2 import pytesseract import numpy as np image_path = "elecBill.jpg" img = cv2.imread(image_path) # Resize (VERY IMPORTANT) img = cv2.resize(img, None, fx=2, fy=2, interpolation=cv2.INTER_CUBIC) # Convert to grayscale gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY) # Increase contrast using CLAHE clahe = cv2.createCLAHE(clipLimit=3.0, tileGridSize=(8,8)) gray = clahe.apply(gray) # Denoise gray = cv2.bilateralFilter(gray, 9, 75, 75) # Strong adaptive threshold thresh = cv2.adaptiveThreshold( gray, 255, cv2.ADAPTIVE_THRESH_GAUSSIAN_C, cv2.THRESH_BINARY, 31, 10 ) # OCR with better config for receipts custom_config = r'--oem 3 --psm 4' text = pytesseract.image_to_string(thresh, config=custom_config) import re def extract_data(text): data = {} clean_text = text.replace('\n', ' ').replace('\r', ' ') # Extract all dates dates = re.findall(r'\d{2}[/-]\d{2}[/-]\d{2,4}', clean_text) if dates: data["dates_found"] = dates # Extract all decimal numbers numbers = re.findall(r'\d+\.\d+', clean_text) if numbers: # Convert to float for comparison float_numbers = [float(n) for n in numbers] # Largest number often = total payable largest = max(float_numbers) data["largest_amount_guess"] = largest # Keyword based search keyword_patterns = { "bill_amount": r'(Bill|Total|Net).{0,20}?(\d+\.\d+)', "units": r'Units.{0,10}?(\d+)' } for key, pattern in keyword_patterns.items(): match = re.search(pattern, clean_text, re.IGNORECASE) if match: data[key] = match.group(match.lastindex) return data extracted = extract_data(text) print("\n------ EXTRACTED DATA ------") print(extracted)
提问者问题
Hey guys here is the complete code. I'm using this program to extract the text from the invoices. And I would like to improve the accuracy to get precise text. How can i do so. Most of the times I get somewhat gibberish,etc.
Any help would be beneficial. And this is my complete code for the program i had made.
提升准确率的实用建议
一、图像预处理再优化
图像是OCR的基础,预处理不到位直接导致识别乱码:
- 调整阈值参数:你当前用的自适应阈值
blockSize=31和C=10强度过高,容易把字符的笔画“磨掉”。可以尝试把blockSize改成11、17这类更小的奇数,C值降到5左右,保留更多文本细节。 - 添加形态学修复:阈值处理后,用形态学运算修复字符的缺口或消除小噪声:
import cv2 # 用小核进行闭运算填补字符缺口,开运算消除噪声点 kernel = np.ones((2,2), np.uint8) thresh = cv2.morphologyEx(thresh, cv2.MORPH_CLOSE, kernel, iterations=1) thresh = cv2.morphologyEx(thresh, cv2.MORPH_OPEN, kernel, iterations=1) - 强制倾斜校正:拍摄的发票几乎都有倾斜,Tesseract对倾斜文本识别率极低。可以用轮廓检测自动校正:
# 检测文本区域的倾斜角度 coords = np.column_stack(np.where(thresh > 0)) angle = cv2.minAreaRect(coords)[-1] # 调整角度范围 if angle < -45: angle = -(90 + angle) else: angle = -angle # 旋转图像校正 (h, w) = thresh.shape[:2] center = (w // 2, h // 2) M = cv2.getRotationMatrix2D(center, angle, 1.0) rotated = cv2.warpAffine(thresh, M, (w, h), flags=cv2.INTER_CUBIC, borderMode=cv2.BORDER_REPLICATE) thresh = rotated
二、Tesseract配置精准调优
Tesseract的参数对识别结果影响极大,别只固定用一套配置:
- 调整PSM模式:你用的
--psm 4适合单列文本,但发票布局复杂,建议测试--psm 6(假设文本是统一块)或--psm 11(保留布局的灵活识别),不同发票适配的模式不一样。 - 切换OEM模式:如果你的Tesseract是4.0+版本,试试
--oem 1(纯LSTM模型),有时候纯LSTM对打印体文本的识别比混合模式更稳定。 - 训练自定义字库:如果你的发票有固定字体或特殊字符(比如行业专用符号),可以用Tesseract的训练工具生成专属字库,或者找现成的票据类预训练模型,能大幅提升特定场景的识别率。
三、数据提取规则细化
即使OCR有小错误,通过规则优化也能拿到正确数据:
- 更细致的文本清洗:除了替换换行,还要过滤非打印字符,避免乱码干扰正则:
import string printable = set(string.printable) clean_text = ''.join(filter(lambda x: x in printable, clean_text)) - 正则规则贴合发票逻辑:你当前的正则太宽泛,比如金额提取可以结合发票的逻辑(总金额通常和“Total”“合计”“应付”绑定,且带货币符号):
日期提取也可以区分开票日期(通常是年-月-日格式)和到期日,用更精准的正则:# 针对人民币发票的总金额提取 total_pattern = r'(?:Total|合计|应付金额|实付金额)[^\d]*(\d+\.\d+)' match = re.search(total_pattern, clean_text, re.IGNORECASE)r'\d{4}[/-]\d{2}[/-]\d{2}'。 - 多轮逻辑验证:比如提取到的总金额应该是所有数字里最大的,且和小计、税额有逻辑关系(总金额=小计+税额),通过这种验证可以排除OCR识别错误的数值。
四、其他小技巧
- 裁剪无关区域:如果发票有大面积空白、logo或广告,先自动/手动裁剪出文本区域,减少Tesseract的识别干扰。
- 优化图像来源:拍摄时保证光线均匀、无阴影,对焦清晰;扫描件用300DPI以上的分辨率,识别效果会好很多。
备注:内容来源于stack exchange,提问作者Sidharth Kumar
相关产品推荐
相关产品推荐

