如何结合Tesseract与OpenCV提取指定bounding box内的OCR文本
解决方案:基于OpenCV+Tesseract的KYC文本结构化提取
一、前置准备
- 安装Tesseract OCR引擎:Linux用
sudo apt install tesseract-ocr,Windows直接下载安装包配置环境变量 - 安装Python依赖:
pip install pytesseract opencv-python
二、核心实现步骤
1. 从Bounding Box提取单区域文本
先遍历OpenCV检测到的所有bbox,裁剪对应区域并通过Tesseract提取文本,建议加预处理提升识别准确率:
import cv2 import pytesseract # 假设已加载KYC图像为img,bbox列表为bboxes(每个元素格式:(x, y, w, h)) extracted_texts = [] for bbox in bboxes: x, y, w, h = bbox # 裁剪ROI区域 roi = img[y:y+h, x:x+w] # 预处理:灰度化+二值化 gray_roi = cv2.cvtColor(roi, cv2.COLOR_BGR2GRAY) _, thresh_roi = cv2.threshold(gray_roi, 0, 255, cv2.THRESH_BINARY + cv2.THRESH_OTSU) # Tesseract提取文本,psm7指定单文本行识别 text = pytesseract.image_to_string(thresh_roi, lang='eng', config='--psm 7').strip() extracted_texts.append((bbox, text))
2. 拼接分散字符为完整姓名
姓名的字符/单词bbox通常处于同一行或垂直相邻区域,通过位置关系分组拼接:
# 先筛选出姓名区域的bbox(可根据KYC布局规则判断,比如靠近"Name"标签) name_bboxes = [item for item in extracted_texts if is_name_region(item[0])] # is_name_region为自定义筛选函数 # 按行分组(y坐标差值小于10像素视为同一行) name_groups = {} for bbox, text in name_bboxes: x, y, w, h = bbox center_y = y + h//2 # 匹配已有分组 matched = False for key in name_groups: if abs(center_y - key) < 10: name_groups[key].append((x, text)) matched = True break if not matched: name_groups[center_y] = [(x, text)] # 同一行按x坐标排序后拼接 full_name = "" for row_y in sorted(name_groups.keys()): row_items = sorted(name_groups[row_y], key=lambda x: x[0]) full_name += " ".join([txt for _, txt in row_items]) + " " full_name = full_name.strip()
3. 提取PAN号
根据PAN号的固定格式(比如12位纯数字,或印度PAN的5字母+4数字+1字母),用正则匹配筛选:
import re # 示例:匹配12位纯数字PAN号 pan_pattern = r'^\d{12}$' pan_number = None for _, text in extracted_texts: if re.fullmatch(pan_pattern, text): pan_number = text break # 如果是字母数字组合PAN,调整正则:比如印度PAN用r'^[A-Z]{5}\d{4}[A-Z]{1}$'
4. 生成结构化结果
将提取内容整理为结构化字典:
structured_result = { "full_name": full_name, "pan_number": pan_number } print(structured_result)
三、优化技巧
- 预处理:对ROI做高斯降噪(
cv2.GaussianBlur)、边缘增强,进一步提升识别率 - bbox筛选:利用KYC文档的固定布局,提前锁定姓名、PAN号的大致区域,减少无效识别
- Tesseract配置:根据文本类型调整
--psm参数,比如识别单个字符用--psm 8,识别整块文本用--psm 6
内容的提问来源于stack exchange,提问作者Levin Jose
相关产品推荐
相关产品推荐

