如何通过OpenCV/OCR为图像中单个字符生成Bounding Box?求助
单个字符Bounding Box识别解决方案
针对EasyOCR的调整
- 将
batch_size设为1,强制模型逐字符处理 - 开启
detail=1,确保返回完整的字符级识别信息 - 把
y_ths和x_ths设为0,完全禁止字符/单词的合并 - 设置
min_size过滤过小的干扰区域,值根据字符实际大小调整
import easyocr as eo reader = eo.Reader(['en'], gpu=True) result = reader.readtext( imgOriginal, y_ths=0.0, x_ths=0.0, paragraph=False, batch_size=1, detail=1, min_size=10 # 可根据图像字符大小微调 ) # 提取单个字符的Bounding Box for detection in result: bbox = detection[0] char = detection[1] confidence = detection[2] print(f"字符: {char}, Bounding Box: {bbox}, 置信度: {confidence}")
针对Tesserocr/Pytesserocr的参数优化
必须使用字符级的PSM模式,推荐以下参数组合:
psm=10:单字符专属识别模式oem=3:混合引擎模式(兼顾准确率和兼容性)- 可选添加
tessedit_char_whitelist,限定识别字符范围,减少干扰
import pytesseract from PIL import Image img = Image.open(imgOriginal) custom_config = r'--oem 3 --psm 10 -c tessedit_char_whitelist=ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789' data = pytesseract.image_to_data(img, config=custom_config, output_type=pytesseract.Output.DICT) # 提取有效字符的Bounding Box for i in range(len(data['text'])): char = data['text'][i].strip() if char: bbox = (data['left'][i], data['top'][i], data['left'][i]+data['width'][i], data['top'][i]+data['height'][i]) confidence = data['conf'][i] print(f"字符: {char}, Bounding Box: {bbox}, 置信度: {confidence}")
预处理建议(提升识别效果)
如果图像存在模糊、倾斜或字符粘连,先做以下处理:
- 二值化:将图像转为黑白,增强字符与背景对比度
- 去噪:使用高斯模糊或中值滤波去除噪点
- 倾斜校正:通过霍夫变换等方法修正字符倾斜
内容的提问来源于stack exchange,提问作者Arnav
相关产品推荐
相关产品推荐

