如何在Tesseract OCR中去除图像内特定高度的文字?
解决方案:去除特定高度无效文字以提升Tesseract OCR准确率
方法一:OpenCV预处理去除指定高度区域
直接定位无效文字所在的高度范围,用背景色填充后再进行OCR识别,适合无效文字位置固定的场景:
import cv2 import pytesseract # 读取图像 img = cv2.imread("target_image.png") # 转为灰度图(简化处理) gray_img = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY) # 自定义要去除的区域高度(根据你的图像调整,比如顶部60像素为无效区) invalid_height = 60 # 用白色填充无效区域(如果图像背景不是白色,替换为对应背景色的灰度值) gray_img[:invalid_height, :] = 255 # 保存处理后的图像 cv2.imwrite("cleaned_image.png", gray_img) # 执行OCR识别 ocr_result = pytesseract.image_to_string("cleaned_image.png", lang="chi_sim") print(ocr_result)
方法二:裁剪有效区域后用Tesseract识别
无需修改原图像,直接裁剪掉包含无效文字的区域,只识别有效部分:
from PIL import Image import pytesseract # 打开图像 img = Image.open("target_image.png") # 定义有效区域坐标:(左, 上, 右, 下),跳过顶部无效区 # 示例中从y=60像素开始向下识别,可根据实际调整 valid_region = (0, 60, img.width, img.height) # 裁剪有效区域 cropped_img = img.crop(valid_region) # 执行OCR识别 ocr_result = pytesseract.image_to_string(cropped_img, lang="chi_sim") print(ocr_result)
进阶优化:自动定位无效文字行
如果无效文字高度不固定但有统一特征,可通过OpenCV轮廓检测自动定位并去除:
import cv2 import numpy as np img = cv2.imread("target_image.png") gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY) # 二值化处理(突出文字) _, thresh = cv2.threshold(gray, 127, 255, cv2.THRESH_BINARY_INV) # 检测文字轮廓 contours, _ = cv2.findContours(thresh, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE) # 遍历轮廓,过滤掉高度符合无效文字的区域 invalid_line_height = 30 # 自定义无效文字的高度阈值 for cnt in contours: x, y, w, h = cv2.boundingRect(cnt) if h <= invalid_line_height: # 用白色填充该轮廓区域 cv2.rectangle(gray, (x, y), (x+w, y+h), (255, 255, 255), -1) # 保存处理后的图像并识别 cv2.imwrite("auto_cleaned.png", gray)
内容的提问来源于stack exchange,提问作者VigneshKumar Selvaraj
相关产品推荐
相关产品推荐

