如何用Python结合OpenCV提取低对比度细字体图片文本?
图片文本提取求助
需求说明
我想从一张图片中提取文本数据,理想输出为:
42 Gandalf le Gris 2,247/2,300 2,035/2,200 1,068/1,100 274,232/285,800 35 Gorbag 7/100 355/1,250 37,530/207,500
输出顺序无关紧要。
尝试过的处理方法
我用Python的OpenCV库做了多种尝试,但都没得到理想结果,具体代码如下:
- 提升图片对比度
lab= cv2.cvtColor(image, cv2.COLOR_BGR2LAB) l_channel, a, b = cv2.split(lab) clahe = cv2.createCLAHE(clipLimit=10.0, tileGridSize=(6,6)) cl = clahe.apply(l_channel) limg = cv2.merge((cl,a,b)) image = cv2.cvtColor(limg, cv2.COLOR_LAB2BGR)
- 遍历阈值做边缘检测
for a in range(1000): i +=3 image = cv2.Canny(image_1, 100 + i, 100 + i) data = pytesseract.image_to_string(image, lang='eng', config='--psm 6')
- 替换特定颜色为白色
for color in colors: rgb = color.split(',') image[np.all(image == (int(rgb[2]), int(rgb[1]), int(rgb[0])), axis=-1)] = (255, 255, 255)
- 灰度转换+反转处理
gray = cv2.cvtColor(image, cv2.COLOR_BGR2GRAY) blur = cv2.GaussianBlur(gray, (1,1), 0) thresh = cv2.threshold(blur, 0, 255, cv2.THRESH_BINARY_INV + cv2.THRESH_OTSU)[1] # Morph open to remove noise and invert image kernel = cv2.getStructuringElement(cv2.MORPH_RECT, (1,1)) opening = cv2.morphologyEx(thresh, cv2.MORPH_OPEN, kernel, iterations=0) invert = 255 - opening # Perform text extraction image = pytesseract.image_to_string(image, lang='eng', config='--psm 6') kernel = pytesseract.image_to_string(kernel, lang='eng', config='--psm 6') invert = pytesseract.image_to_string(invert, lang='eng', config='--psm 6')
核心问题
- 数字的边缘线条过细
- 数字颜色与背景颜色过于接近
请问这种情况是否可以实现文本提取?
内容的提问来源于stack exchange,提问作者JackBow
相关产品推荐
相关产品推荐

