如何用Python+OpenCV实现多行文本图像的行、词、字符分割?
基于OpenCV的多行文本三级分割解决方案
一、整体流程概述
需完成三级文本分割+识别流程:
- 多行文本图像 → 单行文本图像
- 单行文本图像 → 单个单词图像
- 单词图像 → 字符图像+文本识别(基于你已实现的逻辑微调适配)
二、分步实现代码
1. 多行转单行(行分割)
通过横向像素投影定位文本行的上下边界,利用文本行区域像素集中的特性分割出独立行:
import cv2 import numpy as np import matplotlib.pyplot as plt def split_image_to_lines(img): # 转灰度图 gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY) # 二值化(黑底白字,便于投影分析) _, thresh = cv2.threshold(gray, 0, 255, cv2.THRESH_BINARY_INV + cv2.THRESH_OTSU) # 计算横向像素投影:每行的像素总和 row_proj = np.sum(thresh, axis=1) # 筛选出有文本的行索引 non_zero_rows = np.where(row_proj > 0)[0] if len(non_zero_rows) == 0: return [] # 分组连续行,确定每行的上下边界 line_bounds = [] start = non_zero_rows[0] for i in range(1, len(non_zero_rows)): # 行间隙阈值:可根据图像空白大小调整 if non_zero_rows[i] - non_zero_rows[i-1] > 5: line_bounds.append((start, non_zero_rows[i-1])) start = non_zero_rows[i] line_bounds.append((start, non_zero_rows[-1])) # 分割出每行图像,添加少量padding避免裁切文字 line_images = [] padding = 2 for (top, bottom) in line_bounds: top = max(0, top - padding) bottom = min(thresh.shape[0], bottom + padding) line_img = thresh[top:bottom, :] line_images.append(cv2.cvtColor(line_img, cv2.COLOR_GRAY2BGR)) return line_images
2. 单行转单词(单词分割)
通过纵向像素投影定位单词间的空白间隙,分割出单个单词:
def split_line_to_words(line_img): # 转灰度图 gray = cv2.cvtColor(line_img, cv2.COLOR_BGR2GRAY) # 二值化 _, thresh = cv2.threshold(gray, 0, 255, cv2.THRESH_BINARY_INV + cv2.THRESH_OTSU) # 计算纵向像素投影:每列的像素总和 col_proj = np.sum(thresh, axis=0) # 筛选出有文本的列索引 non_zero_cols = np.where(col_proj > 0)[0] if len(non_zero_cols) == 0: return [] # 分组连续列,确定每个单词的左右边界 word_bounds = [] start = non_zero_cols[0] for i in range(1, len(non_zero_cols)): # 单词间隙阈值:可根据单词间距调整 if non_zero_cols[i] - non_zero_cols[i-1] > 10: word_bounds.append((start, non_zero_cols[i-1])) start = non_zero_cols[i] word_bounds.append((start, non_zero_cols[-1])) # 分割出每个单词图像,添加少量padding word_images = [] padding = 2 for (left, right) in word_bounds: left = max(0, left - padding) right = min(thresh.shape[1], right + padding) word_img = thresh[:, left:right] word_images.append(cv2.cvtColor(word_img, cv2.COLOR_GRAY2BGR)) return word_images
3. 适配已有单词转字符函数
将原函数修改为接受图像对象(避免频繁读写文件),并修正字符顺序问题:
# 需提前加载你的loaded_model和class_labels def word_to_text(word_img): extracted_text = "" gray = cv2.cvtColor(word_img, cv2.COLOR_BGR2GRAY) _, thresh = cv2.threshold(gray, 0, 255, cv2.THRESH_BINARY_INV + cv2.THRESH_OTSU) contours, _ = cv2.findContours(thresh, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE) # 按x坐标正序排序轮廓,保证字符顺序正确 contours = sorted(contours, key=lambda ctr: cv2.boundingRect(ctr)[0]) for contour in contours: x, y, w, h = cv2.boundingRect(contour) # 过滤噪声小轮廓 if w < 5 or h < 5: continue letter_img = thresh[y:y + h, x:x + w] # 锐化字符 kernel = np.array([[0, -1, 0], [-1, 5, -1], [0, -1, 0]], np.float32) letter_img = cv2.filter2D(letter_img, -1, kernel) # 添加边框并调整尺寸 letter_img = cv2.copyMakeBorder(letter_img, 4, 4, 4, 4, cv2.BORDER_CONSTANT, value=[0,0,0]) letter_img = cv2.resize(letter_img, (28, 28), interpolation=cv2.INTER_AREA) # 可选:显示字符 # plt.imshow(letter_img, cmap='gray') # plt.axis('off') # plt.show() # 模型预测 letter_img = letter_img.reshape(1, 28, 28, 1).astype('float32') / 255 prediction = loaded_model.predict(letter_img) predicted_class = np.argmax(prediction) extracted_text += class_labels[predicted_class] print('识别单词:', extracted_text) return extracted_text
4. 完整流程调用
def process_multi_line_image(image_path): img = cv2.imread(image_path) if img is None: print("无法读取图像") return # 第一步:分割为单行图像 line_images = split_image_to_lines(img) print(f"分割出 {len(line_images)} 行文本") full_text = [] for line_idx, line_img in enumerate(line_images): print(f"\n处理第 {line_idx+1} 行:") # 第二步:分割为单词图像 word_images = split_line_to_words(line_img) print(f"分割出 {len(word_images)} 个单词") line_text = [] for word_img in word_images: # 第三步:单词转文本 line_text.append(word_to_text(word_img)) full_text.append(' '.join(line_text)) # 输出最终结果 print("\n===== 最终识别结果 =====") for line in full_text: print(line) return full_text # 调用示例 process_multi_line_image(r"C:\Users\student\Desktop\FinalProject\Flask\image\test3IMG.jpg")
三、关键参数调整说明
- 行分割的
行间隙阈值:文本行间距大时,调大5这个数值 - 单词分割的
单词间隙阈值:单词间距大时,调大10这个数值 - 字符过滤的
w < 5 or h < 5:字符较大时,调大阈值过滤更大的噪声
四、优化建议
- 若分割效果差,可在二值化后添加形态学操作增强文本:
kernel = cv2.getStructuringElement(cv2.MORPH_RECT, (2,2)) thresh = cv2.dilate(thresh, kernel, iterations=1)
- 倾斜文本需先做倾斜校正,再执行分割流程,提升准确率。
内容的提问来源于stack exchange,提问作者Rozira
相关产品推荐
相关产品推荐

