You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python+OpenCV实现多行文本图像的行、词、字符分割?

基于OpenCV的多行文本三级分割解决方案

一、整体流程概述

需完成三级文本分割+识别流程:

  1. 多行文本图像 → 单行文本图像
  2. 单行文本图像 → 单个单词图像
  3. 单词图像 → 字符图像+文本识别(基于你已实现的逻辑微调适配)

二、分步实现代码

1. 多行转单行(行分割)

通过横向像素投影定位文本行的上下边界,利用文本行区域像素集中的特性分割出独立行:

import cv2
import numpy as np
import matplotlib.pyplot as plt

def split_image_to_lines(img):
    # 转灰度图
    gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)
    # 二值化(黑底白字,便于投影分析)
    _, thresh = cv2.threshold(gray, 0, 255, cv2.THRESH_BINARY_INV + cv2.THRESH_OTSU)
    
    # 计算横向像素投影:每行的像素总和
    row_proj = np.sum(thresh, axis=1)
    # 筛选出有文本的行索引
    non_zero_rows = np.where(row_proj > 0)[0]
    
    if len(non_zero_rows) == 0:
        return []
    
    # 分组连续行,确定每行的上下边界
    line_bounds = []
    start = non_zero_rows[0]
    for i in range(1, len(non_zero_rows)):
        # 行间隙阈值:可根据图像空白大小调整
        if non_zero_rows[i] - non_zero_rows[i-1] > 5:
            line_bounds.append((start, non_zero_rows[i-1]))
            start = non_zero_rows[i]
    line_bounds.append((start, non_zero_rows[-1]))
    
    # 分割出每行图像,添加少量padding避免裁切文字
    line_images = []
    padding = 2
    for (top, bottom) in line_bounds:
        top = max(0, top - padding)
        bottom = min(thresh.shape[0], bottom + padding)
        line_img = thresh[top:bottom, :]
        line_images.append(cv2.cvtColor(line_img, cv2.COLOR_GRAY2BGR))
    
    return line_images

2. 单行转单词(单词分割)

通过纵向像素投影定位单词间的空白间隙,分割出单个单词:

def split_line_to_words(line_img):
    # 转灰度图
    gray = cv2.cvtColor(line_img, cv2.COLOR_BGR2GRAY)
    # 二值化
    _, thresh = cv2.threshold(gray, 0, 255, cv2.THRESH_BINARY_INV + cv2.THRESH_OTSU)
    
    # 计算纵向像素投影:每列的像素总和
    col_proj = np.sum(thresh, axis=0)
    # 筛选出有文本的列索引
    non_zero_cols = np.where(col_proj > 0)[0]
    
    if len(non_zero_cols) == 0:
        return []
    
    # 分组连续列,确定每个单词的左右边界
    word_bounds = []
    start = non_zero_cols[0]
    for i in range(1, len(non_zero_cols)):
        # 单词间隙阈值:可根据单词间距调整
        if non_zero_cols[i] - non_zero_cols[i-1] > 10:
            word_bounds.append((start, non_zero_cols[i-1]))
            start = non_zero_cols[i]
    word_bounds.append((start, non_zero_cols[-1]))
    
    # 分割出每个单词图像,添加少量padding
    word_images = []
    padding = 2
    for (left, right) in word_bounds:
        left = max(0, left - padding)
        right = min(thresh.shape[1], right + padding)
        word_img = thresh[:, left:right]
        word_images.append(cv2.cvtColor(word_img, cv2.COLOR_GRAY2BGR))
    
    return word_images

3. 适配已有单词转字符函数

将原函数修改为接受图像对象(避免频繁读写文件),并修正字符顺序问题:

# 需提前加载你的loaded_model和class_labels
def word_to_text(word_img):
    extracted_text = ""
    gray = cv2.cvtColor(word_img, cv2.COLOR_BGR2GRAY)
    _, thresh = cv2.threshold(gray, 0, 255, cv2.THRESH_BINARY_INV + cv2.THRESH_OTSU)
    contours, _ = cv2.findContours(thresh, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE)
    
    # 按x坐标正序排序轮廓,保证字符顺序正确
    contours = sorted(contours, key=lambda ctr: cv2.boundingRect(ctr)[0])
    
    for contour in contours:
        x, y, w, h = cv2.boundingRect(contour)
        # 过滤噪声小轮廓
        if w < 5 or h < 5:
            continue
        letter_img = thresh[y:y + h, x:x + w]
        # 锐化字符
        kernel = np.array([[0, -1, 0], [-1, 5, -1], [0, -1, 0]], np.float32)
        letter_img = cv2.filter2D(letter_img, -1, kernel)
        # 添加边框并调整尺寸
        letter_img = cv2.copyMakeBorder(letter_img, 4, 4, 4, 4, cv2.BORDER_CONSTANT, value=[0,0,0])
        letter_img = cv2.resize(letter_img, (28, 28), interpolation=cv2.INTER_AREA)
        
        # 可选:显示字符
        # plt.imshow(letter_img, cmap='gray')
        # plt.axis('off')
        # plt.show()
        
        # 模型预测
        letter_img = letter_img.reshape(1, 28, 28, 1).astype('float32') / 255
        prediction = loaded_model.predict(letter_img)
        predicted_class = np.argmax(prediction)
        extracted_text += class_labels[predicted_class]
    
    print('识别单词:', extracted_text)
    return extracted_text

4. 完整流程调用

def process_multi_line_image(image_path):
    img = cv2.imread(image_path)
    if img is None:
        print("无法读取图像")
        return
    
    # 第一步:分割为单行图像
    line_images = split_image_to_lines(img)
    print(f"分割出 {len(line_images)} 行文本")
    
    full_text = []
    for line_idx, line_img in enumerate(line_images):
        print(f"\n处理第 {line_idx+1} 行:")
        # 第二步:分割为单词图像
        word_images = split_line_to_words(line_img)
        print(f"分割出 {len(word_images)} 个单词")
        
        line_text = []
        for word_img in word_images:
            # 第三步:单词转文本
            line_text.append(word_to_text(word_img))
        
        full_text.append(' '.join(line_text))
    
    # 输出最终结果
    print("\n===== 最终识别结果 =====")
    for line in full_text:
        print(line)
    return full_text

# 调用示例
process_multi_line_image(r"C:\Users\student\Desktop\FinalProject\Flask\image\test3IMG.jpg")

三、关键参数调整说明

  • 行分割的行间隙阈值:文本行间距大时,调大5这个数值
  • 单词分割的单词间隙阈值:单词间距大时,调大10这个数值
  • 字符过滤的w < 5 or h < 5:字符较大时,调大阈值过滤更大的噪声

四、优化建议

  1. 若分割效果差,可在二值化后添加形态学操作增强文本:
kernel = cv2.getStructuringElement(cv2.MORPH_RECT, (2,2))
thresh = cv2.dilate(thresh, kernel, iterations=1)
  1. 倾斜文本需先做倾斜校正,再执行分割流程,提升准确率。

内容的提问来源于stack exchange,提问作者Rozira

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.24 03:52:09