You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Python OpenCV的文本分割及字符处理问题求助

字符分割与图像格式转换解决方案

一、解决"S"和"T"粘连的分割问题

字符粘连通常源于二值化效果不佳、轮廓合并或字符间距过小,可通过以下方法修复:

1. 优化二值化处理

全局阈值易受光照不均影响,改用自适应阈值能更精准保留字符边缘:

gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)
# 自适应高斯阈值,blockSize取奇数,C为亮度偏移常数
binary = cv2.adaptiveThreshold(gray, 255, cv2.ADAPTIVE_THRESH_GAUSSIAN_C, cv2.THRESH_BINARY_INV, 11, 2)

2. 垂直投影分割法(针对粘连字符)

通过计算垂直方向的像素分布,定位字符间的空白区域作为分割线:

import numpy as np

# 计算每列的白色像素数量(垂直投影)
vertical_proj = np.sum(binary == 255, axis=0)
split_points = []
in_gap = False

# 遍历投影值,找到连续空白列的中点作为分割线
for i, val in enumerate(vertical_proj):
    if val == 0 and not in_gap:
        start = i
        in_gap = True
    elif val > 0 and in_gap:
        split_points.append((start + i) // 2)
        in_gap = False

# 用分割线切割粘连的字符区域
# 假设粘连字符的bounding rect为x, y, w, h
for point in split_points:
    if x < point < x + w:
        char1 = binary[y:y+h, x:point]
        char2 = binary[y:y+h, point:x+w]
        # 分别处理两个分割后的字符

3. 形态学操作微调

用垂直方向的小核做腐蚀,断开细小粘连,再膨胀恢复字符原有大小:

kernel = cv2.getStructuringElement(cv2.MORPH_RECT, (1, 2))
eroded = cv2.erode(binary, kernel, iterations=1)
dilated = cv2.dilate(eroded, kernel, iterations=1)
# 从处理后的图像中重新提取轮廓
contours, _ = cv2.findContours(dilated, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE)

二、将字符图像转为28x28数组适配EMNIST模型

需保持字符比例并对齐数据集格式,步骤如下:

1. 裁剪字符并添加padding

提取字符时保留周围留白,避免边缘信息丢失:

# 假设contour为单个字符的轮廓
x, y, w, h = cv2.boundingRect(contour)
padding = 2
# 裁剪并处理边界溢出
char_img = binary[y-padding:y+h+padding, x-padding:x+w+padding]
char_img = np.pad(char_img, 
                  ((max(0, -y+padding), max(0, y+h+padding - binary.shape[0])),
                   (max(0, -x+padding), max(0, x+w+padding - binary.shape[1]))),
                  mode='constant', constant_values=0)

2. 等比例缩放并居中到28x28画布

直接拉伸会导致字符变形,需保持比例后居中放置:

target_size = 28
# 计算缩放比例,确保最大边不超过24(留4像素边距)
scale = min(24 / char_img.shape[0], 24 / char_img.shape[1])
new_w, new_h = int(char_img.shape[1] * scale), int(char_img.shape[0] * scale)
resized_char = cv2.resize(char_img, (new_w, new_h), interpolation=cv2.INTER_AREA)

# 创建28x28空白画布(背景为0,匹配EMNIST黑色背景格式)
canvas = np.zeros((target_size, target_size), dtype=np.uint8)
# 计算居中偏移量
x_offset = (target_size - new_w) // 2
y_offset = (target_size - new_h) // 2
canvas[y_offset:y_offset+new_h, x_offset:x_offset+new_w] = resized_char

3. 转换为模型输入数组

归一化并调整维度以适配模型要求:

# 转为float32并归一化到0-1范围
char_array = canvas.astype(np.float32) / 255.0
# 若模型需要单通道维度,添加通道轴
char_array = np.expand_dims(char_array, axis=-1)

内容的提问来源于stack exchange,提问作者Nuevo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.24 11:41:11