You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于pytesseract优化OCR提取文本的对齐问题

问题描述

提取图像文本时出现文本对齐错乱的情况,推测问题出在格式处理环节而非文本提取环节。想请教是否可以利用边界框坐标来优化文本对齐效果?原图像中价格列的起始位置是统一的,但提取后的输出文本存在对齐错乱问题。

当前使用的代码如下:

def extract_text(self, image):
    gray_image = cv2.cvtColor(image, cv2.COLOR_BGR2GRAY)

    text_regions = []
    retval, labels, stats, centroids = cv2.connectedComponentsWithStats(gray_image)

    section_groups = {}
    section_threshold = 20
    for i in range(1, retval):
        x, y, w, h, area = stats[i]
    section_groups.setdefault(y // section_threshold, []).append((x, y, w, h))

    for section_key, components in section_groups.items():
        section_image = gray_image[components[0][1]:components[-1][1] + components[-1][3],
                        :]
        extracted_text, section_boxes = self._extract_text_and_boxes(section_image)
        text_regions.append((extracted_text, section_boxes))

    return text_regions

def _extract_text_and_boxes(self, component):
    extracted_text = pytesseract.image_to_string(component, config=self.config)
    boxes = self.get_text_boxes(component)
    return extracted_text, boxes

def get_text_boxes(self, image):
    results = pytesseract.image_to_data(image, output_type=Output.DICT)
    boxes = []
    for i in range(0, len(results["text"])):
        x = results["left"][i]
        y = results["top"][i]
        w = results["width"][i]
        h = results["height"][i]

        text = results["text"][i]
        conf = float(results["conf"][i])
        if conf > 0:
            boxes.append((x, y, w, h))

    return boxes
解决方案

完全可以通过边界框坐标优化文本对齐,核心思路是基于每个文本块的x坐标(横向位置)对文本进行分组或对齐排版,而非直接使用Tesseract返回的无格式纯文本。

现有代码的问题

  1. 仅按行分割区域提取文本,未利用边界框位置信息排版,Tesseract的image_to_string输出是无格式纯文本,自然丢失原列对齐结构。
  2. 连通分量分组逻辑存在缩进错误:for循环内的append语句被放在循环外,导致仅最后一个连通分量被加入分组,行分割不准确。

优化步骤

1. 修复连通分量分组的缩进错误

修正extract_text方法的循环缩进,确保每个连通分量正确分组,并按行的y坐标排序保证处理顺序:

def extract_text(self, image):
    gray_image = cv2.cvtColor(image, cv2.COLOR_BGR2GRAY)

    text_regions = []
    retval, labels, stats, centroids = cv2.connectedComponentsWithStats(gray_image)

    section_groups = {}
    section_threshold = 20
    # 修复缩进:将append放入循环内部
    for i in range(1, retval):
        x, y, w, h, area = stats[i]
        section_groups.setdefault(y // section_threshold, []).append((x, y, w, h))

    # 按行的y坐标排序,保证从上到下处理
    sorted_sections = sorted(section_groups.items(), key=lambda item: item[0])
    for section_key, components in sorted_sections:
        # 获取当前行的上下边界
        min_y = min(comp[1] for comp in components)
        max_y = max(comp[1] + comp[3] for comp in components)
        section_image = gray_image[min_y:max_y, :]
        # 提取当前行的文本块与对应边界框
        line_text_boxes = self._extract_text_with_boxes(section_image)
        text_regions.append(line_text_boxes)

    return text_regions

2. 修改文本提取方法,保留文本与边界框的对应关系

替换原_extract_text_and_boxes方法,同时返回每个文本块的内容和位置,并按x坐标排序保证行内顺序:

def _extract_text_with_boxes(self, component):
    results = pytesseract.image_to_data(component, output_type=Output.DICT)
    text_boxes = []
    for i in range(0, len(results["text"])):
        x = results["left"][i]
        y = results["top"][i]
        w = results["width"][i]
        h = results["height"][i]
        text = results["text"][i].strip()
        conf = float(results["conf"][i])
        # 过滤空文本和低置信度结果
        if conf > 0 and text:
            text_boxes.append((x, y, w, h, text))
    # 按x坐标排序,确保同一行内从左到右的顺序
    text_boxes.sort(key=lambda item: item[0])
    return text_boxes

3. 基于边界框实现列对齐排版

通过聚类所有文本块的x坐标确定列基准位置,再对每行文本按基准位置填充空格实现对齐:

def format_aligned_text(self, text_regions):
    # 收集所有文本块的x坐标用于确定列基准
    all_x_coords = []
    for line_boxes in text_regions:
        for box in line_boxes:
            all_x_coords.append(box[0])
    if not all_x_coords:
        return ""

    # 对x坐标聚类,生成列基准位置(相邻x差小于10归为同一列)
    column_threshold = 10
    all_x_coords.sort()
    columns = [all_x_coords[0]]
    for x in all_x_coords[1:]:
        if x - columns[-1] > column_threshold:
            columns.append(x)

    # 按列基准对齐每行文本
    aligned_lines = []
    char_width = 5  # 假设每个字符占5像素宽度,可根据实际调整
    for line_boxes in text_regions:
        line_text = ""
        prev_col_end = 0
        for col_x in columns:
            # 匹配当前列的文本块
            col_text = ""
            for box in line_boxes:
                if abs(box[0] - col_x) <= column_threshold:
                    col_text = box[4]
                    break
            # 计算需要填充的空格数
            space_needed = max(0, (col_x - prev_col_end) // char_width)
            line_text += " " * space_needed + col_text
            prev_col_end = col_x + len(col_text) * char_width
        aligned_lines.append(line_text)
    return "\n".join(aligned_lines)

使用示例

# 提取带边界框的文本区域
text_regions = your_instance.extract_text(your_image)
# 生成对齐后的文本
aligned_text = your_instance.format_aligned_text(text_regions)
print(aligned_text)

内容的提问来源于stack exchange,提问作者code_comm

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.25 00:07:06