You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Tesseract Python封装中SetRectangle后iterate_level报错:无文本返回

问题:修改文本边界框后Tesseract无法返回文本以获取字体属性

我有一张包含几行日文汉字的图片,尝试修改文本边界框后遍历获取字体及其属性。已确认Pyplot能正确显示修改后的边界框,且移除SetRectangle循环后文本可正常返回。

原函数代码

def font_attr(resized):
    img = resized
    image = Image.fromarray(img)

    with PyTessBaseAPI(path='C:\\Users\\mdelal001\\fast_format_assist\\tessdata\\', lang='jpn+msp+hgp', oem=0, psm=3) as api:
        api.SetImage(image)
        
        boxes = api.GetComponentImages(RIL.TEXTLINE, True)
        delta = 5 

        image_array = np.array(image)
        for box in boxes:
            print(box)
            box = box[1]
            x, y, w, h = box['x'] - delta, box['y'] - delta, box['w'] + 2 * delta, box['h'] + 2 * delta
            cv2.line(image_array, (x, y), (x + w, y), (0, 0, 0), 2)
            cv2.line(image_array, (x, y), (x, y + h), (0, 0, 0), 2)
            cv2.line(image_array, (x + w, y), (x + w, y + h), (0, 0, 0), 2)
            cv2.line(image_array, (x, y + h), (x + w, y + h), (0, 0, 0), 2)
        plt.imshow(image_array)
        plt.show()

        for i, (im, box, _, _) in enumerate(boxes):
             print(i,  (im, box, _, _))
             api.SetRectangle(box['x'] - delta, box['y'] - delta, box['w'] + 2 * delta, box['h'] + 2 * delta) 

        api.Recognize()
        ri = api.GetIterator()
        font = []
        attributes = []
        for r in iterate_level(ri, RIL.BLOCK):
            symbol = r.GetUTF8Text(RIL.TEXTLINE)
            conf = r.Confidence(RIL.BLOCK)
            symbol = symbol.replace('\n',' ').replace(' ', '')
            word_attributes = r.WordFontAttributes()
            if not symbol:
                continue
            else:
                font.append([symbol, 'confidence: ',conf])
                attributes.append(word_attributes)
            
            return font, attributes

运行报错信息

Traceback (most recent call last):
  File "c:\Users\m1\fast_format_assist\font_reader.py", line 338, in <module>
    attr = font_attr(resized)
  File "c:\Users\m1\fast_format_assist\font_reader.py", line 216, in font_attr
    symbol = r.GetUTF8Text(RIL.TEXTLINE)
  File "tesserocr.pyx", line 820, in tesserocr._tesserocr.PyLTRResultIterator.GetUTF8Text
RuntimeError: No text returned

问题分析与修复

核心问题点

  1. SetRectangle循环覆盖:循环调用SetRectangle会不断覆盖识别区域,最终仅最后一个边界框生效,且后续迭代用RIL.BLOCK层级,和之前获取的TEXTLINE层级不匹配
  2. 层级不匹配:用RIL.BLOCK迭代,但调用GetUTF8Text时传入RIL.TEXTLINE,层级不一致导致无法获取文本
  3. return位置错误:return在循环内部,第一次迭代就直接返回,无法遍历所有文本框
  4. Recognize调用时机错误:设置所有边界框后才调用Recognize,无法正确识别每个调整后的区域

修复后的代码

def font_attr(resized):
    img = resized
    image = Image.fromarray(img)
    font = []
    attributes = []

    with PyTessBaseAPI(path='C:\\Users\\mdelal001\\fast_format_assist\\tessdata\\', lang='jpn+msp+hgp', oem=0, psm=3) as api:
        api.SetImage(image)
        
        boxes = api.GetComponentImages(RIL.TEXTLINE, True)
        delta = 5 

        # 可视化修改后的边界框(保留原逻辑)
        image_array = np.array(image)
        for box in boxes:
            box_data = box[1]
            x, y, w, h = box_data['x'] - delta, box_data['y'] - delta, box_data['w'] + 2 * delta, box_data['h'] + 2 * delta
            cv2.line(image_array, (x, y), (x + w, y), (0, 0, 0), 2)
            cv2.line(image_array, (x, y), (x, y + h), (0, 0, 0), 2)
            cv2.line(image_array, (x + w, y), (x + w, y + h), (0, 0, 0), 2)
            cv2.line(image_array, (x, y + h), (x + w, y + h), (0, 0, 0), 2)
        plt.imshow(image_array)
        plt.show()

        # 逐个处理每个文本框
        for i, (im, box_data, _, _) in enumerate(boxes):
            # 设置调整后的边界框
            x, y, w, h = box_data['x'] - delta, box_data['y'] - delta, box_data['w'] + 2 * delta, box_data['h'] + 2 * delta
            api.SetRectangle(x, y, w, h)
            
            # 识别当前区域
            api.Recognize()
            ri = api.GetIterator()
            
            # 遍历当前文本框内容(统一用TEXTLINE层级)
            for r in iterate_level(ri, RIL.TEXTLINE):
                symbol = r.GetUTF8Text(RIL.TEXTLINE)
                if not symbol:
                    continue
                # 清理文本格式
                symbol = symbol.replace('\n',' ').strip()
                if not symbol:
                    continue
                # 获取置信度和字体属性
                conf = r.Confidence(RIL.TEXTLINE)
                word_attrs = r.WordFontAttributes()
                
                font.append([symbol, 'confidence: ', conf])
                attributes.append(word_attrs)

    return font, attributes

关键修改说明

  • 将return语句移至函数末尾,确保遍历所有文本框后再返回结果
  • 对每个文本框单独设置SetRectangle并调用Recognize,保证每个调整后的区域都被正确识别
  • 统一迭代层级和文本获取层级为RIL.TEXTLINE,与组件图像层级匹配
  • 优化文本清理逻辑,避免空字符串被加入结果列表

内容的提问来源于stack exchange,提问作者elchapo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.26 04:48:20