You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求Linux下适配Tesseract 5的文本标注GUI工具(生成训练文件)

针对Tesseract 5的行级OCR标注解决方案

1. 适配Tesseract 5的Box Editor改版工具

原qt-box-editor仅支持Tesseract 3,但社区有基于Qt5维护的适配Tesseract 5的分支版本,完全匹配你的需求:

  • 自动基于Tesseract 5的行检测功能拆分页面为单行图像
  • 加载Tesseract 5的识别结果作为初始转录猜测
  • 提供GUI界面直接修正文本内容
  • 导出对应单行的图像文件和标注文本文件,输出格式兼容Tesseract微调流程

该工具基于原项目修改,Linux下可通过Qt5构建工具编译,仅需少量适配调整,无需从零开发标注套件。

2. 轻量自定义方案:PyTesseract + PySimpleGUI

如果找不到现成适配工具,这个方案只需少量代码调整即可快速实现需求:

  • 利用pytesseract.image_to_data()获取行级边界框和识别文本,完成自动行拆分
  • 通过PySimpleGUI快速搭建轻量GUI,加载单行图像并展示Tesseract的初始猜测,支持手动编辑
  • 内置逻辑将修正后的文本与对应行图像保存到指定目录(文件名一一对应,如line_001.png和line_001.txt)

核心代码示例:

import pytesseract
from PIL import Image
import PySimpleGUI as sg
import os

# 加载页面图像并提取行级数据
img = Image.open("input_page.png")
ocr_data = pytesseract.image_to_data(img, output_type=pytesseract.Output.DICT)

# 整理行级信息(合并同一行的字符块)
line_list = []
current_line = None
for idx in range(len(ocr_data["text"])):
    if ocr_data["level"][idx] == 4:  # Tesseract行级别标记
        if not current_line:
            current_line = {
                "left": ocr_data["left"][idx],
                "top": ocr_data["top"][idx],
                "width": ocr_data["width"][idx],
                "height": ocr_data["height"][idx],
                "text": ocr_data["text"][idx]
            }
        else:
            # 合并同一行的边界框和文本
            current_line["width"] = max(current_line["left"] + current_line["width"], ocr_data["left"][idx] + ocr_data["width"][idx]) - current_line["left"]
            current_line["height"] = max(current_line["top"] + current_line["height"], ocr_data["top"][idx] + ocr_data["height"][idx]) - current_line["top"]
            current_line["text"] += f' {ocr_data["text"][idx]}'
    elif ocr_data["level"][idx] < 4 and current_line:
        line_list.append(current_line)
        current_line = None
if current_line:
    line_list.append(current_line)

# 搭建标注GUI
layout = []
for line_idx, line_info in enumerate(line_list):
    # 裁剪单行图像并临时保存
    line_img = img.crop((line_info["left"], line_info["top"], line_info["left"] + line_info["width"], line_info["top"] + line_info["height"]))
    temp_img_path = f"temp_line_{line_idx}.png"
    line_img.save(temp_img_path)
    # 添加GUI行元素:图像+文本输入框
    layout.append([sg.Image(temp_img_path), sg.InputText(line_info["text"], key=f"line_{line_idx}")])
layout.append([sg.Button("保存所有标注")])

window = sg.Window("行级OCR标注工具", layout)

while True:
    event, values = window.read()
    if event == sg.WIN_CLOSED:
        break
    if event == "保存所有标注":
        output_dir = "labeled_lines"
        os.makedirs(output_dir, exist_ok=True)
        # 遍历保存所有行的图像和标注文本
        for line_idx, line_info in enumerate(line_list):
            line_img = img.crop((line_info["left"], line_info["top"], line_info["left"] + line_info["width"], line_info["top"] + line_info["height"]))
            img_save_path = os.path.join(output_dir, f"line_{line_idx:03d}.png")
            txt_save_path = os.path.join(output_dir, f"line_{line_idx:03d}.txt")
            line_img.save(img_save_path)
            with open(txt_save_path, "w", encoding="utf-8") as f:
                f.write(values[f"line_{line_idx}"])
        sg.popup("标注文件已保存完成")
        break

window.close()

这段代码已经实现核心功能,你可根据需求调整UI样式、图像压缩参数或Tesseract识别配置,无需从零构建完整套件。

3. OCRopus 行级标注工具

OCRopus是开源OCR工具链,自带行级标注GUI,支持对接Tesseract 5作为识别后端:

  • 自动拆分页面为单行图像片段
  • 调用Tesseract 5生成初始转录文本
  • 提供直观的GUI界面修正标注内容
  • 导出兼容Tesseract微调的图像-文本文件对

Linux下可通过系统包管理器安装OCRopus,或从源码构建,仅需少量配置即可完成与Tesseract 5的对接。

内容的提问来源于stack exchange,提问作者BradypusRex

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.02 00:40:08