求Linux下适配Tesseract 5的文本标注GUI工具(生成训练文件)
针对Tesseract 5的行级OCR标注解决方案
1. 适配Tesseract 5的Box Editor改版工具
原qt-box-editor仅支持Tesseract 3,但社区有基于Qt5维护的适配Tesseract 5的分支版本,完全匹配你的需求:
- 自动基于Tesseract 5的行检测功能拆分页面为单行图像
- 加载Tesseract 5的识别结果作为初始转录猜测
- 提供GUI界面直接修正文本内容
- 导出对应单行的图像文件和标注文本文件,输出格式兼容Tesseract微调流程
该工具基于原项目修改,Linux下可通过Qt5构建工具编译,仅需少量适配调整,无需从零开发标注套件。
2. 轻量自定义方案:PyTesseract + PySimpleGUI
如果找不到现成适配工具,这个方案只需少量代码调整即可快速实现需求:
- 利用
pytesseract.image_to_data()获取行级边界框和识别文本,完成自动行拆分 - 通过PySimpleGUI快速搭建轻量GUI,加载单行图像并展示Tesseract的初始猜测,支持手动编辑
- 内置逻辑将修正后的文本与对应行图像保存到指定目录(文件名一一对应,如
line_001.png和line_001.txt)
核心代码示例:
import pytesseract from PIL import Image import PySimpleGUI as sg import os # 加载页面图像并提取行级数据 img = Image.open("input_page.png") ocr_data = pytesseract.image_to_data(img, output_type=pytesseract.Output.DICT) # 整理行级信息(合并同一行的字符块) line_list = [] current_line = None for idx in range(len(ocr_data["text"])): if ocr_data["level"][idx] == 4: # Tesseract行级别标记 if not current_line: current_line = { "left": ocr_data["left"][idx], "top": ocr_data["top"][idx], "width": ocr_data["width"][idx], "height": ocr_data["height"][idx], "text": ocr_data["text"][idx] } else: # 合并同一行的边界框和文本 current_line["width"] = max(current_line["left"] + current_line["width"], ocr_data["left"][idx] + ocr_data["width"][idx]) - current_line["left"] current_line["height"] = max(current_line["top"] + current_line["height"], ocr_data["top"][idx] + ocr_data["height"][idx]) - current_line["top"] current_line["text"] += f' {ocr_data["text"][idx]}' elif ocr_data["level"][idx] < 4 and current_line: line_list.append(current_line) current_line = None if current_line: line_list.append(current_line) # 搭建标注GUI layout = [] for line_idx, line_info in enumerate(line_list): # 裁剪单行图像并临时保存 line_img = img.crop((line_info["left"], line_info["top"], line_info["left"] + line_info["width"], line_info["top"] + line_info["height"])) temp_img_path = f"temp_line_{line_idx}.png" line_img.save(temp_img_path) # 添加GUI行元素:图像+文本输入框 layout.append([sg.Image(temp_img_path), sg.InputText(line_info["text"], key=f"line_{line_idx}")]) layout.append([sg.Button("保存所有标注")]) window = sg.Window("行级OCR标注工具", layout) while True: event, values = window.read() if event == sg.WIN_CLOSED: break if event == "保存所有标注": output_dir = "labeled_lines" os.makedirs(output_dir, exist_ok=True) # 遍历保存所有行的图像和标注文本 for line_idx, line_info in enumerate(line_list): line_img = img.crop((line_info["left"], line_info["top"], line_info["left"] + line_info["width"], line_info["top"] + line_info["height"])) img_save_path = os.path.join(output_dir, f"line_{line_idx:03d}.png") txt_save_path = os.path.join(output_dir, f"line_{line_idx:03d}.txt") line_img.save(img_save_path) with open(txt_save_path, "w", encoding="utf-8") as f: f.write(values[f"line_{line_idx}"]) sg.popup("标注文件已保存完成") break window.close()
这段代码已经实现核心功能,你可根据需求调整UI样式、图像压缩参数或Tesseract识别配置,无需从零构建完整套件。
3. OCRopus 行级标注工具
OCRopus是开源OCR工具链,自带行级标注GUI,支持对接Tesseract 5作为识别后端:
- 自动拆分页面为单行图像片段
- 调用Tesseract 5生成初始转录文本
- 提供直观的GUI界面修正标注内容
- 导出兼容Tesseract微调的图像-文本文件对
Linux下可通过系统包管理器安装OCRopus,或从源码构建,仅需少量配置即可完成与Tesseract 5的对接。
内容的提问来源于stack exchange,提问作者BradypusRex
相关产品推荐
相关产品推荐

