You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Tesseract.js生成文本位置偏移问题求助:实现可选中图文转文字工具

解决文本位置偏差的可行方案

1. 修正OCR坐标与文本渲染的基线差异

Tesseract返回的top坐标是文本边界框的顶部,但PIL的draw.text()默认基于文本的基线而非顶部绘制,这是位置偏差的核心原因之一。可以通过字体的 ascent/descent 参数调整y坐标:

from PIL import ImageFont

font = ImageFont.truetype("目标字体路径.ttf", 初始字体大小)
# 获取字体基线到顶部的距离(ascent)和基线到底部的距离(descent)
ascent, descent = font.getmetrics()
# 调整y坐标,让文本顶部对齐OCR边界框顶部
adjusted_y = y + ascent - (ascent + descent)

2. 精准匹配OCR边界框的尺寸

不要直接用OCR返回的height作为字体大小,因为边界框高度包含上下留白。可以用二分法计算适配边界框的字体大小:

def get_fitting_font(text, target_width, target_height, font_path):
    min_size = 1
    max_size = 150
    best_size = min_size
    while min_size <= max_size:
        mid_size = (min_size + max_size) // 2
        font = ImageFont.truetype(font_path, mid_size)
        text_w, text_h = font.getsize(text)
        if text_w <= target_width and text_h <= target_height:
            best_size = mid_size
            min_size = mid_size + 1
        else:
            max_size = mid_size - 1
    return ImageFont.truetype(font_path, best_size)

3. 对齐方式匹配原图文本

根据原图文本的对齐方式(左/中/右)调整绘制坐标:

# 以居中对齐为例
font = get_fitting_font(text, w, h, font_path)
text_w, _ = font.getsize(text)
adjusted_x = x + (w - text_w) // 2  # 让文本水平居中于OCR边界框
draw.text((adjusted_x, adjusted_y), text, font=font, fill=(255,0,0,128), align="center")

4. 统一坐标系统

如果涉及图像缩放、旋转,必须同步调整OCR返回的坐标。比如图像缩放了scale倍,所有x/y/w/h都要乘以scale,避免坐标错位。

5. 优化OCR边界框精度

如果OCR本身返回的边界框不准,可通过图像预处理提升精度:

  • 对图像做二值化、去噪、对比度调整
  • 给Tesseract添加--psm参数,比如单行文本用--psm 7,强制按单行识别

调整后的核心代码片段

import cv2
import pytesseract
from PIL import Image, ImageDraw, ImageFont

def get_fitting_font(text, target_width, target_height, font_path):
    min_size = 1
    max_size = 150
    best_size = min_size
    while min_size <= max_size:
        mid_size = (min_size + max_size) // 2
        font = ImageFont.truetype(font_path, mid_size)
        text_w, text_h = font.getsize(text)
        if text_w <= target_width and text_h <= target_height:
            best_size = mid_size
            min_size = mid_size + 1
        else:
            max_size = mid_size - 1
    return ImageFont.truetype(font_path, best_size)

def adjust_text_position(image_path, font_path):
    img = cv2.imread(image_path)
    rgb_img = cv2.cvtColor(img, cv2.COLOR_BGR2RGB)
    pil_img = Image.fromarray(rgb_img)
    draw = ImageDraw.Draw(pil_img)
    
    results = pytesseract.image_to_data(rgb_img, output_type=pytesseract.Output.DICT, lang="eng")
    
    for i in range(len(results["text"])):
        conf = int(results["conf"][i])
        if conf < 60 or not results["text"][i].strip():
            continue
        
        x, y, w, h = results["left"][i], results["top"][i], results["width"][i], results["height"][i]
        text = results["text"][i].strip()
        
        font = get_fitting_font(text, w, h, font_path)
        ascent, descent = font.getmetrics()
        text_w, text_h = font.getsize(text)
        
        adjusted_y = y + (h - text_h) // 2
        adjusted_x = x + (w - text_w) // 2
        
        draw.text((adjusted_x, adjusted_y), text, font=font, fill=(255, 0, 0, 128))
    
    return pil_img

# 调用示例
adjusted_img = adjust_text_position("测试原图路径.jpg", "字体文件路径.ttf")
adjusted_img.save("效果输出.jpg")

内容的提问来源于stack exchange,提问作者That wolphin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.16 11:00:59