Tesseract.js生成文本位置偏移问题求助:实现可选中图文转文字工具
解决文本位置偏差的可行方案
1. 修正OCR坐标与文本渲染的基线差异
Tesseract返回的top坐标是文本边界框的顶部,但PIL的draw.text()默认基于文本的基线而非顶部绘制,这是位置偏差的核心原因之一。可以通过字体的 ascent/descent 参数调整y坐标:
from PIL import ImageFont font = ImageFont.truetype("目标字体路径.ttf", 初始字体大小) # 获取字体基线到顶部的距离(ascent)和基线到底部的距离(descent) ascent, descent = font.getmetrics() # 调整y坐标,让文本顶部对齐OCR边界框顶部 adjusted_y = y + ascent - (ascent + descent)
2. 精准匹配OCR边界框的尺寸
不要直接用OCR返回的height作为字体大小,因为边界框高度包含上下留白。可以用二分法计算适配边界框的字体大小:
def get_fitting_font(text, target_width, target_height, font_path): min_size = 1 max_size = 150 best_size = min_size while min_size <= max_size: mid_size = (min_size + max_size) // 2 font = ImageFont.truetype(font_path, mid_size) text_w, text_h = font.getsize(text) if text_w <= target_width and text_h <= target_height: best_size = mid_size min_size = mid_size + 1 else: max_size = mid_size - 1 return ImageFont.truetype(font_path, best_size)
3. 对齐方式匹配原图文本
根据原图文本的对齐方式(左/中/右)调整绘制坐标:
# 以居中对齐为例 font = get_fitting_font(text, w, h, font_path) text_w, _ = font.getsize(text) adjusted_x = x + (w - text_w) // 2 # 让文本水平居中于OCR边界框 draw.text((adjusted_x, adjusted_y), text, font=font, fill=(255,0,0,128), align="center")
4. 统一坐标系统
如果涉及图像缩放、旋转,必须同步调整OCR返回的坐标。比如图像缩放了scale倍,所有x/y/w/h都要乘以scale,避免坐标错位。
5. 优化OCR边界框精度
如果OCR本身返回的边界框不准,可通过图像预处理提升精度:
- 对图像做二值化、去噪、对比度调整
- 给Tesseract添加
--psm参数,比如单行文本用--psm 7,强制按单行识别
调整后的核心代码片段
import cv2 import pytesseract from PIL import Image, ImageDraw, ImageFont def get_fitting_font(text, target_width, target_height, font_path): min_size = 1 max_size = 150 best_size = min_size while min_size <= max_size: mid_size = (min_size + max_size) // 2 font = ImageFont.truetype(font_path, mid_size) text_w, text_h = font.getsize(text) if text_w <= target_width and text_h <= target_height: best_size = mid_size min_size = mid_size + 1 else: max_size = mid_size - 1 return ImageFont.truetype(font_path, best_size) def adjust_text_position(image_path, font_path): img = cv2.imread(image_path) rgb_img = cv2.cvtColor(img, cv2.COLOR_BGR2RGB) pil_img = Image.fromarray(rgb_img) draw = ImageDraw.Draw(pil_img) results = pytesseract.image_to_data(rgb_img, output_type=pytesseract.Output.DICT, lang="eng") for i in range(len(results["text"])): conf = int(results["conf"][i]) if conf < 60 or not results["text"][i].strip(): continue x, y, w, h = results["left"][i], results["top"][i], results["width"][i], results["height"][i] text = results["text"][i].strip() font = get_fitting_font(text, w, h, font_path) ascent, descent = font.getmetrics() text_w, text_h = font.getsize(text) adjusted_y = y + (h - text_h) // 2 adjusted_x = x + (w - text_w) // 2 draw.text((adjusted_x, adjusted_y), text, font=font, fill=(255, 0, 0, 128)) return pil_img # 调用示例 adjusted_img = adjust_text_position("测试原图路径.jpg", "字体文件路径.ttf") adjusted_img.save("效果输出.jpg")
内容的提问来源于stack exchange,提问作者That wolphin
相关产品推荐
相关产品推荐

