You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Pillow在图片原文本矩形区域绘制替换后的文本?

问题背景

假设我们有如下图片:
(示例图片内容为两行文本:Write Text on Image 和 using OpenCV)

首先通过以下代码使用Tesseract识别图片文本:

from PIL import Image
import pytesseract
pytesseract.pytesseract.tesseract_cmd = r'C:\Program Files\Tesseract-OCR\tesseract.exe'
text = pytesseract.image_to_string(Image.open('text_image.png'))

随后将识别文本中的“Text”替换为“new Text”,代码如下:

text = text.replace("Text","new Text")
print(text)

替换结果为:

Write new Text on Image
using OpenCV

现在需将该新文本绘制到原图的原文本位置,找到的教程中Pillow的text方法仅传入单点坐标:

d1.text((65, 10), "Sample text", fill =(255, 0, 0),font=myFont)

但原文本对应矩形区域,请问该如何实现此需求?

解决方案

要实现将新文本精准绘制到原文本的矩形区域,核心是先获取原文本的边界框坐标,再基于该区域完成覆盖、文本绘制,具体步骤如下:

1. 获取原文本的边界框信息

使用Tesseract的image_to_data方法,它会返回每个文本片段的位置数据(左上角坐标、宽度、高度等),方便定位原文本区域:

from PIL import Image
import pytesseract
import pandas as pd

pytesseract.pytesseract.tesseract_cmd = r'C:\Program Files\Tesseract-OCR\tesseract.exe'
img = Image.open('text_image.png')

# 获取OCR详细数据并转为DataFrame处理
ocr_data = pytesseract.image_to_data(img, output_type=pytesseract.Output.DATAFRAME)
# 过滤无效的空文本行
ocr_data = ocr_data[ocr_data['text'].notna()]

# 计算整个文本区域的边界框:取所有文本行的最小左上角坐标、最大右下角坐标
min_x = ocr_data['left'].min()
min_y = ocr_data['top'].min()
max_x = ocr_data['left'].max() + ocr_data['width'].max()
max_y = ocr_data['top'].max() + ocr_data['height'].max()
text_bbox = (min_x, min_y, max_x, max_y)

如果只需替换特定的“Text”单词,可单独定位该单词的边界框:

# 找到包含"Text"的文本行,提取其边界框
target_row = ocr_data[ocr_data['text'] == 'Text'].iloc[0]
word_bbox = (target_row['left'], target_row['top'], target_row['left']+target_row['width'], target_row['top']+target_row['height'])

2. 覆盖原文本区域

用与原图背景一致的颜色填充边界框,彻底覆盖原文本:

from PIL import ImageDraw

draw = ImageDraw.Draw(img)
# 假设原图背景为白色,需根据实际背景色调整
draw.rectangle(text_bbox, fill=(255, 255, 255))

3. 在边界框内适配绘制新文本

考虑文本换行、对齐问题,确保新文本完全适配原区域:

from PIL import ImageFont
import textwrap

# 加载匹配原文本风格的字体,字号需根据实际调整
font = ImageFont.truetype('arial.ttf', 16)
new_text = "Write new Text on Image\nusing OpenCV"

# 拆分文本为适配边界框宽度的多行
bbox_width = max_x - min_x
# 根据字体字符宽度计算每行可容纳的字符数
char_width = font.getsize(' ')[0]
wrapped_lines = textwrap.wrap(new_text, width=int(bbox_width / char_width))

# 计算垂直居中的起始位置
line_height = font.getsize('A')[1]
total_text_height = len(wrapped_lines) * line_height
y_start = min_y + (max_y - min_y - total_text_height) // 2
current_y = y_start

for line in wrapped_lines:
    # 左对齐绘制,若需居中可计算水平偏移
    current_x = min_x
    # 水平居中计算:current_x = min_x + (bbox_width - font.getsize(line)[0]) // 2
    draw.text((current_x, current_y), line, fill=(0, 0, 0), font=font)
    current_y += line_height

# 保存修改后的图片
img.save('modified_image.png')

关键注意点

  • 字体的类型、字号要尽量匹配原文本,避免文本超出边界框或风格不协调;
  • 填充颜色必须与原图背景一致,否则会留下明显的修改痕迹;
  • 若替换单个单词,只需针对该单词的边界框操作,无需处理整个文本区域。

内容的提问来源于stack exchange,提问作者neural science

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 05:55:30