如何用Pillow在图片原文本矩形区域绘制替换后的文本?
问题背景
假设我们有如下图片:
(示例图片内容为两行文本:Write Text on Image 和 using OpenCV)
首先通过以下代码使用Tesseract识别图片文本:
from PIL import Image import pytesseract pytesseract.pytesseract.tesseract_cmd = r'C:\Program Files\Tesseract-OCR\tesseract.exe' text = pytesseract.image_to_string(Image.open('text_image.png'))
随后将识别文本中的“Text”替换为“new Text”,代码如下:
text = text.replace("Text","new Text") print(text)
替换结果为:
Write new Text on Image using OpenCV
现在需将该新文本绘制到原图的原文本位置,找到的教程中Pillow的text方法仅传入单点坐标:
d1.text((65, 10), "Sample text", fill =(255, 0, 0),font=myFont)
但原文本对应矩形区域,请问该如何实现此需求?
解决方案
要实现将新文本精准绘制到原文本的矩形区域,核心是先获取原文本的边界框坐标,再基于该区域完成覆盖、文本绘制,具体步骤如下:
1. 获取原文本的边界框信息
使用Tesseract的image_to_data方法,它会返回每个文本片段的位置数据(左上角坐标、宽度、高度等),方便定位原文本区域:
from PIL import Image import pytesseract import pandas as pd pytesseract.pytesseract.tesseract_cmd = r'C:\Program Files\Tesseract-OCR\tesseract.exe' img = Image.open('text_image.png') # 获取OCR详细数据并转为DataFrame处理 ocr_data = pytesseract.image_to_data(img, output_type=pytesseract.Output.DATAFRAME) # 过滤无效的空文本行 ocr_data = ocr_data[ocr_data['text'].notna()] # 计算整个文本区域的边界框:取所有文本行的最小左上角坐标、最大右下角坐标 min_x = ocr_data['left'].min() min_y = ocr_data['top'].min() max_x = ocr_data['left'].max() + ocr_data['width'].max() max_y = ocr_data['top'].max() + ocr_data['height'].max() text_bbox = (min_x, min_y, max_x, max_y)
如果只需替换特定的“Text”单词,可单独定位该单词的边界框:
# 找到包含"Text"的文本行,提取其边界框 target_row = ocr_data[ocr_data['text'] == 'Text'].iloc[0] word_bbox = (target_row['left'], target_row['top'], target_row['left']+target_row['width'], target_row['top']+target_row['height'])
2. 覆盖原文本区域
用与原图背景一致的颜色填充边界框,彻底覆盖原文本:
from PIL import ImageDraw draw = ImageDraw.Draw(img) # 假设原图背景为白色,需根据实际背景色调整 draw.rectangle(text_bbox, fill=(255, 255, 255))
3. 在边界框内适配绘制新文本
考虑文本换行、对齐问题,确保新文本完全适配原区域:
from PIL import ImageFont import textwrap # 加载匹配原文本风格的字体,字号需根据实际调整 font = ImageFont.truetype('arial.ttf', 16) new_text = "Write new Text on Image\nusing OpenCV" # 拆分文本为适配边界框宽度的多行 bbox_width = max_x - min_x # 根据字体字符宽度计算每行可容纳的字符数 char_width = font.getsize(' ')[0] wrapped_lines = textwrap.wrap(new_text, width=int(bbox_width / char_width)) # 计算垂直居中的起始位置 line_height = font.getsize('A')[1] total_text_height = len(wrapped_lines) * line_height y_start = min_y + (max_y - min_y - total_text_height) // 2 current_y = y_start for line in wrapped_lines: # 左对齐绘制,若需居中可计算水平偏移 current_x = min_x # 水平居中计算:current_x = min_x + (bbox_width - font.getsize(line)[0]) // 2 draw.text((current_x, current_y), line, fill=(0, 0, 0), font=font) current_y += line_height # 保存修改后的图片 img.save('modified_image.png')
关键注意点
- 字体的类型、字号要尽量匹配原文本,避免文本超出边界框或风格不协调;
- 填充颜色必须与原图背景一致,否则会留下明显的修改痕迹;
- 若替换单个单词,只需针对该单词的边界框操作,无需处理整个文本区域。
内容的提问来源于stack exchange,提问作者neural science
相关产品推荐
相关产品推荐

