You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python实现图片指定文本坐标获取与高亮填充的优化方案咨询

精准定位并高亮指定文本段落

原问题场景

需要定位图片中指定文本「was the age of wisdom」的坐标,并用浅色填充方式高亮。原代码通过匹配首尾单词「was」和「wisdom,」的方式未得到预期结果,当前错误高亮效果如下:
错误高亮结果
原始图片:
原始图片

原代码问题分析

  • 未确保匹配到的「was」和「wisdom,」属于同一个目标句子,文本中可能存在多个相同单词,导致坐标范围错误
  • 使用边框高亮不符合需求,需要改为浅色填充的高亮方式

改进实现方案

以下代码通过锁定连续单词序列的方式精准定位目标文本,并实现半透明浅色填充高亮:

import pytesseract
import cv2
import numpy as np

# 配置Tesseract路径
pytesseract.pytesseract.tesseract_cmd = "C:\\Program Files\\Tesseract-OCR\\tesseract.exe"

# 读取图片
filename = 'C:\\Users\\vicky\\Downloads\\image.png'
img = cv2.imread(filename)
# 备份原图用于半透明叠加
img_copy = img.copy()

# 获取带坐标的文本数据
from pytesseract import Output
d = pytesseract.image_to_data(img, output_type=Output.DICT)

# 整理单词和对应坐标:过滤空文本,保留有效单词信息
word_info = []
for i in range(len(d['text'])):
    text = d['text'][i].strip()
    if text:
        word_info.append({
            'text': text,
            'left': d['left'][i],
            'top': d['top'][i],
            'width': d['width'][i],
            'height': d['height'][i],
            'right': d['left'][i] + d['width'][i],
            'bottom': d['top'][i] + d['height'][i]
        })

# 目标短语(注意匹配识别出的单词格式,比如末尾是否带标点)
target_phrase = ["was", "the", "age", "of", "wisdom"]

# 查找连续匹配的单词序列
match_indices = []
for i in range(len(word_info) - len(target_phrase) + 1):
    current_sequence = [word['text'].lower() for word in word_info[i:i+len(target_phrase)]]
    # 兼容带逗号的情况,比如"wisdom,"
    current_sequence[-1] = current_sequence[-1].rstrip(',')
    if current_sequence == [w.lower() for w in target_phrase]:
        match_indices = list(range(i, i+len(target_phrase)))
        break

if match_indices:
    # 计算整个目标短语的边界框:取所有单词的最小left、最小top,最大right、最大bottom
    min_left = min(word_info[i]['left'] for i in match_indices)
    min_top = min(word_info[i]['top'] for i in match_indices)
    max_right = max(word_info[i]['right'] for i in match_indices)
    max_bottom = max(word_info[i]['bottom'] for i in match_indices)
    
    # 打印坐标
    print(f"目标文本左上角坐标:({min_left}, {min_top})")
    print(f"目标文本右下角坐标:({max_right}, {max_bottom})")
    
    # 绘制半透明填充高亮
    # 选择浅色(比如淡黄色),设置透明度(alpha=0.3)
    highlight_color = (255, 255, 153)  # 淡黄色
    alpha = 0.3
    # 在备份图上绘制填充矩形
    cv2.rectangle(img_copy, (min_left, min_top), (max_right, max_bottom), highlight_color, -1)
    # 叠加原图和高亮图,实现半透明效果
    img = cv2.addWeighted(img_copy, alpha, img, 1 - alpha, 0)
    
    # 保存结果
    cv2.imwrite('highlight_result.png', img)
else:
    print("未找到目标文本")

关键说明

  1. 精准匹配连续单词:通过遍历单词序列,确保找到的是连续的目标短语,避免跨句子匹配错误
  2. 半透明高亮实现:使用cv2.addWeighted实现图像叠加,既保留原文本清晰度,又实现柔和的填充高亮
  3. 兼容标点差异:处理识别出的单词可能带标点的情况(比如「wisdom,」),确保匹配成功

内容的提问来源于stack exchange,提问作者surendra kumawat

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.30 00:25:12