You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中如何将pytesseract.image_to_data返回字符串转为字典列表?

解析pytesseract.image_to_data返回的TSV数据为字典列表

其实pytesseract本身就支持直接返回结构化数据,不用手动处理制表符分隔的字符串,这是最优的解决方案:

方法一:直接获取结构化字典列表

通过指定output_type参数为Output.DICT,接口会直接返回以字段名为键、对应列数据为值的字典,再通过简单的列表推导就能转成你需要的字典列表:

import pytesseract
from pytesseract import Output

# 直接获取结构化数据
ocr_data = pytesseract.image_to_data("image.png", output_type=Output.DICT)

# 转换为字典列表,方便遍历
ocr_item_list = [dict(zip(ocr_data.keys(), row)) for row in zip(*ocr_data.values())]

# 遍历查找目标文本并提取width/height
target_content = "你要找的文本内容"
for item in ocr_item_list:
    # 跳过空文本的无效条目
    if not item["text"].strip():
        continue
    if target_content in item["text"]:
        print(f"匹配条目:width={item['width']}, height={item['height']}")

方法二:手动解析已获取的TSV字符串

如果已经拿到了接口返回的制表符分隔字符串,可以用Python内置的csv模块来解析,避免自己处理字符串分割可能出现的边界问题:

import csv
from io import StringIO

# 假设已获取TSV格式的返回字符串
tsv_result = pytesseract.image_to_data("image.png")

# 用csv.DictReader解析,自动将首行作为字典键
reader = csv.DictReader(StringIO(tsv_result), delimiter="\t")
ocr_item_list = list(reader)

# 遍历查找(注意这里的字段值是字符串,如需数值类型可自行转换)
target_content = "你要找的文本内容"
for item in ocr_item_list:
    if not item["text"].strip():
        continue
    if target_content in item["text"]:
        width = int(item["width"])
        height = int(item["height"])
        print(f"匹配条目:width={width}, height={height}")

注意事项

  • 方法一是最优选择,直接利用pytesseract的内置功能,减少手动解析的出错概率;
  • 部分旧版本pytesseract可能需要确认Output类是否可用,建议使用较新版本的库;
  • 遍历过程中记得跳过空文本的条目,避免无效匹配。

内容的提问来源于stack exchange,提问作者PythonKiddieScripterX

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 10:25:39