Python中如何将pytesseract.image_to_data返回字符串转为字典列表?
解析pytesseract.image_to_data返回的TSV数据为字典列表
其实pytesseract本身就支持直接返回结构化数据,不用手动处理制表符分隔的字符串,这是最优的解决方案:
方法一:直接获取结构化字典列表
通过指定output_type参数为Output.DICT,接口会直接返回以字段名为键、对应列数据为值的字典,再通过简单的列表推导就能转成你需要的字典列表:
import pytesseract from pytesseract import Output # 直接获取结构化数据 ocr_data = pytesseract.image_to_data("image.png", output_type=Output.DICT) # 转换为字典列表,方便遍历 ocr_item_list = [dict(zip(ocr_data.keys(), row)) for row in zip(*ocr_data.values())] # 遍历查找目标文本并提取width/height target_content = "你要找的文本内容" for item in ocr_item_list: # 跳过空文本的无效条目 if not item["text"].strip(): continue if target_content in item["text"]: print(f"匹配条目:width={item['width']}, height={item['height']}")
方法二:手动解析已获取的TSV字符串
如果已经拿到了接口返回的制表符分隔字符串,可以用Python内置的csv模块来解析,避免自己处理字符串分割可能出现的边界问题:
import csv from io import StringIO # 假设已获取TSV格式的返回字符串 tsv_result = pytesseract.image_to_data("image.png") # 用csv.DictReader解析,自动将首行作为字典键 reader = csv.DictReader(StringIO(tsv_result), delimiter="\t") ocr_item_list = list(reader) # 遍历查找(注意这里的字段值是字符串,如需数值类型可自行转换) target_content = "你要找的文本内容" for item in ocr_item_list: if not item["text"].strip(): continue if target_content in item["text"]: width = int(item["width"]) height = int(item["height"]) print(f"匹配条目:width={width}, height={height}")
注意事项
- 方法一是最优选择,直接利用pytesseract的内置功能,减少手动解析的出错概率;
- 部分旧版本pytesseract可能需要确认
Output类是否可用,建议使用较新版本的库; - 遍历过程中记得跳过空文本的条目,避免无效匹配。
内容的提问来源于stack exchange,提问作者PythonKiddieScripterX
相关产品推荐
相关产品推荐

