Python OCR地址处理:将多行字段合并为每行三字段的格式优化请求
OCR地址输出格式优化方案
当前输出格式
现有tess_address函数生成的address.txt格式如下:
First Name, Address, City State Zip, Second Name, Second Address, Second City State zip,
期望输出格式
需要直接生成每3个字段合并为一行的格式:
First Name, Address, City State Zip, Second Name, Second Address, City State Zip,
优化后的代码
import os import re import pytesseract def tess_address(): files = os.listdir("address") sorted_files = sorted(files) # 一次性打开文件,减少IO操作 with open("address.txt", 'w', encoding='utf8') as output_file: for image in sorted_files: # 用os.path.join适配不同系统路径 image_path = os.path.join("address", image) # OCR识别文本 text = pytesseract.image_to_string(image_path) # 移除原有逗号 clean_text = re.sub(",", "", text) # 分割行并过滤空行(处理OCR可能识别出的无效空行) valid_lines = [line.strip() for line in clean_text.splitlines() if line.strip()] # 每3行一组合并成目标格式 for idx in range(0, len(valid_lines), 3): # 取出当前组的3行内容 address_group = valid_lines[idx:idx+3] # 合并为一行,字段间用", "分隔,末尾加逗号 merged_line = ", ".join(address_group) + "," output_file.write(f"{merged_line}\n")
优化说明
- IO效率提升:使用
with上下文管理器统一处理文件读写,避免循环中反复打开/关闭文件,同时保证文件自动关闭。 - 路径兼容性:替换手动拼接路径为
os.path.join,适配Windows、Linux等不同操作系统的路径规则。 - 空行处理:过滤OCR识别出的空行或空白内容,避免分组逻辑出错。
- 直接生成目标格式:通过切片按3行一组批量合并,无需后续调用额外处理函数,一步到位生成所需格式。
- 代码可读性:简化冗余逻辑,变量命名更清晰,便于后续维护和理解。
内容的提问来源于stack exchange,提问作者kmsgli
相关产品推荐
相关产品推荐

