如何用Python编写脚本解析器将特定格式TXT数据转为竖线分隔格式
需求描述
我有一个TXT数据文件,内容格式如下:
name : jashon m l Address : 2410 Vance Avenue City : Alexandria State : Louisiana Zip code : 71301 Phone : +3187305018 name : rush b Address : address 2 City : city 2 State : state 2 Zip code : 71301 Phone : phone
希望将其转换为竖线分隔的格式:
jashon m l|2410 Vance Avenue|Alexandria|Louisiana|71301|+3187305018 rush b|address 2|city 2|state 2|71301|phone
Python解析脚本
以下是实现需求的Python脚本,逻辑清晰且适配性强:
def convert_txt_to_pipe(input_file_path, output_file_path): current_record = [] # 定义输出字段的固定顺序,需与原文件字段对应 field_order = ["name", "Address", "City", "State", "Zip code", "Phone"] with open(input_file_path, 'r', encoding='utf-8') as infile, \ open(output_file_path, 'w', encoding='utf-8') as outfile: for line in infile: line = line.strip() # 空行代表当前用户记录结束,写入文件并重置缓存 if not line: if current_record: outfile.write('|'.join(current_record) + '\n') current_record = [] continue # 按第一个冒号分割键值对,避免值中含冒号导致解析出错 key, value = line.split(':', 1) key = key.strip() value = value.strip() # 按指定顺序收集字段值 if key in field_order: current_record.append(value) # 处理文件末尾无空行的最后一条记录 if current_record: outfile.write('|'.join(current_record) + '\n') # 运行示例 if __name__ == "__main__": input_path = "input.txt" # 替换为你的输入文件路径 output_path = "output.txt" # 替换为你的输出文件路径 convert_txt_to_pipe(input_path, output_path) print("转换完成!")
脚本说明
- 字段顺序控制:通过
field_order列表指定输出字段的排列顺序,确保与需求一致 - 空行识别:利用空行作为不同用户记录的分隔符,触发写入操作
- 鲁棒性处理:使用
split(':', 1)分割键值对,避免值中包含冒号时出现解析错误 - 编码兼容:指定
utf-8编码,支持更多特殊字符场景
使用时只需替换input_path和output_path为实际文件路径即可执行。
内容的提问来源于stack exchange,提问作者abiz haha
相关产品推荐
相关产品推荐

