如何从CSV文件中提取Name、Address等指定数据?
从格式异常的文本文件中提取指定字段数据
你提供的文件并非标准CSV格式,是字段名与对应值分行排列、夹杂分隔线和行号的错乱文本。以下是两种提取指定数据的方法:
一、手动提取(数据量小时)
直接定位字段名对应的有效行(忽略分隔线和行号前缀):
- Name:
A------- - Gender:
M - Country:
Oman - Date of Birth:
28.10.1995
注意:文本中没有Address相关内容,无法提取。
二、Python脚本批量提取(数据量大时)
用脚本自动过滤无效行,匹配目标字段并抓取对应值:
# 读取文件内容,过滤空行和分隔线行 with open("your_file.txt", "r", encoding="utf-8") as f: lines = [line.strip() for line in f if line.strip() and not line.strip().startswith("---")] # 定义需要提取的目标字段 target_fields = {"Name", "Gender", "Date of Birth", "Address"} extracted_data = {} current_target = None for line in lines: # 判断当前行是否是目标字段 if line in target_fields: current_target = line elif current_target: # 提取行号后的内容(比如"4 A-------"处理成"A-------") value = line.split(maxsplit=1)[1] if " " in line else line extracted_data[current_target] = value current_target = None # 输出提取结果 for field in target_fields: print(f"{field}: {extracted_data.get(field, '未找到对应数据')}")
运行脚本前,将your_file.txt替换为你的实际文件名即可。
内容的提问来源于stack exchange,提问作者Sheryar Malik
相关产品推荐
相关产品推荐

