Python结构化文本转CSV:无法将行映射为列的问题求助
解决结构化员工文本转指定CSV格式的问题
我明白你现在的困扰——原代码只是把每行的键值对拆分成两列,但没办法把同一个员工的所有属性整合到CSV的一行里。咱们来一步步解决这个问题:
问题分析
你的输入文本里,员工信息是以属性行的形式排列的:有的员工以Employee Name:开头,有的则直接从Designation:开始(比如johndoe4到johndoe8)。我们需要把每个员工的所有对应属性(姓名、职位、邮箱、联系方式、资质、专长)映射到CSV的对应列,每个员工占一行,同时还要处理一些脏数据(比如Contact里的Intercom信息、Qualification里的多余符号)。
解决方案代码
下面是经过优化的代码,能完美生成你需要的CSV格式:
import csv # 定义CSV的表头,和你需要的格式完全匹配 csv_headers = ["Employee name", "designation", "email", "contact", "Qualification", "Specialisation"] current_employee = {} employees_list = [] # 清理联系方式:提取第一个有效号码,去掉Intercom等无关内容 def clean_contact(contact_text): for segment in contact_text.split(","): cleaned = segment.strip().replace("Intercom No.", "").rstrip("<").strip() if cleaned: return cleaned return "" # 清理资质信息:去掉多余符号、替换HTML实体、整理格式 def clean_qualification(qual_text): cleaned = qual_text.strip().lstrip(":").strip() cleaned = cleaned.replace("&", "&").replace("|", ", ").rstrip(",") return cleaned # 读取并处理原始文本 with open('test.txt', 'r') as records_file: for line in records_file: line = line.strip() if not line: continue # 拆分键值对,只按第一个冒号分割(避免Qualification里的冒号干扰) key_value_pair = line.split(":", 1) if len(key_value_pair) != 2: continue key, value = key_value_pair key = key.strip() value = value.strip() # 识别新员工的开始:要么遇到Employee Name,要么遇到Designation但当前无员工数据 if key == "Employee Name": # 如果当前已有员工数据,先存入列表 if current_employee: employees_list.append(current_employee) current_employee = {} current_employee["Employee name"] = value elif key == "Designation": # 处理没有Employee Name的员工,默认名字设为Unknown(也可以根据邮箱自定义) if not current_employee: current_employee["Employee name"] = "Unknown" current_employee["designation"] = value elif key == "Email": current_employee["email"] = value elif key == "ContactNo": current_employee["contact"] = clean_contact(value) elif key == "Qualification": current_employee["Qualification"] = clean_qualification(value) elif key == "Area of Interest / Specialisation": current_employee["Specialisation"] = value # 把最后一个员工的数据加入列表 if current_employee: employees_list.append(current_employee) # 写入CSV文件 with open('log.csv', 'w', newline='') as output_file: writer = csv.DictWriter(output_file, fieldnames=csv_headers) writer.writeheader() # 确保每一行都包含所有表头字段,缺失的填空字符串 for employee in employees_list: csv_row = {header: employee.get(header, "") for header in csv_headers} writer.writerow(csv_row)
代码关键点解释
- 用字典收集员工信息:每个员工的属性都存在一个字典里,键对应CSV的表头,这样能轻松映射到列。
- 脏数据处理:专门写了两个清理函数,处理Contact里的Intercom信息、Qualification里的多余冒号、HTML实体(比如
&)等问题。 - 兼容两种员工条目:既处理有
Employee Name开头的员工,也处理直接从Designation开始的员工,给无姓名的员工设置默认值。 - 使用DictWriter:比普通的
csv.writer更适合这种表头对应场景,自动处理列的顺序。
输出效果
运行后生成的CSV会是你需要的格式:每行对应一个员工,所有属性都在指定的列里,脏数据也被清理干净。
内容的提问来源于stack exchange,提问作者ARUN XZA
相关产品推荐
相关产品推荐

