使用Pandas将多份TXT文件转换为DataFrame并保存为CSV的方法
实现逻辑
前置匹配规则
默认所有TXT文件的标题行统一使用标题名:内容的格式,内容支持换行,示例TXT结构如下:
姓名:卡洛斯·阿尔贝托
党派:社会自由党
选区:圣保罗
任期:2023-2027
个人简介:
曾任圣保罗州议会议员,2022年当选联邦议员,
专注于公共卫生、社会保障领域的立法推进。
若实际标题分隔符为英文冒号等其他格式,修改后续代码中的正则规则即可。
完整Python代码
import os import re import pandas as pd # 配置项,可根据实际情况修改 TXT_FOLDER_PATH = "./brazil_congressman_txt" # 存放所有TXT文件的文件夹路径 OUTPUT_CSV_PATH = "./congressman_info.csv" # 标题行匹配正则,若为英文冒号可改为 r'^([^:]+):\s*(.*)$' TITLE_PATTERN = re.compile(r'^([^:]+):\s*(.*)$') FILE_ENCODING = "utf-8" # 若为葡萄牙语编码可改为 "latin-1" 或 "utf-8-sig" if __name__ == "__main__": all_records = [] all_existed_titles = set() # 遍历读取所有TXT文件,提取键值对 for file_name in os.listdir(TXT_FOLDER_PATH): if not file_name.lower().endswith(".txt"): continue file_full_path = os.path.join(TXT_FOLDER_PATH, file_name) current_record = {} current_title = None with open(file_full_path, "r", encoding=FILE_ENCODING) as f: for line in f: line_stripped = line.strip() if not line_stripped: continue # 匹配是否为标题行 match_result = TITLE_PATTERN.match(line_stripped) if match_result: current_title = match_result.group(1).strip() content = match_result.group(2).strip() current_record[current_title] = content all_existed_titles.add(current_title) else: # 非标题行则追加到上一个标题的内容中,处理多行内容 if current_title: current_record[current_title] += f"\n{line_stripped}" all_records.append(current_record) # 筛选所有文件共有的标题作为列名 common_titles = set(all_records[0].keys()) for record in all_records[1:]: common_titles.intersection_update(record.keys()) common_titles = list(common_titles) # 若不需要筛选共有标题,要保留所有出现过的标题,替换上面3行为: # common_titles = list(all_existed_titles) # 生成DataFrame并导出CSV df = pd.DataFrame(all_records, columns=common_titles) df.to_csv(OUTPUT_CSV_PATH, index=False, encoding="utf-8-sig") print(f"导出完成,共处理{len(all_records)}份文件,生成{len(common_titles)}列数据")
输出说明
最终导出的CSV文件表头为所有TXT共有的标题,每一行对应一份TXT文件的内容,换行的内容会保留换行符,用Excel或WPS打开可正常显示。
内容的提问来源于stack exchange,提问作者muharrem bagriyanik
相关产品推荐
相关产品推荐

