You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Pandas将多份TXT文件转换为DataFrame并保存为CSV的方法

实现逻辑

前置匹配规则

默认所有TXT文件的标题行统一使用标题名:内容的格式,内容支持换行,示例TXT结构如下:

姓名:卡洛斯·阿尔贝托
党派:社会自由党
选区:圣保罗
任期:2023-2027
个人简介:
曾任圣保罗州议会议员,2022年当选联邦议员,
专注于公共卫生、社会保障领域的立法推进。

若实际标题分隔符为英文冒号等其他格式,修改后续代码中的正则规则即可。

完整Python代码

import os
import re
import pandas as pd

# 配置项,可根据实际情况修改
TXT_FOLDER_PATH = "./brazil_congressman_txt"  # 存放所有TXT文件的文件夹路径
OUTPUT_CSV_PATH = "./congressman_info.csv"
# 标题行匹配正则,若为英文冒号可改为 r'^([^:]+):\s*(.*)$'
TITLE_PATTERN = re.compile(r'^([^:]+):\s*(.*)$')
FILE_ENCODING = "utf-8"  # 若为葡萄牙语编码可改为 "latin-1" 或 "utf-8-sig"

if __name__ == "__main__":
    all_records = []
    all_existed_titles = set()

    # 遍历读取所有TXT文件,提取键值对
    for file_name in os.listdir(TXT_FOLDER_PATH):
        if not file_name.lower().endswith(".txt"):
            continue
        file_full_path = os.path.join(TXT_FOLDER_PATH, file_name)
        current_record = {}
        current_title = None

        with open(file_full_path, "r", encoding=FILE_ENCODING) as f:
            for line in f:
                line_stripped = line.strip()
                if not line_stripped:
                    continue
                # 匹配是否为标题行
                match_result = TITLE_PATTERN.match(line_stripped)
                if match_result:
                    current_title = match_result.group(1).strip()
                    content = match_result.group(2).strip()
                    current_record[current_title] = content
                    all_existed_titles.add(current_title)
                else:
                    # 非标题行则追加到上一个标题的内容中,处理多行内容
                    if current_title:
                        current_record[current_title] += f"\n{line_stripped}"
        all_records.append(current_record)

    # 筛选所有文件共有的标题作为列名
    common_titles = set(all_records[0].keys())
    for record in all_records[1:]:
        common_titles.intersection_update(record.keys())
    common_titles = list(common_titles)
    # 若不需要筛选共有标题,要保留所有出现过的标题,替换上面3行为:
    # common_titles = list(all_existed_titles)

    # 生成DataFrame并导出CSV
    df = pd.DataFrame(all_records, columns=common_titles)
    df.to_csv(OUTPUT_CSV_PATH, index=False, encoding="utf-8-sig")
    print(f"导出完成,共处理{len(all_records)}份文件,生成{len(common_titles)}列数据")

输出说明

最终导出的CSV文件表头为所有TXT共有的标题,每一行对应一份TXT文件的内容,换行的内容会保留换行符,用Excel或WPS打开可正常显示。

内容的提问来源于stack exchange,提问作者muharrem bagriyanik

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.05 11:57:03