You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python如何将指定格式文本数据转换为带自定义表头的CSV文件

Python键值对格式文本转CSV解决方案

你之前直接用pd.read_csv()按行读取的方式无法处理字段值换行缩进的场景,也无法将键值对映射为表头和对应行值,因此无法得到预期结果。可以按以下方案实现需求:

实现思路

  • 逐行读取原始文本,区分非缩进的键值行和缩进的续值行,将续值行内容拼接到对应上一个字段的值中
  • 以空行作为单条记录的分割标记,将同一条记录的所有键值对整合为一个字典
  • 所有记录处理完成后统一转换为DataFrame,导出为CSV即可

完整实现代码

import pandas as pd

# 存储所有解析完成的记录
records = []
# 存储当前正在解析的单条记录
current_record = {}
# 存储上一个处理的键名,用于拼接缩进的续值
last_key = None

with open("sepsis2015.txt", "r", encoding="utf-8") as f:
    for line in f:
        stripped_line = line.strip()
        # 空行代表当前记录结束
        if not stripped_line:
            if current_record:
                records.append(current_record)
                current_record = {}
                last_key = None
            continue
        
        # 判断是否为缩进的续值行
        if line.startswith((" ", "\t")):
            if last_key:
                current_record[last_key] += " " + stripped_line
            continue
        
        # 非缩进行为新的键值对,按第一个空白分割键和值
        key, value = stripped_line.split(maxsplit=1)
        current_record[key] = value
        last_key = key

# 加入最后一条未入库的记录
if current_record:
    records.append(current_record)

# 转换为DataFrame,可自定义调整列顺序
df = pd.DataFrame(records)
# 如需指定表头输出顺序可取消注释下方代码,替换为实际字段即可
# df = df[["PMID", "STAT", "DA", "CTDT"]]
# 导出为CSV,utf-8-sig编码适配Excel打开
df.to_csv("sepsis_output.csv", index=False, encoding="utf-8-sig")

效果说明

运行代码后会在同级目录生成sepsis_output.csv文件,表头为文本中出现过的所有字段标识,每一行对应原始文本中的一条完整记录,换行缩进的长值会自动合并为完整内容。

内容的提问来源于stack exchange,提问作者MASTANELAL ABDULGAFAR KURESHI

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.06 07:57:04