You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Pandas将CSV中hist列的键值对拆分为多行结构化数据

实现思路

  • 原数据中hist列的每条记录由\n\n分隔为多个独立的属性块,每个属性块内部是key: value格式的键值对
  • 先把每条hist内容解析为字典列表,再通过行展开、字典列拆分的方式得到最终结构化数据

完整实现代码

import pandas as pd

# 1. 读取CSV文件,分隔符为分号
df = pd.read_csv('sample1.csv', sep=';')

# 2. 定义hist列解析函数
def parse_hist(hist_str):
    blocks = hist_str.strip().split('\n\n')
    res = []
    current = {}
    for block in blocks:
        if not block:
            continue
        # 逐行解析键值对
        for line in block.split('\n'):
            if not line:
                continue
            k, v = line.split(':', 1)
            # 处理原数据中id=3的hist里的naction笔误,过滤多余的n前缀
            k = k.strip().lstrip('n')
            current[k.strip()] = v.strip()
        res.append(current.copy())
    return res

# 3. 解析hist并展开为多行
df['hist_parsed'] = df['hist'].apply(parse_hist)
df_exploded = df.explode('hist_parsed', ignore_index=True)

# 4. 把解析后的字典列拆分为独立字段
hist_df = pd.json_normalize(df_exploded['hist_parsed'])

# 5. 合并原字段和新解析字段,空值替换为空字符串匹配输出要求
final_df = pd.concat([df_exploded[['id', 'name', 'desc']], hist_df], axis=1).fillna('')

# 按要求格式输出
print(final_df.to_markdown(index=False))

补充说明

  • 代码默认保留原数据的字段值,如果需要统一action、auto的大小写,可自行添加大小写转换逻辑
  • 如果不需要处理naction的笔误,删除对应lstrip('n')的代码即可

内容的提问来源于stack exchange,提问作者arodrber

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.01 01:48:04