如何使用Pandas将CSV中hist列的键值对拆分为多行结构化数据
实现思路
- 原数据中
hist列的每条记录由\n\n分隔为多个独立的属性块,每个属性块内部是key: value格式的键值对 - 先把每条
hist内容解析为字典列表,再通过行展开、字典列拆分的方式得到最终结构化数据
完整实现代码
import pandas as pd # 1. 读取CSV文件,分隔符为分号 df = pd.read_csv('sample1.csv', sep=';') # 2. 定义hist列解析函数 def parse_hist(hist_str): blocks = hist_str.strip().split('\n\n') res = [] current = {} for block in blocks: if not block: continue # 逐行解析键值对 for line in block.split('\n'): if not line: continue k, v = line.split(':', 1) # 处理原数据中id=3的hist里的naction笔误,过滤多余的n前缀 k = k.strip().lstrip('n') current[k.strip()] = v.strip() res.append(current.copy()) return res # 3. 解析hist并展开为多行 df['hist_parsed'] = df['hist'].apply(parse_hist) df_exploded = df.explode('hist_parsed', ignore_index=True) # 4. 把解析后的字典列拆分为独立字段 hist_df = pd.json_normalize(df_exploded['hist_parsed']) # 5. 合并原字段和新解析字段,空值替换为空字符串匹配输出要求 final_df = pd.concat([df_exploded[['id', 'name', 'desc']], hist_df], axis=1).fillna('') # 按要求格式输出 print(final_df.to_markdown(index=False))
补充说明
- 代码默认保留原数据的字段值,如果需要统一
action、auto的大小写,可自行添加大小写转换逻辑 - 如果不需要处理
naction的笔误,删除对应lstrip('n')的代码即可
内容的提问来源于stack exchange,提问作者arodrber
相关产品推荐
相关产品推荐

