如何使用Pandas DataFrame解析CSV中的嵌套attributes字段并生成指定输出
使用Pandas解析嵌套结构CSV的attributes字段
解决步骤
读取CSV并转换嵌套字段
CSV中的attributes字段是形似字典的字符串,先用ast.literal_eval将其转为可操作的Python字典对象,避免JSON解析因布尔值格式(true)导致的报错。提取嵌套层级数据
从转换后的字典中逐层提取所需字段:- 从
attributes['data']['attributes']中提取created_at、include_percentiles、metric_type、tags - 从
attributes['data']['attributes']['aggregations'][0]中提取space(示例中聚合数组仅一个元素,若有多个可循环或额外展开)
- 从
展开数组字段
使用explode方法将tags数组拆分为多行,让每个标签对应一条记录。
完整代码
import pandas as pd import ast # 读取CSV文件 df = pd.read_csv('your_file.csv') # 将attributes字段的字符串转为Python字典 df['attributes'] = df['attributes'].apply(ast.literal_eval) # 提取嵌套属性并生成新列 df['space'] = df['attributes'].apply(lambda x: x['data']['attributes']['aggregations'][0]['space']) df['created_at'] = df['attributes'].apply(lambda x: x['data']['attributes']['created_at']) df['include_percentiles'] = df['attributes'].apply(lambda x: x['data']['attributes']['include_percentiles']) df['metric_type'] = df['attributes'].apply(lambda x: x['data']['attributes']['metric_type']) df['tags'] = df['attributes'].apply(lambda x: x['data']['attributes']['tags']) # 展开tags数组为多行 df_exploded = df.explode('tags') # 保留需要的列并调整顺序 final_df = df_exploded[['id', 'type', 'space', 'created_at', 'include_percentiles', 'metric_type', 'tags']] # 打印结果 print(final_df.to_string(index=False))
输出结果
id type space created_at include_percentiles metric_type tags 1 xx sum 2020-03-25T09:48:37.463835Z True count app 1 xx sum 2020-03-25T09:48:37.463835Z True count datacenter
内容的提问来源于stack exchange,提问作者rose1110
相关产品推荐
相关产品推荐

