如何将DataFrame中attributes字段拆分为新列并处理嵌套属性值
实现方案
以下基于Python Pandas实现,可直接适配你的需求:
前置准备
首先导入需要的依赖库:
import pandas as pd import json
核心代码实现
1. 编写嵌套字典展平函数
递归处理嵌套结构,将嵌套的key转换为上层属性.下层属性的格式:
def flatten_dict(d, parent_key='', sep='.'): items = [] for k, v in d.items(): new_key = f"{parent_key}{sep}{k}" if parent_key else k # 如果值仍是字典,递归展平 if isinstance(v, dict): items.extend(flatten_dict(v, new_key, sep=sep).items()) else: items.append((new_key, v)) return dict(items)
2. 处理DataFrame生成结果
# 假设你的原始DataFrame变量名为df # 第一步:将attributes列的JSON字符串转为Python字典,如果你的attributes列已经是字典类型,可跳过这一步 df['attributes'] = df['attributes'].apply(lambda x: json.loads(x) if isinstance(x, str) else x) # 第二步:展平所有属性并生成属性列 attr_df = df['attributes'].apply(flatten_dict).apply(pd.Series) # 第三步:缺失属性填充为False attr_df = attr_df.fillna(False) # 第四步:拼接business_id列得到最终结果 result = pd.concat([df['business_id'], attr_df], axis=1)
可选优化
如果你的属性值里有字符串格式的布尔值(比如"True"/"False"),需要转成真正的布尔类型,可在填充缺失值后增加以下处理:
for col in attr_df.columns: attr_df[col] = attr_df[col].replace({"True": True, "False": False})
内容的提问来源于stack exchange,提问作者reign
相关产品推荐
相关产品推荐

