You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将Python中的Dataframe转换为树形结构JSON?

嘿,我来帮你实现从DataFrame生成你要的树形结构JSON!看了你的示例结构,咱们可以用Pandas配合简单的递归或者手动分层构建的方式搞定,下面是具体的实现方案:

实现思路

你的目标结构是多层嵌套:Cat → Band → PE → 日期 → GroupCalcs(包含各类统计字段),核心就是把DataFrame的行按这些层级分组,逐层嵌套成字典后再转换为JSON格式。

前提假设

先假设你的DataFrame结构如下(列顺序对应层级从外到内,最后是统计字段):

CatBandPEDateCountEqWghtRtnIdx_EqWghtRtnIdx_WghtRtnWghtRtn
Cat1Band1PE11995-01382.0-0.01580.984160.975977-0.0240
Cat1Band1PE11995-02382.0-0.01580.984160.975977-0.0240
方法一:递归构建(简洁高效)

这种方法适合层级较多的场景,用递归函数自动逐层分组构建树形:

import pandas as pd
import json

def build_nested_tree(group_df):
    # 当只剩统计字段时,打包成GroupCalcs字典
    stat_cols = ['Count', 'EqWghtRtn', 'Idx_EqWghtRtn', 'Idx_WghtRtn', 'WghtRtn']
    if set(group_df.columns) == set(stat_cols):
        # 取当前分组的唯一统计值(需确保同一层级无重复行)
        return {'GroupCalcs': group_df.iloc[0].to_dict()}
    
    # 取当前最外层的分组列,递归构建子树
    current_level = group_df.columns[0]
    tree = {}
    for key, subgroup in group_df.groupby(current_level):
        # 移除当前分组列,继续处理下一层
        tree[key] = build_nested_tree(subgroup.drop(columns=[current_level]))
    return tree

# 假设你的DataFrame名为df
# 先把日期转成字符串(避免datetime类型在JSON中被转成奇怪的格式)
df['Date'] = df['Date'].astype(str)

# 构建树形字典
tree_structure = build_nested_tree(df)

# 转成格式化的JSON字符串
json_result = json.dumps(tree_structure, indent=4)

# 输出或保存
print(json_result)
方法二:手动分层构建(直观易调试)

如果觉得递归有点绕,可以手动逐层创建字典,逻辑更直观,方便调试:

import pandas as pd
import json

# 假设你的DataFrame名为df
df['Date'] = df['Date'].astype(str)

# 设置多层索引,方便按层级提取数据
df.set_index(['Cat', 'Band', 'PE', 'Date'], inplace=True)

# 先转成「多层索引-统计值」的字典
raw_nested = df.to_dict('index')

# 手动构建目标树形结构
tree_structure = {}
for (cat, band, pe, date), stats in raw_nested.items():
    # 逐层创建字典,不存在则初始化
    if cat not in tree_structure:
        tree_structure[cat] = {}
    if band not in tree_structure[cat]:
        tree_structure[cat][band] = {}
    if pe not in tree_structure[cat][band]:
        tree_structure[cat][band][pe] = {}
    # 把统计值放到GroupCalcs下
    tree_structure[cat][band][pe][date] = {'GroupCalcs': stats}

# 转成格式化的JSON
json_result = json.dumps(tree_structure, indent=4)
print(json_result)
注意事项
  • 层级顺序:不管用哪种方法,都要保证分组层级的顺序是从外到内(Cat → Band → PE → 日期),否则生成的结构会错乱。
  • 日期格式:一定要把日期列转成字符串,不然JSON会把datetime对象转换成ISO格式字符串(比如"1995-01-01T00:00:00"),不符合你的示例要求。
  • 重复值处理:如果同一层级下有重复的键(比如同一个Cat+Band+PE+Date对应多行),两种方法都会取最后一行的值,所以提前要确保DataFrame是按层级去重后的结果。

内容的提问来源于stack exchange,提问作者Dave D

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 08:35:31