You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将层级嵌套结构数据转换为Pandas DataFrame

如何将嵌套层级数据转换为Pandas DataFrame?

问题背景

从网页获取了嵌套结构的数据,获取数据的代码如下:

import pandas as pd
import requests
from bs4 import BeautifulSoup

url = 'https://www.quantalys.com/Recherche' 
soup = BeautifulSoup(requests.get(url).content, "html.parser")
data = soup.find(class_='quantatree-datasource')['value']

原始嵌套数据结构

[{"key":"1","title":"Monétaire",
"children":[{"key":"4","title":"Monétaire Europe",
"children":[{"key":"1","title":"Monétaire euro"}, 
{"key":"2","title":"Monétaire euro dynamique"}, 
{"key":"3","title":"Monétaire autre devise Europe"}]},
{"key":"5","title":"Monétaire hors Europe",
"children":[{"key":"4","title":"Monétaire USD"},
{"key":"5","title":"Monétaire hors Europe autre devise"}]}]}

目标DataFrame格式

希望转换为包含所有层级路径的DataFrame,每一行对应一条从根到节点的路径,列对应层级:

Monétaire     
Monétaire     Monétaire Europe
Monétaire     Monétaire Europe   Monétaire euro
Monétaire     Monétaire Europe   Monétaire euro dynamique
Monétaire     Monétaire Europe   Monétaire autre devise Europe
Monétaire     Monétaire hors Europe
Monétaire     Monétaire hors Europe  Monétaire USD
Monétaire     Monétaire hors Europe  Monétaire hors Europe autre devise

实现方案

通过递归遍历嵌套结构收集所有层级路径,再将路径列表转换为DataFrame,具体步骤如下:

1. 解析JSON数据

首先把获取到的data字符串解析为Python可操作的字典结构:

import json

# 解析JSON字符串
parsed_data = json.loads(data)

2. 递归收集所有路径

编写递归函数,遍历每个节点并记录从根节点到当前节点的完整路径:

def collect_paths(nodes, current_path=None):
    paths = []
    if current_path is None:
        current_path = []
    for node in nodes:
        # 将当前节点的标题加入路径
        new_path = current_path + [node['title']]
        # 保存当前路径
        paths.append(new_path)
        # 如果存在子节点,递归处理子节点
        if 'children' in node and node['children']:
            paths.extend(collect_paths(node['children'], new_path))
    return paths

# 获取所有层级路径
all_paths = collect_paths(parsed_data)

3. 转换为DataFrame

将路径列表转换为DataFrame,对长度不足的路径补空值以对齐列数:

# 确定最大层级深度,作为DataFrame的列数
max_depth = max(len(path) for path in all_paths)
# 补全每个路径的长度,不足部分用空字符串填充
filled_paths = [path + ['']*(max_depth - len(path)) for path in all_paths]
# 创建DataFrame,可自定义列名
df = pd.DataFrame(filled_paths, columns=[f"层级{i+1}" for i in range(max_depth)])

# 若需要和示例一致的无表头输出,执行以下打印语句
print(df.to_string(index=False, header=False))

运行上述代码后,即可得到符合要求的层级化DataFrame。

内容的提问来源于stack exchange,提问作者jacques

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.09 06:15:00