You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何展平DataFrame中的字典并拼接所有结果行

解决方案

问题原因

你遇到的KeyError: 'node'是因为GitHub GraphQL返回的comments和labels是分页结构对象,里面包含nodes(实际的评论/标签列表)和pageInfo(分页信息),直接用apply(pd.Series)解析整个对象会找不到预期字段,得先提取nodes里的列表再展开。

具体实现步骤

假设你从API获取的原始数据结构类似GitHub GraphQL返回格式(模拟示例):

# 模拟GitHub GraphQL返回的Issue数据
raw_data = {
    "data": {
        "repository": {
            "issues": {
                "nodes": [
                    {
                        "number": 1,
                        "title": "测试Issue 1",
                        "body": "这是第一个测试Issue",
                        "comments": {
                            "nodes": [
                                {"id": "C_1", "author": {"login": "user1"}, "body": "评论1"},
                                {"id": "C_2", "author": {"login": "user2"}, "body": "评论2"}
                            ],
                            "pageInfo": {"endCursor": "...", "hasNextPage": False}
                        },
                        "labels": {
                            "nodes": [
                                {"id": "L_1", "name": "bug", "color": "ff0000"},
                                {"id": "L_2", "name": "help wanted", "color": "00ff00"}
                            ],
                            "pageInfo": {"endCursor": "...", "hasNextPage": False}
                        }
                    },
                    {
                        "number": 2,
                        "title": "测试Issue 2",
                        "body": "这是第二个测试Issue",
                        "comments": {
                            "nodes": [{"id": "C_3", "author": {"login": "user3"}, "body": "评论3"}],
                            "pageInfo": {"endCursor": "...", "hasNextPage": False}
                        },
                        "labels": {
                            "nodes": [{"id": "L_3", "name": "enhancement", "color": "0000ff"}],
                            "pageInfo": {"endCursor": "...", "hasNextPage": False}
                        }
                    }
                ]
            }
        }
    }
}

1. 转换初始DataFrame

先提取Issue基础数据,剥离分页信息:

import pandas as pd

# 提取Issue列表
issues_list = raw_data["data"]["repository"]["issues"]["nodes"]
df = pd.DataFrame(issues_list)

# 提取comments和labels中的实际数据列表(丢弃pageInfo)
df["comments"] = df["comments"].apply(lambda x: x["nodes"])
df["labels"] = df["labels"].apply(lambda x: x["nodes"])

2. 展开评论列

将每个评论拆分为单独行,再解析评论的嵌套字段:

# 展开comments列表,每条评论对应一行
df_comments_expanded = df.explode("comments", ignore_index=True)
# 将评论的嵌套字典展开为独立列
df_comments = pd.json_normalize(df_comments_expanded["comments"])
# 合并基础Issue数据与展开的评论字段
df_step1 = pd.concat([df_comments_expanded.drop("comments", axis=1), df_comments], axis=1)

3. 展开标签列

用同样逻辑处理标签字段:

# 展开labels列表,每个标签对应一行
df_labels_expanded = df_step1.explode("labels", ignore_index=True)
# 将标签的嵌套字典展开为独立列
df_labels = pd.json_normalize(df_labels_expanded["labels"])
# 合并得到最终展平后的DataFrame
final_df = pd.concat([df_labels_expanded.drop("labels", axis=1), df_labels], axis=1)

4. 最终效果

final_df会生成每个Issue的每条评论+每个标签的组合行,比如第一个Issue有2条评论×2个标签,会生成4行数据,完全实现嵌套结构的展平。

关键注意点

  • 必须先从comments/labels对象中提取nodes列表,否则直接解析会处理分页信息而非实际数据
  • explode用于将列表类型的列拆分为多行,是实现行扩展的核心
  • pd.json_normalize比apply(pd.Series)更适合解析多层嵌套字典,稳定性更高

内容的提问来源于stack exchange,提问作者acdeicnr

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 01:20:22