如何展平DataFrame中的字典并拼接所有结果行
解决方案
问题原因
你遇到的KeyError: 'node'是因为GitHub GraphQL返回的comments和labels是分页结构对象,里面包含nodes(实际的评论/标签列表)和pageInfo(分页信息),直接用apply(pd.Series)解析整个对象会找不到预期字段,得先提取nodes里的列表再展开。
具体实现步骤
假设你从API获取的原始数据结构类似GitHub GraphQL返回格式(模拟示例):
# 模拟GitHub GraphQL返回的Issue数据 raw_data = { "data": { "repository": { "issues": { "nodes": [ { "number": 1, "title": "测试Issue 1", "body": "这是第一个测试Issue", "comments": { "nodes": [ {"id": "C_1", "author": {"login": "user1"}, "body": "评论1"}, {"id": "C_2", "author": {"login": "user2"}, "body": "评论2"} ], "pageInfo": {"endCursor": "...", "hasNextPage": False} }, "labels": { "nodes": [ {"id": "L_1", "name": "bug", "color": "ff0000"}, {"id": "L_2", "name": "help wanted", "color": "00ff00"} ], "pageInfo": {"endCursor": "...", "hasNextPage": False} } }, { "number": 2, "title": "测试Issue 2", "body": "这是第二个测试Issue", "comments": { "nodes": [{"id": "C_3", "author": {"login": "user3"}, "body": "评论3"}], "pageInfo": {"endCursor": "...", "hasNextPage": False} }, "labels": { "nodes": [{"id": "L_3", "name": "enhancement", "color": "0000ff"}], "pageInfo": {"endCursor": "...", "hasNextPage": False} } } ] } } } }
1. 转换初始DataFrame
先提取Issue基础数据,剥离分页信息:
import pandas as pd # 提取Issue列表 issues_list = raw_data["data"]["repository"]["issues"]["nodes"] df = pd.DataFrame(issues_list) # 提取comments和labels中的实际数据列表(丢弃pageInfo) df["comments"] = df["comments"].apply(lambda x: x["nodes"]) df["labels"] = df["labels"].apply(lambda x: x["nodes"])
2. 展开评论列
将每个评论拆分为单独行,再解析评论的嵌套字段:
# 展开comments列表,每条评论对应一行 df_comments_expanded = df.explode("comments", ignore_index=True) # 将评论的嵌套字典展开为独立列 df_comments = pd.json_normalize(df_comments_expanded["comments"]) # 合并基础Issue数据与展开的评论字段 df_step1 = pd.concat([df_comments_expanded.drop("comments", axis=1), df_comments], axis=1)
3. 展开标签列
用同样逻辑处理标签字段:
# 展开labels列表,每个标签对应一行 df_labels_expanded = df_step1.explode("labels", ignore_index=True) # 将标签的嵌套字典展开为独立列 df_labels = pd.json_normalize(df_labels_expanded["labels"]) # 合并得到最终展平后的DataFrame final_df = pd.concat([df_labels_expanded.drop("labels", axis=1), df_labels], axis=1)
4. 最终效果
final_df会生成每个Issue的每条评论+每个标签的组合行,比如第一个Issue有2条评论×2个标签,会生成4行数据,完全实现嵌套结构的展平。
关键注意点
- 必须先从
comments/labels对象中提取nodes列表,否则直接解析会处理分页信息而非实际数据 explode用于将列表类型的列拆分为多行,是实现行扩展的核心pd.json_normalize比apply(pd.Series)更适合解析多层嵌套字典,稳定性更高
内容的提问来源于stack exchange,提问作者acdeicnr
相关产品推荐
相关产品推荐

