You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中如何将含嵌套origins列表的字典列表转换为DataFrame

嵌套列表转Pandas DataFrame展开origins字段方案

问题说明

现有结构如下的Python列表对象,直接使用pd.DataFrame(myList)转换时,origins列会存储列表类型数据,需要将origins嵌套列表内origin、quantityLeads两个键对应的值展开,和其他外层字段存入同一DataFrame。
注意:原始列表第一个字典的created_date字段后缺失逗号,直接运行会报语法错误,后续代码已补全该问题。

实现方法

方法1:使用pandas explode方法(适合快速实现,pandas 0.25+版本支持)

核心逻辑是先将嵌套的origins列表拆分为多行,再将每行的字典结构展开为独立列,最后删除冗余的原origins列。

import pandas as pd

# 补全语法错误后的原始数据
myList = [
   {
      "id":3105052,
      "title":"Ebook Relatórios Gerenciais",
      "offering":"Institucional",
      "created_date":"2022-06-28",
      "inserted_date":"2022-06-28",
      "channel":"Social",
      "start_date":"2022-06-28",
      "end_date":"2022-06-28",
      "origins":[
         {"origin":"LinkedIn", "quantityLeads":"1"},
         {"origin":"Facebook", "quantityLeads":"1"}
      ]
   },
   {
      "id":3105052,
      "title":"Ebook Relatórios Gerenciais",
      "offering":"Institucional",
      "inserted_date":"2022-06-28",
      "created_date":"2022-06-28",
      "channel":"Direct",
      "start_date":"2022-06-28",
      "end_date":"2022-06-28",
      "origins":[{"origin":"Desconhecida", "quantityLeads":"2"}]
   },
   {
      "id":2918513,
      "title":"Ebook Direct To Consumer",
      "offering":"Supply Chain",
      "created_date":"2022-06-28",
      "inserted_date":"2022-06-28",
      "channel":"Social",
      "start_date":"2022-06-28",
      "end_date":"2022-06-28",
      "origins":[{"origin":"LinkedIn", "quantityLeads":"1"}]
   }
]

df = pd.DataFrame(myList)
# 拆分origins嵌套列表为独立行
df = df.explode('origins', ignore_index=True)
# 将origins字段下的字典拆分为独立列
df[['origin', 'quantityLeads']] = df['origins'].apply(pd.Series)
# 删除冗余的原origins列
df = df.drop(columns=['origins'])

# 可选:将quantityLeads转换为整数类型
df['quantityLeads'] = df['quantityLeads'].astype(int)

方法2:预处理扁平化数据(适合大数据量场景,性能更优)

如果数据量较大,apply(pd.Series)的运行效率较低,可以先在原生Python层把嵌套结构扁平化,再直接传入DataFrame构造函数:

flat_records = []
for record in myList:
    # 提取外层公共字段,排除origins
    base_fields = {k:v for k, v in record.items() if k != 'origins'}
    # 遍历每个来源条目,和公共字段合并为单条记录
    for origin_info in record['origins']:
        flat_records.append({**base_fields, **origin_info})

df = pd.DataFrame(flat_records)
# 可选类型转换
df['quantityLeads'] = df['quantityLeads'].astype(int)

最终输出效果

两种方法得到的DataFrame结构完全一致,共4条记录:

  • 外层字段id/title/offering等会按照origins内的条目数自动重复填充
  • 新增origin和quantityLeads列存储展开后的值,无嵌套结构
idtitleofferingcreated_dateinserted_datechannelstart_dateend_dateoriginquantityLeads
3105052Ebook Relatórios GerenciaisInstitucional2022-06-282022-06-28Social2022-06-282022-06-28LinkedIn1
3105052Ebook Relatórios GerenciaisInstitucional2022-06-282022-06-28Social2022-06-282022-06-28Facebook1
3105052Ebook Relatórios GerenciaisInstitucional2022-06-282022-06-28Direct2022-06-282022-06-28Desconhecida2
2918513Ebook Direct To ConsumerSupply Chain2022-06-282022-06-28Social2022-06-282022-06-28LinkedIn1

内容的提问来源于stack exchange,提问作者Ítalo Magalhães

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.27 19:57:11