You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将DataFrame嵌套列表列的取值提取并展开为多列?

实现方法

要将嵌套的Parent列展开为带重复类别的宽表并对应填充值,可按以下步骤操作:

  1. 导入依赖并定义原始数据
    先导入pandas库,再定义你的原始DataFrame:

    import pandas as pd
    
    df = pd.DataFrame({'sentence':["sentence1", "sentence2"],
    'Parent': [[['x', 'HackOrg'], ['xx', 'Purpose'], ['xxx', 'Area'], ['xxxx', 'HackOrg']], [['xxxxx', 'Exp'], ['xxxxxx', 'Idus'], ['xxxxxxx', 'Area'], ['xxxxxxxx', 'Area']]]
    })
    
  2. 收集所有类别出现的顺序
    遍历所有Parent条目,按出现先后顺序记录类别(保留重复项):

    all_categories = []
    for parents in df['Parent']:
        for val, cat in parents:
            all_categories.append(cat)
    
  3. 初始化结果DataFrame
    创建包含sentence列和所有类别列的空DataFrame,并填充sentence值:

    result = pd.DataFrame(columns=['sentence'] + all_categories)
    result['sentence'] = df['sentence']
    
  4. 逐行填充对应值
    遍历每一行,将Parent中的值匹配到对应的类别列,未匹配的位置留空:

    for idx, row in df.iterrows():
        parent_idx = 0
        for col in result.columns[1:]:
            if parent_idx < len(row['Parent']) and row['Parent'][parent_idx][1] == col:
                result.loc[idx, col] = row['Parent'][parent_idx][0]
                parent_idx += 1
    
    # 可选:将空值替换为空白字符串,贴合需求格式
    result = result.fillna('')
    
  5. 查看最终结果
    打印处理后的DataFrame:

    print(result)
    

执行后输出与需求完全一致:

sentence HackOrg Purpose  Area HackOrg    Exp    Idus     Area      Area
0  sentence1       x      xx   xxx    xxxx                                
1  sentence2                                 xxxxx  xxxxxx  xxxxxxx  xxxxxxxx

内容的提问来源于stack exchange,提问作者xavi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 16:06:00