You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas DataFrame嵌套列表中过滤空Answers元素?

问题:过滤Pandas嵌套DataFrame中的空Answers元素并解决json_normalize报错

现有包含嵌套列表结构的Pandas DataFrame,需过滤其中空的Answers数组。已实现过滤空exResponse.pb_sentence元素,但处理Answers时,执行pd.json_normalize(resp['Answers'])触发报错:AttributeError: 'list' object has no attribute 'keys',推测是空Answers数组导致。

示例数据

{
    "examples": [
        {
            "website": "info",
            "df": [
                {
                    "Question": "What?",
                    "Answers": []
                },
                {
                    "Question": "how?",
                    "Answers": []
                },
                {
                    "Question": "Why?",
                    "Answers": []
                }
            ],
            "whitelisted_url": true,
            "exResponse": {
                "pb_sentence": "",
                "solution_sentence": "",
                "why_sentence": ""
            }
        },            
         {
            "website": "info2",
            "df": [
                {
                    "Question": "What?",
                    "Answers": ["example answer1"]
                },
                {
                    "Question": "how?",
                    "Answers": ["example answer1"]
                },
                {
                    "Question": "Why?",
                    "Answers": []
                }
            ],
            "whitelisted_url": true,
            "exResponse": {
                "pb_sentence": "",
            }
        }
    ]
}

现有过滤函数

def filter(data, name):
   resp = pd.concat([pd.DataFrame(data),
                         pd.json_normalize(data['examples'])],
                        axis=1)

    resp = pd.concat([pd.DataFrame(resp),
                         pd.json_normalize(resp['df'])],
                        axis=1)

    resp['exResponse.pb_sentence'].replace(
        '', np.nan, inplace=True)
    resp.dropna(
        subset=['exResponse.pb_sentence'], inplace=True)
    

    resp.drop(resp[resp['df.Answers'].apply(len) == 0].index, inplace=True)

报错代码段及报错行

resp.rename(
        columns={0: 'Challenge', 1: 'Solution', 2: 'Importance'}, inplace=True)
    # challenge deserializing
    resp = pd.concat([pd.DataFrame(df_resp),
                         pd.json_normalize(resp['Challenge'])],
                        axis=1)
    resp = pd.concat([pd.DataFrame(resp),
                         pd.json_normalize(resp['Answers'])],
                        axis=1)

报错行:

29 resp = pd.concat([pd.DataFrame(resp),
---> 30                      pd.json_normalize(resp['Answers'])],
     31                     axis=1)

解决方案

核心问题分析

pd.json_normalize仅支持展平字典或字典列表结构,而你的Answers列存储的是字符串列表(或空列表),直接传入会触发类型错误。同时原代码未正确展开嵌套的df列表,导致过滤逻辑无法精准作用到每个问答对。

修正后的过滤逻辑

先展开嵌套的df列表,拆分出独立的Question和Answers列,再执行过滤:

import pandas as pd
import numpy as np

def filter_data(data):
    # 1. 加载examples数据并展开嵌套的df列表,每个问答对单独成行
    df_examples = pd.json_normalize(data['examples'])
    df_exploded = df_examples.explode('df', ignore_index=True)
    
    # 2. 拆分df中的Question和Answers为单独列
    df_qa = pd.json_normalize(df_exploded['df'])
    resp = pd.concat([df_exploded.drop('df', axis=1), df_qa], axis=1)
    
    # 3. 过滤空的exResponse.pb_sentence
    resp['exResponse.pb_sentence'] = resp['exResponse.pb_sentence'].replace('', np.nan)
    resp = resp.dropna(subset=['exResponse.pb_sentence'])
    
    # 4. 过滤空的Answers数组(仅保留长度大于0的列表)
    resp = resp[resp['Answers'].apply(len) > 0]
    
    return resp

处理后续的Answers字段

过滤后Answers均为非空字符串列表,若需将其转为字符串列而非列表(示例中为单元素列表),可直接提取列表元素:

# 提取Answers列表的第一个元素转为字符串列
resp['Answer'] = resp['Answers'].str[0]
# 后续直接使用Answer列即可,无需再调用json_normalize

原代码报错原因

原代码中pd.json_normalize(resp['Answers'])传入的是包含空列表/字符串列表的Series,而json_normalize无法处理纯列表结构。通过先展开嵌套列表、过滤空数组,再将列表元素转为字符串列,即可彻底避免该错误。


内容的提问来源于stack exchange,提问作者Serkan Gün

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 07:31:06