如何在Pandas DataFrame嵌套列表中过滤空Answers元素?
问题:过滤Pandas嵌套DataFrame中的空Answers元素并解决json_normalize报错
现有包含嵌套列表结构的Pandas DataFrame,需过滤其中空的Answers数组。已实现过滤空exResponse.pb_sentence元素,但处理Answers时,执行pd.json_normalize(resp['Answers'])触发报错:AttributeError: 'list' object has no attribute 'keys',推测是空Answers数组导致。
示例数据
{ "examples": [ { "website": "info", "df": [ { "Question": "What?", "Answers": [] }, { "Question": "how?", "Answers": [] }, { "Question": "Why?", "Answers": [] } ], "whitelisted_url": true, "exResponse": { "pb_sentence": "", "solution_sentence": "", "why_sentence": "" } }, { "website": "info2", "df": [ { "Question": "What?", "Answers": ["example answer1"] }, { "Question": "how?", "Answers": ["example answer1"] }, { "Question": "Why?", "Answers": [] } ], "whitelisted_url": true, "exResponse": { "pb_sentence": "", } } ] }
现有过滤函数
def filter(data, name): resp = pd.concat([pd.DataFrame(data), pd.json_normalize(data['examples'])], axis=1) resp = pd.concat([pd.DataFrame(resp), pd.json_normalize(resp['df'])], axis=1) resp['exResponse.pb_sentence'].replace( '', np.nan, inplace=True) resp.dropna( subset=['exResponse.pb_sentence'], inplace=True) resp.drop(resp[resp['df.Answers'].apply(len) == 0].index, inplace=True)
报错代码段及报错行
resp.rename( columns={0: 'Challenge', 1: 'Solution', 2: 'Importance'}, inplace=True) # challenge deserializing resp = pd.concat([pd.DataFrame(df_resp), pd.json_normalize(resp['Challenge'])], axis=1) resp = pd.concat([pd.DataFrame(resp), pd.json_normalize(resp['Answers'])], axis=1)
报错行:
29 resp = pd.concat([pd.DataFrame(resp), ---> 30 pd.json_normalize(resp['Answers'])], 31 axis=1)
解决方案
核心问题分析
pd.json_normalize仅支持展平字典或字典列表结构,而你的Answers列存储的是字符串列表(或空列表),直接传入会触发类型错误。同时原代码未正确展开嵌套的df列表,导致过滤逻辑无法精准作用到每个问答对。
修正后的过滤逻辑
先展开嵌套的df列表,拆分出独立的Question和Answers列,再执行过滤:
import pandas as pd import numpy as np def filter_data(data): # 1. 加载examples数据并展开嵌套的df列表,每个问答对单独成行 df_examples = pd.json_normalize(data['examples']) df_exploded = df_examples.explode('df', ignore_index=True) # 2. 拆分df中的Question和Answers为单独列 df_qa = pd.json_normalize(df_exploded['df']) resp = pd.concat([df_exploded.drop('df', axis=1), df_qa], axis=1) # 3. 过滤空的exResponse.pb_sentence resp['exResponse.pb_sentence'] = resp['exResponse.pb_sentence'].replace('', np.nan) resp = resp.dropna(subset=['exResponse.pb_sentence']) # 4. 过滤空的Answers数组(仅保留长度大于0的列表) resp = resp[resp['Answers'].apply(len) > 0] return resp
处理后续的Answers字段
过滤后Answers均为非空字符串列表,若需将其转为字符串列而非列表(示例中为单元素列表),可直接提取列表元素:
# 提取Answers列表的第一个元素转为字符串列 resp['Answer'] = resp['Answers'].str[0] # 后续直接使用Answer列即可,无需再调用json_normalize
原代码报错原因
原代码中pd.json_normalize(resp['Answers'])传入的是包含空列表/字符串列表的Series,而json_normalize无法处理纯列表结构。通过先展开嵌套列表、过滤空数组,再将列表元素转为字符串列,即可彻底避免该错误。
内容的提问来源于stack exchange,提问作者Serkan Gün
相关产品推荐
相关产品推荐

