如何将含字典列表的Pandas DataFrame行展开并保留关联字段?
问题描述
现有一个包含date、person等字段的Pandas DataFrame,person列的值为字典列表,格式如下:
date person 0 2002-09-04 [{'name':'anna', 'weight':'2.9', 'hospital':'x'}, {'name': 'jacob', ...}, ...] 1 2002-10-16 [{'name':'lynn', 'weight':'3.0', 'hospital':'y'}, {'name': 'tony', ...}, ...]
希望转换为如下格式的新DataFrame:
date name weight hospital 0 2002-09-04 anna 2.9 x 1 2002-09-04 jacob ... ... n 2002-10-16 lynn 3.0 y n1 2002-10-16 tony ... ...
目前已通过以下代码将person列的字典列表展开为新DataFrame,但无法保留对应的date等关联字段:
df_person = pd.DataFrame() for row, _ in enumerate(df['person']): df_person = df_person.append(df['person'][row], ignore_index = True, sort = False)
解决方案
方法一:使用explode + json_normalize(推荐)
这是Pandas 0.25+版本支持的高效写法,无需循环,代码简洁且性能更优:
# 将person列的列表拆分成多行,每行对应一个字典 df_exploded = df.explode('person') # 将字典列解析为单独字段,并关联原date列 result = pd.json_normalize(df_exploded['person']).join(df_exploded['date']) # 调整列顺序,把date放到首位 result = result[['date', 'name', 'weight', 'hospital']]
方法二:改进循环代码
如果坚持用循环实现,需要在每次循环时给每个person字典添加对应的date字段,再合并:
dfs = [] for idx, row in df.iterrows(): # 给当前行的每个person字典追加date字段 person_list = [dict(p, date=row['date']) for p in row['person']] dfs.append(pd.DataFrame(person_list)) # 合并所有子DataFrame df_person = pd.concat(dfs, ignore_index=True) # 调整列顺序 df_person = df_person[['date', 'name', 'weight', 'hospital']]
注意:原代码中使用的
append方法在Pandas 2.0+已标记为过时,改用pd.concat能获得更好的性能。
内容的提问来源于stack exchange,提问作者Jana
相关产品推荐
相关产品推荐

