如何将DataFrame的prediction列转换为指定结构的新DataFrame
解决方法:将字典列表列转换为结构化DataFrame
先构造示例测试数据
import pandas as pd data = { 'prediction': [ [ {'answer': 'My name is Andrew.', 'text': 'What is your name ?', 'id': 1}, {'answer': 'I live at California.', 'text': 'Where do you live ?', 'id': 2}, {'answer': 'I want to become a doctor.', 'text': 'What do you want to do ?', 'id': 3} ], [ {'answer': 'My name is Julie.', 'text': 'What is your name ?', 'id': 1}, {'answer': 'I live at NY.', 'text': 'Where do you live ?', 'id': 2}, {'answer': 'I want to be a cook.', 'text': 'What do you want to do ?', 'id': 3} ] ] } df = pd.DataFrame(data)
方法一:使用explode + pivot(适合大数据量)
# 把每行的字典列表拆分为单独行 exploded_df = df.explode('prediction') # 将字典展开为独立列 expanded_df = exploded_df['prediction'].apply(pd.Series) # 透视转换为目标结构,重置索引 result_df = expanded_df.pivot( index=expanded_df.index, columns='text', values='answer' ).reset_index(drop=True)
方法二:逐行处理字典列表(代码更简洁)
# 对每行的字典列表生成单行DataFrame,再拼接所有行 result_df = pd.concat( df['prediction'].apply( lambda lst: pd.DataFrame(lst).set_index('text')['answer'].to_frame().T ), ignore_index=True )
两种方法最终都会得到你需要的结构:
What is your name ? Where do you live ? What do you want to do ? 0 My name is Andrew. I live at California. I want to become a doctor. 1 My name is Julie. I live at NY. I want to be a cook.
内容的提问来源于stack exchange,提问作者d_b
相关产品推荐
相关产品推荐

