合并PySpark转Pandas DataFrame中含数组的列并规整数据
PySpark转Pandas后的数据规整方案
实现步骤
1. 拆分处理answer与question字段组
由于answer和question的字段是并行数组结构(其中paragraphlabels为嵌套数组),我们可以分别处理这两部分,再合并结果:
处理answer部分
先提取answer相关字段并重命名,再分两次展开数组——先处理外层数组确保Line与对应Label、Text组关联,再展开嵌套数组让每个标签和文本对应单独一行:
import pandas as pd # 提取answer字段并改名 answer_df = df[['answer.end', 'answer.paragraphlabels', 'answer.text']].rename( columns={ 'answer.end': 'Line', 'answer.paragraphlabels': 'Label', 'answer.text': 'Text' } ) # 第一次展开:把外层数组拆成行,保证Line和对应的Label/Text组对应 answer_df = answer_df.explode(['Line', 'Label', 'Text']) # 第二次展开:把嵌套的Label和Text数组拆成单独行 answer_df = answer_df.explode(['Label', 'Text'])
处理question部分
用同样的逻辑处理question字段:
# 提取question字段并改名 question_df = df[['question.end', 'question.paragraphlabels', 'question.text']].rename( columns={ 'question.end': 'Line', 'question.paragraphlabels': 'Label', 'question.text': 'Text' } ) # 分步展开数组 question_df = question_df.explode(['Line', 'Label', 'Text']) question_df = question_df.explode(['Label', 'Text'])
2. 合并结果并整理
将处理后的两部分数据合并,重置索引后就得到你需要的规整表格:
# 合并answer和question的结果 final_df = pd.concat([answer_df, question_df], ignore_index=True) # 调整列顺序并重置索引 final_df = final_df[['Line', 'Label', 'Text']].reset_index(drop=True)
方案说明
- 为什么不用
np.concatenate直接拼列:这种方式会把数组强行拉平,但丢失了Line与Label、Text的对应关系,必然导致数据混乱。 - 分步
explode的优势:先关联外层数组的对应关系,再展开嵌套数组,能彻底避免数据错位,保证每一行的Line、Label、Text都是严格对应的。 - 这个方案能完整保留所有原始数据,不管answer和question的数组长度是否一致,都能正确展开。
内容的提问来源于stack exchange,提问作者Мария Кашталян
相关产品推荐
相关产品推荐

