You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

合并PySpark转Pandas DataFrame中含数组的列并规整数据

PySpark转Pandas后的数据规整方案

实现步骤

1. 拆分处理answer与question字段组

由于answer和question的字段是并行数组结构(其中paragraphlabels为嵌套数组),我们可以分别处理这两部分,再合并结果:

处理answer部分

先提取answer相关字段并重命名,再分两次展开数组——先处理外层数组确保Line与对应Label、Text组关联,再展开嵌套数组让每个标签和文本对应单独一行:

import pandas as pd

# 提取answer字段并改名
answer_df = df[['answer.end', 'answer.paragraphlabels', 'answer.text']].rename(
    columns={
        'answer.end': 'Line',
        'answer.paragraphlabels': 'Label',
        'answer.text': 'Text'
    }
)

# 第一次展开:把外层数组拆成行,保证Line和对应的Label/Text组对应
answer_df = answer_df.explode(['Line', 'Label', 'Text'])

# 第二次展开:把嵌套的Label和Text数组拆成单独行
answer_df = answer_df.explode(['Label', 'Text'])

处理question部分

用同样的逻辑处理question字段:

# 提取question字段并改名
question_df = df[['question.end', 'question.paragraphlabels', 'question.text']].rename(
    columns={
        'question.end': 'Line',
        'question.paragraphlabels': 'Label',
        'question.text': 'Text'
    }
)

# 分步展开数组
question_df = question_df.explode(['Line', 'Label', 'Text'])
question_df = question_df.explode(['Label', 'Text'])

2. 合并结果并整理

将处理后的两部分数据合并,重置索引后就得到你需要的规整表格:

# 合并answer和question的结果
final_df = pd.concat([answer_df, question_df], ignore_index=True)

# 调整列顺序并重置索引
final_df = final_df[['Line', 'Label', 'Text']].reset_index(drop=True)

方案说明

  • 为什么不用np.concatenate直接拼列:这种方式会把数组强行拉平,但丢失了Line与Label、Text的对应关系,必然导致数据混乱。
  • 分步explode的优势:先关联外层数组的对应关系,再展开嵌套数组,能彻底避免数据错位,保证每一行的Line、Label、Text都是严格对应的。
  • 这个方案能完整保留所有原始数据,不管answer和question的数组长度是否一致,都能正确展开。

内容的提问来源于stack exchange,提问作者Мария Кашталян

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 01:55:29