如何在PySpark中合并连续的多行数据?
如何合并数据集中的连续n行并保留列结构
我有如下示例数据集:
| Column A | Column B | Column C | Column D |
|---|---|---|---|
| Cell A1 | Cell B1 | Cell C1 | Cell D1 |
| Cell A2 | Cell B2 | Cell C2 | Cell D2 |
| Cell A3 | Cell B3 | Cell C3 | Cell D3 |
| Cell A4 | Cell B4 | Cell C4 | Cell D4 |
想要将连续的n行进行合并(比如n=2),保留原有列结构,得到如下结果:
| Column A | Column B | Column C | Column D |
|---|---|---|---|
| Cell A1, A2 | Cell B1, B2 | Cell C1, C2 | Cell D1, D2 |
| Cell A3, A4 | Cell B3, B4 | Cell C3, C4 | Cell D3, D4 |
解决方案(以Python pandas为例)
通过分组+聚合即可实现,核心逻辑是给每n行分配相同的分组标识,再对每个分组内的列值进行拼接合并。
步骤1:加载数据
import pandas as pd # 构造示例数据集 data = { 'Column A': ['Cell A1', 'Cell A2', 'Cell A3', 'Cell A4'], 'Column B': ['Cell B1', 'Cell B2', 'Cell B3', 'Cell B4'], 'Column C': ['Cell C1', 'Cell C2', 'Cell C3', 'Cell C4'], 'Column D': ['Cell D1', 'Cell D2', 'Cell D3', 'Cell D4'] } df = pd.DataFrame(data)
步骤2:分组聚合实现合并
设置合并行数n=2,生成分组键后对每列进行字符串拼接:
n = 2 # 生成分组键:每n行一组,如0,0,1,1... group_keys = df.index // n # 按分组键聚合,拼接每组内的列值 merged_df = df.groupby(group_keys).agg(lambda x: ', '.join(x))
执行后得到的结果:
Column A Column B Column C Column D 0 Cell A1, Cell A2 Cell B1, Cell B2 Cell C1, Cell C2 Cell D1, Cell D2 1 Cell A3, Cell A4 Cell B3, Cell B4 Cell C3, Cell C4 Cell D3, Cell D4
如果需要和示例一样仅保留单元格后缀拼接(如Cell A1, A2),可先处理单元格内容再聚合:
# 提取每个单元格的后缀部分 df_processed = df.applymap(lambda x: x.split(' ')[1]) # 聚合时拼接后缀并添加前缀 merged_df_custom = df_processed.groupby(group_keys).agg(lambda x: f'Cell {", ".join(x)}')
最终结果完全匹配需求:
Column A Column B Column C Column D 0 Cell A1, A2 Cell B1, B2 Cell C1, C2 Cell D1, D2 1 Cell A3, A4 Cell B3, B4 Cell C3, C4 Cell D3, D4
Excel实现思路
无需代码的话,可通过INDEX+CONCAT组合公式实现:
在新工作表的A1单元格输入公式:
="Cell "&INDEX(原表!A:A,ROW()*2-1)&", "&INDEX(原表!A:A,ROW()*2)
向右、向下填充公式,即可完成每2行的合并。
内容的提问来源于stack exchange,提问作者Basir Mahmood
相关产品推荐
相关产品推荐

