You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于指定列的空单元格合并DataFrame的对应相邻行

这个需求完全可以实现,你可以通过Pandas的分组聚合逻辑完成,两种实现方案如下:

基础实现(适用于Col1非空值无重复的场景)

import pandas as pd
import numpy as np

# 构造你的样例数据
df = pd.DataFrame({
    "Index": [1,2,3,4],
    "Col1": [10, np.nan, np.nan, 20],
    "Col2": ["x", "y", "z", "a"],
    "Col3": [40, 50, 60, 30]
})

# 若原数据Col1空值是空字符串而非NaN,先执行下面一行转换
# df['Col1'] = df['Col1'].replace('', np.nan)

# 空值向前填充,关联到上方最近的非空Col1值
df['Col1'] = df['Col1'].ffill()

# 按Col1分组聚合
res = df.groupby('Col1', as_index=False).agg(
    Index=('Index', 'first'), # 取每组第一个Index值
    Col2=('Col2', ','.join), # Col2值用逗号拼接
    Col3=('Col3', lambda x: ','.join(x.astype(str))) # 数字列先转字符串再拼接
)[['Index', 'Col1', 'Col2', 'Col3']] # 调整列顺序与原表一致

print(res)

输出结果:

Index  Col1   Col2      Col3
0      1    10  x,y,z  40,50,60
1      4    20      a        30

严谨实现(兼容Col1存在重复非空值的场景)

如果你的数据中Col1可能出现重复的非空值(比如后续再次出现值为10的行),可以用分组标记的方式避免不连续的同值行被错误合并:

import pandas as pd
import numpy as np

df = pd.DataFrame({
    "Index": [1,2,3,4],
    "Col1": [10, np.nan, np.nan, 20],
    "Col2": ["x", "y", "z", "a"],
    "Col3": [40, 50, 60, 30]
})

# 生成分组标记:Col1非空时计数+1,空值继承上一个计数
df['group_id'] = df['Col1'].notna().cumsum()

res = df.groupby('group_id', as_index=False).agg(
    Index=('Index', 'first'),
    Col1=('Col1', 'first'),
    Col2=('Col2', ','.join),
    Col3=('Col3', lambda x: ','.join(x.astype(str)))
).drop('group_id', axis=1) # 删除临时分组列

print(res)

内容的提问来源于stack exchange,提问作者Pearl

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.28 00:24:01