You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中拼接DataFrame并拆分group列的实现方法求助

解决多任务学习数据集拼接问题

核心思路

先为两个DataFrame的group列赋予不同名称,再通过纵向拼接实现需求,避免merge的过度扩展或直接concat的列冲突问题。

具体实现代码

import pandas as pd

# 原始数据
df1 = pd.DataFrame({
    'sample': ['A', 'B', 'C', 'D'],
    'group': [1,0,1,0],
    'value': [123, 64, 534, 873]
})

df2 = pd.DataFrame({
    'sample': ['A', 'D', 'E'],
    'group': [1,1,0],
    'value': [372, 981, 23]
})

# 1. 重命名group列,避免拼接时冲突
df1_renamed = df1.rename(columns={'group': 'group_x'})
df2_renamed = df2.rename(columns={'group': 'group_y'})

# 2. 纵向拼接两个数据集,重置索引
df3 = pd.concat([df1_renamed, df2_renamed], axis=0, ignore_index=True)

# 查看结果
print(df3)

输出结果

sample  group_x  group_y  value
0      A      1.0      NaN    123
1      B      0.0      NaN     64
2      C      1.0      NaN    534
3      D      0.0      NaN    873
4      A      NaN      1.0    372
5      D      NaN      1.0    981
6      E      NaN      0.0     23

方法说明

  • 直接pd.concat未重命名时,会将两个group列合并为一列,无法区分来源;重命名后拼接则会保留各自的分组标识,空缺值自动填充NaN。
  • pd.merge会基于sample等列做匹配连接,导致相同样本的行被合并或扩展,不符合保留原始行结构的需求。

内容的提问来源于stack exchange,提问作者Ssong

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.21 17:48:11