You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何合并两个不同长度的Pandas DataFrame生成指定结构的新DataFrame

问题描述

现有两个DataFrame:
第一个DataFrame df1:

import pandas as pd
index = [0,1,2,3,4,5]
s = pd.Series([1,2,2,3,3,6], index=index)
t = pd.Series([2,4,6,8,10,12], index=index)
df1 = pd.DataFrame(s, columns=["MUL1"])
df1["MUL2"] = t

输出结构:

MUL1  MUL2
0     1     2
1     2     4
2     2     6
3     3     8
4     3    10
5     6    12

第二个DataFrame df2:

u = pd.Series([1,2,3,6], index=[0,1,2,3])
v = pd.Series([2,8,10,12], index=[0,1,2,3])
df2 = pd.DataFrame(u, columns=["MUL3"])
df2["MUL4"] = v

输出结构:

MUL3  MUL4
0     1     2
1     2     8
2     3    10
3     6    12

需要合并得到如下结构的新DataFrame:

MUL6  MUL7
0     1     2
1     2     8
2     2     8
3     3    10
4     3    10
5     6    12

尝试了以下方法,但生成的新DataFrame长度不符合df1的长度:

X1 = df1.to_numpy()
X2 = df2.to_numpy()

list = []
for i in range(X1.shape[0]):
  for j in range(X2.shape[0]):
    if X1[i, -1] == X2[j, -1]:
      list.append(X2[X1[i, -1]==X2[j, -1], -1])
解决方案

你的需求本质是根据df1的MUL1匹配df2的MUL3,提取对应的MUL4值,不需要转numpy数组遍历,用pandas内置的方法就能轻松实现:

方法1:使用map(简洁高效)

先把df2转换成字典(MUL3为键,MUL4为值),再用df1["MUL1"]映射得到MUL7,最后重命名列即可:

# 构建映射字典
mul3_to_mul4 = df2.set_index("MUL3")["MUL4"].to_dict()
# 生成新DataFrame
result = df1[["MUL1"]].rename(columns={"MUL1": "MUL6"})
result["MUL7"] = result["MUL6"].map(mul3_to_mul4)

方法2:使用merge(适合更复杂的匹配场景)

通过MUL1和MUL3进行左连接,保留df1的所有行,再重命名列:

result = pd.merge(df1[["MUL1"]], df2, left_on="MUL1", right_on="MUL3", how="left")
result = result[["MUL1", "MUL4"]].rename(columns={"MUL1": "MUL6", "MUL4": "MUL7"})

两种方法都能得到你想要的结果,且长度和df1完全一致。

内容的提问来源于stack exchange,提问作者breez

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 22:55:48