You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何无需两次合并,将单数据集index匹配至双边数据集的两国别列

问题

我有一个双边数据集,包含Country和Partner Country两列国别字段,现有数据集df和df2如下:

df

CountryPartner Countryyearvar1
TurkeySpain2019320

df2

Countryyearindex
Turkey20190
Spain20191

我希望得到如下结果,且无需执行两次合并操作:

期望输出

Countryindex_countryPartner Countryindex_partneryearvar1
Turkey0Spain12019320
解决方案

可以将df2转换为基于(Country, year)的索引映射,直接在df中生成所需的索引列,避免两次合并:

方法1:字典映射 + apply

import pandas as pd

# 构建(国家, 年份)到index的映射字典
index_mapper = df2.set_index(['Country', 'year'])['index'].to_dict()

# 直接添加两个索引列
df_result = df.assign(
    index_country=df.apply(lambda x: index_mapper[(x['Country'], x['year'])], axis=1),
    index_partner=df.apply(lambda x: index_mapper[(x['Partner Country'], x['year'])], axis=1)
)

# 调整列顺序匹配期望结果
df_result = df_result[['Country', 'index_country', 'Partner Country', 'index_partner', 'year', 'var1']]

方法2:索引定位(更高效)

如果数据集较大,推荐这种无需apply的方法,性能更优:

import pandas as pd

# 构建多索引的Series
index_series = df2.set_index(['Country', 'year'])['index']

# 通过zip匹配多索引,获取对应值
df_result = df.assign(
    index_country=index_series.loc[zip(df['Country'], df['year'])].values,
    index_partner=index_series.loc[zip(df['Partner Country'], df['year'])].values
)

# 调整列顺序
df_result = df_result[['Country', 'index_country', 'Partner Country', 'index_partner', 'year', 'var1']]

这两种方法都只需要一次数据预处理,不需要执行两次merge操作,就能得到目标结果。

内容的提问来源于stack exchange,提问作者user19562955

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 12:45:48