You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中Pandas多列条件匹配的Index Match功能实现方法

多列条件匹配生成DataFrame新列实现方案

原有实现存在三个核心问题:

  • 映射规则和需求不符:需求的匹配条件是年份+type,但你定义的映射表把id作为匹配维度,导致id为cc的行全部匹配失败
  • 缺少年份提取逻辑:原date列是带季度前缀的字符串(如q1 2021),直接拿2021做完整匹配无法命中
  • 方法选择错误:replace仅适用于单列值替换,不支持多列联合匹配,同时代码中混用df和df1两个变量名也会直接报错

方案1:np.select(性能最优,适合大数据量)

import pandas as pd
import numpy as np

# 原始数据读取后直接使用即可,这里是示例构造代码
df = pd.DataFrame({
    'id': ['aa', 'bb', 'cc', 'cc', 'cc'],
    'date': ['q1 2021', 'q1 2022', 'q1 2021', 'q1 2022', 'q1 2023'],
    'type': ['aa', 'aa', 'aa', 'aa', 'bb']
})

# 提取date列末尾的年份
df['year'] = df['date'].str.split().str[-1]

# 按需求定义匹配条件和对应取值
conditions = [
    (df['year'] == '2021') & (df['type'] == 'aa'),
    (df['year'] == '2022') & (df['type'] == 'aa'),
    (df['year'] == '2023') & (df['type'] == 'bb')
]
values = [10, 20, 50]

# 生成source列,未匹配到的默认返回0,可自行修改default参数
df['source'] = np.select(conditions, values, default=0)

# 删除临时生成的year列
df.drop('year', axis=1, inplace=True)

print(df)

方案2:apply自定义函数(写法直观,适合小数据量)

def map_source(row):
    # 逐行提取年份
    year = row['date'].split()[-1]
    if year == '2021' and row['type'] == 'aa':
        return 10
    elif year == '2022' and row['type'] == 'aa':
        return 20
    elif year == '2023' and row['type'] == 'bb':
        return 50
    # 未匹配的返回值可自行修改
    return None

df['source'] = df.apply(map_source, axis=1)

两种方案运行后都可以得到你期望的输出结果。


内容的提问来源于stack exchange,提问作者Lynn

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.26 06:15:03