You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于DataFrame列子串匹配为另一DataFrame新列赋值?

问题:基于DataFrame的名称匹配添加新列

原始数据

两个DataFrame定义如下:

import pandas as pd

df1 = pd.DataFrame(list(zip(['name1, Name2, name5', 'name4, name3', 'name6xx'],
                            [150, 230, 'name6xx'])),
                    columns=['name', 'compound1'])

df2 = pd.DataFrame(list(zip(['name1 ppl', 'PPL name2', 'Name3 PPL', 'name4 ppl', 'name5 ppl', 'name6xx'])), columns=['name'])

目标结果

需要为df2添加compound列,最终得到:

df2 = pd.DataFrame(list(zip(['name1 ppl', 'PPL name2', 'Name3 PPL', 'name4 ppl', 'name5 ppl', 'name6xx'],
                           [150,150,230,230,150,'name6xx'])),
                    columns=['name', 'compound'])

尝试的错误代码

以下代码未达到预期效果:

def match(name):
    cond = df1['name'].str.contains(name, case=False)[0]
    if cond:
        return df1.loc[cond, 'compound1'].values[0]
    else:
        return name
df2.compound_new = df2.name.apply(match)

正确实现方案

原代码仅检查df1的第一行,且未处理df2名称中附带的ppl等额外字符。正确思路是先构建名称到compound值的映射字典,再从df2的名称中匹配对应值:

# 1. 构建名称映射字典:将df1中逗号分隔的名称统一转小写,关联对应compound1值
name_map = {}
for _, row in df1.iterrows():
    # 拆分名称并去除空格、转小写
    names = [n.strip().lower() for n in row['name'].split(',')]
    for name in names:
        name_map[name] = row['compound1']

# 2. 为df2匹配compound值
def get_compound_value(df2_name):
    lower_name = df2_name.lower()
    # 遍历映射字典,检查名称是否包含在df2的名称中
    for key in name_map:
        if key in lower_name:
            return name_map[key]
    # 无匹配时返回原名称
    return df2_name

# 3. 添加新列
df2['compound'] = df2['name'].apply(get_compound_value)

执行后即可得到目标结果,该方案支持不区分大小写匹配,且能处理df2名称中附带的额外字符。

内容的提问来源于stack exchange,提问作者Joe

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 10:35:22