You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas中实现多条件匹配结果拼接生成新列的方法

解决np.select仅返回单个匹配项的问题

当前使用np.select生成reference列时,若codes_desc包含多个匹配代码,只会返回最后一个匹配的描述。要实现收集所有匹配描述并按序号拼接的需求,可参考以下两种方法:

方法一:逐行遍历(易理解,适合小数据量)

先将代码与对应描述整理为字典,再对每行文本检查匹配情况:

import pandas as pd
import numpy as np

# 代码与对应描述的映射字典,便于维护
code_mapping = {
    'R27': 'The person is a Ninja',
    'R38': 'The person is a Pirate',
    'R52': 'The person is a Doctor',
    'R62': 'The person is a Samurai',
    'R21': 'The person is a Admiral',
    'R22': 'The person is a Police',
    'R23': 'The person is a Teacher',
    'R57': 'The person is a Singer',
    'R82': 'The person is a Guitarist',
    'R86': 'The person is a Chef',
    'R20': 'The person is a Runner',
    'R98': 'The person is a Wizard'
}

col = 'codes_desc'

# 定义处理单文本的函数
def get_matched_references(text):
    matched = []
    # 按字典顺序遍历,记录匹配项并添加序号
    for idx, (code, desc) in enumerate(code_mapping.items(), start=1):
        if pd.notna(text) and code.lower() in text.lower():
            matched.append(f"{idx}. '{desc}'")
    return '\n'.join(matched) if matched else 'Reason Unknown'

# 应用到DataFrame生成reference列
df_merged["reference"] = df_merged[col].apply(get_matched_references)

方法二:向量化处理(效率更高,适合大数据量)

通过生成布尔矩阵批量匹配,再格式化结果:

import pandas as pd
import numpy as np

code_mapping = {
    'R27': 'The person is a Ninja',
    'R38': 'The person is a Pirate',
    'R52': 'The person is a Doctor',
    'R62': 'The person is a Samurai',
    'R21': 'The person is a Admiral',
    'R22': 'The person is a Police',
    'R23': 'The person is a Teacher',
    'R57': 'The person is a Singer',
    'R82': 'The person is a Guitarist',
    'R86': 'The person is a Chef',
    'R20': 'The person is a Runner',
    'R98': 'The person is a Wizard'
}

col = 'codes_desc'

# 生成布尔匹配矩阵:每列对应一个描述,值为该行是否匹配对应代码
matches_df = pd.DataFrame({
    desc: df_merged[col].str.contains(code, case=False, na=False)
    for code, desc in code_mapping.items()
})

# 格式化每行的匹配结果
def format_matches(row):
    # 筛选出当前行匹配的描述,添加序号
    matched_items = [
        f"{idx+1}. '{desc}'" 
        for idx, desc in enumerate(row.index[row])
    ]
    return '\n'.join(matched_items) if matched_items else 'Reason Unknown'

df_merged["reference"] = matches_df.apply(format_matches, axis=1)

两种方法最终都会生成类似以下格式的结果:

  1. 'The person is a Ninja'
  2. 'The person is a Police'

内容的提问来源于stack exchange,提问作者Ash

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 22:25:26