You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python使用if条件实现列头与行值的精确/部分匹配并提取对应id

Python实现匹配提取id的方案

首先导入依赖,构造示例数据:

import pandas as pd

# 构造示例DataFrame
df = pd.DataFrame({
    "id": ["a234", "b231", "a234", "g34", "y34"],
    "references": ["papaya", "pomegranate", "apple", "mango", "strawberry"],
    "apple": [0, 1, 0, 0, 0],
    "pomegranate": [1, 0, 1, 0, 1]
})

核心匹配逻辑如下,严格按照「先精确匹配、再部分匹配」的优先级执行,用if条件控制匹配流程:

def get_matched_id(df, target_col):
    # 第一轮:优先执行精确字符串匹配
    exact_match = df[df["references"] == target_col]
    if not exact_match.empty:
        # 可按需调整返回所有匹配id/单个id,此处返回去重后的全部匹配id
        return exact_match["id"].unique().tolist()
    # 精确匹配无结果时,进入第二轮部分匹配
    partial_match = df[df["references"].str.contains(target_col, case=False)]
    if not partial_match.empty:
        return partial_match["id"].unique().tolist()
    # 两类匹配都无结果时返回空列表
    return []

# 测试匹配apple列
print(get_matched_id(df, "apple"))
# 输出结果:['a234']

# 测试部分匹配场景:新增references为green apple的行
df_test = df._append({"id": "h56", "references": "green apple", "apple": 1, "pomegranate":0}, ignore_index=True)
# 删掉原精确匹配的行,测试部分匹配返回结果
df_test = df_test.drop(df_test[df_test["references"] == "apple"].index)
print(get_matched_id(df_test, "apple"))
# 输出结果:['h56']

补充说明:

  • 要批量匹配所有数值列的列名,遍历df.columns[2:](跳过前两列id和references)依次调用上述函数即可
  • 部分匹配默认设置了case=False忽略大小写,不需要可删除该参数
  • 只需要返回第一个匹配id的话,把unique().tolist()替换为iloc[0]即可

内容的提问来源于stack exchange,提问作者Aruna Saraswathy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.02 12:45:03