You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何跨DataFrame执行模糊匹配,计算元素得分并实现职业分类

需求实现:基于模糊匹配的职位归类

现有数据结构

df1(职位基准表)

MarketingSalesIT
marketing managersales leadsoftware eng
marketing specsales mgrdata scientist

df2(待归类职位表)

ProfessionJob Title
ITdata science manager
MarketingMarketing manager

需求说明

将df2中的每个Job Title与df1的每一行、每一列元素做模糊匹配,计算匹配得分;为每个Job Title生成对应得分表,找出得分≥90的最高分所属列(职业类别),以此将该Job Title归类到对应职业。

实现方案

使用fuzzywuzzy库的fuzz.ratio()计算模糊匹配得分,Python代码实现如下:

import pandas as pd
from fuzzywuzzy import fuzz

# 初始化数据
df1 = pd.DataFrame({
    'Marketing': ['marketing manager', 'marketing spec'],
    'Sales': ['sales lead', 'sales mgr'],
    'IT': ['software eng', 'data scientist']
})

df2 = pd.DataFrame({
    'Profession': ['IT', 'Marketing'],
    'Job Title': ['data science manager', 'Marketing manager']
})

# 计算模糊匹配得分
def get_match_scores(job_title, base_df):
    return base_df.applymap(lambda x: fuzz.ratio(job_title.lower(), x.lower()))

# 遍历处理每个待归类职位
for _, row in df2.iterrows():
    job_title = row['Job Title']
    print(f"### {job_title} 匹配得分表")
    score_table = get_match_scores(job_title, df1)
    print(score_table.to_markdown(index=False))
    
    # 筛选并确定归类职业
    col_max_scores = score_table.max()
    qualified_cols = col_max_scores[col_max_scores >= 90]
    if not qualified_cols.empty:
        target_profession = qualified_cols.idxmax()
        print(f"该职位归类到:**{target_profession}**")
    else:
        print("无得分≥90的匹配项,无法归类")
    print("\n")

输出结果示例

data science manager 匹配得分表

MarketingSalesIT
50030
0091

该职位归类到:IT

Marketing manager 匹配得分表

MarketingSalesIT
10000
5500

该职位归类到:Marketing

内容的提问来源于stack exchange,提问作者Suraj Bhala

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.25 07:45:51