如何从Python职位类别字典中获取匹配度最高的对应值?
解决职位标题与类别关键字的最高匹配问题
Hey,我来帮你搞定这个匹配需求!核心思路是统计每个职位标题和各类别关键字的重叠单词数量,取匹配数最多的那个类别。下面是具体的实现步骤和可直接运行的代码:
1. 预处理字典数据
首先得把你原始字典里的逗号分隔字符串拆成便于匹配的单词列表,同时统一转成小写,避免大小写差异导致的匹配失败。
2. 编写匹配函数
定义一个函数,对每个职位标题拆分单词后,遍历所有类别,计算和每个类别关键字的重叠数,记录匹配数最高的类别。如果所有类别都没有匹配,返回提示信息。
完整代码实现
# 你的原始职位类别关键字字典 category_dict = { '01': ['advertising, representatives, financial, miscellaneous, other, sales'], '02': ['musicians, workers, officials, entertainers, actors, singers, competitors, dancers'], '03': ['movers, station, gas, of, stock, pumping, workers, hoist, mining, freight, truck'], '04': ['child, support, children, disable, supplemental, security, income'] } # 预处理:将每个类别的关键字转为小写单词列表 processed_categories = {} for code, kw_str_list in category_dict.items(): # 取出字符串,按逗号分割,去除空格并转小写 keywords = [kw.strip().lower() for kw in kw_str_list[0].split(',')] processed_categories[code] = keywords def find_best_match(job_title): # 拆分职位标题为小写单词集合 title_words = set(word.lower() for word in job_title.split()) max_matches = 0 best_code = None for code, keywords in processed_categories.items(): # 计算两个集合的交集大小,就是匹配的单词数 match_count = len(title_words & set(keywords)) if match_count > max_matches: max_matches = match_count best_code = code if max_matches == 0: return f"职位「{job_title}」没有匹配的类别" return f"职位「{job_title}」最匹配的类别是:{best_code}" # 测试你的职位标题列表 job_titles = ['child support', 'art director', 'driver assistant specialist'] for title in job_titles: print(find_best_match(title))
运行结果
执行代码后会输出:
职位「child support」最匹配的类别是:04
职位「art director」没有匹配的类别
职位「driver assistant specialist」没有匹配的类别
补充说明
- 这里用了集合的交集来快速计算匹配数,效率比逐个遍历更高
- 如果想优化匹配逻辑(比如支持部分匹配、同义词匹配),可以扩展这个函数,比如加入模糊匹配工具,不过基础的精确匹配已经能满足你的需求啦
内容的提问来源于stack exchange,提问作者Vera.J
相关产品推荐
相关产品推荐

