You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python的RAKE库精准提取招聘描述中的技术关键词

问题描述

需要从以下LinkedIn招聘描述中提取Numpy、Scipy、HTML这类明确的工具、框架类技术关键词:

input_text = "In-depth understanding of the Python software development stacks, ecosystems, frameworks and tools such as Numpy, Scipy, Pandas, Dask, spaCy, NLTK, sci-kit-learn and PyTorch.Experience with front-end development using HTML, CSS, and JavaScript.
Familiarity with database technologies such as SQL and NoSQL.Excellent problem-solving ability with solid communication and collaboration skills.
Preferred Skills And QualificationsExperience with popular Python frameworks such as Django, Flask or Pyramid."

当前使用RAKE的代码如下,但输出包含“frameworks”“experience”等无关词汇:

from rake_nltk import Rake

r = Rake()
r.extract_keywords_from_text(input_text)
keywords = r.get_ranked_phrases_with_scores()

for score, keyword in keywords:
    if len(keyword.split()) == 1:  # Check if the keyword is one word
        print(f"{keyword}: {score}")
解决方案

方法一:优化RAKE提取逻辑,精准过滤非技术词

通过自定义停用词+词性标注过滤,排除常见非技术类词汇,只保留名词类的单技术词:

from rake_nltk import Rake
import nltk
from nltk.corpus import stopwords
from nltk.tag import pos_tag

# 首次运行需下载NLTK资源
nltk.download('stopwords')
nltk.download('averaged_perceptron_tagger')

# 自定义停用词:添加招聘文本中常见的非技术词汇
custom_stopwords = set(stopwords.words('english'))
custom_stopwords.update(['frameworks', 'experience', 'familiarity', 'skills', 'qualifications', 'ability', 'technologies'])

# 初始化RAKE并使用自定义停用词
r = Rake(stopwords=custom_stopwords)

r.extract_keywords_from_text(input_text)
keywords = r.get_ranked_phrases_with_scores()

# 过滤规则:保留单字词+名词词性+不在停用词列表中
for score, keyword in keywords:
    if len(keyword.split()) == 1:
        word_tag = pos_tag([keyword])[0]
        # 仅保留名词类词性(NN/NNP等)
        if word_tag[1] in ['NN', 'NNP', 'NNPS', 'NNS'] and keyword.lower() not in custom_stopwords:
            print(f"{keyword}: {score}")

方法二:用预设技术列表过滤RAKE输出

如果需要100%精准的结果,直接用预设的技术词汇列表匹配RAKE输出是最可靠的方式:

实现代码

from rake_nltk import Rake

# 预设技术词汇列表
tech_keywords = {
    'Python', 'Numpy', 'Scipy', 'Pandas', 'Dask', 'spaCy', 'NLTK', 'sci-kit-learn', 
    'PyTorch', 'HTML', 'CSS', 'JavaScript', 'SQL', 'NoSQL', 'Django', 'Flask', 'Pyramid'
}

r = Rake()
r.extract_keywords_from_text(input_text)
keywords = r.get_ranked_phrases_with_scores()

# 仅保留出现在预设列表中的关键词
for score, keyword in keywords:
    if keyword in tech_keywords:
        print(f"{keyword}: {score}")

构建全面技术列表的方法

  • 按领域分类收集:将技术分为后端语言/框架、前端技术、数据科学工具、数据库等类别,逐个领域整理主流技术栈;
  • 参考招聘平台高频词:整理同类型岗位招聘描述中反复出现的技术词汇;
  • 依托开源社区资源:参考GitHub热门仓库标签、Stack Overflow技术标签,或者Awesome系列汇总项目(如Awesome Python);
  • 参考行业报告:提取Stack Overflow开发者调查报告、TIOBE编程语言排行榜中的热门技术词汇;
  • 自动化扩展+人工校验:用WordNet等语义工具,基于核心技术词扩展相关词汇,再人工去重筛选。

内容的提问来源于stack exchange,提问作者Fatemeh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.24 21:30:19