如何用Python的RAKE库精准提取招聘描述中的技术关键词
问题描述
需要从以下LinkedIn招聘描述中提取Numpy、Scipy、HTML这类明确的工具、框架类技术关键词:
input_text = "In-depth understanding of the Python software development stacks, ecosystems, frameworks and tools such as Numpy, Scipy, Pandas, Dask, spaCy, NLTK, sci-kit-learn and PyTorch.Experience with front-end development using HTML, CSS, and JavaScript. Familiarity with database technologies such as SQL and NoSQL.Excellent problem-solving ability with solid communication and collaboration skills. Preferred Skills And QualificationsExperience with popular Python frameworks such as Django, Flask or Pyramid."
当前使用RAKE的代码如下,但输出包含“frameworks”“experience”等无关词汇:
from rake_nltk import Rake r = Rake() r.extract_keywords_from_text(input_text) keywords = r.get_ranked_phrases_with_scores() for score, keyword in keywords: if len(keyword.split()) == 1: # Check if the keyword is one word print(f"{keyword}: {score}")
解决方案
方法一:优化RAKE提取逻辑,精准过滤非技术词
通过自定义停用词+词性标注过滤,排除常见非技术类词汇,只保留名词类的单技术词:
from rake_nltk import Rake import nltk from nltk.corpus import stopwords from nltk.tag import pos_tag # 首次运行需下载NLTK资源 nltk.download('stopwords') nltk.download('averaged_perceptron_tagger') # 自定义停用词:添加招聘文本中常见的非技术词汇 custom_stopwords = set(stopwords.words('english')) custom_stopwords.update(['frameworks', 'experience', 'familiarity', 'skills', 'qualifications', 'ability', 'technologies']) # 初始化RAKE并使用自定义停用词 r = Rake(stopwords=custom_stopwords) r.extract_keywords_from_text(input_text) keywords = r.get_ranked_phrases_with_scores() # 过滤规则:保留单字词+名词词性+不在停用词列表中 for score, keyword in keywords: if len(keyword.split()) == 1: word_tag = pos_tag([keyword])[0] # 仅保留名词类词性(NN/NNP等) if word_tag[1] in ['NN', 'NNP', 'NNPS', 'NNS'] and keyword.lower() not in custom_stopwords: print(f"{keyword}: {score}")
方法二:用预设技术列表过滤RAKE输出
如果需要100%精准的结果,直接用预设的技术词汇列表匹配RAKE输出是最可靠的方式:
实现代码
from rake_nltk import Rake # 预设技术词汇列表 tech_keywords = { 'Python', 'Numpy', 'Scipy', 'Pandas', 'Dask', 'spaCy', 'NLTK', 'sci-kit-learn', 'PyTorch', 'HTML', 'CSS', 'JavaScript', 'SQL', 'NoSQL', 'Django', 'Flask', 'Pyramid' } r = Rake() r.extract_keywords_from_text(input_text) keywords = r.get_ranked_phrases_with_scores() # 仅保留出现在预设列表中的关键词 for score, keyword in keywords: if keyword in tech_keywords: print(f"{keyword}: {score}")
构建全面技术列表的方法
- 按领域分类收集:将技术分为后端语言/框架、前端技术、数据科学工具、数据库等类别,逐个领域整理主流技术栈;
- 参考招聘平台高频词:整理同类型岗位招聘描述中反复出现的技术词汇;
- 依托开源社区资源:参考GitHub热门仓库标签、Stack Overflow技术标签,或者Awesome系列汇总项目(如Awesome Python);
- 参考行业报告:提取Stack Overflow开发者调查报告、TIOBE编程语言排行榜中的热门技术词汇;
- 自动化扩展+人工校验:用WordNet等语义工具,基于核心技术词扩展相关词汇,再人工去重筛选。
内容的提问来源于stack exchange,提问作者Fatemeh
相关产品推荐
相关产品推荐

