Gensim keywords函数pos_filter参数无效问题及正确用法咨询
问题:Gensim keywords() 的 pos_filter 参数未按预期生效
我帮你梳理了问题的核心原因,以及对应的解决方案:
首先,你的代码存在两个关键语法/参数格式错误
- 第二个
keywords()调用里,pos_filter('NN','JJ')少了赋值符号=,应该写成pos_filter=('NN','JJ'),否则代码会报错,根本无法正常执行。 - 当传递单个POS标签时,你写的
pos_filter=('NN')会被Python识别成字符串,而非元组——正确的写法是pos_filter=('NN',)(注意末尾的逗号),这才是函数期望的参数格式。
这两个错误是导致参数完全不生效的主要原因。
修正后的验证代码
先把参数问题修复,再测试效果:
from gensim.summarization import keywords import requests url = 'https://www.nytimes.com/2018/05/16/opinion/ramadan-spirit-america.html' text = requests.get(url).text # 仅保留名词(注意元组格式) print("=== 仅保留名词 ===") print(keywords(text, words=15, pos_filter=('NN',), lemmatize=True, scores=True)) # 保留名词和形容词 print("\n=== 保留名词+形容词 ===") print(keywords(text, words=15, pos_filter=('NN', 'JJ'), lemmatize=True, scores=True)) # 仅保留形容词 print("\n=== 仅保留形容词 ===") print(keywords(text, words=15, pos_filter=('JJ',), lemmatize=True, scores=True))
额外提醒:Gensim summarization 模块已被弃用
gensim.summarization 属于旧版模块,已经被官方标记为「已弃用」,后续可能不再维护,也存在一些已知的POS过滤逻辑bug。如果需要更稳定、精准的关键词提取,推荐使用 spaCy 这类现代NLP工具:
import spacy from collections import Counter import requests nlp = spacy.load("en_core_web_sm") url = 'https://www.nytimes.com/2018/05/16/opinion/ramadan-spirit-america.html' text = requests.get(url).text doc = nlp(text) # 提取名词(含普通名词、专有名词) nouns = [token.lemma_ for token in doc if token.pos_ in ['NOUN', 'PROPN'] and not token.is_stop and not token.is_punct] print("=== 名词关键词 ===") print(Counter(nouns).most_common(15)) # 提取形容词 adjectives = [token.lemma_ for token in doc if token.pos_ == 'ADJ' and not token.is_stop and not token.is_punct] print("\n=== 形容词关键词 ===") print(Counter(adjectives).most_common(15))
这个方案能更精准地控制词性过滤,且标注准确率更高。
内容的提问来源于stack exchange,提问作者Nandani
相关产品推荐
相关产品推荐

