Python中是否有库可生成与指定词汇搭配的动词列表?
可行的解决方案
有几个库或组合方案可以实现提取特定名词搭配动词的需求,以下是具体实现方式:
1. 使用spaCy(推荐)
spaCy的预训练语言模型支持依存句法分析和词向量语义匹配,能精准识别与目标名词搭配的动词。
示例代码:
import spacy # 加载中等规模预训练模型,大规模模型en_core_web_lg效果更优 nlp = spacy.load("en_core_web_md") target_noun = "car" # 构建包含目标名词的示例文本,也可替换为大规模语料 doc = nlp(f"A {target_noun} can be driven, crashed, repaired, washed, parked, or stolen.") verb_collocations = set() for token in doc: if token.text.lower() == target_noun: # 从依存关系中提取搭配动词(主动词或被动词) for child in token.children: if child.dep_ in ("nsubj", "dobj"): verb_collocations.add(child.head.lemma_) # 结合词向量语义匹配补充相关动词 similar_verbs = [ token.text for token in nlp.vocab if token.pos_ == "VERB" and nlp(target_noun).similarity(token) > 0.5 ] verb_collocations.update(similar_verbs) # 格式化输出为"to + 动词原形"格式 print([f"to {v}" for v in verb_collocations if v])
2. 使用Pattern库
Pattern是专注于英语NLP的轻量库,内置搭配提取功能,可直接获取与名词关联的动词。
示例代码:
from pattern.en import lemma, collocations target_noun = "car" # 提取与目标名词搭配的动词(pos="VB"指定动词类) colloc_list = collocations(target_noun, pos="VB") # 格式化输出为动词原形格式 print([f"to {lemma(v)}" for v in colloc_list])
3. 结合NLTK语料库自定义统计
如果偏好NLTK,可以利用其自带语料库统计名词与动词的共现频率,筛选高频搭配。
示例代码:
import nltk from nltk.corpus import brown from collections import defaultdict # 下载所需语料和标签集 nltk.download('brown') nltk.download('universal_tagset') target_noun = "car" verb_counts = defaultdict(int) # 遍历Brown语料库,统计与car搭配的动词 for sentence in brown.tagged_sents(tagset='universal'): for idx, (word, tag) in enumerate(sentence): if word.lower() == target_noun and tag == 'NOUN': # 统计前置动词(car作宾语) if idx > 0: prev_word, prev_tag = sentence[idx-1] if prev_tag == 'VERB': verb_counts[prev_word.lower()] += 1 # 统计后置动词(car作主语) if idx < len(sentence)-1: next_word, next_tag = sentence[idx+1] if next_tag == 'VERB': verb_counts[next_word.lower()] += 1 # 按频率取前10个搭配动词 top_verbs = sorted(verb_counts.items(), key=lambda x: x[1], reverse=True)[:10] print([f"to {v}" for v, _ in top_verbs])
内容的提问来源于stack exchange,提问作者Jonas De Boeck
相关产品推荐
相关产品推荐

