You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中是否有库可生成与指定词汇搭配的动词列表?

可行的解决方案

有几个库或组合方案可以实现提取特定名词搭配动词的需求,以下是具体实现方式:

1. 使用spaCy(推荐)

spaCy的预训练语言模型支持依存句法分析和词向量语义匹配,能精准识别与目标名词搭配的动词。

示例代码:

import spacy

# 加载中等规模预训练模型,大规模模型en_core_web_lg效果更优
nlp = spacy.load("en_core_web_md")

target_noun = "car"
# 构建包含目标名词的示例文本,也可替换为大规模语料
doc = nlp(f"A {target_noun} can be driven, crashed, repaired, washed, parked, or stolen.")

verb_collocations = set()
for token in doc:
    if token.text.lower() == target_noun:
        # 从依存关系中提取搭配动词(主动词或被动词)
        for child in token.children:
            if child.dep_ in ("nsubj", "dobj"):
                verb_collocations.add(child.head.lemma_)
        # 结合词向量语义匹配补充相关动词
        similar_verbs = [
            token.text for token in nlp.vocab
            if token.pos_ == "VERB" and nlp(target_noun).similarity(token) > 0.5
        ]
        verb_collocations.update(similar_verbs)

# 格式化输出为"to + 动词原形"格式
print([f"to {v}" for v in verb_collocations if v])

2. 使用Pattern库

Pattern是专注于英语NLP的轻量库,内置搭配提取功能,可直接获取与名词关联的动词。

示例代码:

from pattern.en import lemma, collocations

target_noun = "car"
# 提取与目标名词搭配的动词(pos="VB"指定动词类)
colloc_list = collocations(target_noun, pos="VB")
# 格式化输出为动词原形格式
print([f"to {lemma(v)}" for v in colloc_list])

3. 结合NLTK语料库自定义统计

如果偏好NLTK,可以利用其自带语料库统计名词与动词的共现频率,筛选高频搭配。

示例代码:

import nltk
from nltk.corpus import brown
from collections import defaultdict

# 下载所需语料和标签集
nltk.download('brown')
nltk.download('universal_tagset')

target_noun = "car"
verb_counts = defaultdict(int)

# 遍历Brown语料库,统计与car搭配的动词
for sentence in brown.tagged_sents(tagset='universal'):
    for idx, (word, tag) in enumerate(sentence):
        if word.lower() == target_noun and tag == 'NOUN':
            # 统计前置动词(car作宾语)
            if idx > 0:
                prev_word, prev_tag = sentence[idx-1]
                if prev_tag == 'VERB':
                    verb_counts[prev_word.lower()] += 1
            # 统计后置动词(car作主语)
            if idx < len(sentence)-1:
                next_word, next_tag = sentence[idx+1]
                if next_tag == 'VERB':
                    verb_counts[next_word.lower()] += 1

# 按频率取前10个搭配动词
top_verbs = sorted(verb_counts.items(), key=lambda x: x[1], reverse=True)[:10]
print([f"to {v}" for v, _ in top_verbs])

内容的提问来源于stack exchange,提问作者Jonas De Boeck

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 22:02:06