如何用Python的NLTK或同类工具实现动词词原形(lemma)变位及词性转换?
基于动词原形的变位及形态获取方案
1. 使用inflect库
inflect是轻量级语言处理库,支持动词变位、名词单复数转换,适配Python3.9+版本。
安装命令:
pip install inflect
使用示例:
import inflect p = inflect.engine() # 获取动词的全变位形式 verb_lemma = "run" all_conjugations = p.conjugate(verb_lemma, all=True) print(all_conjugations) # 生成特定时态/人称的变位 print(p.conjugate("run", tense="present", person=3, number="singular")) # 输出:runs print(p.conjugate("run", tense="past")) # 输出:ran print(p.conjugate("run", tense="past", aspect="perfect")) # 输出:has run
2. 使用pyconjug库
pyconjug是专门针对英语动词变位的库,可直接返回动词所有时态、语态形式,API直观易上手。
安装命令:
pip install pyconjug
使用示例:
from pyconjug import Conjugator conjugator = Conjugator(lang='en') verb_obj = conjugator.conjugate("run") # 获取所有变位形式的字典 print(verb_obj.all_conjugations) # 单独获取特定形态 print(verb_obj.present_tense) # 现在时形式 print(verb_obj.past_tense) # 过去式形式 print(verb_obj.present_participle) # 现在分词 print(verb_obj.past_participle) # 过去分词
3. 使用spaCy搭配spacy-inflect插件
借助spaCy的语法分析能力,结合spacy-inflect插件可精准生成动词指定词性形态,适合需结合NLP pipeline的场景。
安装命令:
pip install spacy spacy-inflect python -m spacy download en_core_web_sm
使用示例:
import spacy from spacy_inflect import Inflect nlp = spacy.load("en_core_web_sm") inflect = Inflect(nlp) doc = nlp("run") verb_token = doc[0] # 根据Penn Treebank标签生成对应形态 print(inflect.inflect(verb_token, tag="VBZ")) # 第三人称单数现在时:runs print(inflect.inflect(verb_token, tag="VBD")) # 过去式:ran print(inflect.inflect(verb_token, tag="VBG")) # 现在分词:running print(inflect.inflect(verb_token, tag="VBN")) # 过去分词:run
4. 适配Python3的pattern3分支
原版pattern库对高版本Python支持不佳,但社区维护的pattern3分支解决了兼容性问题,可在Python3.9/3.10中正常使用。
安装命令:
pip install pattern3
使用示例:
from pattern.en import conjugate # 生成不同时态、语态的变位 print(conjugate("run", tense="present", person=3)) # runs print(conjugate("run", tense="past")) # ran print(conjugate("run", tense="future")) # will run print(conjugate("run", aspect="progressive")) # am running
内容的提问来源于stack exchange,提问作者Patrick Autilio
相关产品推荐
相关产品推荐

