如何让NLTK结合上下文进行词形还原?
如何让NLTK结合上下文对单词进行词形还原?
你观察得特别准——默认情况下WordNetLemmatizer把leaves还原成leaf,是因为它默认把所有单词当成名词处理。但像你举的两个例子,leaves在不同语境里可能是动词(第三人称单数)或名词(复数),这时候就得先搞定「词性识别」,再让词形还原工具结合词性工作,才能得到正确结果。
下面是具体的实现思路和代码:
核心步骤:先做词性标注,再映射格式
WordNetLemmatizer只认四种词性参数:'n'(名词)、'v'(动词)、'a'(形容词)、'r'(副词),但NLTK自带的词性标注器输出的标签是更细分的(比如VBZ代表动词第三人称单数、NNS代表复数名词),所以我们需要先把标注结果映射成工具能识别的格式。
完整示例代码
from nltk.stem.wordnet import WordNetLemmatizer from nltk.tag import pos_tag from nltk.tokenize import word_tokenize # 初始化词形还原工具 lem = WordNetLemmatizer() def lemmatize_with_context(text): # 第一步:把文本拆成单个单词 tokens = word_tokenize(text) # 第二步:给每个单词打词性标签 tagged_tokens = pos_tag(tokens) lemmatized_result = [] for token, tag in tagged_tokens: # 把NLTK的POS标签映射成WordNet支持的格式 if tag.startswith('VB'): # 所有动词类标签(VB/VBD/VBG/VBN/VBP/VBZ) pos_type = 'v' elif tag.startswith('NN'): # 所有名词类标签(NN/NNS/NNP/NNPS) pos_type = 'n' elif tag.startswith('JJ'): # 所有形容词类标签(JJ/JJR/JJS) pos_type = 'a' elif tag.startswith('RB'): # 所有副词类标签(RB/RBR/RBS) pos_type = 'r' else: # 其他词性默认按名词处理 pos_type = 'n' # 第三步:结合词性做词形还原 processed_word = lem.lemmatize(token, pos=pos_type) lemmatized_result.append(processed_word) return ' '.join(lemmatized_result) # 测试你给出的两个例子 print(lemmatize_with_context("Mary leaves the room")) # 输出:Mary leave the room print(lemmatize_with_context("Dew drops fall from the leaves")) # 输出:Dew drop fall from the leaf
额外说明
- 词性标注的准确性会直接影响最终结果,NLTK的默认标注器在日常文本场景下表现已经足够稳定;
- 如果遇到特殊领域的文本,可能需要训练自定义的词性标注器,但大部分普通场景用上面的代码就够啦。
内容的提问来源于stack exchange,提问作者James Ko
相关产品推荐
相关产品推荐

