You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何让NLTK结合上下文进行词形还原?

如何让NLTK结合上下文对单词进行词形还原?

你观察得特别准——默认情况下WordNetLemmatizer把leaves还原成leaf,是因为它默认把所有单词当成名词处理。但像你举的两个例子,leaves在不同语境里可能是动词(第三人称单数)或名词(复数),这时候就得先搞定「词性识别」,再让词形还原工具结合词性工作,才能得到正确结果。

下面是具体的实现思路和代码:

核心步骤:先做词性标注,再映射格式

WordNetLemmatizer只认四种词性参数:'n'(名词)、'v'(动词)、'a'(形容词)、'r'(副词),但NLTK自带的词性标注器输出的标签是更细分的(比如VBZ代表动词第三人称单数、NNS代表复数名词),所以我们需要先把标注结果映射成工具能识别的格式。

完整示例代码

from nltk.stem.wordnet import WordNetLemmatizer
from nltk.tag import pos_tag
from nltk.tokenize import word_tokenize

# 初始化词形还原工具
lem = WordNetLemmatizer()

def lemmatize_with_context(text):
    # 第一步:把文本拆成单个单词
    tokens = word_tokenize(text)
    # 第二步:给每个单词打词性标签
    tagged_tokens = pos_tag(tokens)
    
    lemmatized_result = []
    for token, tag in tagged_tokens:
        # 把NLTK的POS标签映射成WordNet支持的格式
        if tag.startswith('VB'):  # 所有动词类标签(VB/VBD/VBG/VBN/VBP/VBZ)
            pos_type = 'v'
        elif tag.startswith('NN'):  # 所有名词类标签(NN/NNS/NNP/NNPS)
            pos_type = 'n'
        elif tag.startswith('JJ'):  # 所有形容词类标签(JJ/JJR/JJS)
            pos_type = 'a'
        elif tag.startswith('RB'):  # 所有副词类标签(RB/RBR/RBS)
            pos_type = 'r'
        else:
            # 其他词性默认按名词处理
            pos_type = 'n'
        
        # 第三步:结合词性做词形还原
        processed_word = lem.lemmatize(token, pos=pos_type)
        lemmatized_result.append(processed_word)
    
    return ' '.join(lemmatized_result)

# 测试你给出的两个例子
print(lemmatize_with_context("Mary leaves the room"))  # 输出:Mary leave the room
print(lemmatize_with_context("Dew drops fall from the leaves"))  # 输出:Dew drop fall from the leaf

额外说明

  • 词性标注的准确性会直接影响最终结果,NLTK的默认标注器在日常文本场景下表现已经足够稳定;
  • 如果遇到特殊领域的文本,可能需要训练自定义的词性标注器,但大部分普通场景用上面的代码就够啦。

内容的提问来源于stack exchange,提问作者James Ko

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 11:08:50