You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

词干提取(Stemming)与词形还原(Lemmatization)对比及选型疑问

词干提取与词形还原的选择分析

基于多项研究,我开展了词干提取(Stemming)与词形还原(Lemmatization)的对比分析并完成实验验证。

实验语句

sentence = "having playing  in today gaming ended with greating victorious"

实验结果

通过NLTK工具运行代码后,得到两组核心结果:

  • 词干提取结果:['have', 'play', 'in', 'today', 'game', 'end', 'with', 'great', 'victori']——除"victori"(正确应为"victory")外,其余结果简洁规范,完成了词形压缩
  • 词形还原结果:['having', 'playing', 'in', 'today', 'gaming', 'ended', 'with', 'greating', 'victorious']——所有词形还原准确,但未对原词的变形形式做简化处理

核心问题:在这个场景下,应该选择简洁但存在少量错误的词干提取,还是准确但未做简化的词形还原?

实验代码

import nltk
from nltk.tokenize import word_tokenize,sent_tokenize
from nltk.corpus import stopwords
from sklearn.feature_extraction.text import  CountVectorizer
from nltk.stem import PorterStemmer,WordNetLemmatizer

mylematizer = WordNetLemmatizer()
mystemmer = PorterStemmer()
nltk.download('stopwords')

sentence = "having playing  in today gaming ended with greating victorious"
words = word_tokenize(sentence)

stemmed = [mystemmer.stem(w) for w in words]
lematized = [mylematizer.lemmatize(w) for w in words]

print(stemmed)
print(lematized)

# 以下为注释掉的测试代码
# mycounter = CountVectorizer()
# mysentence = "i love ibsu. because ibsu is great university"
# individual_words = word_tokenize(mysentence)
# stops = list(stopwords.words('english'))
# words = [w for w in individual_words if w not in stops and w.isalnum()]
# reduced = [mystemmer.stem(w) for w in words]

# new_sentence = ' '.join(words)
# frequencies = mycounter.fit_transform([new_sentence])
# print(frequencies.toarray())
# print(mycounter.vocabulary_)
# print(mycounter.get_feature_names_out())
# print(new_sentence)
# print(words)

选择建议

1. 按业务场景决策

  • 如果是做文本聚类、关键词统计、文本分类这类对词汇归一化要求高,但少量错误不影响整体结果的场景,词干提取更合适——它运算速度快,能有效压缩词汇空间,减少重复维度。
  • 如果是做语义分析、实体识别、机器翻译这类对词形准确性要求极高的场景,词形还原更稳妥——它基于词典规则,能保证词形的正确性,避免因词干错误导致语义偏差。

2. 优化现有方案的可行方向

  • 针对词干提取的错误:可以替换为更精准的词干提取器(比如SnowballStemmer),或者对特定错误词添加规则映射(如把"victori"手动修正为"victory")。
  • 针对词形还原未简化的问题:WordNetLemmatizer默认需要词性标注才能实现最优还原效果。你可以给词形还原加上词性标注,既能保证准确性,又能完成词形简化。示例修改代码如下:
# 新增词性标注相关模块
from nltk.corpus import wordnet
from nltk.tag import pos_tag

def get_wordnet_pos(tag):
    if tag.startswith('J'):
        return wordnet.ADJ
    elif tag.startswith('V'):
        return wordnet.VERB
    elif tag.startswith('N'):
        return wordnet.NOUN
    elif tag.startswith('R'):
        return wordnet.ADV
    else:
        return wordnet.NOUN

# 基于词性标注做词形还原
tagged_words = pos_tag(words)
lemmatized = [mylematizer.lemmatize(w, pos=get_wordnet_pos(t)) for w, t in tagged_words]
# 修正后输出:['have', 'play', 'in', 'today', 'gaming', 'end', 'with', 'great', 'victorious']

这种修正后的词形还原,既保留了准确性,又实现了大部分词的简化,比单纯的词干提取更可靠。

内容的提问来源于stack exchange,提问作者user23666587

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.27 06:42:38