You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何获取目标词的词形变体与语义关联词?NLP技术咨询

解决你的两个NLP需求

需求1:获取目标词的所有词形变体(比如phone → Phones/phones/Phone's等)

你遇到的问题本质是词形标准化和生成屈折变体,不用复杂的Word2Vec,用轻量的NLP工具就能搞定:

方案1:标准化输入用于匹配字典

如果核心需求是让phones/Phone's这类输入能匹配字典里的phone键,先对输入做标准化处理即可:

  1. 统一转为小写,消除大小写差异
  2. 去掉所有格后缀('s或s')
  3. 用词形还原工具把复数转成单数

用NLTK和正则就能实现,代码示例:

from nltk.stem import WordNetLemmatizer
import re

# 初始化词形还原器
lemmatizer = WordNetLemmatizer()

def normalize_input(input_word):
    # 转小写
    lower_word = input_word.lower()
    # 移除所有格后缀
    normalized = re.sub(r'(\'s|s\')$', '', lower_word)
    # 词形还原(复数转单数,指定名词类型)
    return lemmatizer.lemmatize(normalized, pos='n')

# 测试各种变体
test_inputs = ['phone', 'Phones', 'phones', 'Phone\'s', 'phone\'s']
for word in test_inputs:
    print(f"输入{word} → 标准化后{normalize_input(word)}")
# 所有输入都会输出'phone',直接用这个结果去匹配字典键就行

方案2:生成所有可能的词形变体

如果是要主动生成phone的所有变体(比如展示给用户),可以用inflect库,它专门处理英语的屈折变化:

import inflect

p = inflect.engine()
base_word = 'phone'

# 生成各类变体
variants = [
    base_word,  # 小写原型
    base_word.capitalize(),  # 首字母大写
    p.plural(base_word),  # 复数小写
    p.plural(base_word).capitalize(),  # 复数首字母大写
    f"{base_word}'s",  # 所有格小写
    f"{base_word.capitalize()}'s"  # 所有格首字母大写
]

print(variants)
# 输出: ['phone', 'Phone', 'phones', 'Phones', "phone's", "Phone's"]

需求2:获取同语义关联词(比如Dog → Puppy/Kitty等)

你用WordNet只拿到synsets是因为没提取里面的具体词,另外WordNet还能帮你找到下位词(更具体的词,比如Dog的下位词Puppy)、上位词,甚至同类语义的词。如果要更灵活的语义关联,也可以用预训练词向量模型。

方案1:用WordNet提取同义词和相关词

修改你的代码,从synsets里取出具体的词,再加上下位词/同类词:

from nltk.corpus import wordnet as wn

def get_related_words(target_word):
    related = set()
    for syn in wn.synsets(target_word, pos=wn.NOUN):  # 限定名词类型
        # 添加同义词
        related.update(syn.lemma_names())
        # 添加下位词(更具体的子类,比如Dog → Puppy)
        for hypo in syn.hyponyms():
            related.update(hypo.lemma_names())
        # 添加上位词的下位词(同类别的词,比如Dog和Cat同属animal子类)
        for hyper in syn.hypernyms():
            for hypo_of_hyper in hyper.hyponyms():
                related.update(hypo_of_hyper.lemma_names())
    # 处理大小写,添加首字母大写形式
    related.update([word.capitalize() for word in related])
    # 把下划线换成空格(比如domestic_dog → domestic dog)
    related = {word.replace('_', ' ') for word in related}
    return related

# 测试Dog
print(get_related_words('dog'))
# 会包含'Dog', 'dog', 'Puppy', 'puppy', 'Cat', 'kitty'等相关词

方案2:用预训练词向量模型(Word2Vec/FastText)

如果想要更贴合语境的语义关联,比如Dog和Puppy这种强相关词,可以用Gensim加载预训练的Word2Vec模型,比如Google的News预训练模型:

from gensim.models import KeyedVectors

# 加载预训练模型(需先下载GoogleNews-vectors-negative300.bin)
model = KeyedVectors.load_word2vec_format('GoogleNews-vectors-negative300.bin', binary=True)

# 获取Top10相似词
similar_words = model.most_similar('Dog', topn=10)
# 提取词(忽略相似度分数)
result = [word for word, score in similar_words]
print(result)
# 输出类似:['Puppy', 'dog', 'Dogs', 'Puppies', 'Canine', 'Yorkshire_terrier', ...]

你也可以用FastText模型,它对生僻词和拼写错误的处理效果更好。

内容的提问来源于stack exchange,提问作者user9158931

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 06:47:38