如何通过同义词映射优化KMeans句子聚类?求相关工具包
同义词映射工具与聚类代码优化方案
要解决聚类未识别同义词的问题,有多种工具和方法可实现同义词映射,以下是几种实用方案:
一、自定义同义词映射表(最适合明确指定同义词的场景)
这种方法直接简单,完全按需求手动指定同义词映射规则,无需依赖外部语料库。
修改后的代码实现:
from sklearn.feature_extraction.text import TfidfVectorizer from sklearn.cluster import KMeans import pandas as pd # 定义同义词映射表 synonym_map = { "capability": "ability", "power": "ability" } # 同义词替换预处理函数 def replace_synonyms(sentence): words = sentence.split() replaced_words = [synonym_map.get(word, word) for word in words] return " ".join(replaced_words) sentences = ["I have the ability", "I have the weakness", "I have the capability", "I have the power"] # 对句子进行同义词替换预处理 processed_sentences = [replace_synonyms(sent) for sent in sentences] tfidf = TfidfVectorizer(max_features=300) tfidf.fit(processed_sentences) X = tfidf.transform(processed_sentences) k = 2 model = KMeans(n_clusters=k, random_state=1) model.fit(X) # 输出结果 result = pd.DataFrame(columns=["original_sentence", "processed_sentence"], data=zip(sentences, processed_sentences)) result = result.join(pd.DataFrame(columns=["cluster"], data=model.labels_)) print(result)
运行输出:
| index | original_sentence | processed_sentence | cluster |
|---|---|---|---|
| 0 | I have the ability | I have the ability | 0 |
| 1 | I have the weakness | I have the weakness | 1 |
| 2 | I have the capability | I have the ability | 0 |
| 3 | I have the power | I have the ability | 0 |
可见ability、capability、power对应的句子被归为同一簇,weakness单独成簇,符合需求。
二、基于NLTK WordNet的同义词扩展
如果需要自动识别同义词,NLTK的WordNet语料库可获取单词的同义词集合,也可结合自定义映射补充:
示例代码片段:
from nltk.corpus import wordnet import nltk # 首次运行需下载WordNet nltk.download('wordnet') def get_synonyms(word): synonyms = set() for syn in wordnet.synsets(word): for lemma in syn.lemmas(): synonyms.add(lemma.name()) # 手动补充自定义同义词 if word == "power": synonyms.add("ability") return synonyms # 可扩展替换逻辑,用同义词集合中的基准词统一替换
三、其他可用工具
- spaCy:通过自定义
Matcher或结合预训练词向量,实现同义词识别与替换,适配复杂文本场景。 - gensim:利用预训练的Word2Vec/GloVe词向量,通过计算词相似度识别同义词,再做归一化处理。
内容的提问来源于stack exchange,提问作者james pow
相关产品推荐
相关产品推荐

