Gensim Word2Vec训练后仅生成单字符词汇,求问题排查与解决
问题诊断与解决
嘿,这个坑我之前也踩过!核心原因是你给Word2Vec的输入格式完全错了——Gensim的Word2Vec模型要求输入必须是**「句子组成的列表」,每个句子又是一个「单词组成的列表」**;而你直接传了一个扁平的单词列表,模型会把每个单词字符串当成可迭代对象,自动拆成单个字符来训练,自然就得到了字符级的词汇表。
为什么会这样?
举个直观的例子:如果你传入的是["potential", "xyz"],模型会把每个字符串拆成["p","o","t","e","n","t","i","a","l"]和["x","y","z"]来处理,最终词汇表就全是单个字符了。
快速修正方案
你只需要把预处理好的单词列表包裹在一个外层列表里,让模型识别这是一个完整的句子(如果你的文本是单句的话):
# word Embedding from gensim.models import Word2Vec # 重点修改:把单词列表放到外层列表中,作为单个句子传入 train_corpus = [token_list5] # train model model = Word2Vec(train_corpus, min_count=2) # summarize the loaded model print("The model is :") print(model,"\n") # summarize vocabulary words = list(model.wv.vocab) print("The learned vocabulary words are : \n",words)
更规范的多句处理方式
如果你的原始文本包含多个句子,建议先拆分句子再逐个预处理,最后组成标准的句子列表输入模型,这样训练效果会更好:
# 先拆分句子(替代之前直接对整段文本分词) sentences = nltk.sent_tokenize(raw_text) processed_sentences = [] lemmatizer = WordNetLemmatizer() stop_words = set(stopwords.words("english")) # 提前转成集合,提升查询效率 for sent in sentences: # 单句预处理流程 tokens = nltk.word_tokenize(sent) # 去标点 tokens = list(filter(lambda token : punkt.PunktToken(token).is_non_punct, tokens)) # 转小写 tokens = [word.lower() for word in tokens] # 去停用词(用集合查询更快) tokens = [token for token in tokens if token not in stop_words] # 词形还原 tokens = [lemmatizer.lemmatize(word) for word in tokens] processed_sentences.append(tokens) # 用标准的句子列表训练模型 model = Word2Vec(processed_sentences, min_count=2)
额外小建议
预处理后建议先打印token_list5(或processed_sentences)的内容,确认里面都是完整的单词,避免分词、去标点等环节出现隐性错误。
内容的提问来源于stack exchange,提问作者amanuel_ng
相关产品推荐
相关产品推荐

