如何用Gensim生成葡萄牙语词嵌入?及适配问题咨询
Hey there! Let's break down why your Portuguese phrase similarity isn't working as expected, and how to fix it. Gensim does fully support Portuguese—the issue here is with your preprocessing pipeline and model setup, not language support.
1. Your code has messy, language-mismatched preprocessing
Looking at your code, you're mixing English stopwords with Portuguese text, and your preprocessing steps are inconsistent (you first generate one set of texts, then overwrite it with tokenized text without applying stopword filtering). This means your model is cluttered with meaningless Portuguese function words like de, do, da that should be removed, and it can't properly learn semantic relationships between your target phrases.
2. You're using LSI, not Word2Vec (your plot title is misleading)
Your code uses Latent Semantic Indexing (LSI)—a topic modeling algorithm based on TF-IDF—instead of Word2Vec (a word embedding model). Both work for similarity tasks, but they require different preprocessing, and Word2Vec is better suited for capturing fine-grained semantic similarities between phrases.
Fixes to get your Portuguese similarity working
Let's adjust your code with proper Portuguese-specific preprocessing, and show both improved LSI and Word2Vec approaches:
First, install/download required tools
pip install gensim nltk scikit-learn
Then download NLTK's Portuguese language resources:
import nltk nltk.download('punkt') nltk.download('stopwords') nltk.download('rslp') # Portuguese stemmer
Option 1: Improved LSI with proper preprocessing
This fixes your original pipeline with correct Portuguese cleanup:
import logging logging.basicConfig(format='%(asctime)s : %(levelname)s : %(message)s', level=logging.INFO) import matplotlib.pyplot as plt from gensim import corpora, models, similarities from nltk.tokenize import word_tokenize from nltk.corpus import stopwords from nltk.stem import RSLPStemmer import numpy as np # Portuguese preprocessing tools stop_words = set(stopwords.words('portuguese')) stemmer = RSLPStemmer() # Reduces words to their root form # Your original documents documents = [ "Interface máquina humana para aplicações computacionais de laboratório abc", "Um levantamento da opinião do usuário sobre o tempo de resposta do sistema informático", "O sistema de gerenciamento de interface do usuário EPS", "Sistema e testes de engenharia de sistemas humanos de EPS", "Relação do tempo de resposta percebido pelo usuário para a medição de erro", "A geração de árvores não ordenadas binárias aleatórias", "O gráfico de interseção dos caminhos nas árvores", "Gráfico de menores IV Largura de árvores e bem quase encomendado", "Gráficos menores Uma pesquisa" ] # Unified preprocessing function for all text def preprocess(text): # Tokenize Portuguese text tokens = word_tokenize(text.lower(), language='portuguese') # Keep only alphabetic tokens, remove stopwords, apply stemming return [stemmer.stem(token) for token in tokens if token.isalpha() and token not in stop_words] # Preprocess all documents texts = [preprocess(doc) for doc in documents] # Build dictionary and corpus dictionary = corpora.Dictionary(texts) corpus = [dictionary.doc2bow(text) for text in texts] # Train TF-IDF and LSI models tfidf = models.TfidfModel(corpus) corpus_tfidf = tfidf[corpus] lsi = models.LsiModel(corpus_tfidf, id2word=dictionary, num_topics=2) # Process your target phrase new_doc = "Tempo de resposta e medição de erro" new_vec_bow = dictionary.doc2bow(preprocess(new_doc)) new_vec_lsi = lsi[tfidf[new_vec_bow]] # Calculate and sort similarities index = similarities.MatrixSimilarity(lsi[corpus_tfidf]) sims = index[new_vec_lsi] sims_sorted = sorted(enumerate(sims), key=lambda x: -x[1]) print("LSI Similarity Results:") for idx, score in sims_sorted: print(f"{score:.4f} - {documents[idx]}") # (You can keep your visualization code here, just update the data sources to use the cleaned texts/models)
Option 2: Word2Vec for semantic phrase similarity
If you want true word embedding-based similarity (like your plot title mentions), use Word2Vec and average word vectors to get phrase vectors:
from gensim.models import Word2Vec from sklearn.metrics.pairwise import cosine_similarity # Train Word2Vec on your preprocessed texts w2v_model = Word2Vec( sentences=texts, vector_size=100, window=5, # Context window size min_count=1, # Keep rare words workers=4 ) # Function to get a phrase vector by averaging word vectors def get_phrase_vector(phrase): tokens = preprocess(phrase) # Only use words present in the Word2Vec model valid_vecs = [w2v_model.wv[token] for token in tokens if token in w2v_model.wv] if not valid_vecs: return np.zeros(w2v_model.vector_size) return np.mean(valid_vecs, axis=0) # Calculate vectors for your target phrase and all documents new_vec = get_phrase_vector(new_doc) doc_vecs = [get_phrase_vector(doc) for doc in documents] # Compute cosine similarity sims = cosine_similarity([new_vec], doc_vecs)[0] sims_sorted = sorted(enumerate(sims), key=lambda x: -x[1]) print("\nWord2Vec Similarity Results:") for idx, score in sims_sorted: print(f"{score:.4f} - {documents[idx]}")
Key Takeaways
- Gensim supports Portuguese perfectly—no extra configuration needed beyond language-specific preprocessing.
- Always use language-matched stopwords, tokenizers, and stemmers/lemmatizers (for Portuguese, NLTK's
RSLPStemmeror spaCy's Portuguese model work great). - Choose the right model: LSI is good for topic-based similarity, while Word2Vec/Doc2Vec is better for semantic, word-level similarity.
内容的提问来源于stack exchange,提问作者razimbres

