You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Gensim生成葡萄牙语词嵌入?及适配问题咨询

Hey there! Let's break down why your Portuguese phrase similarity isn't working as expected, and how to fix it. Gensim does fully support Portuguese—the issue here is with your preprocessing pipeline and model setup, not language support.

1. Your code has messy, language-mismatched preprocessing

Looking at your code, you're mixing English stopwords with Portuguese text, and your preprocessing steps are inconsistent (you first generate one set of texts, then overwrite it with tokenized text without applying stopword filtering). This means your model is cluttered with meaningless Portuguese function words like de, do, da that should be removed, and it can't properly learn semantic relationships between your target phrases.

2. You're using LSI, not Word2Vec (your plot title is misleading)

Your code uses Latent Semantic Indexing (LSI)—a topic modeling algorithm based on TF-IDF—instead of Word2Vec (a word embedding model). Both work for similarity tasks, but they require different preprocessing, and Word2Vec is better suited for capturing fine-grained semantic similarities between phrases.

Fixes to get your Portuguese similarity working

Let's adjust your code with proper Portuguese-specific preprocessing, and show both improved LSI and Word2Vec approaches:

First, install/download required tools

pip install gensim nltk scikit-learn

Then download NLTK's Portuguese language resources:

import nltk
nltk.download('punkt')
nltk.download('stopwords')
nltk.download('rslp') # Portuguese stemmer

Option 1: Improved LSI with proper preprocessing

This fixes your original pipeline with correct Portuguese cleanup:

import logging
logging.basicConfig(format='%(asctime)s : %(levelname)s : %(message)s', level=logging.INFO)
import matplotlib.pyplot as plt
from gensim import corpora, models, similarities
from nltk.tokenize import word_tokenize
from nltk.corpus import stopwords
from nltk.stem import RSLPStemmer
import numpy as np

# Portuguese preprocessing tools
stop_words = set(stopwords.words('portuguese'))
stemmer = RSLPStemmer() # Reduces words to their root form

# Your original documents
documents = [
    "Interface máquina humana para aplicações computacionais de laboratório abc",
    "Um levantamento da opinião do usuário sobre o tempo de resposta do sistema informático",
    "O sistema de gerenciamento de interface do usuário EPS",
    "Sistema e testes de engenharia de sistemas humanos de EPS",
    "Relação do tempo de resposta percebido pelo usuário para a medição de erro",
    "A geração de árvores não ordenadas binárias aleatórias",
    "O gráfico de interseção dos caminhos nas árvores",
    "Gráfico de menores IV Largura de árvores e bem quase encomendado",
    "Gráficos menores Uma pesquisa"
]

# Unified preprocessing function for all text
def preprocess(text):
    # Tokenize Portuguese text
    tokens = word_tokenize(text.lower(), language='portuguese')
    # Keep only alphabetic tokens, remove stopwords, apply stemming
    return [stemmer.stem(token) for token in tokens if token.isalpha() and token not in stop_words]

# Preprocess all documents
texts = [preprocess(doc) for doc in documents]

# Build dictionary and corpus
dictionary = corpora.Dictionary(texts)
corpus = [dictionary.doc2bow(text) for text in texts]

# Train TF-IDF and LSI models
tfidf = models.TfidfModel(corpus)
corpus_tfidf = tfidf[corpus]
lsi = models.LsiModel(corpus_tfidf, id2word=dictionary, num_topics=2)

# Process your target phrase
new_doc = "Tempo de resposta e medição de erro"
new_vec_bow = dictionary.doc2bow(preprocess(new_doc))
new_vec_lsi = lsi[tfidf[new_vec_bow]]

# Calculate and sort similarities
index = similarities.MatrixSimilarity(lsi[corpus_tfidf])
sims = index[new_vec_lsi]
sims_sorted = sorted(enumerate(sims), key=lambda x: -x[1])

print("LSI Similarity Results:")
for idx, score in sims_sorted:
    print(f"{score:.4f} - {documents[idx]}")

# (You can keep your visualization code here, just update the data sources to use the cleaned texts/models)

Option 2: Word2Vec for semantic phrase similarity

If you want true word embedding-based similarity (like your plot title mentions), use Word2Vec and average word vectors to get phrase vectors:

from gensim.models import Word2Vec
from sklearn.metrics.pairwise import cosine_similarity

# Train Word2Vec on your preprocessed texts
w2v_model = Word2Vec(
    sentences=texts,
    vector_size=100,
    window=5, # Context window size
    min_count=1, # Keep rare words
    workers=4
)

# Function to get a phrase vector by averaging word vectors
def get_phrase_vector(phrase):
    tokens = preprocess(phrase)
    # Only use words present in the Word2Vec model
    valid_vecs = [w2v_model.wv[token] for token in tokens if token in w2v_model.wv]
    if not valid_vecs:
        return np.zeros(w2v_model.vector_size)
    return np.mean(valid_vecs, axis=0)

# Calculate vectors for your target phrase and all documents
new_vec = get_phrase_vector(new_doc)
doc_vecs = [get_phrase_vector(doc) for doc in documents]

# Compute cosine similarity
sims = cosine_similarity([new_vec], doc_vecs)[0]
sims_sorted = sorted(enumerate(sims), key=lambda x: -x[1])

print("\nWord2Vec Similarity Results:")
for idx, score in sims_sorted:
    print(f"{score:.4f} - {documents[idx]}")

Key Takeaways

  • Gensim supports Portuguese perfectly—no extra configuration needed beyond language-specific preprocessing.
  • Always use language-matched stopwords, tokenizers, and stemmers/lemmatizers (for Portuguese, NLTK's RSLPStemmer or spaCy's Portuguese model work great).
  • Choose the right model: LSI is good for topic-based similarity, while Word2Vec/Doc2Vec is better for semantic, word-level similarity.

内容的提问来源于stack exchange,提问作者razimbres

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 08:02:27