使用Sklearn TF-IDF提取文档关键词及余弦相似度技术问询
TF-IDF (via TfidfVectorizer) is exactly what you need for both extracting relevant keywords from your documents and calculating cosine similarity with query texts. The sparse matrix output you’re seeing is totally normal and actually efficient for this kind of work—let me break down what’s going on and how to get the results you need.
Why you’re seeing a sparse matrix
Most terms don’t appear in most documents, so storing the entire dense matrix (filled with lots of zeros) would waste memory. The Compressed Sparse Row (CSR) format only stores non-zero values, which is the standard for text processing tasks. You don’t need to convert it to a dense matrix for most operations (like cosine similarity), but if you want to inspect the raw scores for small datasets, you can do:
# Convert to dense array (not recommended for large datasets!) dense_matrix = sklearn_representation.toarray()
Extracting top keywords per document
To get the most relevant keywords for each document based on TF-IDF scores, map the highest-scoring indices back to your vocabulary terms. Here’s a quick code snippet:
# Get the vocabulary (all unique terms) from the vectorizer feature_names = sklearn_tfidf.get_feature_names_out() # For each document, retrieve top 3 keywords for doc_idx in range(len(rawDocument)): # Get TF-IDF scores for the current document tfidf_scores = sklearn_representation[doc_idx].toarray().flatten() # Get indices of the top N highest scores (reverse to get descending order) top_term_indices = tfidf_scores.argsort()[-3:][::-1] # Map indices to actual terms top_keywords = [feature_names[idx] for idx in top_term_indices] print(f"Document {doc_idx+1} top keywords: {', '.join(top_keywords)}")
This will give you the terms with the highest TF-IDF values—these are your most relevant keywords for each document.
Calculating cosine similarity with a query document
You don’t need to convert the sparse matrix to compute similarity. Use sklearn.metrics.pairwise.cosine_similarity directly, which works seamlessly with sparse formats. Here’s how:
from sklearn.metrics.pairwise import cosine_similarity # Example query document (already filtered for stopwords) query = ["british american sports english"] # Transform the query using the same vectorizer you fitted on your raw documents query_tfidf = sklearn_tfidf.transform(query) # Compute cosine similarity between the query and all your documents similarity_scores = cosine_similarity(query_tfidf, sklearn_representation)[0] # Print similarity scores (0 = no overlap, 1 = identical content) for doc_idx, score in enumerate(similarity_scores): print(f"Similarity with Document {doc_idx+1}: {score:.4f}")
Optional parameter tweaks for better results
Your current TfidfVectorizer parameters are solid, but you might want to adjust these for more precise outputs:
max_df: Exclude terms that appear in too many documents (e.g.,max_df=0.8to ignore terms present in 80%+ of docs)ngram_range: Include multi-word phrases if relevant (e.g.,ngram_range=(1,2)to capture unigrams and bigrams like "british english")min_df: Ignore terms that appear in too few documents (e.g.,min_df=2to exclude one-off typos or rare terms)
Content of the question originates from Stack Exchange, asked by Federico Cuozzo

