You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Sklearn TF-IDF提取文档关键词及余弦相似度技术问询

Great news—you’re using the right tool!

TF-IDF (via TfidfVectorizer) is exactly what you need for both extracting relevant keywords from your documents and calculating cosine similarity with query texts. The sparse matrix output you’re seeing is totally normal and actually efficient for this kind of work—let me break down what’s going on and how to get the results you need.

Why you’re seeing a sparse matrix

Most terms don’t appear in most documents, so storing the entire dense matrix (filled with lots of zeros) would waste memory. The Compressed Sparse Row (CSR) format only stores non-zero values, which is the standard for text processing tasks. You don’t need to convert it to a dense matrix for most operations (like cosine similarity), but if you want to inspect the raw scores for small datasets, you can do:

# Convert to dense array (not recommended for large datasets!)
dense_matrix = sklearn_representation.toarray()

Extracting top keywords per document

To get the most relevant keywords for each document based on TF-IDF scores, map the highest-scoring indices back to your vocabulary terms. Here’s a quick code snippet:

# Get the vocabulary (all unique terms) from the vectorizer
feature_names = sklearn_tfidf.get_feature_names_out()

# For each document, retrieve top 3 keywords
for doc_idx in range(len(rawDocument)):
    # Get TF-IDF scores for the current document
    tfidf_scores = sklearn_representation[doc_idx].toarray().flatten()
    # Get indices of the top N highest scores (reverse to get descending order)
    top_term_indices = tfidf_scores.argsort()[-3:][::-1]
    # Map indices to actual terms
    top_keywords = [feature_names[idx] for idx in top_term_indices]
    print(f"Document {doc_idx+1} top keywords: {', '.join(top_keywords)}")

This will give you the terms with the highest TF-IDF values—these are your most relevant keywords for each document.

Calculating cosine similarity with a query document

You don’t need to convert the sparse matrix to compute similarity. Use sklearn.metrics.pairwise.cosine_similarity directly, which works seamlessly with sparse formats. Here’s how:

from sklearn.metrics.pairwise import cosine_similarity

# Example query document (already filtered for stopwords)
query = ["british american sports english"]

# Transform the query using the same vectorizer you fitted on your raw documents
query_tfidf = sklearn_tfidf.transform(query)

# Compute cosine similarity between the query and all your documents
similarity_scores = cosine_similarity(query_tfidf, sklearn_representation)[0]

# Print similarity scores (0 = no overlap, 1 = identical content)
for doc_idx, score in enumerate(similarity_scores):
    print(f"Similarity with Document {doc_idx+1}: {score:.4f}")

Optional parameter tweaks for better results

Your current TfidfVectorizer parameters are solid, but you might want to adjust these for more precise outputs:

  • max_df: Exclude terms that appear in too many documents (e.g., max_df=0.8 to ignore terms present in 80%+ of docs)
  • ngram_range: Include multi-word phrases if relevant (e.g., ngram_range=(1,2) to capture unigrams and bigrams like "british english")
  • min_df: Ignore terms that appear in too few documents (e.g., min_df=2 to exclude one-off typos or rare terms)

Content of the question originates from Stack Exchange, asked by Federico Cuozzo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 07:51:05