请求Doc2Vec单文档句向量生成及聚类实现代码(Python)
Doc2Vec for Sentence Vectors & Clustering: Step-by-Step Solutions
Hey there! Let's break down your two needs into clear, newbie-friendly code and explanations. I'll build on the bits of code you already tried to keep things consistent.
1. Getting Sentence Vectors from a Single Document
Doc2Vec works by treating each sentence as its own "document" during training. Here's how to extract vectors for every sentence in your single text file:
First, Install Required Tools
Run this in your terminal if you haven't already:
pip install gensim nltk pandas scikit-learn matplotlib
Complete Code
import nltk from gensim.models.doc2vec import Doc2Vec, TaggedDocument from nltk.tokenize import word_tokenize, sent_tokenize # Download NLTK's tokenizer data (run once) nltk.download('punkt') # Load your single document (replace with your file path or raw text) with open("your_document.txt", "r", encoding="utf-8") as f: document_text = f.read() # Split the document into individual sentences sentences = sent_tokenize(document_text) # Preprocess each sentence: tokenize, lowercase, clean up processed_sentences = [] for idx, sentence in enumerate(sentences): # Split sentence into words and lowercase them tokens = word_tokenize(sentence.lower()) # Keep only alphanumeric words (filters out punctuation) tokens = [token for token in tokens if token.isalnum()] # Create a TaggedDocument (required for Doc2Vec training) processed_sentences.append(TaggedDocument(tokens, [idx])) # Train the Doc2Vec model model = Doc2Vec( vector_size=100, # Size of the output vector (adjust based on your needs) min_count=1, # Ignore words that appear less than once epochs=50, # Number of training loops dm=1 # Use "distributed memory" mode (good for sentence context) ) # Build vocabulary from our processed sentences model.build_vocab(processed_sentences) # Train the model on our sentence-documents model.train(processed_sentences, total_examples=model.corpus_count, epochs=model.epochs) # Generate vectors for each sentence sentence_vectors = [model.infer_vector(doc.words) for doc in processed_sentences] # Check the results print(f"Generated {len(sentence_vectors)} sentence vectors!") print("Sample vector for first sentence:", sentence_vectors[0])
Quick Explanation
- We split your document into sentences using NLTK's
sent_tokenizetool. - Each sentence gets converted into a
TaggedDocument(a required format for Doc2Vec) with a unique tag (the sentence index). - The model learns to map each sentence to a numerical vector that captures its meaning.
infer_vectorgenerates the final vector for each sentence—even though we trained on these sentences, this step ensures consistent, usable vectors.
2. Clustering Sentences from kkk.csv
Let's process your CSV file, generate sentence vectors, and cluster similar sentences using KMeans. We'll even add a visualization to see how clusters group together.
Complete Code
import pandas as pd import nltk from gensim.models.doc2vec import Doc2Vec, TaggedDocument from nltk.tokenize import word_tokenize from sklearn.cluster import KMeans from sklearn.decomposition import PCA import matplotlib.pyplot as plt nltk.download('punkt') # Load your CSV file (adjust column name if yours isn't 'sentence') df = pd.read_csv("kkk.csv") sentences = df['sentence'].tolist() # Replace 'sentence' with your actual column name # Preprocess sentences (same as before) processed_sentences = [] for idx, sentence in enumerate(sentences): # Convert to string to avoid errors with missing values tokens = word_tokenize(str(sentence).lower()) tokens = [token for token in tokens if token.isalnum()] processed_sentences.append(TaggedDocument(tokens, [idx])) # Train Doc2Vec model model = Doc2Vec( vector_size=100, min_count=1, epochs=50, dm=1 ) model.build_vocab(processed_sentences) model.train(processed_sentences, total_examples=model.corpus_count, epochs=model.epochs) # Get sentence vectors sentence_vectors = [model.infer_vector(doc.words) for doc in processed_sentences] # Find the optimal number of clusters (elbow method) inertia = [] k_range = range(1, 11) for k in k_range: kmeans = KMeans(n_clusters=k, random_state=42) kmeans.fit(sentence_vectors) inertia.append(kmeans.inertia_) # Plot elbow curve to pick your k plt.figure(figsize=(8,4)) plt.plot(k_range, inertia, 'bo-') plt.xlabel('Number of Clusters (k)') plt.ylabel('Inertia') plt.title('Elbow Method to Choose Optimal k') plt.show() # Choose your k (based on the elbow plot—try 3-5 for 101 sentences) num_clusters = 3 kmeans = KMeans(n_clusters=num_clusters, random_state=42) cluster_labels = kmeans.fit_predict(sentence_vectors) # Add cluster labels to your original dataframe df['cluster'] = cluster_labels # Print sentences grouped by cluster for cluster in range(num_clusters): print(f"\n--- Cluster {cluster} ---") cluster_sentences = df[df['cluster'] == cluster]['sentence'].tolist() for sent in cluster_sentences: print(f"- {sent}") # Visualize clusters in 2D using PCA pca = PCA(n_components=2) reduced_vectors = pca.fit_transform(sentence_vectors) plt.figure(figsize=(10,6)) scatter = plt.scatter(reduced_vectors[:,0], reduced_vectors[:,1], c=cluster_labels, cmap='viridis') plt.legend(handles=scatter.legend_elements()[0], labels=[f"Cluster {i}" for i in range(num_clusters)]) plt.title('Sentence Clusters (2D PCA Projection)') plt.show()
Key Tips for Newbies
- CSV Column Name: If your CSV doesn't use
sentenceas the column header, replace that string with your actual column name (e.g.,df['text']). - Choosing k: The elbow plot shows where adding more clusters stops improving the grouping (the "elbow" point). For 101 sentences, k=3-5 is a safe starting point.
- Missing Values: If your CSV has empty rows, add
df = df.dropna(subset=['sentence'])before processing to remove them.
内容的提问来源于stack exchange,提问作者Rekha Sharma
相关产品推荐
相关产品推荐

