You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

请求Doc2Vec单文档句向量生成及聚类实现代码(Python)

Doc2Vec for Sentence Vectors & Clustering: Step-by-Step Solutions

Hey there! Let's break down your two needs into clear, newbie-friendly code and explanations. I'll build on the bits of code you already tried to keep things consistent.


1. Getting Sentence Vectors from a Single Document

Doc2Vec works by treating each sentence as its own "document" during training. Here's how to extract vectors for every sentence in your single text file:

First, Install Required Tools

Run this in your terminal if you haven't already:

pip install gensim nltk pandas scikit-learn matplotlib

Complete Code

import nltk
from gensim.models.doc2vec import Doc2Vec, TaggedDocument
from nltk.tokenize import word_tokenize, sent_tokenize

# Download NLTK's tokenizer data (run once)
nltk.download('punkt')

# Load your single document (replace with your file path or raw text)
with open("your_document.txt", "r", encoding="utf-8") as f:
    document_text = f.read()

# Split the document into individual sentences
sentences = sent_tokenize(document_text)

# Preprocess each sentence: tokenize, lowercase, clean up
processed_sentences = []
for idx, sentence in enumerate(sentences):
    # Split sentence into words and lowercase them
    tokens = word_tokenize(sentence.lower())
    # Keep only alphanumeric words (filters out punctuation)
    tokens = [token for token in tokens if token.isalnum()]
    # Create a TaggedDocument (required for Doc2Vec training)
    processed_sentences.append(TaggedDocument(tokens, [idx]))

# Train the Doc2Vec model
model = Doc2Vec(
    vector_size=100,  # Size of the output vector (adjust based on your needs)
    min_count=1,      # Ignore words that appear less than once
    epochs=50,        # Number of training loops
    dm=1              # Use "distributed memory" mode (good for sentence context)
)

# Build vocabulary from our processed sentences
model.build_vocab(processed_sentences)

# Train the model on our sentence-documents
model.train(processed_sentences, total_examples=model.corpus_count, epochs=model.epochs)

# Generate vectors for each sentence
sentence_vectors = [model.infer_vector(doc.words) for doc in processed_sentences]

# Check the results
print(f"Generated {len(sentence_vectors)} sentence vectors!")
print("Sample vector for first sentence:", sentence_vectors[0])

Quick Explanation

  • We split your document into sentences using NLTK's sent_tokenize tool.
  • Each sentence gets converted into a TaggedDocument (a required format for Doc2Vec) with a unique tag (the sentence index).
  • The model learns to map each sentence to a numerical vector that captures its meaning.
  • infer_vector generates the final vector for each sentence—even though we trained on these sentences, this step ensures consistent, usable vectors.

2. Clustering Sentences from kkk.csv

Let's process your CSV file, generate sentence vectors, and cluster similar sentences using KMeans. We'll even add a visualization to see how clusters group together.

Complete Code

import pandas as pd
import nltk
from gensim.models.doc2vec import Doc2Vec, TaggedDocument
from nltk.tokenize import word_tokenize
from sklearn.cluster import KMeans
from sklearn.decomposition import PCA
import matplotlib.pyplot as plt

nltk.download('punkt')

# Load your CSV file (adjust column name if yours isn't 'sentence')
df = pd.read_csv("kkk.csv")
sentences = df['sentence'].tolist()  # Replace 'sentence' with your actual column name

# Preprocess sentences (same as before)
processed_sentences = []
for idx, sentence in enumerate(sentences):
    # Convert to string to avoid errors with missing values
    tokens = word_tokenize(str(sentence).lower())
    tokens = [token for token in tokens if token.isalnum()]
    processed_sentences.append(TaggedDocument(tokens, [idx]))

# Train Doc2Vec model
model = Doc2Vec(
    vector_size=100,
    min_count=1,
    epochs=50,
    dm=1
)
model.build_vocab(processed_sentences)
model.train(processed_sentences, total_examples=model.corpus_count, epochs=model.epochs)

# Get sentence vectors
sentence_vectors = [model.infer_vector(doc.words) for doc in processed_sentences]

# Find the optimal number of clusters (elbow method)
inertia = []
k_range = range(1, 11)
for k in k_range:
    kmeans = KMeans(n_clusters=k, random_state=42)
    kmeans.fit(sentence_vectors)
    inertia.append(kmeans.inertia_)

# Plot elbow curve to pick your k
plt.figure(figsize=(8,4))
plt.plot(k_range, inertia, 'bo-')
plt.xlabel('Number of Clusters (k)')
plt.ylabel('Inertia')
plt.title('Elbow Method to Choose Optimal k')
plt.show()

# Choose your k (based on the elbow plot—try 3-5 for 101 sentences)
num_clusters = 3
kmeans = KMeans(n_clusters=num_clusters, random_state=42)
cluster_labels = kmeans.fit_predict(sentence_vectors)

# Add cluster labels to your original dataframe
df['cluster'] = cluster_labels

# Print sentences grouped by cluster
for cluster in range(num_clusters):
    print(f"\n--- Cluster {cluster} ---")
    cluster_sentences = df[df['cluster'] == cluster]['sentence'].tolist()
    for sent in cluster_sentences:
        print(f"- {sent}")

# Visualize clusters in 2D using PCA
pca = PCA(n_components=2)
reduced_vectors = pca.fit_transform(sentence_vectors)

plt.figure(figsize=(10,6))
scatter = plt.scatter(reduced_vectors[:,0], reduced_vectors[:,1], c=cluster_labels, cmap='viridis')
plt.legend(handles=scatter.legend_elements()[0], labels=[f"Cluster {i}" for i in range(num_clusters)])
plt.title('Sentence Clusters (2D PCA Projection)')
plt.show()

Key Tips for Newbies

  • CSV Column Name: If your CSV doesn't use sentence as the column header, replace that string with your actual column name (e.g., df['text']).
  • Choosing k: The elbow plot shows where adding more clusters stops improving the grouping (the "elbow" point). For 101 sentences, k=3-5 is a safe starting point.
  • Missing Values: If your CSV has empty rows, add df = df.dropna(subset=['sentence']) before processing to remove them.

内容的提问来源于stack exchange,提问作者Rekha Sharma

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 08:23:34