如何对数百页大文本进行保连续性的段落级聚类(章节划分)
Hey there! Let’s break down how you can tackle this paragraph-level clustering for long texts while keeping continuity (aka splitting into logical chapters). It’s a niche problem, so it makes sense standard solutions don’t fit perfectly—here’s a practical starting path:
- Start by splitting your full text into logical paragraphs (not just line breaks—use rules like double newlines, or regex patterns like
\n\n+for plain text to catch proper paragraph separators). - Strip out noise: remove extra whitespace, standardize punctuation, and filter out repetitive boilerplate (like page headers/footers if it’s a scanned or formatted document).
- Assign each paragraph a sequential index—this will be critical for enforcing continuity later on.
Isolated paragraph embeddings won’t cut it here—you need features that capture how each paragraph connects to its neighbors:
- Use a pre-trained language model (like BERT, RoBERTa, or the lightweight
all-MiniLM-L6-v2) to generate embeddings. Instead of embedding just the single paragraph, create a "windowed embedding": concatenate the current paragraph’s embedding with the embeddings of the previous 1-2 paragraphs (tweak the window size based on how context-dependent your text is). - For simpler structured texts (like academic papers), combine topic features (TF-IDF on paragraph text) with positional features (e.g., normalized paragraph index, to weight where it falls in the overall text).
- Another trick: Generate embeddings for sliding chunks of 3-5 consecutive paragraphs, then map each individual paragraph to its chunk’s embedding—this preserves the flow of ideas across adjacent paragraphs.
Standard clustering (K-means, DBSCAN) ignores order, so you need methods that prioritize keeping consecutive paragraphs together:
- Hierarchical Agglomerative Clustering (HAC) with a custom distance metric: Instead of using vanilla Euclidean distance, define a metric that penalizes splitting sequential paragraphs. For example, if two paragraphs are adjacent, reduce their distance by a fixed factor; or add a penalty if a cluster would separate a consecutive block.
- Dynamic Time Warping (DTW) + Clustering: Treat the sequence of paragraph embeddings as a time series. DTW measures similarity between sequences, so you can cluster segments with similar patterns—this naturally keeps consecutive paragraphs grouped.
- Constraint-based K-means: Add must-link constraints (adjacent paragraphs should stay in the same cluster unless there’s a strong topic shift) and cannot-link constraints (non-adjacent paragraphs with wildly different embeddings shouldn’t be grouped). While
scikit-learndoesn’t have built-in constrained K-means, you can implement a simple version by adjusting cluster assignments to respect sequence, or use specialized libraries likeCluSTer.
Initial clusters will need cleanup to feel like natural chapters:
- Merge small, isolated clusters that sit between larger, topically similar clusters (a single-paragraph cluster between two clusters about the same topic is almost certainly a mistake).
- Identify clear topic shift points: Calculate cosine similarity between each paragraph’s embedding and the next, then set a threshold for what counts as a "significant shift"—these are your natural chapter breaks.
- Manually review a sample of clusters to tweak your parameters (like window size or distance threshold)—since "chapter" is subjective, human feedback will help the model align with your needs.
Here’s a minimal example to kick things off with windowed embeddings and constrained HAC:
import numpy as np from sklearn.cluster import AgglomerativeClustering from sentence_transformers import SentenceTransformer # Load a lightweight pre-trained embedding model model = SentenceTransformer('all-MiniLM-L6-v2') # Assume `paragraphs` is your ordered list of text paragraphs paragraphs = ["Paragraph 1 content...", "Paragraph 2 content...", ...] # Generate windowed embeddings (current + previous paragraph) windowed_embeddings = [] for idx, para in enumerate(paragraphs): curr_emb = model.encode(para) if idx == 0: # First paragraph uses only its own embedding windowed_embeddings.append(curr_emb) else: # Combine current and previous paragraph embeddings prev_emb = model.encode(paragraphs[idx-1]) combined_emb = np.concatenate([curr_emb, prev_emb]) windowed_embeddings.append(combined_emb) # Custom distance matrix that penalizes non-consecutive splits def custom_distance_matrix(X): num_paras = len(X) dist_matrix = np.zeros((num_paras, num_paras)) for i in range(num_paras): for j in range(num_paras): # Base cosine distance between embeddings cosine_dist = 1 - np.dot(X[i], X[j])/(np.linalg.norm(X[i])*np.linalg.norm(X[j])) # Add penalty for non-adjacent pairs to discourage splitting sequence if abs(i - j) > 1: cosine_dist += 0.2 # Adjust penalty weight based on your text dist_matrix[i][j] = cosine_dist return dist_matrix # Run hierarchical clustering with custom distance clustering = AgglomerativeClustering( n_clusters=None, distance_threshold=0.5, # Adjust to control number of chapters linkage='average', affinity='precomputed' ) chapter_labels = clustering.fit_predict(custom_distance_matrix(windowed_embeddings)) # `chapter_labels` now assigns each paragraph to a chapter cluster
- Window size for embeddings (1-3 previous paragraphs works well for most texts)
- Penalty weight in the custom distance metric
- Distance threshold in HAC (higher = fewer chapters, lower = more granular splits)
- Embedding models (try
all-mpnet-base-v2for better accuracy if speed isn’t a constraint)
内容的提问来源于stack exchange,提问作者ashcrok

