You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何对数百页大文本进行保连续性的段落级聚类(章节划分)

Hey there! Let’s break down how you can tackle this paragraph-level clustering for long texts while keeping continuity (aka splitting into logical chapters). It’s a niche problem, so it makes sense standard solutions don’t fit perfectly—here’s a practical starting path:

1. First: Preprocess Your Text Cleanly
  • Start by splitting your full text into logical paragraphs (not just line breaks—use rules like double newlines, or regex patterns like \n\n+ for plain text to catch proper paragraph separators).
  • Strip out noise: remove extra whitespace, standardize punctuation, and filter out repetitive boilerplate (like page headers/footers if it’s a scanned or formatted document).
  • Assign each paragraph a sequential index—this will be critical for enforcing continuity later on.
2. Extract Context-Aware Paragraph Features

Isolated paragraph embeddings won’t cut it here—you need features that capture how each paragraph connects to its neighbors:

  • Use a pre-trained language model (like BERT, RoBERTa, or the lightweight all-MiniLM-L6-v2) to generate embeddings. Instead of embedding just the single paragraph, create a "windowed embedding": concatenate the current paragraph’s embedding with the embeddings of the previous 1-2 paragraphs (tweak the window size based on how context-dependent your text is).
  • For simpler structured texts (like academic papers), combine topic features (TF-IDF on paragraph text) with positional features (e.g., normalized paragraph index, to weight where it falls in the overall text).
  • Another trick: Generate embeddings for sliding chunks of 3-5 consecutive paragraphs, then map each individual paragraph to its chunk’s embedding—this preserves the flow of ideas across adjacent paragraphs.
3. Use Constrained Clustering to Respect Sequence

Standard clustering (K-means, DBSCAN) ignores order, so you need methods that prioritize keeping consecutive paragraphs together:

  • Hierarchical Agglomerative Clustering (HAC) with a custom distance metric: Instead of using vanilla Euclidean distance, define a metric that penalizes splitting sequential paragraphs. For example, if two paragraphs are adjacent, reduce their distance by a fixed factor; or add a penalty if a cluster would separate a consecutive block.
  • Dynamic Time Warping (DTW) + Clustering: Treat the sequence of paragraph embeddings as a time series. DTW measures similarity between sequences, so you can cluster segments with similar patterns—this naturally keeps consecutive paragraphs grouped.
  • Constraint-based K-means: Add must-link constraints (adjacent paragraphs should stay in the same cluster unless there’s a strong topic shift) and cannot-link constraints (non-adjacent paragraphs with wildly different embeddings shouldn’t be grouped). While scikit-learn doesn’t have built-in constrained K-means, you can implement a simple version by adjusting cluster assignments to respect sequence, or use specialized libraries like CluSTer.
4. Post-Process to Refine Chapter Boundaries

Initial clusters will need cleanup to feel like natural chapters:

  • Merge small, isolated clusters that sit between larger, topically similar clusters (a single-paragraph cluster between two clusters about the same topic is almost certainly a mistake).
  • Identify clear topic shift points: Calculate cosine similarity between each paragraph’s embedding and the next, then set a threshold for what counts as a "significant shift"—these are your natural chapter breaks.
  • Manually review a sample of clusters to tweak your parameters (like window size or distance threshold)—since "chapter" is subjective, human feedback will help the model align with your needs.
Quick Starter Code Snippet (Python)

Here’s a minimal example to kick things off with windowed embeddings and constrained HAC:

import numpy as np
from sklearn.cluster import AgglomerativeClustering
from sentence_transformers import SentenceTransformer

# Load a lightweight pre-trained embedding model
model = SentenceTransformer('all-MiniLM-L6-v2')

# Assume `paragraphs` is your ordered list of text paragraphs
paragraphs = ["Paragraph 1 content...", "Paragraph 2 content...", ...]

# Generate windowed embeddings (current + previous paragraph)
windowed_embeddings = []
for idx, para in enumerate(paragraphs):
    curr_emb = model.encode(para)
    if idx == 0:
        # First paragraph uses only its own embedding
        windowed_embeddings.append(curr_emb)
    else:
        # Combine current and previous paragraph embeddings
        prev_emb = model.encode(paragraphs[idx-1])
        combined_emb = np.concatenate([curr_emb, prev_emb])
        windowed_embeddings.append(combined_emb)

# Custom distance matrix that penalizes non-consecutive splits
def custom_distance_matrix(X):
    num_paras = len(X)
    dist_matrix = np.zeros((num_paras, num_paras))
    for i in range(num_paras):
        for j in range(num_paras):
            # Base cosine distance between embeddings
            cosine_dist = 1 - np.dot(X[i], X[j])/(np.linalg.norm(X[i])*np.linalg.norm(X[j]))
            # Add penalty for non-adjacent pairs to discourage splitting sequence
            if abs(i - j) > 1:
                cosine_dist += 0.2  # Adjust penalty weight based on your text
            dist_matrix[i][j] = cosine_dist
    return dist_matrix

# Run hierarchical clustering with custom distance
clustering = AgglomerativeClustering(
    n_clusters=None,
    distance_threshold=0.5,  # Adjust to control number of chapters
    linkage='average',
    affinity='precomputed'
)
chapter_labels = clustering.fit_predict(custom_distance_matrix(windowed_embeddings))

# `chapter_labels` now assigns each paragraph to a chapter cluster
Key Parameters to Experiment With
  • Window size for embeddings (1-3 previous paragraphs works well for most texts)
  • Penalty weight in the custom distance metric
  • Distance threshold in HAC (higher = fewer chapters, lower = more granular splits)
  • Embedding models (try all-mpnet-base-v2 for better accuracy if speed isn’t a constraint)

内容的提问来源于stack exchange,提问作者ashcrok

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 09:55:20