You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为Affinity Propagation编码词元及文本数据自然聚类方案咨询

Text Feature Conversion for Affinity Propagation & Algorithm Suitability

Great question! Let's break this down into two key parts: converting your lemmatized text (plus other data types) into a format Affinity Propagation can handle, and whether this algorithm is actually the best choice for your task.

Part 1: Converting Text & Mixed Data to Numerical Features

Affinity Propagation works with numerical feature matrices (or precomputed similarity matrices), so we need to turn your text and other columns into structured numerical data. Here are the most effective approaches:

1. TF-IDF Vectorization (Most Common & Practical)

Since your text is already lemmatized, you can skip the lemmatization step in vectorizers. TF-IDF is ideal because it weights terms by their importance (downplaying common words like "the" that don't add clustering value).

  • Use sklearn.feature_extraction.text.TfidfVectorizer to convert each lemmatized text entry into a numerical vector.
  • Tweak parameters like max_features to limit the number of terms (prevents overfitting) or ngram_range to capture phrase patterns.

2. Word Embeddings (For Semantic Clustering)

If you want to capture deeper semantic meaning (e.g., grouping texts about "car" and "automobile" together), use word embeddings:

  • Train a custom Word2Vec/GloVe model on your lemmatized text, or use pre-trained embeddings (like spaCy's en_core_web_md vectors).
  • Aggregate word vectors for each document (e.g., take the mean of all word vectors in the paragraph) to get a single document-level vector.

3. Integrate Non-Text Features

Don't ignore your int/float/datetime columns! Convert them to numerical values and combine with text features:

  • For datetime columns: Convert to timestamps, or extract components like year/month/day as integers.
  • Use sklearn.compose.ColumnTransformer to seamlessly combine text vectorization, numerical scaling, and datetime processing into a single pipeline.

Example Pipeline Code

import pandas as pd
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import StandardScaler, FunctionTransformer
from sklearn.cluster import AffinityPropagation

# Sample DataFrame setup (adjust column names to match yours)
df = pd.DataFrame({
    "lemmatized_text": ["this is sample lemmatized text", "another example paragraph"],
    "int_col": [10, 20],
    "float_col": [3.14, 2.71],
    "date_col": pd.to_datetime(["2023-01-01", "2023-02-01"])
})

# Helper to convert datetime to timestamp
def process_datetime(data):
    return data["date_col"].apply(lambda x: x.timestamp()).to_frame()

# Build preprocessing pipeline
preprocessor = ColumnTransformer(
    transformers=[
        ("text", TfidfVectorizer(max_features=5000), "lemmatized_text"),
        ("numeric", StandardScaler(), ["int_col", "float_col"]),
        ("datetime", FunctionTransformer(process_datetime, validate=False), ["date_col"])
    ]
)

# Generate final feature matrix
X = preprocessor.fit_transform(df)

# Run Affinity Propagation
ap = AffinityPropagation(damping=0.5, random_state=42)
cluster_labels = ap.fit_predict(X)

Bonus: Use Precomputed Similarity Matrices

Affinity Propagation accepts precomputed similarity matrices (set affinity='precomputed'). For text, cosine similarity is more meaningful than Euclidean distance, so you can compute this directly from your text features:

from sklearn.metrics.pairwise import cosine_similarity

# Get TF-IDF features first
tfidf = TfidfVectorizer(max_features=5000)
text_features = tfidf.fit_transform(df["lemmatized_text"])

# Compute cosine similarity matrix
similarity_matrix = cosine_similarity(text_features)

# Run Affinity Propagation with precomputed similarity
ap = AffinityPropagation(affinity="precomputed", damping=0.5, random_state=42)
cluster_labels = ap.fit_predict(similarity_matrix)

Part 2: Is Affinity Propagation the Optimal Choice?

It depends on your dataset size and goals:

  • When it works well: If you have a small-to-medium dataset (hundreds to low thousands of samples) and don't want to pre-specify the number of clusters, Affinity Propagation can be a good fit. It automatically identifies cluster centers based on "messages" between samples.
  • When to avoid it:
    • Large datasets: It has O(n²) time complexity, so it becomes extremely slow with tens of thousands of samples.
    • Stability concerns: It's sensitive to the damping parameter and can produce inconsistent cluster results across runs.
    • Dense clusters: If your text data has clear, dense clusters, algorithms like K-Means (faster) or HDBSCAN (better for irregular cluster shapes) may perform more reliably.

Alternatives to Consider

  • K-Means: Fast, scalable, and widely used for text clustering (you'll need to choose n_clusters, but tools like the elbow method can help).
  • HDBSCAN: No need to specify cluster count, handles outliers well, and works with density-based structures.
  • DBSCAN: Great for datasets with irregularly shaped clusters, but requires tuning eps and min_samples.

内容的提问来源于stack exchange,提问作者LMGagne

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 04:02:14