为Affinity Propagation编码词元及文本数据自然聚类方案咨询
Great question! Let's break this down into two key parts: converting your lemmatized text (plus other data types) into a format Affinity Propagation can handle, and whether this algorithm is actually the best choice for your task.
Part 1: Converting Text & Mixed Data to Numerical Features
Affinity Propagation works with numerical feature matrices (or precomputed similarity matrices), so we need to turn your text and other columns into structured numerical data. Here are the most effective approaches:
1. TF-IDF Vectorization (Most Common & Practical)
Since your text is already lemmatized, you can skip the lemmatization step in vectorizers. TF-IDF is ideal because it weights terms by their importance (downplaying common words like "the" that don't add clustering value).
- Use
sklearn.feature_extraction.text.TfidfVectorizerto convert each lemmatized text entry into a numerical vector. - Tweak parameters like
max_featuresto limit the number of terms (prevents overfitting) orngram_rangeto capture phrase patterns.
2. Word Embeddings (For Semantic Clustering)
If you want to capture deeper semantic meaning (e.g., grouping texts about "car" and "automobile" together), use word embeddings:
- Train a custom Word2Vec/GloVe model on your lemmatized text, or use pre-trained embeddings (like spaCy's
en_core_web_mdvectors). - Aggregate word vectors for each document (e.g., take the mean of all word vectors in the paragraph) to get a single document-level vector.
3. Integrate Non-Text Features
Don't ignore your int/float/datetime columns! Convert them to numerical values and combine with text features:
- For datetime columns: Convert to timestamps, or extract components like year/month/day as integers.
- Use
sklearn.compose.ColumnTransformerto seamlessly combine text vectorization, numerical scaling, and datetime processing into a single pipeline.
Example Pipeline Code
import pandas as pd from sklearn.feature_extraction.text import TfidfVectorizer from sklearn.compose import ColumnTransformer from sklearn.preprocessing import StandardScaler, FunctionTransformer from sklearn.cluster import AffinityPropagation # Sample DataFrame setup (adjust column names to match yours) df = pd.DataFrame({ "lemmatized_text": ["this is sample lemmatized text", "another example paragraph"], "int_col": [10, 20], "float_col": [3.14, 2.71], "date_col": pd.to_datetime(["2023-01-01", "2023-02-01"]) }) # Helper to convert datetime to timestamp def process_datetime(data): return data["date_col"].apply(lambda x: x.timestamp()).to_frame() # Build preprocessing pipeline preprocessor = ColumnTransformer( transformers=[ ("text", TfidfVectorizer(max_features=5000), "lemmatized_text"), ("numeric", StandardScaler(), ["int_col", "float_col"]), ("datetime", FunctionTransformer(process_datetime, validate=False), ["date_col"]) ] ) # Generate final feature matrix X = preprocessor.fit_transform(df) # Run Affinity Propagation ap = AffinityPropagation(damping=0.5, random_state=42) cluster_labels = ap.fit_predict(X)
Bonus: Use Precomputed Similarity Matrices
Affinity Propagation accepts precomputed similarity matrices (set affinity='precomputed'). For text, cosine similarity is more meaningful than Euclidean distance, so you can compute this directly from your text features:
from sklearn.metrics.pairwise import cosine_similarity # Get TF-IDF features first tfidf = TfidfVectorizer(max_features=5000) text_features = tfidf.fit_transform(df["lemmatized_text"]) # Compute cosine similarity matrix similarity_matrix = cosine_similarity(text_features) # Run Affinity Propagation with precomputed similarity ap = AffinityPropagation(affinity="precomputed", damping=0.5, random_state=42) cluster_labels = ap.fit_predict(similarity_matrix)
Part 2: Is Affinity Propagation the Optimal Choice?
It depends on your dataset size and goals:
- When it works well: If you have a small-to-medium dataset (hundreds to low thousands of samples) and don't want to pre-specify the number of clusters, Affinity Propagation can be a good fit. It automatically identifies cluster centers based on "messages" between samples.
- When to avoid it:
- Large datasets: It has O(n²) time complexity, so it becomes extremely slow with tens of thousands of samples.
- Stability concerns: It's sensitive to the
dampingparameter and can produce inconsistent cluster results across runs. - Dense clusters: If your text data has clear, dense clusters, algorithms like K-Means (faster) or HDBSCAN (better for irregular cluster shapes) may perform more reliably.
Alternatives to Consider
- K-Means: Fast, scalable, and widely used for text clustering (you'll need to choose
n_clusters, but tools like the elbow method can help). - HDBSCAN: No need to specify cluster count, handles outliers well, and works with density-based structures.
- DBSCAN: Great for datasets with irregularly shaped clusters, but requires tuning
epsandmin_samples.
内容的提问来源于stack exchange,提问作者LMGagne

