You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于scikit-learn:Pipeline+GridSearchCV的KMeans文本聚类参数调优

Hey there! Let's dive into your text clustering pipeline setup and how to make GridSearchCV work effectively for tuning those preprocessing parameters. I’ll break down the key parts, fixes, and optimizations you need:

1. Understanding Your Current Pipeline

First, let's unpack each component of your pipeline to spot strengths and areas to refine:

  • vect1: A initial CountVectorizer that turns raw text into word counts. This feeds into your custom lemmatization step.
  • MyLemmanization: Your custom transformer for lemmatizing tokens—critical note: this class must implement scikit-learn's fit() and transform() methods to play nice with the pipeline.
  • vect2: A second vectorizer that cleans up the lemmatized tokens by filtering out overly common/rare terms (max_df/min_df) and limiting vocabulary size (max_features).
  • tfidf: Converts raw word counts to TF-IDF weights, which usually improves clustering performance by downplaying frequent, low-information words.
  • clust: KMeans clustering with 10 clusters, using k-means++ initialization to get better starting centroids (a smart choice!).
2. Making the Pipeline GridSearchCV-Ready

The biggest hurdle here is ensuring your custom MyLemmanization class is scikit-learn compatible. Here's a polished version that handles sparse matrices from CountVectorizer:

from sklearn.base import BaseEstimator, TransformerMixin
import nltk
from nltk.stem import WordNetLemmatizer

class MyLemmanization(BaseEstimator, TransformerMixin):
    def __init__(self, lemmatize=True, leave_other_words=True):
        self.lemmatize = lemmatize
        self.leave_other_words = leave_other_words
        self.lemmatizer = WordNetLemmatizer()
    
    def fit(self, X, y=None):
        # No fitting needed for lemmatization—just return self
        return self
    
    def transform(self, X):
        transformed_docs = []
        # Convert sparse matrix from vect1 back to token lists
        for doc in X:
            tokens = self.vect1.inverse_transform(doc)[0]
            lemmatized = []
            for token in tokens:
                if self.lemmatize:
                    lemmatized.append(self.lemmatizer.lemmatize(token))
                else:
                    lemmatized.append(token)
            # Add logic for leave_other_words here (e.g., filter stopwords)
            if not self.leave_other_words:
                lemmatized = [t for t in lemmatized if t not in nltk.corpus.stopwords.words('english')]
            transformed_docs.append(' '.join(lemmatized))
        return transformed_docs

Pro tip: Using two CountVectorizer steps is redundant. A cleaner approach is to combine tokenization and lemmatization into a single custom tokenizer for one vectorizer—we’ll cover that later!

Next, define your parameter grid and hook up GridSearchCV. Since clustering is unsupervised, we’ll use silhouette score (a metric that measures cluster separation) to rank parameter combinations:

from sklearn.model_selection import GridSearchCV
from sklearn.metrics import silhouette_score, make_scorer

# Define parameters to tune across preprocessing steps
param_grid = {
    # vect1 parameters
    'vect1__ngram_range': [(1,1), (1,2)],  # Test unigrams vs unigrams+bigrams
    'vect1__stop_words': [None, 'english'],  # Remove stopwords or not
    # Custom lemmatizer parameters
    'myfun__lemmatize': [True, False],
    'myfun__leave_other_words': [True, False],
    # vect2 parameters
    'vect2__max_df': [0.9, 0.95, 0.99],
    'vect2__min_df': [1, 2, 3],
    'vect2__max_features': [1000, 2000, 3000],
    # TF-IDF parameters
    'tfidf__use_idf': [True, False],
    # Optional: Tune KMeans itself
    'clust__n_clusters': [8, 10, 12]
}

# Create a scorer for unsupervised learning
silhouette_scorer = make_scorer(silhouette_score, metric='euclidean')

# Initialize GridSearchCV
grid_search = GridSearchCV(
    estimator=text_clf,
    param_grid=param_grid,
    scoring=silhouette_scorer,
    cv=5,  # 5-fold cross-validation
    n_jobs=-1,  # Use all CPU cores for speed
    verbose=2
)

# Fit to your raw text corpus (replace X with your data)
grid_search.fit(X)

# Get the best results
print("Best preprocessing parameters:", grid_search.best_params_)
print("Best silhouette score:", grid_search.best_score_)
3. Key Optimizations to Simplify & Improve
  • Simplify the Pipeline: Ditch the two vectorizers! Use a single CountVectorizer with a custom lemmatizing tokenizer:

    def lemmatized_tokenizer(text):
        tokens = nltk.word_tokenize(text)
        lemmatizer = WordNetLemmatizer()
        return [lemmatizer.lemmatize(token) for token in tokens]
    
    # Revised pipeline
    text_clf = Pipeline([
        ('vect', CountVectorizer(
            analyzer="word",
            tokenizer=lemmatized_tokenizer,
            max_df=0.95,
            min_df=2,
            max_features=2000,
            stop_words='english'
        )),
        ('tfidf', TfidfTransformer()),
        ('clust', KMeans(n_clusters=10, init='k-means++', max_iter=100, n_init=10, verbose=1))
    ])
    

    This cuts down on complexity and avoids sparse matrix conversion headaches.

  • Tune KMeans Stability: Your current n_init=1 is risky—KMeans can get stuck in local minima. Bump this to n_init=10 (the default) for more consistent cluster results.

  • Choose the Right Scoring Metric: Silhouette score (1 = perfect separation, -1 = poor separation) is intuitive, but Calinski-Harabasz score (higher = better clusters) is faster to compute for large datasets. Pick based on your data size.

4. Post-GridSearch Next Steps

Once you have your best pipeline:

  1. Extract it with best_clf = grid_search.best_estimator_
  2. Generate cluster labels with labels = best_clf.predict(X)
  3. Analyze cluster quality by checking top terms in each cluster (use best_clf.named_steps['vect'].get_feature_names_out() and KMeans centroids)

内容的提问来源于stack exchange,提问作者Lukas

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 07:42:37