基于scikit-learn:Pipeline+GridSearchCV的KMeans文本聚类参数调优
Hey there! Let's dive into your text clustering pipeline setup and how to make GridSearchCV work effectively for tuning those preprocessing parameters. I’ll break down the key parts, fixes, and optimizations you need:
First, let's unpack each component of your pipeline to spot strengths and areas to refine:
vect1: A initialCountVectorizerthat turns raw text into word counts. This feeds into your custom lemmatization step.MyLemmanization: Your custom transformer for lemmatizing tokens—critical note: this class must implement scikit-learn'sfit()andtransform()methods to play nice with the pipeline.vect2: A second vectorizer that cleans up the lemmatized tokens by filtering out overly common/rare terms (max_df/min_df) and limiting vocabulary size (max_features).tfidf: Converts raw word counts to TF-IDF weights, which usually improves clustering performance by downplaying frequent, low-information words.clust: KMeans clustering with 10 clusters, usingk-means++initialization to get better starting centroids (a smart choice!).
The biggest hurdle here is ensuring your custom MyLemmanization class is scikit-learn compatible. Here's a polished version that handles sparse matrices from CountVectorizer:
from sklearn.base import BaseEstimator, TransformerMixin import nltk from nltk.stem import WordNetLemmatizer class MyLemmanization(BaseEstimator, TransformerMixin): def __init__(self, lemmatize=True, leave_other_words=True): self.lemmatize = lemmatize self.leave_other_words = leave_other_words self.lemmatizer = WordNetLemmatizer() def fit(self, X, y=None): # No fitting needed for lemmatization—just return self return self def transform(self, X): transformed_docs = [] # Convert sparse matrix from vect1 back to token lists for doc in X: tokens = self.vect1.inverse_transform(doc)[0] lemmatized = [] for token in tokens: if self.lemmatize: lemmatized.append(self.lemmatizer.lemmatize(token)) else: lemmatized.append(token) # Add logic for leave_other_words here (e.g., filter stopwords) if not self.leave_other_words: lemmatized = [t for t in lemmatized if t not in nltk.corpus.stopwords.words('english')] transformed_docs.append(' '.join(lemmatized)) return transformed_docs
Pro tip: Using two CountVectorizer steps is redundant. A cleaner approach is to combine tokenization and lemmatization into a single custom tokenizer for one vectorizer—we’ll cover that later!
Next, define your parameter grid and hook up GridSearchCV. Since clustering is unsupervised, we’ll use silhouette score (a metric that measures cluster separation) to rank parameter combinations:
from sklearn.model_selection import GridSearchCV from sklearn.metrics import silhouette_score, make_scorer # Define parameters to tune across preprocessing steps param_grid = { # vect1 parameters 'vect1__ngram_range': [(1,1), (1,2)], # Test unigrams vs unigrams+bigrams 'vect1__stop_words': [None, 'english'], # Remove stopwords or not # Custom lemmatizer parameters 'myfun__lemmatize': [True, False], 'myfun__leave_other_words': [True, False], # vect2 parameters 'vect2__max_df': [0.9, 0.95, 0.99], 'vect2__min_df': [1, 2, 3], 'vect2__max_features': [1000, 2000, 3000], # TF-IDF parameters 'tfidf__use_idf': [True, False], # Optional: Tune KMeans itself 'clust__n_clusters': [8, 10, 12] } # Create a scorer for unsupervised learning silhouette_scorer = make_scorer(silhouette_score, metric='euclidean') # Initialize GridSearchCV grid_search = GridSearchCV( estimator=text_clf, param_grid=param_grid, scoring=silhouette_scorer, cv=5, # 5-fold cross-validation n_jobs=-1, # Use all CPU cores for speed verbose=2 ) # Fit to your raw text corpus (replace X with your data) grid_search.fit(X) # Get the best results print("Best preprocessing parameters:", grid_search.best_params_) print("Best silhouette score:", grid_search.best_score_)
Simplify the Pipeline: Ditch the two vectorizers! Use a single
CountVectorizerwith a custom lemmatizing tokenizer:def lemmatized_tokenizer(text): tokens = nltk.word_tokenize(text) lemmatizer = WordNetLemmatizer() return [lemmatizer.lemmatize(token) for token in tokens] # Revised pipeline text_clf = Pipeline([ ('vect', CountVectorizer( analyzer="word", tokenizer=lemmatized_tokenizer, max_df=0.95, min_df=2, max_features=2000, stop_words='english' )), ('tfidf', TfidfTransformer()), ('clust', KMeans(n_clusters=10, init='k-means++', max_iter=100, n_init=10, verbose=1)) ])This cuts down on complexity and avoids sparse matrix conversion headaches.
Tune KMeans Stability: Your current
n_init=1is risky—KMeans can get stuck in local minima. Bump this ton_init=10(the default) for more consistent cluster results.Choose the Right Scoring Metric: Silhouette score (1 = perfect separation, -1 = poor separation) is intuitive, but Calinski-Harabasz score (higher = better clusters) is faster to compute for large datasets. Pick based on your data size.
Once you have your best pipeline:
- Extract it with
best_clf = grid_search.best_estimator_ - Generate cluster labels with
labels = best_clf.predict(X) - Analyze cluster quality by checking top terms in each cluster (use
best_clf.named_steps['vect'].get_feature_names_out()and KMeans centroids)
内容的提问来源于stack exchange,提问作者Lukas

