文档聚类优化Pandas列相似推荐字符串合并的技术咨询
Got it, let's tackle this problem step by step. Dealing with 188k similar recommendation strings is tricky, especially since fuzzy matching is too slow and your initial TF-IDF+KMeans setup isn't grouping semantically similar phrases like "replace" and "change" together. Here are some practical, actionable solutions to fix this:
1. Fix Your TF-IDF + KMeans Pipeline First
Your current setup isn't capturing semantic similarity (e.g., "replace" vs "change") because TF-IDF relies on exact word matches, not meaning. Let's tweak it:
a. Preprocess to Remove Noise & Focus on Core Content
Most of your strings share the fixed prefix "It is recommended to"—this adds no value for clustering. Strip it out, and extract the core action + object instead:
# Strip the fixed prefix and clean text def clean_recommendation(text): prefix = "It is recommended to " if text.startswith(prefix): return text[len(prefix):].strip() return text hardware['core_text'] = hardware['resolution_modified'].apply(clean_recommendation)
b. Replace Synonyms with High-Frequency Terms
Since you want to keep the most frequent string as the cluster representative, pre-map synonyms to their highest-occurrence counterparts first:
from collections import Counter import spacy nlp = spacy.load("en_core_web_sm") # Extract all root verbs from core texts to find synonyms verbs = [] for text in hardware['core_text']: doc = nlp(text) for token in doc: if token.dep_ == 'ROOT' and token.pos_ == 'VERB': verbs.append(token.lemma_) # Get the most common verbs to build synonym mappings verb_counts = Counter(verbs) # Example mapping: adjust based on your actual data synonym_map = { "change": "replace", "swap": "replace", "install": "replace", "image": "reimage" # If "image scanner" means reimaging it } def replace_synonyms(text): doc = nlp(text) tokens = [] for token in doc: tokens.append(synonym_map.get(token.lemma_, token.text)) return " ".join(tokens) hardware['processed_text'] = hardware['core_text'].apply(replace_synonyms)
c. Tune KMeans Cluster Count
Setting n_clusters=30 arbitrarily is likely hurting results. Use the elbow method to find the optimal number of clusters:
from sklearn.cluster import KMeans import matplotlib.pyplot as plt # Recompute TF-IDF with cleaned text v = TfidfVectorizer(max_features=6000, ngram_range=(1,3), stop_words='english', strip_accents='ascii') x = v.fit_transform(hardware['processed_text']) # Test cluster counts and plot SSE sse = [] for k in range(50, 200, 20): kmeans = KMeans(n_clusters=k, random_state=42) kmeans.fit(x) sse.append(kmeans.inertia_) plt.plot(range(50, 200, 20), sse) plt.xlabel("Number of Clusters") plt.ylabel("SSE (Sum of Squared Errors)") plt.show()
Pick the cluster count where the SSE curve starts to flatten (the "elbow" point).
2. Switch to Semantic Embeddings (Better for Short Text)
TF-IDF fails with synonyms because it doesn't understand word meaning. Use sentence-level embeddings like Sentence-BERT—they're designed to capture semantic similarity for short texts:
a. Generate Sentence Embeddings
from sentence_transformers import SentenceTransformer # Use a lightweight, fast model for large datasets model = SentenceTransformer('all-MiniLM-L6-v2') embeddings = model.encode(hardware['processed_text'], show_progress_bar=True)
b. Cluster with KMeans or DBSCAN
KMeans: Use the elbow method on the embeddings to find optimal clusters, then assign each cluster to its most frequent original string:
# Assume optimal k is 150 from elbow analysis kmeans = KMeans(n_clusters=150, random_state=42) hardware['cluster_id'] = kmeans.fit_predict(embeddings) # Map each cluster to its most frequent recommendation cluster_reps = {} for cluster in hardware['cluster_id'].unique(): cluster_texts = hardware[hardware['cluster_id'] == cluster]['resolution_modified'] cluster_reps[cluster] = cluster_texts.value_counts().index[0] hardware['unified_recommendation'] = hardware['cluster_id'].map(cluster_reps)DBSCAN: If you don't want to predefine cluster counts, DBSCAN groups dense regions. For Sentence-BERT embeddings (normalized), start with
eps=0.5andmin_samples=5:from sklearn.cluster import DBSCAN dbscan = DBSCAN(eps=0.5, min_samples=5, metric='cosine') hardware['cluster_id'] = dbscan.fit_predict(embeddings) # Handle noise points (cluster_id = -1) by assigning them to their closest cluster or keeping as-is
3. Fix Memory Issues with Hierarchical Clustering
AgglomerativeClustering is O(n²), which is impossible for 188k rows. Instead:
- Pre-reduce dimensionality: Use PCA to shrink Sentence-BERT embeddings to 20-50 dimensions before clustering.
- Use MiniBatchKMeans: It's faster and uses less memory than standard KMeans, with nearly identical results.
4. Quick Win: Deduplicate First
Before any clustering, remove exact duplicates to reduce your dataset size:
hardware = hardware.drop_duplicates(subset='resolution_modified')
This can cut down your 188k rows significantly, making all subsequent steps faster.
内容的提问来源于stack exchange,提问作者justanewb

