如何去除DataFrame重复样本并生成带权重新表?DBSCAN内存优化咨询
Great question! Removing duplicate samples and leveraging their occurrence counts is an excellent solution to fix DBSCAN's memory issues with large datasets—and it won’t compromise your clustering results. Here’s a breakdown of why this works and how to implement it step-by-step:
Why This Fix Works
DBSCAN’s memory bottleneck comes from its need to calculate pairwise distances (or run nearest neighbor searches) across all samples. With 800k samples, this leads to unsustainable O(n²) memory usage. By keeping only unique samples (and tracking how often each repeats), you drastically reduce the number of samples n (e.g., if 50% of your data is duplicates, n drops to 400k, or even lower if duplicates are more frequent).
Since duplicate samples have identical features, they will always fall into the same cluster (or be marked as noise) in DBSCAN. This means we can safely cluster only the unique samples, then map the results back to the full dataset without losing any meaningful information.
Step 1: Remove Duplicates and Track Occurrences
First, we’ll identify unique samples in your DataFrame, assign each a unique ID for mapping, and count how many times each unique sample appears (useful for algorithms like K-Means that support weights):
# Define which columns are your features (skip the first column assuming it's a label) feature_cols = df.columns[1:] # Add a unique ID for each group of duplicate samples df['duplicate_id'] = df.groupby(feature_cols.tolist()).ngroup() # Extract only unique samples (keep the first occurrence of each duplicate group) df_unique = df.drop_duplicates(subset=feature_cols).sort_values('duplicate_id') # Count occurrences (weight) of each unique sample df_unique['weight'] = df.groupby('duplicate_id').size().values # Prepare feature array for unique samples train_unique = np.array(df_unique[feature_cols])
Step 2: Cluster Unique Samples and Map Results
Run DBSCAN on the reduced unique dataset (this will use a fraction of the original memory), then map the cluster labels back to every sample in your original dataset:
def Cluster(data, algo="DBSCAN"): if algo == "DBSCAN": dbscan = DBSCAN(eps=0.9, min_samples=1000, metric="euclidean", n_jobs=1) y_result = dbscan.fit_predict(data) centroidi = "In DBSCAN there are not Centroids" return y_result, centroidi # Add K-Means support here if needed (see note below) # Run DBSCAN on unique samples y_unique, centroidi = Cluster(data=train_unique, algo="DBSCAN") # Create a mapping from duplicate ID to cluster label id_to_label = dict(zip(df_unique['duplicate_id'], y_unique)) # Map labels back to the full original dataset y_dbscan = df['duplicate_id'].map(id_to_label).values
Bonus: Optimizing K-Means with Weights
Unlike DBSCAN, scikit-learn’s KMeans supports a sample_weight parameter. For K-Means, you can train directly on the unique samples using their occurrence counts as weights—this produces identical results to training on the full dataset but uses far less memory:
from sklearn.cluster import KMeans def Cluster(data, weights=None, algo="KMeans"): if algo == "KMeans": # Adjust n_clusters to match your task requirements kmeans = KMeans(n_clusters=5, n_init='auto') y_result = kmeans.fit_predict(data, sample_weight=weights) centroidi = kmeans.cluster_centers_ return y_result, centroidi # Run K-Means with weights y_kmeans, centroids = Cluster(data=train_unique, weights=df_unique['weight'].values, algo="KMeans") # Map results back to full dataset if needed id_to_kmeans_label = dict(zip(df_unique['duplicate_id'], y_kmeans)) y_kmeans_full = df['duplicate_id'].map(id_to_kmeans_label).values
Final Notes
- This approach will completely resolve your memory issues because you’re only clustering a fraction of the original data.
- For DBSCAN, the mapped results are identical to clustering the full dataset—no information is lost.
- For K-Means, using weights ensures the centroid calculations account for duplicate samples correctly.
内容的提问来源于stack exchange,提问作者Davide Aureli

