无预设聚类数的图像视觉特征聚类及2D可视化方案咨询
Got it, let's walk through exactly how to pull this off—you want to group images by raw visual traits (color, shape, gradients), no predefined categories needed, get that clean 2D clustered plot like the t-SNE example you referenced, and even organize images into folders by their clusters. Here's an actionable, code-friendly approach:
1. Extract Combined Visual Features
First, turn each image into a numerical vector that captures its key visual properties. Combine multiple feature types for a robust representation:
- HOG (Gradients/Shape): Perfect for capturing edge structures and shape details. Use scikit-image or OpenCV to compute this.
- Color Histograms: Capture color distribution (HSV histograms work great for color invariance, better than RGB).
- Optional: Simple shape stats (like aspect ratio, edge density) if you want extra granularity.
Example code snippet:
import cv2 import numpy as np from skimage.feature import hog from sklearn.preprocessing import StandardScaler def extract_features(img_path): # Load and resize image (consistent input size is critical) img = cv2.imread(img_path) img = cv2.resize(img, (64, 64)) # HOG features for shape/gradients gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY) hog_features = hog(gray, orientations=9, pixels_per_cell=(8,8), cells_per_block=(2,2), block_norm='L2-Hys') # HSV color histogram for color distribution hsv = cv2.cvtColor(img, cv2.COLOR_BGR2HSV) hist = cv2.calcHist([hsv], [0,1,2], None, [8,8,8], [0,256,0,256,0,256]) hist = cv2.normalize(hist, hist).flatten() # Combine all features into a single vector combined = np.hstack([hog_features, hist]) return combined # Extract features for all your images image_paths = ["img1.jpg", "img2.jpg", ...] # Replace with your image paths features = np.array([extract_features(path) for path in image_paths]) # Standardize features (critical for consistent clustering/dimensionality reduction) scaler = StandardScaler() scaled_features = scaler.fit_transform(features)
2. Reduce Dimensions to 2D for Visualization
Use t-SNE (or UMAP, which is faster for large datasets) to project high-dimensional features into a 2D space. This preserves local similarity, so visually similar images will cluster close together:
from sklearn.manifold import TSNE # Run t-SNE projection tsne = TSNE(n_components=2, perplexity=30, random_state=42) tsne_embeddings = tsne.fit_transform(scaled_features)
Note: Adjust perplexity based on your dataset size—aim for 5-50, higher values work better with more images.
3. Cluster with DBSCAN (No Predefined Cluster Count)
Since you don't know how many clusters exist, DBSCAN is ideal—it groups dense regions of points and ignores outliers, no need to set a cluster number upfront:
from sklearn.cluster import DBSCAN # Run DBSCAN dbscan = DBSCAN(eps=0.5, min_samples=5) cluster_labels = dbscan.fit_predict(tsne_embeddings) # Filter out outliers (labeled as -1) unique_clusters = np.unique(cluster_labels[cluster_labels != -1])
Tuning tips:
- Plot a k-distance graph (sorted distances from each point to its k-th nearest neighbor) to pick an elbow point for
eps. - Adjust
min_samplesbased on cluster tightness—higher values mean fewer, more tightly grouped clusters.
4. Create the 2D Clustered Visualization
Plot the t-SNE embeddings, color-coded by cluster, and optionally overlay small image thumbnails to match the example you liked:
import matplotlib.pyplot as plt from matplotlib.offsetbox import OffsetImage, AnnotationBbox plt.figure(figsize=(12,12)) # Plot each cluster with a unique color for cluster in unique_clusters: mask = cluster_labels == cluster plt.scatter(tsne_embeddings[mask, 0], tsne_embeddings[mask, 1], label=f"Cluster {cluster}", alpha=0.7) # Optional: Add image thumbnails for clarity def get_thumbnail(path): img = cv2.imread(path) img = cv2.cvtColor(img, cv2.COLOR_BGR2RGB) return cv2.resize(img, (32,32)) for i, (x, y) in enumerate(tsne_embeddings): if cluster_labels[i] != -1: img = get_thumbnail(image_paths[i]) ab = AnnotationBbox(OffsetImage(img, zoom=0.5), (x, y), frameon=False) plt.gca().add_artist(ab) plt.legend() plt.title("Image Clustering by Visual Features (t-SNE + DBSCAN)") plt.show()
5. Organize Images into Cluster Folders
Finally, use Python's os and shutil modules to copy images into folders named after their cluster:
import os import shutil # Create folders for each cluster for cluster in unique_clusters: cluster_dir = f"cluster_{cluster}" if not os.path.exists(cluster_dir): os.makedirs(cluster_dir) # Copy images to their respective cluster folders for path, label in zip(image_paths, cluster_labels): if label != -1: dest_path = os.path.join(f"cluster_{label}", os.path.basename(path)) shutil.copy(path, dest_path) # Handle outliers (optional) outlier_dir = "outliers" if not os.path.exists(outlier_dir): os.makedirs(outlier_dir) for path, label in zip(image_paths, cluster_labels): if label == -1: dest_path = os.path.join(outlier_dir, os.path.basename(path)) shutil.copy(path, dest_path)
Pro Tips
- Preprocess Consistently: Always resize images to the same dimensions before feature extraction to ensure consistent feature vectors.
- Speed Up with PCA: If your feature vector is too long, run PCA first to reduce dimensions before t-SNE—this cuts down computation time.
- UMAP for Large Datasets: For 1000+ images, UMAP is faster than t-SNE while preserving similar clustering quality.
内容的提问来源于stack exchange,提问作者Bartek Wójcik

