寻求适用于无标签三维特征数据的无监督异常检测模型
Hey there! Let's walk through practical, actionable approaches for your unsupervised anomaly detection task—since you've already handled normalization, t-SNE dimensionality reduction, and re-normalization, we can focus directly on methods that fit your unlabeled spatial + speed data.
Top Methods to Try
1. Isolation Forest
This is my go-to for quick, effective unsupervised anomaly detection. It works by isolating outliers through random tree splits, which is great for datasets with spatial components (like your latitude/longitude) plus speed.
from sklearn.ensemble import IsolationForest import numpy as np # Assume your preprocessed data is stored in X (shape: [n_samples, 2] if t-SNE to 2D, or [n_samples,3] if keeping original features) isolation_forest = IsolationForest( n_estimators=100, contamination="auto", # Let the model estimate anomaly ratio, or set manually (e.g., 0.05 for 5% anomalies) random_state=42 ) predicted_labels = isolation_forest.fit_predict(X) # Extract anomalies (-1 marks outliers, 1 marks normal points) anomalies = X[predicted_labels == -1]
2. DBSCAN (Density-Based Spatial Clustering of Applications with Noise)
Perfect for your spatial data! DBSCAN groups dense regions and labels low-density points as anomalies. You'll need to tune eps (maximum distance between neighbors) and min_samples (minimum points to form a cluster) based on your data's density.
from sklearn.cluster import DBSCAN # Adjust eps and min_samples based on your t-SNE normalized data dbscan = DBSCAN(eps=0.3, min_samples=5) cluster_labels = dbscan.fit_predict(X) # Points labeled -1 are anomalies anomalies = X[cluster_labels == -1]
3. Local Outlier Factor (LOF)
LOF compares the local density of a point to its neighbors—points with much lower density than their surroundings are flagged as outliers. It's great for detecting "local" anomalies that might not stand out globally.
from sklearn.neighbors import LocalOutlierFactor lof = LocalOutlierFactor( n_neighbors=20, # Number of neighbors to compare density against contamination="auto" ) predicted_labels = lof.fit_predict(X) anomalies = X[predicted_labels == -1] # You can also check the outlier score (higher = more anomalous) outlier_scores = lof.negative_outlier_factor_
4. Autoencoder (Deep Learning Approach)
If you want to capture complex patterns in your data, an autoencoder learns to reconstruct normal data points—points with high reconstruction error are anomalies. This works well if your speed/location patterns have subtle, non-linear relationships.
import tensorflow as tf from tensorflow.keras.models import Model from tensorflow.keras.layers import Input, Dense import numpy as np # Define the autoencoder architecture input_dim = X.shape[1] input_layer = Input(shape=(input_dim,)) encoder = Dense(8, activation="relu")(input_layer) encoder = Dense(4, activation="relu")(encoder) decoder = Dense(8, activation="relu")(encoder) output_layer = Dense(input_dim, activation="linear")(decoder) autoencoder = Model(inputs=input_layer, outputs=output_layer) autoencoder.compile(optimizer="adam", loss="mse") # Train the model on your data autoencoder.fit( X, X, epochs=50, batch_size=32, validation_split=0.1, shuffle=True ) # Calculate reconstruction error reconstructions = autoencoder.predict(X) mse_errors = np.mean(np.power(X - reconstructions, 2), axis=1) # Set a threshold (e.g., 95th percentile of errors) to flag anomalies threshold = np.percentile(mse_errors, 95) anomalies = X[mse_errors > threshold]
How to Validate & Refine
Since you don't have labels, use these strategies to assess your results:
Visualization: Plot your t-SNE data with anomalies highlighted—this is the quickest way to spot if the model is picking up meaningful outliers:
import matplotlib.pyplot as plt plt.scatter(X[:, 0], X[:, 1], c="blue", alpha=0.6, label="Normal Points") plt.scatter(anomalies[:, 0], anomalies[:, 1], c="red", label="Anomalies") plt.legend() plt.title("Anomaly Detection Results") plt.show()Internal Metrics: Use these scores to compare different models:
from sklearn.metrics import silhouette_score, davies_bouldin_score # Convert labels to 0 (anomaly) and 1 (normal) for metric calculation normalized_labels = np.where(predicted_labels == -1, 0, 1) silhouette = silhouette_score(X, normalized_labels) davies_bouldin = davies_bouldin_score(X, normalized_labels) print(f"Silhouette Score (closer to 1 = better): {silhouette:.2f}") print(f"Davies-Bouldin Index (closer to 0 = better): {davies_bouldin:.2f}")
Quick Tips for Your Dataset
- Tune t-SNE Parameters: If your t-SNE output isn't capturing anomaly patterns, experiment with different
perplexityvalues (try 5-50) to ensure the dimensionality reduction preserves outlier structure. - Cross-Validate Methods: Run 2-3 of the above methods and check for overlapping anomalies—these are likely the most reliable outliers in your data.
- Adjust Contamination: If you have a rough estimate of how many anomalies exist (e.g., 3-5% of the dataset), set
contaminationto that value instead of "auto" for more precise results.
内容的提问来源于stack exchange,提问作者M.Arıcı

