如何对比多组时间序列数据并基于Python实现轮胎扭矩异常检测
Hey there! Let's break down how to tackle your problem—comparing 10 variable-length tire torque time series and building anomaly detection for them using Python. I'll cover both comparison strategies and anomaly detection approaches, including sklearn tools and some handy alternatives.
1. First: Preprocess Your Time Series (Critical for Variable Lengths)
Since each tire's data has different numbers of points, the first step is to get them onto a consistent time grid. Resampling is your friend here. Using pandas, you can align all series to a fixed frequency (e.g., every second, minute—whatever makes sense for your torque data).
import pandas as pd # Load your combined dataset df = pd.read_csv("tire_torque_data.csv") df['timestamp'] = pd.to_datetime(df['timestamp']) # Resample each tire's data to a 1-second frequency (adjust as needed) resampled_dfs = [] for tire_id in df['tire_id'].unique(): tire_df = df[df['tire_id'] == tire_id].set_index('timestamp') # Use mean to fill gaps (or forward fill if that fits your data's behavior) resampled = tire_df['sig_value'].resample('1S').mean().reset_index() resampled['tire_id'] = tire_id resampled_dfs.append(resampled) # Combine back into a single aligned DataFrame aligned_df = pd.concat(resampled_dfs)
Now all your time series have matching time intervals, making comparison and anomaly detection far more straightforward.
2. Comparing the Time Series Datasets
Once aligned, you can compare the tires in a few practical ways:
a. Feature-Based Comparison
Extract meaningful statistical or time-series features from each tire's data, then compare these features across tires. This pairs perfectly with sklearn for clustering or similarity scoring.
Common features to compute:
- Basic stats: mean, median, standard deviation, min/max of
sig_value - Temporal features: autocorrelation at lag 1, rolling mean/std (e.g., 10-second window)
- Frequency-domain features: FFT peak frequency, power spectral density
Example code to extract features:
import numpy as np def extract_features(tire_df): sig = tire_df['sig_value'].dropna() return { 'tire_id': tire_df['tire_id'].iloc[0], 'mean_sig': sig.mean(), 'std_sig': sig.std(), 'median_sig': sig.median(), 'autocorr_lag1': sig.autocorr(lag=1), 'rolling_mean_10s': sig.rolling(window=10).mean().mean(), # Add more features as needed for your use case } # Extract features for all tires feature_df = pd.DataFrame([extract_features(df) for _, df in aligned_df.groupby('tire_id')]) # Compare similarity using cosine similarity from sklearn.metrics.pairwise import cosine_similarity feature_matrix = feature_df.drop('tire_id', axis=1).values similarity_matrix = cosine_similarity(feature_matrix) # Print similarity between tire 0 and tire 1 print(f"Similarity between tire 0 and 1: {similarity_matrix[0][1]:.4f}")
b. Dynamic Time Warping (DTW)
For direct shape comparison (even if lengths were still unaligned), DTW measures similarity by warping the time axis to match patterns. Use fastdtw for efficient computation:
from fastdtw import fastdtw from scipy.spatial.distance import euclidean # Get two tire series tire_0 = aligned_df[aligned_df['tire_id'] == 0]['sig_value'].dropna().values tire_1 = aligned_df[aligned_df['tire_id'] == 1]['sig_value'].dropna().values distance, _ = fastdtw(tire_0, tire_1, dist=euclidean) print(f"DTW distance between tire 0 and 1: {distance:.4f}")
Lower DTW distance means more similar time series shapes.
3. Anomaly Detection Across Multiple Time Series
Now for detecting anomalies—like a tire behaving differently from the rest. Here are a few robust approaches:
a. Unsupervised Anomaly Detection with Sklearn
Use sklearn's unsupervised models on the extracted features we made earlier. Isolation Forest is fast and handles high-dimensional data well:
from sklearn.ensemble import IsolationForest # Prepare feature matrix (drop tire_id) X = feature_df.drop('tire_id', axis=1).values # Train Isolation Forest (contamination = expected anomaly rate) clf = IsolationForest(contamination=0.1, random_state=42) feature_df['anomaly_flag'] = clf.fit_predict(X) # -1 indicates anomaly, 1 indicates normal anomalous_tires = feature_df[feature_df['anomaly_flag'] == -1]['tire_id'] print(f"Anomalous tires: {list(anomalous_tires)}")
One-Class SVM is another option, but it's slower for larger datasets.
b. Time-Series Specific Anomaly Detection
If you want to detect anomalies within individual time series (not just cross-tire), or capture complex temporal patterns:
i. Prophet (Facebook)
Prophet is excellent for time series forecasting and built-in anomaly detection, even with missing data:
from prophet import Prophet def detect_anomalies_with_prophet(tire_df): # Prophet expects ds (datetime) and y (value) columns prophet_df = tire_df.rename(columns={'timestamp': 'ds', 'sig_value': 'y'}) model = Prophet(interval_width=0.95) # 95% confidence interval model.fit(prophet_df) forecast = model.predict(prophet_df) # Mark points outside the confidence interval as anomalies forecast['anomaly'] = np.where( (prophet_df['y'] < forecast['y_lower']) | (prophet_df['y'] > forecast['y_upper']), 1, 0 ) return forecast # Run for a single tire tire_0_df = aligned_df[aligned_df['tire_id'] == 0] forecast = detect_anomalies_with_prophet(tire_0_df) print(f"Number of anomalies in tire 0: {forecast['anomaly'].sum()}")
ii. LSTM Autoencoder
For capturing complex temporal patterns, an LSTM autoencoder learns to reconstruct normal time series—high reconstruction error indicates anomalies. Use Keras/TensorFlow:
import tensorflow as tf from tensorflow.keras.models import Model from tensorflow.keras.layers import Input, LSTM, RepeatVector, TimeDistributed # Create sequences for LSTM (e.g., 10-time-step windows) def create_sequences(data, seq_length): X = [] for i in range(len(data) - seq_length): X.append(data[i:i+seq_length]) return np.array(X) # Train on "normal" tire data (e.g., tire 0) tire_0_data = aligned_df[aligned_df['tire_id'] == 0]['sig_value'].dropna().values.reshape(-1, 1) seq_length = 10 X_train = create_sequences(tire_0_data, seq_length) # Build autoencoder input_layer = Input(shape=(seq_length, 1)) encoder = LSTM(32, activation='relu', return_sequences=False)(input_layer) decoder = RepeatVector(seq_length)(encoder) decoder = LSTM(32, activation='relu', return_sequences=True)(decoder) output_layer = TimeDistributed(tf.keras.layers.Dense(1))(decoder) autoencoder = Model(inputs=input_layer, outputs=output_layer) autoencoder.compile(optimizer='adam', loss='mse') # Train the model autoencoder.fit(X_train, X_train, epochs=50, batch_size=16, validation_split=0.1) # Calculate reconstruction error threshold X_recon = autoencoder.predict(X_train) mse = np.mean(np.power(X_train - X_recon, 2), axis=(1,2)) threshold = np.percentile(mse, 95) # 95th percentile as anomaly cutoff # Test on another tire tire_1_data = aligned_df[aligned_df['tire_id'] == 1]['sig_value'].dropna().values.reshape(-1,1) X_test = create_sequences(tire_1_data, seq_length) X_test_recon = autoencoder.predict(X_test) test_mse = np.mean(np.power(X_test - X_test_recon, 2), axis=(1,2)) anomalies = test_mse > threshold print(f"Anomalous windows in tire 1: {np.sum(anomalies)}")
c. Toolkit Recommendation: tsfresh
tsfresh automatically extracts hundreds of time-series features, saving you from manual feature engineering. Pair it with sklearn models for quick results:
from tsfresh import extract_features from tsfresh.utilities.dataframe_functions import impute # Extract features using tsfresh extracted_features = extract_features( aligned_df, column_id='tire_id', column_sort='timestamp', column_value='sig_value' ) # Impute missing values extracted_features = impute(extracted_features) # Use with Isolation Forest clf = IsolationForest(contamination=0.1) anomaly_scores = clf.fit_predict(extracted_features)
Final Notes
- Start with feature-based comparison and Isolation Forest if you want a quick, sklearn-native solution.
- Use DTW if shape similarity is more important than statistical features.
- For temporal anomalies within a tire, Prophet or LSTM autoencoders are better choices.
- tsfresh is a huge time-saver for feature extraction—definitely check it out if you don't want to handcraft features.
内容的提问来源于stack exchange,提问作者alwaysaskingquestions

