如何比较多组时间序列数据?基于Python的轮胎扭矩数据异常检测
Great question—time series anomaly detection for sensor data like tire torque is super common in industrial use cases, and sklearn is a solid starting point (though we’ll cover some specialized tools that might work better too). Let’s break this down step by step, from prepping your data to running detection and validating results.
Before jumping into modeling, you need to get your data in shape. Here’s what to focus on:
- Fix timestamps: Convert your
timestampcolumn to datetime format so pandas can handle it properly:import pandas as pd df['timestamp'] = pd.to_datetime(df['timestamp']) - Handle missing values: Sensor data often has gaps. For time series, linear interpolation is usually safe if gaps are small:
df['sig_value'] = df.groupby('tire_id')['sig_value'].transform(lambda x: x.interpolate(method='linear')) - Align time axes (optional but helpful): If your 10 tires have different sampling frequencies, resample them to a consistent interval (e.g., 1-minute averages) to make comparisons fair:
df_resampled = df.groupby('tire_id').resample('1min', on='timestamp')['sig_value'].mean().reset_index()
Sklearn’s core algorithms don’t work directly on raw time series—you need to convert each tire’s signal into a set of meaningful features. Here are the most useful ones for torque data:
- Summary statistics: Mean, standard deviation, min/max, median, skewness, kurtosis (these capture overall signal behavior)
- Rolling window stats: 5-minute or 10-minute rolling averages/variances (captures short-term fluctuations)
- Frequency domain features: Use FFT to get dominant frequencies (good for detecting unusual vibration patterns)
Here’s a quick snippet to compute summary stats for each tire:
grouped = df_resampled.groupby('tire_id') tire_features = grouped['sig_value'].agg([ 'mean', 'std', 'max', 'min', 'median', 'skew', 'kurt' ]).reset_index()
For rolling stats, you can calculate them and then take their own summary stats (e.g., average of rolling variances) to get a single value per tire:
rolling_stats = grouped['sig_value'].rolling(window=5).agg(['mean', 'std']).reset_index() rolling_summary = rolling_stats.groupby('tire_id')[['mean', 'std']].agg(['mean', 'max']) # Flatten the column names for sklearn rolling_summary.columns = ['_'.join(col) for col in rolling_summary.columns] # Merge with your summary stats tire_features = pd.merge(tire_features, rolling_summary, on='tire_id')
Since you know there are 2 anomalous tires but not which ones, unsupervised methods are your best bet.
Sklearn’s Unsupervised Algorithms
These work well with the tabular features we just created:
- Isolation Forest: Fast, efficient, and designed specifically for outlier detection. Perfect for this use case because it handles high-dimensional data well.
from sklearn.ensemble import IsolationForest # Drop tire_id to get feature matrix X = tire_features.drop('tire_id', axis=1).values # Set contamination to the known anomaly rate (2/10 = 0.2) clf = IsolationForest(contamination=0.2, random_state=42) predictions = clf.fit_predict(X) # -1 = anomaly, 1 = normal tire_features['is_anomaly'] = predictions anomalous_tires = tire_features[tire_features['is_anomaly'] == -1]['tire_id'].tolist() - Local Outlier Factor (LOF): Great if anomalies are "locally" unusual (e.g., one tire has a similar mean to others but way higher variance). It measures how isolated a data point is from its neighbors.
from sklearn.neighbors import LocalOutlierFactor clf = LocalOutlierFactor(n_neighbors=3, contamination=0.2) predictions = clf.fit_predict(X) tire_features['is_anomaly'] = predictions
Specialized Time Series Tools (Better for Raw Signal)
If you want to work directly with the raw time series (instead of features), these tools are tailored for the job:
- ADTK: A dedicated time series anomaly detection library that wraps many algorithms. For example, using Isolation Forest directly on the time series:
from adtk.data import validate_series from adtk.detector import IsolationForestAD # Convert each tire's data to a validated time series ts_data = df_resampled.set_index('timestamp')['sig_value'].groupby('tire_id').apply(validate_series) # Detect anomalies in each time series detector = IsolationForestAD(contamination=0.05) # 5% of points as anomalous anomaly_points = detector.fit_detect(ts_data) # Count anomaly points per tire—tires with high counts are anomalous anomaly_counts = anomaly_points.groupby('tire_id').sum() anomalous_tires = anomaly_counts[anomaly_counts > len(df_resampled)/10 * 0.1].index.tolist() # 10% of points anomalous - Prophet: Facebook’s time series forecasting library. You can use it to predict expected torque values, then flag tires where the actual signal deviates heavily from predictions:
from prophet import Prophet anomaly_scores = [] for tire_id in df_resampled['tire_id'].unique(): # Prep data for Prophet tire_df = df_resampled[df_resampled['tire_id'] == tire_id][['timestamp', 'sig_value']] tire_df.columns = ['ds', 'y'] # Fit model m = Prophet() m.fit(tire_df) # Predict on existing data forecast = m.predict(tire_df) # Calculate residual (actual vs predicted) tire_df['residual'] = abs(tire_df['y'] - forecast['yhat']) # Use mean residual as an anomaly score anomaly_scores.append((tire_id, tire_df['residual'].mean())) # Sort by score and pick top 2 as anomalies anomaly_scores.sort(key=lambda x: x[1], reverse=True) anomalous_tires = [tire for tire, score in anomaly_scores[:2]]
- Visualize: Plot each tire’s torque signal to manually check if your detected anomalies make sense. Anomalous tires might have bigger spikes, higher overall values, or weird patterns:
import matplotlib.pyplot as plt fig, axes = plt.subplots(5, 2, figsize=(16, 20)) axes = axes.flatten() for idx, tire_id in enumerate(df_resampled['tire_id'].unique()): ax = axes[idx] data = df_resampled[df_resampled['tire_id'] == tire_id] ax.plot(data['timestamp'], data['sig_value']) ax.set_title(f"Tire {tire_id}") if tire_id in anomalous_tires: ax.set_facecolor('#fff3e0') # Highlight anomalous tires with light orange plt.tight_layout() plt.show() - Tweak: If your model isn’t picking up the right anomalies, try adjusting features (add frequency domain stats!) or tuning algorithm parameters (e.g.,
contaminationin Isolation Forest).
内容的提问来源于stack exchange,提问作者alwaysaskingquestions

