You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于SciPy least_squares()的鲁棒曲线拟合:一阶LTI系统电机异常值处理

Awesome, let's tackle this robustness issue for your first-order LTI step response fitting. You've already got a good baseline with standard least squares, so we'll build on that with methods specifically designed to suppress outliers in your motor speed data.

First, let's recap your model for clarity:

( v(t) = K \cdot (1 - \exp(-t/T)) )
where ( K ) is the steady-state speed and ( T ) is the time constant of your motor system.

Here are three practical, Python-based approaches to boost fitting robustness:


1. Use Robust Loss Functions with Nonlinear Least Squares

Standard least squares squares all residuals, which gives outliers disproportionate influence. Instead, use robust loss functions (like Huber or Cauchy) that penalize large residuals more gently. The scipy.optimize.least_squares function has built-in support for this.

Implementation Code:

import numpy as np
from scipy.optimize import least_squares
import matplotlib.pyplot as plt

# Define your step response model
def step_response(t, K, T):
    return K * (1 - np.exp(-t / T))

# Define residual function for fitting
def residual(params, t, v):
    K, T = params
    return v - step_response(t, K, T)

# Assume your raw data is stored in t_data (time array) and v_data (speed array)
# Start with a reasonable initial guess: K ≈ max speed, T ≈ 1 (adjust based on your data)
initial_guess = [np.max(v_data), 1.0]

# Fit with Huber loss (balances squared and linear penalty for residuals)
huber_fit = least_squares(
    residual, 
    initial_guess, 
    args=(t_data, v_data), 
    loss='huber', 
    f_scale=0.1  # Adjust this based on your noise level (smaller = more robust)
)
K_huber, T_huber = huber_fit.x

# Fit with Cauchy loss (even more robust to extreme outliers)
cauchy_fit = least_squares(
    residual, 
    initial_guess, 
    args=(t_data, v_data), 
    loss='cauchy', 
    f_scale=0.1
)
K_cauchy, T_cauchy = cauchy_fit.x

Key Notes:

  • Huber Loss: Uses squared error for small residuals, linear error for residuals larger than f_scale—great for mild outliers.
  • Cauchy Loss: Has a slower-growing penalty for large residuals—ideal if you have extreme, far-out outliers.
  • Tune f_scale to match your data's typical noise (e.g., if your speed measurements have ±0.5 RPM noise, set f_scale=0.5).

2. RANSAC Algorithm for Outlier Rejection

RANSAC (Random Sample Consensus) iteratively selects subsets of "inliers" (data that fits the model) and fits only those, automatically ignoring outliers. This works especially well if outliers make up a significant portion of your data.

Since your model is nonlinear, we'll wrap it in a custom scikit-learn-compatible estimator to use with sklearn.linear_model.RANSACRegressor.

Implementation Code:

from sklearn.base import BaseEstimator, RegressorMixin
from sklearn.utils.validation import check_X_y, check_array, check_is_fitted
from sklearn.linear_model import RANSACRegressor

# Custom regressor for your step response model
class StepResponseRegressor(BaseEstimator, RegressorMixin):
    def __init__(self, initial_guess=[1.0, 1.0]):
        self.initial_guess = initial_guess
        self.K_ = None
        self.T_ = None
    
    def fit(self, X, y):
        # Validate input data
        X, y = check_X_y(X, y)
        t = X.flatten()  # Ensure time data is 1D
        
        # Fit using nonlinear least squares
        fit_result = least_squares(residual, self.initial_guess, args=(t, y))
        self.K_, self.T_ = fit_result.x
        return self
    
    def predict(self, X):
        # Validate input and check if model is fitted
        check_is_fitted(self)
        X = check_array(X)
        t = X.flatten()
        
        # Generate predictions
        return step_response(t, self.K_, self.T_)

# Initialize RANSAC with your custom regressor
ransac_reg = RANSACRegressor(
    base_estimator=StepResponseRegressor(initial_guess=[np.max(v_data), 1.0]),
    min_samples=0.7,  # Use 70% of data for each trial (adjust based outlier ratio)
    residual_threshold=0.5,  # Max residual for a point to be considered an inlier
    max_trials=1000  # More trials = better chance of finding a good inlier set
)

# Fit RANSAC (note: X must be 2D for scikit-learn)
ransac_reg.fit(t_data.reshape(-1, 1), v_data)

# Extract fitted parameters and inlier/outlier masks
K_ransac, T_ransac = ransac_reg.estimator_.K_, ransac_reg.estimator_.T_
inlier_mask = ransac_reg.inlier_mask_
outlier_mask = np.logical_not(inlier_mask)

Key Notes:

  • min_samples: Adjust based on your estimated outlier ratio (e.g., if 30% of data is outliers, set to 0.7).
  • residual_threshold: Set to your typical measurement noise (e.g., ±0.5 RPM for speed data).
  • RANSAC will explicitly separate inliers and outliers, which is useful for debugging your data collection pipeline.

3. Weighted Least Squares with Robust Residual Thresholding

If you want more control over which points get downweighted, you can compute weights based on robust residual statistics (like Median Absolute Deviation, MAD) and run a weighted least squares fit.

Implementation Code:

# First run a standard least squares fit to get initial residuals
ols_fit = least_squares(residual, initial_guess, args=(t_data, v_data))
K_ols, T_ols = ols_fit.x

# Calculate residuals from the initial fit
residuals = v_data - step_response(t_data, K_ols, T_ols)

# Use Median Absolute Deviation (MAD) to identify outliers (robust to outliers itself)
mad = np.median(np.abs(residuals - np.median(residuals)))
# Define weights: 1 for inliers (residual < 3*MAD), 0.01 for outliers (small weight)
weights = np.where(np.abs(residuals) < 3 * mad, 1.0, 0.01)

# Define weighted residual function
def weighted_residual(params, t, v, weights):
    return weights * (v - step_response(t, params[0], params[1]))

# Run weighted least squares
weighted_fit = least_squares(
    weighted_residual, 
    [K_ols, T_ols],  # Use OLS result as initial guess
    args=(t_data, v_data, weights)
)
K_weighted, T_weighted = weighted_fit.x

Key Notes:

  • MAD is a robust alternative to standard deviation—outliers don't skew its value.
  • The 3*MAD threshold is a common rule of thumb for identifying outliers, but you can adjust it based on your data.
  • This method is great if you want to retain some influence from outliers (instead of fully ignoring them) but minimize their impact.

Visualize the Results

To verify your robust fits, plot the raw data, outliers, and fitted curves:

plt.figure(figsize=(10, 6))
plt.scatter(t_data, v_data, label='Raw Data', alpha=0.5)
plt.scatter(t_data[outlier_mask], v_data[outlier_mask], color='red', label='Detected Outliers')
plt.plot(t_data, step_response(t_data, K_huber, T_huber), color='blue', label='Huber Loss Fit')
plt.plot(t_data, step_response(t_data, K_ransac, T_ransac), color='green', label='RANSAC Fit')
plt.xlabel('Time (s)')
plt.ylabel('Motor Speed')
plt.title('Robust Step Response Fitting')
plt.legend()
plt.grid(True)
plt.show()

Final Recommendations:

  • If you have mild, sparse outliers: Use Huber loss (simple and effective).
  • If you have extreme or frequent outliers: Use RANSAC (explicitly rejects outliers).
  • If you want fine-grained control over weights: Use weighted least squares with MAD.

内容的提问来源于stack exchange,提问作者Yannick

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 10:34:27