如何在Python中向量化SMA预测与卡尔曼滤波的两个循环?
Hey Juan, let's tackle those performance bottlenecks one by one—both loops are definitely candidates for optimization, and we can eliminate most of the slow Python-level looping by leaning into NumPy's vectorized operations.
1. Vectorizing the SMA Prediction Loop
Your first loop is iterating over each index to grab three specific SMA values, fit a quadratic polynomial, and predict the last value of vector_X. The main issue here is that we're doing this one iteration at a time instead of processing all windows at once.
Here's how to vectorize this entirely:
import numpy as np import talib # Original parameters window_sma = 200 sma_index = 500 offset = 50 vector_X = [1, 2, 3, 15] target_pred_idx = len(vector_X) - 1 # We only care about the last prediction value # Compute SMA once (as you did before) SMA = talib.SMA(values, timeperiod=window_sma) # Generate all i, j, k indices in vectorized form i_array = np.arange(sma_index, len(SMA)) j_array = i_array - offset k_array = i_array - (offset // 2) # Offset is 50, so this gives i-25 # Grab the three SMA values for each window: shape (number_of_windows, 3) sma_windows = np.stack([SMA[j_array], SMA[k_array], SMA[i_array]], axis=1) # Fit quadratic polynomials for all windows at once # np.polyfit accepts a 2D y array, returns coefficients with shape (3, number_of_windows) coeffs = np.polyfit([1, 2, 3], sma_windows.T, 2) # Predict for vector_X, then extract the last value of each prediction y_hats = np.polyval(coeffs.T, vector_X) sma_predicted = y_hats[:, target_pred_idx]
Why this works faster:
This replaces the Python loop with NumPy's optimized C-level operations. Stacking SMA windows into a 2D array lets np.polyfit process all fits in one go, which is orders of magnitude faster than looping through each window individually.
2. Optimizing the Kalman Filter Loop (Eliminate Dynamic Array Appends)
Your second loop's biggest performance killer is x = np.append(x, new_x_col, axis=1). NumPy arrays are fixed-size, so every append creates a new array and copies all existing data—this gets exponentially slower as x grows. We can fix this by pre-allocating the x array upfront, since we know exactly how many columns it will have.
Also, note that your original code used element-wise multiplication (*) for matrix operations, which is incorrect for Kalman filter math. We'll replace that with proper matrix multiplication using the @ operator.
Here's the optimized version:
from filterpy.kalman import KalmanFilter # Assuming you're using filterpy's implementation # Kalman Filter setup (fixed parameters) km = KalmanFilter(dim_x=2, dim_z=1) km.F = np.array([[1., 1.], [0., 1.]]) # State transition matrix km.H = np.array([[1., 0.]]) # Measurement function dt = 0.0001 a = 1.5 # Process noise covariance Q (proper matrix multiplication) km.Q = (a**2) * np.array([[dt**4/4, dt**3/2], [dt**3/2, dt**2]]) km.R = 1000 # Measurement noise variance I = np.eye(2) # Identity matrix (cleaner than hardcoding) km.Z = np.array(sma_predicted) # Measurements # Pre-allocate x array: shape (2, number_of_steps) n_steps = len(sma_predicted) x = np.zeros((2, n_steps)) x[:, 0] = np.array([sma_predicted[0], 0]) # Initial state km.P = np.array([[1000, 0], [0, 1000]]) # Initial covariance # Kalman Filter loop with in-place updates for i in range(n_steps - 1): # Prediction step x[:, i+1] = km.F @ x[:, i] # Matrix multiplication using @ operator km.P = km.F @ km.P @ km.F.T + km.Q # Correction step S = km.H @ km.P @ km.H.T + km.R K = km.P @ km.H.T / S residual = km.Z[i+1] - km.H @ x[:, i+1] x[:, i+1] += K @ residual km.P = (I - K @ km.H) @ km.P
Key optimizations here:
- Pre-allocation: We create
xwith its final shape upfront, so no more expensive array copies fromnp.append. - Correct Matrix Math: Replaced element-wise
*with@for matrix operations, ensuring the Kalman filter calculations are accurate while using NumPy's optimized routines. - Readable, Efficient Steps: Broke down complex calculations into simpler lines without sacrificing speed, making the code easier to debug.
Additional Tips for Overall Performance
- Data Type Optimization: Use smaller data types (e.g.,
float32instead offloat64) if precision allows—this reduces memory usage and speeds up operations. - Talib SMA Alternative: If
talib.SMAfeels slow for 400k points, try a vectorized SMA implementation withnp.convolve(though Talib is already highly optimized, it's worth testing). - Batch Processing: For extremely large datasets, consider splitting the data into batches and processing them in parallel with libraries like
Dask, but the vectorized code above should handle 400k points easily.
内容的提问来源于stack exchange,提问作者Juan Solana

