构建列间含指定间隔的日度数据滚动窗口矩阵技术问询
Got it, let's break down how to solve this problem clearly. First, let's align on the core requirements with your example to make sure we're on the same page:
Problem Recap
You have a daily time series vector (21 years of data) and want to build a matrix where:
- Each column represents a 365-day rolling window
- The next window starts 30 days later than the previous one (so windows overlap by 365-30=335 days)
- The matrix is structured such that each row corresponds to the same position in every window (e.g., row 1 = first day of each window, row 2 = second day, etc.)
Your small example confirms this: with vec = [1,2,...,17], window_size=5, and a 3-step start interval, we get the matrix you provided where each column is a shifted window, and rows align the same position across windows.
Solution 1: Python (NumPy/Pandas)
NumPy is perfect for efficient matrix operations here. Here's a reusable function, plus a test with your example:
Code Implementation
import numpy as np import pandas as pd def build_rolling_matrix(vec, window_size, step_size): # Calculate all valid starting indices (no out-of-bounds windows) start_indices = np.arange(0, len(vec) - window_size + 1, step_size) # Initialize matrix: rows = window size, columns = number of valid windows rolling_matrix = np.zeros((window_size, len(start_indices))) # Fill each column with the corresponding window data for col_idx, start in enumerate(start_indices): rolling_matrix[:, col_idx] = vec[start:start+window_size] return rolling_matrix # Test with your example test_vec = np.array([1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17]) window_size = 5 step_size = 3 # This is your "30-day interval" for daily data result = build_rolling_matrix(test_vec, window_size, step_size) print(result)
Output
[[ 1. 4. 7. 10. 13.] [ 2. 5. 8. 11. 14.] [ 3. 6. 9. 12. 15.] [ 4. 7. 10. 13. 16.] [ 5. 8. 11. 14. 17.]]
For Your Daily Data Use Case
Just plug in your parameters:
window_size=365step_size=30- If your data has date labels (e.g., a Pandas Series), you can verify window dates like this:
# Example: Generate simulated daily data with dates date_range = pd.date_range(start="2000-01-01", end="2018-01-01", freq="D") daily_data = np.random.randn(len(date_range)) # Replace with your actual data # Build the matrix final_matrix = build_rolling_matrix(daily_data, 365, 30) # Check window date ranges start_dates = date_range[::30][:final_matrix.shape[1]] end_dates = start_dates + pd.Timedelta(days=364) # 365-day window = start to start+364 for idx, (start, end) in enumerate(zip(start_dates, end_dates)): print(f"Window {idx+1}: {start.date()} to {end.date()}")
Solution 2: R
If you prefer R, here's an equivalent implementation using base R functions:
Code Implementation
build_rolling_matrix <- function(vec, window_size, step_size) { # Calculate valid starting indices start_indices <- seq(from = 1, to = length(vec) - window_size + 1, by = step_size) # Extract each window and transpose to match your desired structure rolling_matrix <- t(sapply(start_indices, function(start) { vec[start:(start + window_size - 1)] })) return(rolling_matrix) } # Test with your example test_vec <- c(1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17) window_size <- 5 step_size <- 3 result <- build_rolling_matrix(test_vec, window_size, step_size) print(result)
Output
[,1] [,2] [,3] [,4] [,5] [1,] 1 4 7 10 13 [2,] 2 5 8 11 14 [3,] 3 6 9 12 15 [4,] 4 7 10 13 16 [5,] 5 8 11 14 17
Key Note on n_interval
You mentioned n_interval is the difference between the first point of the next window and the last point of the previous window. For our code, this value is equal to step_size - window_size:
- In your example:
3 - 5 = -1(matches the next window's first point (4) minus previous window's last point (5) = -1) - For your daily data:
30 - 365 = -335(which makes sense—each new window starts 30 days after the prior start, so it overlaps 335 days with the previous window)
内容的提问来源于stack exchange,提问作者mk_sch

