如何高效计算起始于首行的可变窗口宽度滚动分位数?
It looks like you're trying to calculate expanding window quantiles (where each window starts at row 1 and extends to the current row) instead of fixed-size rolling windows. Your initial rollapply syntax was on the right track, but the width parameter needed adjustment to target this expanding behavior. Here are a few efficient ways to solve this:
Method 1: Using zoo::rollapply (most straightforward for xts)
Since your data is an xts object (which inherits from zoo), you can use rollapply with a dynamic width sequence that grows with each row. This tells the function to use all rows from the start up to the current position:
library(zoo) # Calculate expanding 95th percentile expanding_95q <- rollapply( x[,"random"], width = seq_along(x), # Width increases from 1 to nrow(x) FUN = quantile, probs = 0.95, align = "right" # Ensures window ends at current row (default, but explicit here) )
Method 2: Using slider::slide_index (cleaner for time-indexed data)
The slider package is designed for flexible window calculations, and works seamlessly with xts's time index. Setting .before = Inf ensures we include all prior rows up to the current one:
library(slider) expanding_95q <- slide_index( x[,"random"], .i = index(x), # Use the xts time index .f = ~ quantile(.x, probs = 0.95), .before = Inf # Include all data from the start to current index )
Method 3: Base R (for small datasets)
If you prefer not to use additional packages, a sapply approach works for smaller datasets (though it's less efficient for large data):
# Calculate quantiles for each expanding window quantile_vals <- sapply(1:nrow(x), function(i) { quantile(x[1:i, "random"], probs = 0.95) }) # Convert back to xts to preserve the time index expanding_95q <- xts(quantile_vals, order.by = index(x))
Verifying the Result
You can confirm the calculation works by checking a specific row (e.g., row 35) against a manual calculation:
# Manual 95th percentile for rows 1-35 manual_q <- quantile(x[1:35, "random"], probs = 0.95) # Compare to the computed value expanding_95q[35]
These two values should match exactly.
Notes
- If your data contains missing values, add
na.rm = TRUEto thequantilecall to handle them. - All methods will return an xts object with the same time index as your original data, so you can easily combine or plot the results alongside your raw data.
内容的提问来源于stack exchange,提问作者mynameisJEFF

