针对多不同时间段公司的月度股票收益滚动标准差计算需求问询
Got it, let's break down how to solve this rolling standard deviation problem correctly—no cross-company window leakage, exactly as you need. Since your data's already sorted by company, year, and month, we can use grouping operations to isolate each company's time series.
Python (Pandas) Approach
Pandas makes this straightforward with groupby() paired with rolling(). Here's a step-by-step implementation:
Optional (but safe): Verify sorting
Even if you said the data's sorted, it never hurts to double-check to avoid unexpected behavior:import pandas as pd df = df.sort_values(['company', 'year', 'month'], ascending=[True, True, True])Calculate rolling 24-month standard deviation per company
We'll group by company, then apply a rolling window of size 24, only computing the std when we have a full 24 observations (so the first 23 rows per company will beNaN):df['rolling_24m_std'] = df.groupby('company')['monthly_return'] \ .rolling(window=24, min_periods=24) \ .std() \ .reset_index(level=0, drop=True)groupby('company'): Ensures each company's time series is processed independently.rolling(window=24, min_periods=24): Sets a 24-month window, and only returns a value when all 24 observations are present.reset_index(level=0, drop=True): Drops the group index so the result aligns correctly with the original DataFrame.
R (dplyr + slider) Approach
If you're working in R, the slider package is perfect for precise rolling window operations, paired with dplyr for grouping:
Verify sorting
library(dplyr) df <- df %>% arrange(company, year, month)Compute rolling std per company
library(slider) df <- df %>% group_by(company) %>% mutate(rolling_24m_std = slide_dbl( .x = monthly_return, .f = sd, .before = 23, # Current row + previous 23 = 24 total observations .complete = TRUE # Only return value if window is full )) %>% ungroup().before = 23: Defines the window as the current row plus the 23 prior rows (total 24 months)..complete = TRUE: Ensures we only get a result when the window has all 24 observations—otherwise returnsNA.
Key Notes
- Companies with fewer than 24 monthly observations will have
NaN/NAin therolling_24m_stdcolumn, which aligns with your requirement to start calculating from the 24th observation. - The grouping operations strictly prevent cross-company window leakage—each company's rolling calculation is entirely isolated from others.
内容的提问来源于stack exchange,提问作者Fred Olsen

