协整时间序列的多期领先与滞后缺失值插补技术问询
Great question! The issue with na.interpolation() and na.interp() is that they're univariate interpolation methods—they only look at FB's own past/future values, completely ignoring the cointegrating relationships with MSFT and SPY that you want to leverage. Let's fix that by building a cointegration-aware imputation workflow:
Step 1: Validate Cointegration First
Before we jump into imputation, we need to confirm that FB, MSFT, and SPY actually share a long-term cointegrating relationship. We'll use the Johansen test (via the urca package) for this:
# Load required packages library(xts) library(urca) library(tsDyn) library(ggplot2) library(reshape2) library(data.table) library(quantmod) # Recreate your initial data prep (from your code) tickers <- c("MSFT", "SPY","FB") dataEnv <- new.env() getSymbols(tickers, from="2010-06-30", to="2015-06-30", env=dataEnv) plist <- eapply(dataEnv, Ad) pframe <- as.data.frame(do.call(merge, plist)) pframe <- as.data.table(pframe,keep.rownames = T) pframe$FB.Adjusted[ (nrow(pframe) - 150):nrow(pframe) ] <- NA # Convert to xts for easier time series handling pxts <- xts(pframe[,-1], order.by = as.Date(pframe$rn)) # Run Johansen cointegration test joh_test <- ca.jo(pxts[, c("FB.Adjusted", "MSFT.Adjusted", "SPY.Adjusted")], type = "trace", K = 2, ecdet = "none") summary(joh_test)
Check the test results: if the trace statistic exceeds the critical value at your chosen significance level (e.g., 5%), you've confirmed a cointegrating relationship. Note the number of cointegrating relations (r) for the next step.
Step 2: Impute with a Vector Error Correction Model (VECM)
Cointegrated time series are best modeled with a VECM, which captures both their long-term equilibrium and short-term adjustments. We'll use this model to predict FB's missing values using the complete MSFT and SPY data:
# Split data into training (non-missing FB) and missing segments train_data <- pxts[1:(nrow(pxts)-151), ] missing_dates <- index(pxts[(nrow(pxts)-150):nrow(pxts), ]) # Fit VECM on training data (use `r` from Johansen test, e.g., r=1) vecm_model <- VECM(train_data[, c("FB.Adjusted", "MSFT.Adjusted", "SPY.Adjusted")], lag = 2, r = 1) # Extract actual MSFT/SPY values for the missing period (we need these to predict FB) actual_msft_spy <- pxts[missing_dates, c("MSFT.Adjusted", "SPY.Adjusted")] # Predict FB's missing values using the VECM fb_imputations <- predict(vecm_model, newdata = actual_msft_spy, n.ahead = 151) # Replace missing FB values in the original dataset pxts_imputed <- pxts pxts_imputed[missing_dates, "FB.Adjusted"] <- fb_imputations[, "FB.Adjusted"] # Convert back to data.table for plotting pframe_imputed <- data.table(rn = index(pxts_imputed), coredata(pxts_imputed)) # Visualize the result meltdf <- melt(pframe_imputed, id="rn") ggplot(meltdf, aes(x=rn, y=value, colour=variable, group=variable)) + geom_line() + ggtitle("FB Imputation Using Cointegration with MSFT & SPY") + theme(axis.text.x = element_text(angle=45, hjust=1))
Step 3: Optional: Multiple Imputation for Uncertainty
If you want to account for imputation uncertainty (e.g., for downstream statistical analysis), you can use mice with a custom VECM-based imputation function:
library(mice) # Define a custom VECM imputation function vecm_impute <- function(y, x, ...) { # y = missing variable (FB), x = complete variables (MSFT, SPY) combined <- cbind(y, x) complete_cases <- na.omit(combined) # Fallback if too few complete observations if(nrow(complete_cases) < 20) return(y) # Fit VECM and predict missing values vecm_fit <- VECM(complete_cases, lag=2, r=1) missing_idx <- which(is.na(y)) if(length(missing_idx) == 0) return(y) pred <- predict(vecm_fit, newdata = x[missing_idx, ], n.ahead=length(missing_idx)) y[missing_idx] <- pred[, colnames(y)] return(y) } # Run multiple imputation (5 datasets by default) imp <- mice(pframe[, -1], method = list(FB.Adjusted = "vecm_impute", MSFT.Adjusted = "pmm", SPY.Adjusted = "pmm"), m = 5) # Extract one imputed dataset (or pool results with `pool()` for analysis) pframe_mice_imputed <- complete(imp, 1) pframe_mice_imputed$rn <- pframe$rn # Plot meltdf <- melt(pframe_mice_imputed, id="rn") ggplot(meltdf, aes(x=rn, y=value, colour=variable, group=variable)) + geom_line() + ggtitle("Multiple Imputation of FB (VECM-Based)") + theme(axis.text.x = element_text(angle=45, hjust=1))
Key Notes
- Always validate cointegration first—this approach only makes sense if the series share a long-term equilibrium.
- Adjust the lag order (
lag=2) based on AIC/BIC (useVARselect()from thevarspackage to find the optimal lag). - The VECM imputation uses the actual observed values of MSFT and SPY during the missing period, ensuring FB's imputed values stay aligned with the group's cointegrating relationship.
内容的提问来源于stack exchange,提问作者BlockChained

