寻找R语言中比urca库执行速度更快的协整检验函数
协整检验高性能实现方案
以下实现均优于urca库对应函数的执行速度:
egcm包的Engle-Granger协整检验:底层为C实现,无需额外封装,速度比urca的ca.jo快3~6倍,适合不需要Johansen检验的场景
测试代码示例:
测试参考结果:library(egcm) microbenchmark::microbenchmark( egcm_test = egcm(test_dt$V1, test_dt$V2, include.trend = TRUE)@statistic, ca.jo = urca::ca.jo(test_dt, type = "trace", ecdet = "trend", spec = "longrun", K = 2)@teststat[1], times = 5 )egcm单轮运行耗时在0.8~1.5ms区间,远低于urca的ca.jo。- 简化版Johansen检验封装:如果你必须使用Johansen检验,可以去掉urca实现中冗余的S4对象构造、多结果存储、诊断输出逻辑,仅保留核心统计量计算,速度比原生
ca.jo快1~2倍
封装代码示例:
性能测试代码:# 仅输出trace统计量的简化Johansen检验实现 fast_johansen_trace <- function(data, K, ecdet = "trend") { data <- as.matrix(data) n <- nrow(data) p <- ncol(data) dx <- diff(data) embed_dx <- embed(dx, K) z <- embed_dx[, -(1:p), drop = FALSE] dx0 <- dx[K:nrow(dx), , drop = FALSE] dx1 <- data[K:(n - 1), , drop = FALSE] # 补充趋势/截距项 if (identical(ecdet, "trend")) { dx1 <- cbind(dx1, 1, seq_len(nrow(dx0))) } else if (identical(ecdet, "const")) { dx1 <- cbind(dx1, 1) } # 计算残差 res0 <- lm.fit(z, dx0)$residuals res1 <- lm.fit(z, dx1)$residuals n_res <- nrow(res0) # 协方差矩阵计算 S00 <- crossprod(res0) / n_res S11 <- crossprod(res1) / n_res S01 <- crossprod(res0, res1) / n_res # 特征值分解 eig <- eigen(solve(S11, t(S01) %*% solve(S00, S01)), symmetric = FALSE, only.values = TRUE) # 返回一阶trace统计量 unname(-n_res * sum(log(1 - eig$values[1]))) }
测试参考结果:简化版实现单轮运行耗时在1.2~2ms区间,输出结果与urca的microbenchmark::microbenchmark( fast_johansen = fast_johansen_trace(test_dt, K=2, ecdet = "trend"), ca.jo = urca::ca.jo(test_dt, type = "trace", ecdet = "trend", spec = "longrun", K = 2)@teststat[1], times = 5 )ca.jo完全一致。
小提示:如果你的使用场景需要批量跑大量协整检验,还可以将核心矩阵运算逻辑用Rcpp改写,速度还能再提升2~5倍
内容的提问来源于stack exchange,提问作者Saurabh
相关产品推荐
相关产品推荐

