You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在逻辑回归中处理时间自相关?——基于R语言非随机效应方法解决黑猩猩行为时间序列数据问题

Hey there! Let's work through how to handle that strong first-order autocorrelation (ACF = 0.865) in your chimpanzee behavior dataset—without relying on random effects. Below are three practical R approaches tailored to your binomial logistic regression setup:

1. Add a Lagged Response Variable to the GLM

A simple, intuitive way to control first-order autocorrelation is to include the previous observation's response as a predictor. This directly models how the chimp's prior behavior influences its current choice of route.

# Load dplyr for lagging (base R's lag() works too)
library(dplyr)

# Create a lagged version of your response variable (onHR)
xdata$onHR_lag1 <- lag(xdata$onHR, n = 1)

# Remove the first row (since it has no lagged value)
xdata_adj <- xdata[-1, ]

# Fit the updated logistic regression
model_lagged <- glm(onHR ~ z.no_indep_log + z.MRatio + sin(date.rad) + cos(date.rad) + rain*I(z.time^2) + on_off_trail + onHR_lag1, 
                    data = xdata_adj, family = binomial("logit"))

summary(model_lagged)

Notes:

  • This controls for immediate temporal dependence but may introduce minor endogeneity if unmeasured variables affect both current and past behavior.
  • You’ll lose one data point due to the lag, which is negligible if your dataset is large.

2. Use Generalized Estimating Equations (GEE) with an AR(1) Correlation Structure

GEE is built for repeated-measures/time-series data and lets you explicitly specify an autocorrelation structure without random effects. It estimates population-averaged effects, which aligns well with your goal of understanding how group traits impact route use probabilities overall.

# Load the geepack package
library(geepack)

# Assign a single ID (since you're working with one chimp) and define time waves
xdata$id <- 1
xdata$wave <- 1:nrow(xdata)

# Fit the GEE model with AR(1) correlation
model_gee <- geeglm(onHR ~ z.no_indep_log + z.MRatio + sin(date.rad) + cos(date.rad) + rain*I(z.time^2) + on_off_trail, 
                    data = xdata, family = binomial("logit"),
                    id = id, corstr = "ar1", waves = wave)

summary(model_gee)

Notes:

  • The corstr = "ar1" parameter tells GEE to model first-order autocorrelation, matching your ACF result.
  • GEE accounts for the correlation structure when calculating standard errors, so your coefficient significance tests will be more reliable than a standard GLM.

3. Fit an Explicit AR(1) Logistic Model via Maximum Likelihood

For a more formal modeling of the autocorrelation mechanism, you can directly fit a logistic regression where the latent propensity for route use follows an AR(1) process. This uses maximum likelihood to estimate both your predictor coefficients and the autocorrelation parameter.

# Load maxLik for likelihood maximization
library(maxLik)

# Define the log-likelihood function for an AR(1) logistic model
ar1_logit_loglik <- function(par, X, y) {
  # Split parameters: beta coefficients + autocorrelation phi
  beta <- par[1:ncol(X)]
  phi <- par[ncol(X) + 1]
  n <- length(y)
  
  # Initialize log-likelihood with the first observation
  mu_first <- X[1, ] %*% beta
  loglik <- dbinom(y[1], 1, plogis(mu_first), log = TRUE)
  
  # Iterate through remaining observations to calculate conditional likelihood
  prev_latent <- qlogis(plogis(mu_first))
  for(t in 2:n) {
    current_latent <- X[t, ] %*% beta + phi * prev_latent
    current_prob <- plogis(current_latent)
    loglik <- loglik + dbinom(y[t], 1, current_prob, log = TRUE)
    prev_latent <- current_latent
  }
  
  return(loglik)
}

# Prepare design matrix and response vector
X <- model.matrix(~ z.no_indep_log + z.MRatio + sin(date.rad) + cos(date.rad) + rain*I(z.time^2) + on_off_trail, data = xdata)
y <- xdata$onHR

# Initial parameter values: use standard GLM coefficients + initial phi guess
init_par <- c(coef(glm(y ~ X - 1, family = binomial)), 0.5)

# Maximize the likelihood
model_ar1_logit <- maxLik(ar1_logit_loglik, start = init_par, X = X, y = y)

summary(model_ar1_logit)

Notes:

  • This method models conditional effects (how predictors influence route use given the previous state) and gives you an explicit estimate of the autocorrelation strength (phi).
  • The code is more complex, but it’s the most statistically rigorous option if you want to formalize the temporal dependence structure.

Quick Recommendation:

Start with the GEE approach—it’s simple, robust, and perfectly suited to your goal of estimating average effects while controlling autocorrelation. If you want to focus on how past behavior predicts current choices, the lagged-variable GLM is a fast alternative.

内容的提问来源于stack exchange,提问作者Mirte

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 14:04:05