You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

分层线性回归(Hierarchical Linear Regression)建模出现NaN问题求助

分层线性回归模型出现NaN的原因及解决方法

问题重现

你编写的分层线性回归代码中,step3加入全部4个预测变量后,模型结果出现大量NaN,但移除其中一个变量(Adherence或Competence)后模型可正常运行。代码及数据如下:

# HLR
setwd("c:/users/miria/desktop")
patient1 <- read.csv("patient1.csv")
library(readr)
Statistics <- read_csv("patient1.csv")

# step 1
step1 <- lm(Symptoms~ Meeting + Alliance, data= Statistics)
summary(step1)

# step 2
step2 <- lm(Symptoms~ Meeting + Alliance + Adherence, data= Statistics)
summary(step2)

# step 3 出现NaN的模型
step3 <- lm(Symptoms~ Meeting + Alliance + Adherence + Competence, data= Statistics)
summary(step3)

数据结构:

# patient1数据
structure(list(Meeting = 1:5, Competence = c(4.75, 4.44, 3.33, 4.4, 3.8), 
               Adherence = c(0.23, 1.65, 0.32, 1.54, 1.16), 
               Alliance = c(12L, 2L, 5L, 6L, 7L), 
               Symptoms = c(37L, 46L, 47L, 48L, 40L)), 
          class = "data.frame", row.names = c(NA, -5L))

核心原因

你的样本量只有5条观测,但step3的模型需要估计的参数包括:截距项 + 4个预测变量的系数,一共5个参数。此时模型的残差自由度为 n - k = 5 - 5 = 0,没有剩余自由度计算标准误、t统计量、p值等指标,这些统计量因此变成NaN。
简单来说,5个数据点刚好能完美拟合5个参数的模型,但无法评估模型的统计显著性和稳定性。

解决办法

  • 增加样本量:这是最彻底的解决方案。收集更多患者的观测数据,让样本量远大于参数个数(一般建议至少是参数数的5-10倍),模型就能正常计算所有统计量。
  • 精简预测变量:
    1. 用cor(Statistics[,c("Adherence","Competence","Meeting","Alliance")])查看变量间相关性,移除高度相关的冗余变量;
    2. 参考step1、step2的模型结果,保留对Symptoms解释力更强的变量(比如看R²变化、系数显著性),去掉解释力弱的变量。
  • 使用正则化回归:如果必须保留所有4个预测变量,可采用岭回归或LASSO这类正则化方法,在小样本、参数过多的情况下给出稳定的参数估计。示例代码(用glmnet包):
library(glmnet)
# 构建矩阵
x <- model.matrix(Symptoms~ Meeting + Alliance + Adherence + Competence, data=Statistics)[,-1]
y <- Statistics$Symptoms
# 岭回归
ridge_model <- glmnet(x, y, alpha=0)
# 查看结果
print(ridge_model)

内容的提问来源于stack exchange,提问作者miriam

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 18:50:41