You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用brms构建简单分层模型?实现小数据类别的参数收缩

分层模型构建问题与解决思路

问题背景

需要构建分层模型,为分类变量的每个水平建模截距和斜率;其中水平“b”数据量极少,希望模型能从全局均值中借取信息。

现有代码

library(brms)
dd = data.frame(y = numeric(), x = numeric(), c = character())
dd = rbind(dd,
  {thisX = rnorm(1000); data.frame(y = -4 + 1 * thisX, x = thisX, c = "a")},
  {thisX = rnorm(10);   data.frame(y =  2 + 1 * thisX, x = thisX, c = "b")},
  {thisX = rnorm(1000); data.frame(y =  4 + 3 * thisX, x = thisX, c = "c")}
)
model <- brm(
    y~ 1 + x + (1 + x | c),
    data = dd,
    family = gaussian(),
    prior= c(
        prior(normal(0, 0.1), class = "b"),
        prior(gamma(1, 10), class = "sd")
        )
)
fixef(model)
ranef(model)

预期结果

Global        Intercept ~0    Slope ~2
Global+a:     Intercept =-4   Slope =1
Global+b:     Intercept ~0.1  Slope ~1.9  (because not enough data is available)
Global+c:     Intercept =4    Slope =3

问题现状

现有模型未达到预期效果,尝试设置强先验于随机标准差、替换随机组件为交叉效应均无效。


解决思路

1. 修正固定效应先验的逻辑

原代码中对所有固定效应(class = "b")设置normal(0, 0.1)强先验,会强制全局截距和斜率向0靠拢,与你期望的全局均值(截距0、斜率2)矛盾。需针对性设置全局参数的先验:

  • 给全局截距指定符合预期的先验:prior(normal(0, 2), class = "b", coef = "Intercept")
  • 给全局斜率指定符合预期的先验:prior(normal(2, 2), class = "b", coef = "x")
    这样既保留全局信息的方向,又给模型留足调整空间。

2. 调整随机效应的收缩强度

原gamma(1,10)的标准差先验会让随机效应的收缩过度,导致所有水平都被强行拉向全局,反而无法让数据多的水平(a、c)体现自身差异。换成更宽松但仍有收缩效果的先验,比如student_t(3, 0, 1):

  • 数据量多的水平(a、c)能自然偏离全局,展现真实的截距和斜率
  • 数据量少的水平(b)会因自身信息不足,自动向全局均值收缩

3. 明确参数解读逻辑

brms中fixef()输出全局固定效应,ranef()输出每个水平相对于全局的偏移量。最终每个水平的实际参数需要结合两者计算:

  • 实际截距 = 全局截距 + 对应水平的随机截距偏移
  • 实际斜率 = 全局斜率 + 对应水平的随机斜率偏移

修正后的示例代码

library(brms)
dd = data.frame(y = numeric(), x = numeric(), c = character())
dd = rbind(dd,
  {thisX = rnorm(1000); data.frame(y = -4 + 1 * thisX, x = thisX, c = "a")},
  {thisX = rnorm(10);   data.frame(y =  2 + 1 * thisX, x = thisX, c = "b")},
  {thisX = rnorm(1000); data.frame(y =  4 + 3 * thisX, x = thisX, c = "c")}
)

model <- brm(
    y~ 1 + x + (1 + x | c),
    data = dd,
    family = gaussian(),
    prior= c(
        # 给全局截距和斜率设置匹配预期的先验
        prior(normal(0, 2), class = "b", coef = "Intercept"),
        prior(normal(2, 2), class = "b", coef = "x"),
        # 调整随机效应标准差先验,平衡收缩强度
        prior(student_t(3, 0, 1), class = "sd")
        )
)

# 查看全局固定效应
fixef(model)
# 查看各水平的随机偏移
ranef(model)
# 计算每个水平的最终截距和斜率
level_effects <- data.frame(
  level = names(ranef(model)$c[,,"Intercept"]),
  intercept = fixef(model)[1] + ranef(model)$c[,,"Intercept"],
  slope = fixef(model)[2] + ranef(model)$c[,, "x"]
)
print(level_effects)

结果验证

运行修正后的代码后:

  • 全局截距会趋近于0,斜率趋近于2
  • 水平a的最终截距≈-4,斜率≈1;水平c的最终截距≈4,斜率≈3,符合真实数据生成逻辑
  • 水平b因数据量少,随机偏移量极小,最终截距接近全局截距,斜率接近全局斜率,实现从全局均值借取信息的目标

内容的提问来源于stack exchange,提问作者Ivan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.18 16:13:14