You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R中使用logistf包运行Firth模型时出现不收敛及卡顿问题

Firth逻辑回归模型不收敛且卡住的问题解决

问题重现

数据集处理代码

# Save the data frame
saveRDS(train_proc_rds, "train_proc_example1.rds")
fwrite(train_proc_rds, "train_proc_example2.csv")

# reload data frame
train_proc_df1 <- readRDS("train_proc_example1.rds")
train_proc_df2 <- fread("train_proc_example2.csv")

模型定义与训练代码

library(caret)
library(logistf)
library(data.table)

# Define training control
train.control <- trainControl(method = "repeatedcv", 
                              number = 3, repeats = 3,
                              savePredictions = TRUE,
                              classProbs = TRUE)

# Define the custom model function
firth_model <- list(
  type = "Classification",
  library = "logistf",
  loop = NULL,
  parameters = data.frame(parameter = c("none"), class = c("character"), label = c("none")),
  grid = function(x, y, len = NULL, search = "grid") {
    data.frame(none = "none")
  },
  fit = function(x, y, wts, param, lev, last, classProbs, ...) {
    data <- as.data.frame(x)
    data$group <- y
    logistf(group ~ ., data = data, control = logistf.control(maxit = 100), ...)
  },
  predict = function(modelFit, newdata, submodels = NULL) {
    as.factor(ifelse(predict(modelFit, newdata, type = "response") > 0.5, "AD", "control"))
  },
  prob = function(modelFit, newdata, submodels = NULL) {
    preds <- predict(modelFit, newdata, type = "response")
    data.frame(control = 1 - preds, AD = preds)
  }
)

# Training the model
set.seed(123)
model1 <- train(train_proc_df1[, .SD, .SDcols = !c("AD", "group")],
                train_proc_df1$group,
                method = firth_model,
                trControl = train.control)
print(model1)

报错与卡壳现象

运行后反复出现不收敛警告,调整maxit=100无效,模型在交叉验证过程中卡住,CPU使用率达99%:

model2 <- train(train_proc_df2[, .SD, .SDcols = !c("AD", "group")],
                train_proc_df2$group,
                method = firth_model,
                trControl = train.control)
+ Fold1.Rep1: none=none 
Warning in logistf(group ~ ., data = data, control = logistf.control(maxit = 100),  :
  Nonconverged PL confidence limits: maximum number of iterations for variables: (Intercept), x1, x2, x3, x4, x5, x6, x7, x8, x9, x10, x11, x12, x13, x14, x15, x16, x17, x18, x19, x20, x21, x22, x23, x24 exceeded. Try to increase the number of iterations by passing 'logistpl.control(maxit=...)' to parameter plcontrol
- Fold1.Rep1: none=none 
+ Fold2.Rep1: none=none 
Warning in logistf(group ~ ., data = data, control = logistf.control(maxit = 100),  :
  Nonconverged PL confidence limits: maximum number of iterations for variables: (Intercept), x1, x2, x3, x4, x5, x6, x7, x8, x9, x10, x11, x12, x13, x14, x15, x16, x17, x18, x19, x20, x21, x22, x23, x24 exceeded. Try to increase the number of iterations by passing 'logistpl.control(maxit=...)' to parameter plcontrol
- Fold2.Rep1: none=none 
+ Fold3.Rep1: none=none 
Warning in logistf(group ~ ., data = data, control = logistf.control(maxit = 100),  :
  Nonconverged PL confidence limits: maximum number of iterations for variables: (Intercept), x1, x2, x3, x4, x5, x6, x7, x8, x9, x10, x11, x12, x13, x14, x15, x16, x17, x18, x19, x20, x21, x22, x23, x24 exceeded. Try to increase the number of iterations by passing 'logistpl.control(maxit=...)' to parameter plcontrol
- Fold3.Rep1: none=none 
+ Fold1.Rep2: none=none 
Warning in logistf(group ~ ., data = data, control = logistf.control(maxit = 100),  :
  logistf.fit: Maximum number of iterations for full model exceeded. Try to increase the number of iterations or alter step size by passing 'logistf.control(maxit=..., maxstep=...)' to parameter control

环境信息

  • R版本:
R.version
               _                           
platform       aarch64-apple-darwin20      
arch           aarch64                     
os             darwin20                    
system         aarch64, darwin20            
status                                       
major          4                           
minor          3.2                         
year           2023                        
month          10                          
day            31                          
svn rev        85441                       
language       R                           
version.string R version 4.3.2 (2023-10-31)
nickname       Eye Holes   
  • 包版本:
> packageVersion("caret")
[1] ‘6.0.94’
> packageVersion("logistf")
[1] ‘1.26.0’
> packageVersion("data.table")
[1] ‘1.14.10’

解决建议

1. 调整迭代与置信限参数

同时增加模型迭代次数和PL置信限的迭代次数,调整步长避免迭代震荡:

# 在fit函数中更新logistf的控制参数
logistf(group ~ ., data = data, 
        control = logistf.control(maxit = 500, maxstep = 0.5),
        plcontrol = logistpl.control(maxit = 500), ...)

2. 排查数据问题

  • 检查是否存在完全/准完全分离:
# 对单个数据集检查分离情况
sep_check <- checksep(group ~ ., data = as.data.frame(cbind(x, y)))
print(sep_check)
  • 检查多重共线性:使用car包计算VIF值,移除VIF过高的特征:
library(car)
vif_values <- vif(glm(group ~ ., data = as.data.frame(cbind(x, y)), family = binomial))
print(vif_values)

3. 简化验证流程

先跳过交叉验证,直接拟合单个模型验证收敛性:

# 直接拟合模型测试
test_data <- as.data.frame(cbind(train_proc_df1[, .SD, .SDcols = !c("AD", "group")], group = train_proc_df1$group))
firth_test <- logistf(group ~ ., data = test_data, 
                      control = logistf.control(maxit = 500),
                      plcontrol = logistpl.control(maxit = 500))
summary(firth_test)

如果单个模型能收敛,再逐步增加交叉验证的复杂度。

4. 兼容性与版本检查

  • 更新logistf到最新版本,或尝试在x86架构的R环境中运行,排查arm架构的兼容性问题。
  • 在logistf中添加verbose=TRUE参数,查看迭代过程的详细输出,定位卡壳环节。

内容的提问来源于stack exchange,提问作者LCheng

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 22:10:05