R循环提取回归系数与标准误:基于hsb2数据集的技术求助
解决提取回归模型系数与标准误的问题
Hey there! Let's sort out this coefficient extraction for you. Your initial attempt has a couple of small issues—cov isn't defined in your code, and you weren't linking the variable names from varlist to each model's results. Here's how to get that clean three-column output you're aiming for:
修正后的代码(Base R 版本)
First, let's make sure we're mapping each variable name to its corresponding model results. We can do this easily with sapply to pull the values we need, then combine everything into a data frame:
hsb2 <- read.csv("https://stats.idre.ucla.edu/stat/data/hsb2.csv") varlist <- names(hsb2)[8:11] models <- lapply(varlist, function(x) { lm(substitute(read ~ i, list(i = as.name(x))), data = hsb2) }) # 提取变量名、系数和标准误,生成目标数据框 results_df <- data.frame( variable = varlist, coefficient = sapply(models, function(mod) coef(summary(mod))[2, 1]), std_error = sapply(models, function(mod) coef(summary(mod))[2, 2]), stringsAsFactors = FALSE ) # 查看结果 print(results_df)
代码解释
variable = varlist: 直接用我们定义的varlist作为第一列的变量名,保证每个结果对应正确的自变量。sapply(models, function(mod) coef(summary(mod))[2, 1]): 遍历每个模型,提取汇总表中第二行第一列的值(也就是自变量的回归系数)。sapply(models, function(mod) coef(summary(mod))[2, 2]): 同理,提取第二行第二列的值(自变量系数的标准误)。
示例输出
运行代码后,你会得到类似这样的结果:
variable coefficient std_error 1 write 0.5521847 0.0727684 2 math 0.6439474 0.0666910 3 science 0.5889194 0.0701747 4 socst 0.5302468 0.0747758
如果你更喜欢用lapply结合do.call(rbind, ...)的方式(更灵活处理复杂情况),这里有另一种写法:
# 给models列表命名,对应自变量名 names(models) <- varlist results_df <- do.call(rbind, lapply(names(models), function(var) { mod_summary <- summary(models[[var]]) coef_vals <- coef(mod_summary)[2, ] data.frame( variable = var, coefficient = coef_vals[1], std_error = coef_vals[2], stringsAsFactors = FALSE ) }))
这两种方法都能帮你得到想要的三列输出,挑你觉得更顺手的就行!
内容的提问来源于stack exchange,提问作者Oscar
相关产品推荐
相关产品推荐

