You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于R语言线性回归:将独热编码特征系数合并为单一系数

解决R语言线性回归独热编码系数合并问题

步骤1:提取模型系数并关联原始特征列

先从训练好的线性回归模型中提取系数,再通过独热编码生成的特征名,拆分出对应的原始列名(比如color.blue拆分出color)。

# 提取模型系数(排除截距项)
coef_df <- as.data.frame(summary(model)$coefficients[-1, , drop = FALSE])
coef_df$feature_name <- rownames(coef_df)

# 从独热编码特征名中拆分原始列名
coef_df$original_feature <- gsub("\\..*", "", coef_df$feature_name)

步骤2:按原始列名聚合系数

根据需求选择聚合逻辑(平均值、中位数、求和等),这里以取平均值为例:

library(dplyr)

# 按原始特征分组计算平均系数,同时加入截距项
aggregated_coef <- coef_df %>%
  group_by(original_feature) %>%
  summarise(Estimate = mean(Estimate, na.rm = TRUE)) %>%
  bind_rows(tibble(original_feature = "(Intercept)", Estimate = coef(model)[1])) %>%
  arrange(match(original_feature, c("(Intercept)", unique(coef_df$original_feature))))

# 输出合并后的结果
print(aggregated_coef)

步骤3:处理参考水平的无效系数

独热编码的参考水平系数为0且带有NA值,聚合时用na.rm = TRUE可自动忽略这些无效值,确保计算的是有效编码特征的系数均值。

完整示例代码

# 模拟带多分类特征的数据集
set.seed(123)
data <- data.frame(
  y = rnorm(100),
  color = factor(sample(c("red", "blue", "green"), 100, replace = TRUE)),
  size = factor(sample(c("S", "M", "L"), 100, replace = TRUE))
)

# 用caret做独热编码
library(caret)
dmy <- dummyVars(" ~ .", data = data)
data_encoded <- data.frame(predict(dmy, newdata = data))

# 训练线性回归模型
model <- lm(y ~ ., data = data_encoded)

# 提取并合并系数
library(dplyr)
coef_df <- as.data.frame(summary(model)$coefficients[-1, , drop = FALSE])
coef_df$feature_name <- rownames(coef_df)
coef_df$original_feature <- gsub("\\..*", "", coef_df$feature_name)

aggregated_coef <- coef_df %>%
  group_by(original_feature) %>%
  summarise(Estimate = mean(Estimate, na.rm = TRUE)) %>%
  bind_rows(tibble(original_feature = "(Intercept)", Estimate = coef(model)[1])) %>%
  arrange(match(original_feature, c("(Intercept)", unique(coef_df$original_feature))))

# 查看最终合并结果
aggregated_coef

示例输出

# A tibble: 3 × 2
  original_feature Estimate
  <chr>               <dbl>
1 (Intercept)        0.0779
2 color             -0.0870
3 size               0.0442

补充说明

  • 若需要其他聚合方式,将mean替换为median、sum等函数即可
  • 拆分原始列名的正则表达式适用于原始列名.类别的特征名格式,若使用其他编码工具生成的格式不同,需调整正则规则
  • 用model.matrix做独热编码时,处理逻辑完全一致,因为特征名格式相同

内容的提问来源于stack exchange,提问作者Subhashree Kar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 05:30:47