You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将sparklyr拟合管道的系数值映射至预测变量名称

在Azure Databricks中用sparklyr关联GLM系数与预测变量名称

在Azure Databricks笔记本中使用sparklyr拟合广义线性模型(GLM)后,需要将模型系数值映射到对应的预测变量名称。以下是拟合模型并提取系数的示例代码:

library(sparklyr)

sc <- spark_connect(method = "databricks")

data <- copy_to(sc, mtcars, "mtcars", overwrite = TRUE)

pipeline <- ml_pipeline(sc) %>%
  ft_r_formula(vs ~ cyl + carb) %>%
  ml_generalized_linear_regression(family = "binomial")

partitioned_data <- sdf_random_split(data, train = 0.80, test = 0.20, seed = 42)

fitted_pipeline <- ml_fit(pipeline, partitioned_data$train)

glrm_transformer <- ml_stage(fitted_pipeline, length(fitted_pipeline$stages))

with(glrm_transformer, c(intercept, coefficients))

解决方法

要关联系数与变量名称,关键是从ft_r_formula阶段提取处理后的特征名称,这些名称的顺序与模型系数的顺序完全对应:

  1. 获取公式处理阶段的特征信息
    从拟合完成的流水线中取出第一个阶段(即ft_r_formula),提取其生成的特征名称:

    # 获取公式处理阶段
    formula_stage <- ml_stage(fitted_pipeline, 1)
    # 提取所有预测变量的名称
    feature_names <- formula_stage$features$names
    
  2. 合并截距与系数,关联变量名
    将模型的截距项单独作为"(Intercept)",与系数、变量名合并为数据框:

    # 提取截距和系数
    model_coefficients <- with(glrm_transformer, c(intercept, coefficients))
    # 组合变量名(包含截距)
    all_variable_names <- c("(Intercept)", feature_names)
    # 生成对应的数据框
    coefficient_mapping <- data.frame(
      Variable = all_variable_names,
      Coefficient = model_coefficients,
      stringsAsFactors = FALSE
    )
    
    # 查看结果
    print(coefficient_mapping)
    

说明

  • ft_r_formula会自动处理公式中的变量转换(比如因子哑编码、交互项等),features$names返回的是处理后的最终特征名称,确保与模型系数的顺序完全匹配。
  • 最终生成的coefficient_mapping数据框清晰展示了每个变量(包括截距)对应的系数值,方便后续分析。

内容的提问来源于stack exchange,提问作者the-mad-statter

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.13 11:41:10