You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R中通过tidyverse实现one-hot encoding 不使用caret、mltools等包

Tidyverse下实现n-1格式独热编码的方法

首先加载基础依赖和测试数据集:

# 加载所需包
library(tidyverse)
library(vcd)

# 载入测试数据集
data(Arthritis)

方法1:仅用dplyr+tidyr核心tidyverse包实现

不需要安装任何tidyverse之外的第三方包,依赖R原生的model.matrix默认生成n-1格式编码的特性实现,代码如下:

arthritis_encoded <- Arthritis %>%
  # 对所有分类列生成n-1独热编码,跳过数值列Age和ID
  mutate(
    across(where(is.factor) & !c(Age, ID), 
           ~model.matrix(~.x)[, -1, drop = FALSE] %>% as_tibble(.name_repair = "unique"))
  ) %>%
  # 展开生成的编码列
  unnest(where(is_tibble), names_sep = "_") %>%
  # 清理冗余列名字符
  rename_with(~str_remove(.x, "\\.x_"), ends_with(c("Treated", "Male", "Some", "Marked")))

执行后查看前5行结果验证:

head(arthritis_encoded, 5)

输出示例:

ID Treatment Sex Age Improved Treatment_Treated Sex_Male Improved_Some Improved_Marked
1 57   Treated Male  27     Some                 1        1             1               0
2 46   Treated Male  29     None                 1        1             0               0
3 77   Treated Male  30     None                 1        1             0               0
4 17   Treated Male  32   Marked                 1        1             0               1
5 36   Treated Male  46   Marked                 1        1             0               1

可以看到二分类的Sex仅保留1个编码列,三分类的Improved保留2个编码列,Age完全保留未参与编码,完全符合需求。


方法2:使用tidyverse生态recipes包实现

如果需要将编码融入机器学习预处理流水线,推荐使用tidymodels体系下的recipes包,代码更简洁可复用,默认行为就是生成n-1格式编码:

library(recipes)

# 构建预处理配方
encoding_rec <- recipe(~ ., data = Arthritis) %>%
  # 指定对所有分类列做哑变量编码,默认输出n-1个特征
  step_dummy(all_nominal_predictors()) %>%
  # 执行预处理
  prep() %>%
  # 提取处理后的数据集
  juice()

内容的提问来源于stack exchange,提问作者Eisen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.05 21:45:02