You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何为未知类别数量的分类变量生成哑变量(R语言)

在R中生成类别数减1的哑变量

问题场景

假设我们有如下数据框:

df <- data.frame(type = c("A","B","AB","O","O","B","A"))

该type列包含4种类别,但实际场景中无法提前知晓类别数量,需要生成数量比类别数少1的哑变量(此例中为3个),预期输出如下:

df <- data.frame(type = c("A","B","AB","O","O","B","A"),
                 A = c(1,0,0,0,0,0,1),
                 B = c(0,1,0,0,0,1,0),
                 AB = c(0,0,1,0,0,0,0))

具体选择哪些类别作为哑变量不做要求,核心是实现未知类别情况下的自动生成。

解决方案

方法一:原生R实现(无需额外包)

使用model.matrix()直接生成哑变量矩阵,再与原数据合并:

# 将type转为因子(确保识别所有类别)
df$type <- factor(df$type)
# 生成所有类别哑变量,再移除最后一列(保证数量为类别数-1)
dummy_matrix <- model.matrix(~ type - 1, data = df)[, -nlevels(df$type)]
# 合并原数据与哑变量
df_with_dummies <- cbind(df, dummy_matrix)

方法二:fastDummies包快速实现

如果偏好tidyverse风格,fastDummies包提供了更简洁的接口:

# 首次使用需安装包
# install.packages("fastDummies")
library(fastDummies)

df$type <- factor(df$type)
# 生成type列的哑变量,保留原列
df_with_dummies <- dummy_cols(df, select_columns = "type", remove_selected_columns = FALSE)
# 移除最后一个类别对应的哑变量列
last_level_col <- paste0("type_", levels(df$type)[nlevels(df$type)])
df_with_dummies <- df_with_dummies[, !names(df_with_dummies) %in% last_level_col]

方法三:dplyr自定义实现

依赖dplyr包,通过逻辑判断生成哑变量:

library(dplyr)

df$type <- factor(df$type)
# 获取除最后一个类别外的所有目标类别
target_levels <- levels(df$type)[-nlevels(df$type)]
# 逐个生成哑变量列
df_with_dummies <- df %>%
  mutate(across(all_of(target_levels), ~ as.integer(type == .x)))

内容的提问来源于stack exchange,提问作者Bae

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.24 17:09:45