You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将DataFrame中以Disease_开头的列转换为0/1二进制值?

处理DataFrame中特定前缀列的二值化转换

原始数据

> dput(df)
structure(list(Person_id = c(123L, 32432L, 34534L, 546L, 567L
), City = c("New York", "LA", "Boston", "New York", "New York"
), Disease_Heart_Failure = c(4L, 0L, 1L, 0L, 8L), Disease_Covid = c(7L, 
2L, 1L, 9L, 0L)), class = "data.frame", row.names = c(NA, -5L
))

需求

针对所有以Disease_开头的列,将列中值为1或大于1的转换为1,值为0的保持不变。

解决方案

方法1:Base R 实现

先筛选出目标列,再批量处理二值化:

# 筛选以Disease_开头的列名
disease_cols <- grep("^Disease_", names(df), value = TRUE)
# 对目标列转换:≥1的转为1,0保持不变
df[disease_cols] <- lapply(df[disease_cols], function(x) as.integer(x >= 1))

方法2:dplyr(tidyverse)实现

用across函数批量处理指定前缀的列:

library(dplyr)

df <- df %>%
  mutate(across(starts_with("Disease_"), ~ as.integer(.x >= 1)))

处理后的结果

> dput(df)
structure(list(Person_id = c(123L, 32432L, 34534L, 546L, 567L
), City = c("New York", "LA", "Boston", "New York", "New York"
), Disease_Heart_Failure = c(1L, 0L, 1L, 0L, 1L), Disease_Covid = c(1L, 
1L, 1L, 1L, 0L)), class = "data.frame", row.names = c(NA, -5L
))

内容的提问来源于stack exchange,提问作者Jamie

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 20:32:48