You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R语言中将文本变量拆分为二元变量的实现方法

需求说明

我有一份患者级数据,其中category_name_en列存储了患者特征。实际仅存在17种独立特征,但由于多数患者同时具备多个特征(用|分隔),导致该列出现了600余种唯一组合。我需要将这些特征拆分为单独的列,用Yes/No二元变量标记患者是否具备对应特征,解决组合过多的问题。

示例数据框(sample_df)

structure(list(category_name_en = c("Chronic Heart Disease | Genetic Blood Disorders (including Sickle Cell Anemia) | Diabetes", 
"Chronic Heart Disease | Diabetes | Pregnancy", "Chronic Heart Disease | Chronic Liver Disease | Diabetes", 
"Elderly ≥ 55 years old | Chronic Obstructive Pulmonary Disease (including Asthma) | Acquired or Genetic Immunodeficiency  Disorders", 
"(Dialysis (Failure Kid | Diabetes | Chronic Obstructive Pulmonary Disease (including Asthma)", 
"Chronic Obstructive Pulmonary Disease (including Asthma) | Elderly ≥ 55 years old | Diabetes", 
"Genetic Blood Disorders (including Sickle Cell Anemia) | Chronic Heart Disease | Elderly ≥ 55 years old", 
"Chronic Heart Disease | Genetic Blood Disorders (including Sickle Cell Anemia) | Diabetes | Elderly ≥ 55 years old | Neurological Disease", 
"Elderly = 55 years old | Diabetes | Genetic Blood Disorders (including Sickle Cell Anemia)", 
"Pregnancy | Chronic Obstructive Pulmonary Disease (including Asthma)", 
"Healthcare practitioners | Healthy Client | Diabetes", "Chronic Heart Disease | Elderly ≥ 55 years old |  (Dialysis (Failure Kid", 
"Healthy Client | Elderly = 55 years old | Pregnancy", "Chronic Obstructive Pulmonary Disease (including Asthma) | Acquired or Genetic Immunodeficiency  Disorders", 
"Chronic Heart Disease | Elderly ≥ 55 years old | Diabetes", 
"Elderly ≥ 55 years old | Chronic Heart Disease | Diabetes | Chronic Obstructive Pulmonary Disease (including Asthma)", 
"Elderly ≥ 55 years old | Chronic Obstructive Pulmonary Disease (including Asthma) | Chronic Heart Disease | Chronic Liver Disease | Diabetes", 
"Chronic Heart Disease | Diabetes | Healthcare practitioners", 
"Cancer | Elderly ≥ 55 years old | Acquired or Genetic Immunodeficiency  Disorders", 
"Chronic Obstructive Pulmonary Disease (including Asthma) | Cancer | Diabetes | Chronic Heart Disease"
), patid = c(428L, 625L, 149L, 356L, 393L, 444L, 618L, 622L, 
488L, 289L, 130L, 335L, 522L, 284L, 39L, 187L, 391L, 663L, 653L, 
563L)), class = "data.frame", row.names = c(NA, -20L))

期望输出格式

每个唯一特征作为单独列,用Yes/No标记患者是否具备该特征:

patid慢性心脏病遗传性血液疾病(含镰状细胞贫血)糖尿病妊娠慢性肝病55岁及以上慢性阻塞性肺疾病(含哮喘)获得性或遗传性免疫缺陷疾病透析(肾功能衰竭)神经系统疾病医护人员健康人群55岁癌症
428YesYesYesNoNoNoNoNoNoNoNoNoNoNo
625YesNoYesYesNoNoNoNoNoNoNoNoNoNo
149YesNoYesNoYesNoNoNoNoNoNoNoNoNo
.............................................

解决方案(R语言)

使用tidyverse工具集的separate_rows和pivot_wider函数即可实现需求,步骤如下:

  1. 加载依赖包
library(tidyverse)
  1. 拆分特征并转换为宽格式
# 拆分|分隔的特征,每行保留单个特征,同时去除特征前后空格
split_df <- sample_df %>%
  separate_rows(category_name_en, sep = "\\|") %>%
  mutate(category_name_en = str_trim(category_name_en)) %>%
  mutate(has_feature = "Yes")

# 转换为宽格式,缺失特征标记为No
result_df <- split_df %>%
  pivot_wider(
    id_cols = patid,
    names_from = category_name_en,
    values_from = has_feature,
    values_fill = list(has_feature = "No")
  )
  1. 可选:简化特征名称
    如果需要去掉特征名称中的冗余说明(如括号内的补充内容),可以在拆分后添加名称处理步骤:
split_df <- sample_df %>%
  separate_rows(category_name_en, sep = "\\|") %>%
  mutate(category_name_en = str_trim(category_name_en)) %>%
  # 移除括号及内部内容
  mutate(category_name_en = str_remove(category_name_en, "\\s*\\(.*\\)")) %>%
  mutate(has_feature = "Yes")

运行上述代码后,result_df即为符合要求的输出格式。

内容的提问来源于stack exchange,提问作者abrar_r

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.30 22:00:15