You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言中如何实现多选变量的spread展开拆分为多列

R实现多选国家列拆分宽表方案

原始country列为逗号分隔的多选存储格式,存在尾逗号、分隔符带空格的情况,先做简单清洗再转宽格式即可,提供两种常用实现方式:

方法1:tidyverse 实现(推荐)

先加载依赖包,构造和示例一致的测试数据:

library(tidyverse)

# 构造示例原始数据
df <- tibble(
  row_id = 1:3,
  country = c("1, 2, 3", "1,", "2")
)

核心处理逻辑:先清洗拆分多选内容为列表,拆为长表后转宽表填充对应值:

df_res <- df %>%
  mutate(
    # 清洗字段:去掉末尾多余逗号、去除首尾空格,按", "拆分出所有选中的国家
    country_selected = str_split(str_trim(str_remove(country, ",$")), ", ")
  ) %>%
  # 把列表形式的选中项拆成单独行
  unnest(country_selected) %>%
  mutate(
    col_name = paste0("country_", country_selected),
    col_value = as.numeric(country_selected)
  ) %>%
  # 长表转宽表,未选中的位置默认填充NA
  pivot_wider(names_from = col_name, values_from = col_value) %>%
  # 移除中间处理列和原始country列
  select(-country, -country_selected)

运行后得到的结果和目标格式完全一致:

# A tibble: 3 × 4
 row_id country_1 country_2 country_3
  <int>     <dbl>     <dbl>     <dbl>
1      1         1         2         3
2      2         1        NA        NA
3      3        NA         2        NA

如果需要未选中的单元格留空而非显示NA,在pivot_wider()中加入参数values_fill = list(col_value = "")即可。

方法2:基础R实现(无需安装额外包)

# 清洗并拆分每个样本的选中国家
country_split <- strsplit(trimws(gsub(",$", "", df$country)), ", ")
# 提取所有出现过的国家选项,生成新列名
all_countries <- sort(unique(unlist(country_split)))
new_col_names <- paste0("country_", all_countries)
# 初始化空的结果矩阵
res_matrix <- matrix(
  NA, 
  nrow = nrow(df), 
  ncol = length(new_col_names),
  dimnames = list(NULL, new_col_names)
)
# 逐行填充选中的国家值
for (i in seq_along(country_split)) {
  selected_id <- as.numeric(country_split[[i]])
  res_matrix[i, paste0("country_", selected_id)] <- selected_id
}
# 合并其他原始列得到最终结果
df_res_base <- cbind(df[, setdiff(names(df), "country")], res_matrix)

内容的提问来源于stack exchange,提问作者Nick

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.28 21:06:33