R语言批量修改多列值:修正model.matrix()实现同ID国家列赋值
R数据框按ID批量标记对应国家列问题
问题背景
- 初始测试数据的
dput()输出:
structure(list(id = c(1, 1, 2, 4, 4), country = c("USA", "Japan", "Germany", "Germany", "USA"), USA = c(0, 0, 0, 0, 0), Germany = c(0, 0, 0, 0, 0), Japan = c(0, 0, 0, 0, 0)), class = "data.frame", row.names = c(NA, -5L))
- 核心需求:利用
df$country列存储的ID关联国家信息,将数据框中和国家名同名的列对应位置置为1。注意ID存在重复值,同一个ID关联过的所有国家,都需要在该ID的所有行的对应国家列标记为1。方法需要适配10万行以上规模数据集,且禁止使用pivot类宽长表转换函数,避免批量处理多数据集时出现列结构不统一的问题。 - 目标输出的
dput()结果:
structure(list(id = c(1, 1, 2, 4, 4), country = c("USA", "Japan", "Germany", "Germany", "USA"), USA = c(1, 1, 0, 1, 1), Germany = c(0, 0, 1, 1, 1), Japan = c(1, 1, 0, 0, 0)), class = "data.frame", row.names = c(NA, -5L))
- 现有错误方案:使用
model.matrix逐行生成独热编码,仅给当前行country对应的列赋值1,没有按ID聚合关联国家,输出结果不符合预期,代码和错误输出如下:
df[levels(factor(df$country))] = model.matrix(~country - 1, df)
structure(list(id = c(1, 1, 2, 4, 4), country = c("USA", "Japan", "Germany", "Germany", "USA"), USA = c(1, 0, 0, 0, 1), Germany = c(0, 0, 1, 1, 0), Japan = c(0, 1, 0, 0, 0)), row.names = c(NA, -5L ), class = "data.frame")
解决方法
原有代码的问题很直接:model.matrix是逐行做独热转换,完全没处理同ID共享国家标记的逻辑。只要先提取每个ID关联的所有国家集合,再批量给对应列赋值即可,全程不会改动原有列结构,不需要pivot操作,性能足够支撑10万行以上数据:
# 提取所有国家列名 country_cols <- levels(factor(df$country)) # 生成ID和关联国家的唯一映射,去重避免重复计算 id_country_rel <- unique(df[, c("id", "country")]) # 遍历每个国家列,给对应关联ID的行赋值1 for (col in country_cols) { match_ids <- id_country_rel$id[id_country_rel$country == col] df[[col]] <- as.integer(df$id %in% match_ids) }
如果处理百万行级别的数据想进一步提升速度,可以换成向量化实现,省去循环开销:
country_cols <- levels(factor(df$country)) id_country_rel <- unique(df[, c("id", "country")]) # 批量匹配ID和对应国家的关联关系 match_result <- outer(df$id, country_cols, function(cur_ids, cur_col) { cur_ids %in% id_country_rel$id[id_country_rel$country == cur_col] }) df[, country_cols] <- as.integer(match_result) dim(df[, country_cols]) <- dim(match_result)
运行上述代码后输出的结果和目标结构完全一致,国家列会自动适配每个数据集里country列的实际取值,不会出现缺列、多列导致的结构不统一问题。
内容的提问来源于stack exchange,提问作者ZZ Top
相关产品推荐
相关产品推荐

