You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何利用部分匹配的列名更便捷地创建新列?

批量创建新变量的R语言便捷方案

原需求是通过mutate批量生成a_new、b_new这类变量,避免手动重复编写* = * * *_increase的逻辑,以下是几种tidyverse生态下的高效实现方式:

原代码参考

library(tidyverse)  

test_data <- data.frame(a=c(1:3),
                        b=c(3:5),
                        c=c(1:3),
                        a_increase=c(1:3)/100,
                        b_increase=c(3:5)/100,
                        c_increase=c(6:8)/100)

test_data %>% mutate(
  a_new =a* a_increase,
  b_new =b* b_increase,
  c_new =c* c_increase
)

方案1:使用across()(推荐)

利用dplyr::across()遍历目标变量,结合字符串拼接自动匹配对应_increase列,无需手动逐个编写计算式:

手动指定基础变量

如果明确知道基础变量名(如a、b、c):

test_data %>%
  mutate(
    across(c(a, b, c),  # 或用all_of(c("a", "b", "c"))处理字符串向量
           ~ .x * get(str_c(cur_column(), "_increase")),
           .names = "{.col}_new")  # 新变量命名规则:原变量名+_new
  )

自动识别基础变量

如果基础变量和_increase变量命名规则统一,可自动提取基础变量(排除带_increase后缀的列):

# 提取所有不含_increase后缀的列作为基础变量
base_vars <- setdiff(names(test_data), str_subset(names(test_data), "_increase"))

test_data %>%
  mutate(
    across(all_of(base_vars),
           ~ .x * get(str_c(cur_column(), "_increase")),
           .names = "{.col}_new")
  )

方案2:使用purrr::map_dfc

通过purrr的映射函数批量生成新列,再合并到原数据:

test_data %>%
  bind_cols(
    map_dfc(base_vars, ~ {
      # 计算单个新列并命名
      test_data[[.x]] * test_data[[str_c(.x, "_increase")]] %>%
        set_names(str_c(.x, "_new"))
    })
  )

以上两种方案都能适配批量变量场景,只要基础变量与对应_increase变量的命名规则统一,不管是a到z还是更多变量,都能一键完成计算。

内容的提问来源于stack exchange,提问作者anderwyang

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.24 18:23:09