You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R中使用mutate(across())动态处理{.col}生成*_se列报错求助

问题排查:mutate(across())计算标准误的符号错误

需求说明

为数据框df中的6组变量(female_*、birth_year_*等)分别新增*_se列,计算逻辑为:对应变量的*_sd列值除以*_n列值的平方根。尝试用mutate(across())实现时触发报错,需排查原因并修正。

数据代码

df <- structure(list(treat = structure(1:4, levels = c("Control", "Treatment 1", "Treatment 2", "Treatment 3"), class = "factor"), female_n = c(314709L, 10456L, 10481L, 10455L), female_mean = c(0.506, 0.506, 0.504, 0.5), female_sd = c(0.5, 0.5, 0.5, 0.5), birth_year_n = c(314709L, 10456L, 10481L, 10455L), birth_year_mean = c(1973.74, 1973.654, 1973.486, 1973.766), birth_year_sd = c(16.867, 16.997, 16.869, 16.89), provided_phone_no_n = c(314709L, 10456L, 10481L, 10455L), provided_phone_no_mean = c(0.656, 0.666, 0.663, 0.647), provided_phone_no_sd = c(0.475, 0.472, 0.473, 0.478), dem_n = c(314709L, 10456L, 10481L, 10455L), dem_mean = c(0.48, 0.474, 0.482, 0.478), dem_sd = c(0.5, 0.499, 0.5, 0.5), rep_n = c(314709L, 10456L, 10481L, 10455L), rep_mean = c(0.136, 0.141, 0.142, 0.138), rep_sd = c(0.343, 0.348, 0.349, 0.345), uaf_n = c(314709L, 10456L, 10481L, 10455L), uaf_mean = c(0.363, 0.365, 0.357, 0.363), uaf_sd = c(0.481, 0.481, 0.479, 0.481)), class = c("tbl_df", "tbl", "data.frame"), row.names = c(NA, -4L))

尝试代码

df %>%
    mutate(
    across(ends_with("_sd"),
           list(
            se = ~.x / sqrt(!!ensym("{str_replace(.col, '_sd', '_n')}"))
           )
    )

报错信息

Error in ensym():
! arg must be a symbol
Backtrace:

  1. ... %>% ...
  2. rlang::abort(message = message)

问题原因

  1. 语法解析错误:代码中用引号包裹{str_replace(.col, '_sd', '_n')},导致它被当作字符串字面量而非表达式解析,ensym()无法将字符串直接转换为符号。
  2. 函数使用不当:ensym()要求传入的是符号(如变量名),而.col在across()的函数体内是列名的字符串,需要先将字符串转换为符号,而非直接用ensym()处理带引号的表达式。

修正方案

方案1:使用sym()转换字符串为符号

library(dplyr)
library(stringr)

df %>%
  mutate(
    across(ends_with("_sd"),
           list(se = ~ .x / sqrt(!!sym(str_replace(.col, "_sd", "_n"))))
    )
  )

方案2:使用.data代词(更简洁安全)

无需手动转换符号,直接通过字符串索引数据框列:

library(dplyr)
library(stringr)

df %>%
  mutate(
    across(ends_with("_sd"),
           list(se = ~ .x / sqrt(.data[[str_replace(.col, "_sd", "_n")]]))
    )
  )

方案3:简化命名格式

利用across()的命名参数自动生成列名,无需list(se=...):

library(dplyr)
library(stringr)

df %>%
  mutate(
    across(ends_with("_sd"),
           ~ .x / sqrt(.data[[str_replace(cur_column(), "_sd", "_n")]]),
           .names = "{str_remove(.col, '_sd')}_se"
    )
  )

内容的提问来源于stack exchange,提问作者C.Robin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 12:44:52