You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R语言中计算data.frame各列样本差值的方法求助

计算两个样本的数值列差值(dplyr实现)

问题描述

现有如下数据框:

df <- data.frame(name = c('sample1', 'sample2'),
      a = c(1.12, 3.18),
      b = c(2.23, 6.29),
      c = c(3.62, 1.72),
      d = c(11.56, 6.22),
      e = c(4.11, 19.38))

我可以用以下代码计算各数值列的均值:

library(dplyr)
df_mean <- df %>% summarise(across(where(is.numeric), ~mean(.x, na.rm=TRUE)))

但想要生成结构类似的data.frame,用sample2与sample1的差值替代均值,尝试了下面的代码但无法运行:

df_difference <- df %>% summarise(across(where(is.numeric), ~.y-.x))

预期结果为:

df_difference <- data.frame(a=2.06, b=4.06, c=-1.9, d=-5.34, e=15.27)

解决方案

正确代码实现

由于数据仅包含两个样本,直接对每个数值列取第二行(sample2)减去第一行(sample1)的值即可。在across的匿名函数中,.x代表当前处理的列向量,利用索引取值计算差值:

library(dplyr)

df_difference <- df %>% 
  summarise(across(where(is.numeric), ~.x[2] - .x[1]))

运行后输出结果与预期完全一致:

> df_difference
     a    b    c     d     e
1 2.06 4.06 -1.9 -5.34 15.27

错误原因说明

之前的代码中使用的.y在across的函数上下文里没有定义,across的匿名函数默认仅提供.x参数(当前列的向量),因此调用未定义的.y会导致对象不存在的报错。

内容的提问来源于stack exchange,提问作者Sylvia Rodriguez

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.03 10:52:21