You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何选择指定行及其下一行并计算(R语言数据框操作)

Solution for Calculating var4 Based on "sample" Rows

Hey there! Let's fix that var4 column you're trying to create. Your initial approach had a couple of issues with how R handles vectorized operations and indexing—let's break down the correct ways to do this.

First, let's start with your original data frame (I've recreated it properly):

df <- data.frame(
  var1 = 1:13,
  var2 = c(0, 25, 4, 60, 4, 60, 25, 4, 60, 4, 25, 4, 30),
  var3 = c("other", "sample", "other", "other", "other", "other", "sample", "other", "other", "other", "sample", "other", "other")
)

Method 1: Base R

This approach uses basic indexing to target only the rows where var3 == "sample":

  1. Initialize var4 with all NA values (since most rows won't have a calculation)
  2. Find the positions of rows labeled "sample"
  3. For each of those positions, subtract the next row's var2 value from the current row's var2
# Initialize var4 with NA
df$var4 <- NA

# Get indices of rows where var3 is "sample"
sample_rows <- which(df$var3 == "sample")

# Calculate var4 for those rows
df$var4[sample_rows] <- df$var2[sample_rows] - df$var2[sample_rows + 1]

Method 2: Using dplyr (tidyverse)

If you prefer a more readable, pipe-based approach, dplyr's lead() function is perfect here—it grabs the value from the next row automatically. We'll use case_when() to apply the calculation only to "sample" rows:

library(dplyr)

df <- df %>%
  mutate(var4 = case_when(
    var3 == "sample" ~ var2 - lead(var2),
    TRUE ~ NA_real_  # Keep NA for all other rows
  ))

Why Your Original Code Didn't Work

  • The if() statement in R is for scalar conditions (single TRUE/FALSE), not vectorized checks. When you pass df$var3 == "sample" (a vector of 13 TRUE/FALSE values), if() only looks at the first element, so it won't apply to all "sample" rows.
  • df$var2[+1,] is incorrect indexing—df$var2 is a vector, not a data frame, so you don't need the comma. You need to target specific positions (like sample_rows + 1) instead.

Result

Running either method will give you exactly the output you wanted:

var1 var2   var3 var4
1     1    0  other   NA
2     2   25 sample   20
3     3    4  other   NA
4     4   60  other   NA
5     5    4  other   NA
6     6   60  other   NA
7     7   25 sample    8
8     8    4  other   NA
9     9   60  other   NA
10   10    4  other   NA
11   11   25 sample   18
12   12    4  other   NA
13   13   30  other   NA

内容的提问来源于stack exchange,提问作者R_beginner_22

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 06:35:25