You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于分组特定行列值创建新变量Var3的R语言技术求助

Solution to Generate Var3 by Grouping ID with Your Specified Rules

Let's tackle this problem step by step. The core requirement is to calculate Var3 by grouping records by ID, and referencing the row where Var1 == 1 within each group for the third rule. We'll use the dplyr package from the tidyverse, which makes grouped operations straightforward and readable.

First, Fix the Original Data Input

Your original read.table code had formatting issues (the data wasn't properly split into rows). Here's the corrected version to load the data correctly:

dd <- read.table(text= "ID Var1 Year Month Var2
1 0 2008 2 -4.17
1 0 2009 2 -3.17
1 0 2010 2 -2.17
1 0 2011 2 -1.17
1 1 2013 2 0.17
1 0 2014 10 2.84
2 0 2008 2 -3.33
2 0 2010 2 -1.33
2 1 2011 2 0.33
2 0 2013 2 5.43
2 0 2015 11 6.43
3 1 2010 2 0.85
3 0 2011 2 2.33
3 0 2013 2 4.86
3 0 2015 11 6.33
3 0 2017 10 4.12", header=TRUE)

Solution Code Using dplyr

library(dplyr)

dd_new <- dd %>%
  group_by(ID) %>%
  # Extract reference values from the row where Var1 == 1 for each group
  mutate(
    Var2_ref = first(Var2[Var1 == 1]),
    Year_ref = first(Year[Var1 == 1]),
    Month_ref = first(Month[Var1 == 1])
  ) %>%
  # Calculate Var3 using your three rules
  mutate(Var3 = case_when(
    # Rule 1: If Var1 == 1, Var3 equals Var2
    Var1 == 1 ~ Var2,
    # Rule 2: If Var2 < 0, Var3 equals Var2
    Var2 < 0 ~ Var2,
    # Rule 3: For Var2 >= 0, compute using reference values
    TRUE ~ Var2_ref + (Year - Year_ref) + (Month - Month_ref)/12
  )) %>%
  # Remove temporary reference columns (optional, clean up the output)
  select(-Var2_ref, -Year_ref, -Month_ref) %>%
  ungroup()

# View the final result
print(dd_new, digits = 6)

How This Works

  • Grouping by ID: group_by(ID) ensures all calculations are isolated to each individual's records.
  • Reference Values: first(Var2[Var1 == 1]) grabs the Var2 value from the single row where Var1 == 1 in each group (we assume each ID has exactly one such row, which matches your sample data). We replicate this for Year and Month to get the reference date point.
  • Case Logic: case_when() applies your rules in priority order—earlier conditions override later ones, which aligns with your requirements (e.g., a row that meets Rule 1 won't be evaluated against Rule 2 or 3).

Verification Against Your Target

Running this code will match your dd_new example closely. For example:

  • For ID=1, the 6th row (2014-10): 0.17 + (2014-2013) + (10-2)/12 = 0.17 + 1 + 8/12 ≈ 1.836667, which matches your target value.
  • Note: Your sample target lists ID=3's Var1=1 row as having Var2=0.67, but your original data uses 0.85—this explains the minor discrepancy in the final row's Var3 value, which our code calculates correctly based on your input data.

Notes

  • If an ID has multiple rows where Var1 ==1, adjust the code to pick the correct reference row (e.g., use last() instead of first(), or add filtering logic).
  • If you prefer base R, you can use split() + lapply() to process each group, but dplyr is far more readable for this type of grouped operation.

内容的提问来源于stack exchange,提问作者Mary B.

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 11:07:33