基于分组特定行列值创建新变量Var3的R语言技术求助
Solution to Generate Var3 by Grouping ID with Your Specified Rules
Let's tackle this problem step by step. The core requirement is to calculate Var3 by grouping records by ID, and referencing the row where Var1 == 1 within each group for the third rule. We'll use the dplyr package from the tidyverse, which makes grouped operations straightforward and readable.
First, Fix the Original Data Input
Your original read.table code had formatting issues (the data wasn't properly split into rows). Here's the corrected version to load the data correctly:
dd <- read.table(text= "ID Var1 Year Month Var2 1 0 2008 2 -4.17 1 0 2009 2 -3.17 1 0 2010 2 -2.17 1 0 2011 2 -1.17 1 1 2013 2 0.17 1 0 2014 10 2.84 2 0 2008 2 -3.33 2 0 2010 2 -1.33 2 1 2011 2 0.33 2 0 2013 2 5.43 2 0 2015 11 6.43 3 1 2010 2 0.85 3 0 2011 2 2.33 3 0 2013 2 4.86 3 0 2015 11 6.33 3 0 2017 10 4.12", header=TRUE)
Solution Code Using dplyr
library(dplyr) dd_new <- dd %>% group_by(ID) %>% # Extract reference values from the row where Var1 == 1 for each group mutate( Var2_ref = first(Var2[Var1 == 1]), Year_ref = first(Year[Var1 == 1]), Month_ref = first(Month[Var1 == 1]) ) %>% # Calculate Var3 using your three rules mutate(Var3 = case_when( # Rule 1: If Var1 == 1, Var3 equals Var2 Var1 == 1 ~ Var2, # Rule 2: If Var2 < 0, Var3 equals Var2 Var2 < 0 ~ Var2, # Rule 3: For Var2 >= 0, compute using reference values TRUE ~ Var2_ref + (Year - Year_ref) + (Month - Month_ref)/12 )) %>% # Remove temporary reference columns (optional, clean up the output) select(-Var2_ref, -Year_ref, -Month_ref) %>% ungroup() # View the final result print(dd_new, digits = 6)
How This Works
- Grouping by ID:
group_by(ID)ensures all calculations are isolated to each individual's records. - Reference Values:
first(Var2[Var1 == 1])grabs theVar2value from the single row whereVar1 == 1in each group (we assume each ID has exactly one such row, which matches your sample data). We replicate this forYearandMonthto get the reference date point. - Case Logic:
case_when()applies your rules in priority order—earlier conditions override later ones, which aligns with your requirements (e.g., a row that meets Rule 1 won't be evaluated against Rule 2 or 3).
Verification Against Your Target
Running this code will match your dd_new example closely. For example:
- For ID=1, the 6th row (2014-10):
0.17 + (2014-2013) + (10-2)/12 = 0.17 + 1 + 8/12 ≈ 1.836667, which matches your target value. - Note: Your sample target lists ID=3's Var1=1 row as having
Var2=0.67, but your original data uses0.85—this explains the minor discrepancy in the final row's Var3 value, which our code calculates correctly based on your input data.
Notes
- If an ID has multiple rows where
Var1 ==1, adjust the code to pick the correct reference row (e.g., uselast()instead offirst(), or add filtering logic). - If you prefer base R, you can use
split()+lapply()to process each group, butdplyris far more readable for this type of grouped operation.
内容的提问来源于stack exchange,提问作者Mary B.
相关产品推荐
相关产品推荐

