关于验证选定变量滞后值与当前值是否相等及批量处理data.frame变量差异判断的技术问询
Automated Solution to Check if All Selected Variables Match Previous Row
Hey there! Let's solve this problem so you don't have to hardcode variable names ever again. You want a flexible way to generate that same column, regardless of how many variables you have or what they're named—here are two straightforward approaches:
Using dplyr (Tidyverse)
This method is clean and readable, perfect if you're already using the tidyverse ecosystem:
library(dplyr) # Your original dataset df <- data.frame(id = c(1,2,3,4), var1 = c(1,1,2,2), var2 = c(1,1,2,3)) # Generate the `same` column automatically df <- df %>% # Step 1: For every column except `id`, check if current row equals the previous row mutate(across(-id, ~ . == lag(.))) %>% # Step 2: For each row, check if ALL those comparisons are TRUE rowwise() %>% mutate(same = all(c_across(-id))) %>% ungroup() %>% # Step 3: Clean up temporary columns and set first row to NA select(-all_of(setdiff(names(.), c("id", "var1", "var2", "same")))) %>% mutate(same = if_else(row_number() == 1, NA, same))
How it works:
across(-id, ~ . == lag(.)): Creates temporary logical columns for each variable (excludingid) where each value isTRUEif it matches the row above.rowwise() + all(c_across(-id)): Checks every row to see if all the temporary logical values areTRUE(meaning every variable matches the previous row).- The final steps clean up the temporary columns and set the first row's
samevalue toNA(since there's no prior row to compare against).
Using Base R
If you prefer not to load extra libraries, this base R approach works just as well:
# Your original dataset df <- data.frame(id = c(1,2,3,4), var1 = c(1,1,2,2), var2 = c(1,1,2,3)) # Automatically get the columns to compare (all except `id`) compare_cols <- setdiff(names(df), "id") # Compare each row to the one above, then check if all matches are TRUE row_matches <- apply(df[-1, compare_cols] == df[-nrow(df), compare_cols], 1, all) # Add the `same` column (first row is NA, followed by our match results) df$same <- c(NA, row_matches)
How it works:
setdiff(names(df), "id"): Dynamically grabs all variable names exceptid—no hardcoding needed, even if you add/remove variables later.df[-1, compare_cols] == df[-nrow(df), compare_cols]: Creates a matrix where each cell isTRUEif the current row's variable matches the previous row's.apply(..., 1, all): Checks each row of the matrix to see if every value isTRUE(all variables match the prior row).c(NA, row_matches): Adds theNAfor the first row and appends our match results for the rest.
Both methods will output exactly the target dataset you wanted:
> df id var1 var2 same 1 1 1 1 NA 2 2 1 1 TRUE 3 3 2 2 FALSE 4 4 2 3 FALSE
内容的提问来源于stack exchange,提问作者Jirka Čep
相关产品推荐
相关产品推荐

