基于R语言按姓名分组排查同名不同银行账户的员工数据问题
Fixing Your R Code to Find Names with Different Bank Accounts
The core issue with your current code is that you're deduplicating bank accounts across all names instead of checking for duplicates within each name group. That's why Chris's 1280 account got removed—because Cassy's 1280 appeared first in your sorted data.
Here's the corrected approach using dplyr that groups by name, filters for names with multiple distinct bank accounts, and keeps unique account entries per name:
library(dplyr) # Your original sample data sample <- data.frame( Emp_id = c("123","134","143","143","235","433","231","120","135","150","150","900","900"), Name = c("Joan","Karyn","Larry","Larry","Larry","Larry","Larry","Amy","Amy","Chris","Chris","Cassy","Cassy"), Bank_Account = c("6758","1244","4900","5201","5201","5201","5201","7890","7890","1280","6565","1280","9873") ) # Step-by-step solution result <- sample %>% # Group rows by employee name group_by(Name) %>% # Only keep groups where there are 2+ distinct bank accounts filter(n_distinct(Bank_Account) > 1) %>% # Keep one unique entry per (Name, Bank_Account) pair (preserving Emp_id) distinct(Bank_Account, .keep_all = TRUE) %>% # Remove grouping and sort by name for readability ungroup() %>% arrange(Name) # View the result print(result)
Expected Output:
Emp_id Name Bank_Account 1 900 Cassy 1280 2 900 Cassy 9873 3 150 Chris 1280 4 150 Chris 6565 5 143 Larry 4900 6 143 Larry 5201
Key Fixes Explained:
- Grouping by Name: Ensures we only analyze bank accounts within each employee's name group, not across all names.
- Filtering for Multiple Accounts: Uses
n_distinct(Bank_Account) > 1to keep only names that have different bank accounts. - Distinct per Group:
distinct(Bank_Account, .keep_all = TRUE)keeps one row for each unique bank account per name, avoiding cross-name deduplication.
内容的提问来源于stack exchange,提问作者Jazz
相关产品推荐
相关产品推荐

