R语言:如何识别数据框同一行内的重复分组列数据
Got it, let's figure out how to spot duplicate (nameX, typeX, numX) groups within each row of your dataframe. First, I'll use a complete version of your sample data (since the original cuts off mid-entry) to demonstrate the solution clearly:
# Complete sample dataframe df <- data.frame( key = c('1', '2', '3', '4', '5'), name1 = c('black','black','black','red','red'), type1 = c('chair','chair','sofa','sofa','plate'), num1 = c(4,5,12,4,3), name2 = c('black', 'red', 'black', 'green', 'blue'), type2 = c('chair','chair','sofa','bed','plate'), num2 = c(4,7,12,3,1), name3 = c('blue', 'green', 'black', 'blue', 'red'), type3 = c('chair','chair','sofa','bed','plate'), num3 = c(4,7,12,3,2), stringsAsFactors = FALSE )
Basic Solution (Fixed Number of Groups)
If you know exactly how many (name/type/num) groups you have, you can use a simple function to check for duplicates per row:
# Function to detect duplicate groups in a single row find_duplicate_groups <- function(row) { # Combine each group's values into a single string for easy comparison groups <- list( group1 = paste(row["name1"], row["type1"], row["num1"], sep = "|"), group2 = paste(row["name2"], row["type2"], row["num2"], sep = "|"), group3 = paste(row["name3"], row["type3"], row["num3"], sep = "|") ) # Find all groups that appear more than once duplicates <- groups[duplicated(groups) | duplicated(groups, fromLast = TRUE)] # Return group names if duplicates exist, else NA if (length(duplicates) > 0) { return(names(duplicates)) } else { return(NA) } } # Apply the function to every row and add results as a new column df$duplicate_groups <- apply(df, 1, find_duplicate_groups) # View the output print(df)
When you run this, you'll get a new column duplicate_groups that lists which groups are duplicates in each row (e.g., row 1 will show group1, group2 since those two groups are identical).
Flexible Solution (Any Number of Groups)
If you might have more groups (like name4/type4/num4, etc.), use this dynamic version that automatically detects all existing groups:
# Generalized function for any number of (nameX, typeX, numX) groups find_duplicate_groups_general <- function(row) { # Extract all group numbers from column names (e.g., 1,2,3 from name1, name2, name3) group_nums <- unique(sub("name(\\d+)", "\\1", grep("name\\d+", names(row), value = TRUE))) # Dynamically create groups for each number groups <- lapply(group_nums, function(x) { paste(row[paste0("name", x)], row[paste0("type", x)], row[paste0("num", x)], sep = "|") }) names(groups) <- paste0("group", group_nums) # Find duplicate groups duplicates <- groups[duplicated(groups) | duplicated(groups, fromLast = TRUE)] return(if (length(duplicates) > 0) names(duplicates) else NA) } # Apply the generalized function df$duplicate_groups <- apply(df, 1, find_duplicate_groups_general)
This version works no matter how many (name/type/num) groups you have in your dataframe—it automatically scans for all columns matching the nameX pattern and builds the corresponding groups.
内容的提问来源于stack exchange,提问作者Adam_S

