R语言:提取两列字符串匹配与不匹配部分的实现需求
Efficiently Extract Matching and Unique String Parts in R
Hey there! I get it—manually splitting cells over and over is a total drag. Let's use a clean, scalable approach that handles all the string manipulation and set operations in one go, using the stringr package (which you're likely already familiar with from using str_detect).
Step-by-Step Solution
First, load the required package (install it first if you haven't with install.packages("stringr")):
library(stringr)
Next, let's define a reusable function to handle the string comparison logic. This will work for single rows or entire data frames:
# Your original data x <- c("apple, banana, pine nuts, almond") y <- c("orange, apple, almond, grapes, carrots") j <- data.frame(x, y, stringsAsFactors = FALSE) # Helper function to compare two comma-separated strings compare_string_sets <- function(x_str, y_str) { # Split strings into individual items, trimming extra spaces x_items <- str_trim(str_split(x_str, ",")[[1]]) y_items <- str_trim(str_split(y_str, ",")[[1]]) # Calculate intersections and unique elements shared_items <- intersect(x_items, y_items) x_unique <- setdiff(x_items, shared_items) y_unique <- setdiff(y_items, shared_items) # Collapse back into comma-separated strings list( both = str_c(shared_items, collapse = ", "), xonly = str_c(x_unique, collapse = ", "), yonly = str_c(y_unique, collapse = ", ") ) }
Now apply this function to your data frame and assemble the final result:
# Apply the function to each row of your data frame comparison_results <- t(apply(j, 1, function(row) compare_string_sets(row["x"], row["y"]))) # Combine original data with results to create data frame k k <- cbind(j, as.data.frame(comparison_results, stringsAsFactors = FALSE))
What This Does
str_trimensures we don't have mismatches from extra spaces (like" apple"vs"apple").intersectandsetdiffhandle the set operations cleanly—no manual checking needed.- The function is scalable: if your data frame
jhas multiple rows, this will process every row automatically.
Final Output
When you run print(k), you'll get exactly what you wanted:
> k x y both yonly xonly 1 apple, banana, pine nuts, almond orange, apple, almond, grapes, carrots apple, almond orange, grapes, carrots banana, pine nuts
内容的提问来源于stack exchange,提问作者Maya
相关产品推荐
相关产品推荐

