You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言:提取两列字符串匹配与不匹配部分的实现需求

Efficiently Extract Matching and Unique String Parts in R

Hey there! I get it—manually splitting cells over and over is a total drag. Let's use a clean, scalable approach that handles all the string manipulation and set operations in one go, using the stringr package (which you're likely already familiar with from using str_detect).

Step-by-Step Solution

First, load the required package (install it first if you haven't with install.packages("stringr")):

library(stringr)

Next, let's define a reusable function to handle the string comparison logic. This will work for single rows or entire data frames:

# Your original data
x <- c("apple, banana, pine nuts, almond")
y <- c("orange, apple, almond, grapes, carrots")
j <- data.frame(x, y, stringsAsFactors = FALSE)

# Helper function to compare two comma-separated strings
compare_string_sets <- function(x_str, y_str) {
  # Split strings into individual items, trimming extra spaces
  x_items <- str_trim(str_split(x_str, ",")[[1]])
  y_items <- str_trim(str_split(y_str, ",")[[1]])
  
  # Calculate intersections and unique elements
  shared_items <- intersect(x_items, y_items)
  x_unique <- setdiff(x_items, shared_items)
  y_unique <- setdiff(y_items, shared_items)
  
  # Collapse back into comma-separated strings
  list(
    both = str_c(shared_items, collapse = ", "),
    xonly = str_c(x_unique, collapse = ", "),
    yonly = str_c(y_unique, collapse = ", ")
  )
}

Now apply this function to your data frame and assemble the final result:

# Apply the function to each row of your data frame
comparison_results <- t(apply(j, 1, function(row) compare_string_sets(row["x"], row["y"])))

# Combine original data with results to create data frame k
k <- cbind(j, as.data.frame(comparison_results, stringsAsFactors = FALSE))

What This Does

  • str_trim ensures we don't have mismatches from extra spaces (like " apple" vs "apple").
  • intersect and setdiff handle the set operations cleanly—no manual checking needed.
  • The function is scalable: if your data frame j has multiple rows, this will process every row automatically.

Final Output

When you run print(k), you'll get exactly what you wanted:

> k
                                      x                                          y            both               yonly             xonly
1 apple, banana, pine nuts, almond orange, apple, almond, grapes, carrots apple, almond orange, grapes, carrots banana, pine nuts

内容的提问来源于stack exchange,提问作者Maya

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 09:25:33