You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在数据框每行的两列字符串中提取共同词汇?

Hey there! Let's figure out how to extract common words between two string columns row-wise in your R data frame—especially since you're dealing with a large dataset and need an efficient solution. You're right that intersect() alone won't work here because it operates on vectors, not the individual string rows we need to process. Let's break down the best approaches:

First, let's set up your sample data

# Your sample input data frame
df <- data.frame(
  C1 = c("Roy goes to Japan", "I go to Japan"),
  C2 = c("Roy goes to Australia", "He goes to Japan"),
  stringsAsFactors = FALSE
)

Efficient Solution 1: Tidyverse (dplyr + purrr + stringr)

If you prefer the tidyverse ecosystem, this approach uses str_split() to break each string into word lists, then map2_chr() to run intersect() row-wise and collapse the results back into a single string:

library(dplyr)
library(purrr)
library(stringr)

df <- df %>%
  mutate(
    # Split each string into a list of words
    C1_words = str_split(C1, "\\s+"),
    C2_words = str_split(C2, "\\s+"),
    # Find common words per row and collapse to a space-separated string
    Result = map2_chr(C1_words, C2_words, ~paste(intersect(.x, .y), collapse = " "))
  ) %>%
  # Remove intermediate columns if you don't need them
  select(-C1_words, -C2_words)

Efficient Solution 2: data.table + stringi (Best for Large Datasets)

For huge datasets, data.table is far more memory-efficient and faster than base R loops or even some tidyverse operations. Pair it with stringi (which has faster C++-backed string operations than stringr) for optimal performance:

library(data.table)
library(stringi)

# Convert to data.table for fast operations
setDT(df)

# Split strings, find common words, then clean up intermediate columns
df[, `:=`(
  C1_words = stri_split_regex(C1, "\\s+"),
  C2_words = stri_split_regex(C2, "\\s+")
)][, Result := mapply(function(x, y) paste(intersect(x, y), collapse = " "), C1_words, C2_words)][, c("C1_words", "C2_words") := NULL]

Optional: Handle Edge Cases

  • Extra spaces: If your strings have leading/trailing spaces or multiple spaces between words, add str_trim() or stri_trim() before splitting:
    str_split(str_trim(C1), "\\s+")
    
  • Case insensitivity: If you want to match words regardless of case (e.g., "Roy" and "roy"), convert strings to lowercase first, then map back to original case if needed:
    df <- df %>%
      mutate(
        C1_lower = str_to_lower(str_trim(C1)),
        C2_lower = str_to_lower(str_trim(C2)),
        C1_words = str_split(C1_lower, "\\s+"),
        C2_words = str_split(C2_lower, "\\s+"),
        Result_lower = map2_chr(C1_words, C2_words, ~paste(intersect(.x, .y), collapse = " ")),
        # Preserve original case in the result
        Result = map2_chr(C1, Result_lower, ~{
          original_words <- str_split(.x, "\\s+")[[1]]
          common_lower <- str_split(.y, "\\s+")[[1]]
          paste(original_words[tolower(original_words) %in% common_lower], collapse = " ")
        })
      ) %>%
      select(C1, C2, Result)
    

Key Notes for Performance

  • Avoid base R for loops at all costs—they're extremely slow for large datasets.
  • stringi functions are generally faster than their stringr counterparts because they're implemented in C++.
  • data.table minimizes memory copies, which is critical when working with big data.

内容的提问来源于stack exchange,提问作者Siddd

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 12:47:45