如何在数据框每行的两列字符串中提取共同词汇?
Hey there! Let's figure out how to extract common words between two string columns row-wise in your R data frame—especially since you're dealing with a large dataset and need an efficient solution. You're right that intersect() alone won't work here because it operates on vectors, not the individual string rows we need to process. Let's break down the best approaches:
First, let's set up your sample data
# Your sample input data frame df <- data.frame( C1 = c("Roy goes to Japan", "I go to Japan"), C2 = c("Roy goes to Australia", "He goes to Japan"), stringsAsFactors = FALSE )
Efficient Solution 1: Tidyverse (dplyr + purrr + stringr)
If you prefer the tidyverse ecosystem, this approach uses str_split() to break each string into word lists, then map2_chr() to run intersect() row-wise and collapse the results back into a single string:
library(dplyr) library(purrr) library(stringr) df <- df %>% mutate( # Split each string into a list of words C1_words = str_split(C1, "\\s+"), C2_words = str_split(C2, "\\s+"), # Find common words per row and collapse to a space-separated string Result = map2_chr(C1_words, C2_words, ~paste(intersect(.x, .y), collapse = " ")) ) %>% # Remove intermediate columns if you don't need them select(-C1_words, -C2_words)
Efficient Solution 2: data.table + stringi (Best for Large Datasets)
For huge datasets, data.table is far more memory-efficient and faster than base R loops or even some tidyverse operations. Pair it with stringi (which has faster C++-backed string operations than stringr) for optimal performance:
library(data.table) library(stringi) # Convert to data.table for fast operations setDT(df) # Split strings, find common words, then clean up intermediate columns df[, `:=`( C1_words = stri_split_regex(C1, "\\s+"), C2_words = stri_split_regex(C2, "\\s+") )][, Result := mapply(function(x, y) paste(intersect(x, y), collapse = " "), C1_words, C2_words)][, c("C1_words", "C2_words") := NULL]
Optional: Handle Edge Cases
- Extra spaces: If your strings have leading/trailing spaces or multiple spaces between words, add
str_trim()orstri_trim()before splitting:str_split(str_trim(C1), "\\s+") - Case insensitivity: If you want to match words regardless of case (e.g., "Roy" and "roy"), convert strings to lowercase first, then map back to original case if needed:
df <- df %>% mutate( C1_lower = str_to_lower(str_trim(C1)), C2_lower = str_to_lower(str_trim(C2)), C1_words = str_split(C1_lower, "\\s+"), C2_words = str_split(C2_lower, "\\s+"), Result_lower = map2_chr(C1_words, C2_words, ~paste(intersect(.x, .y), collapse = " ")), # Preserve original case in the result Result = map2_chr(C1, Result_lower, ~{ original_words <- str_split(.x, "\\s+")[[1]] common_lower <- str_split(.y, "\\s+")[[1]] paste(original_words[tolower(original_words) %in% common_lower], collapse = " ") }) ) %>% select(C1, C2, Result)
Key Notes for Performance
- Avoid base R
forloops at all costs—they're extremely slow for large datasets. stringifunctions are generally faster than theirstringrcounterparts because they're implemented in C++.data.tableminimizes memory copies, which is critical when working with big data.
内容的提问来源于stack exchange,提问作者Siddd

