使用R语言统计相似字符串出现次数的优化方案求助
Got it, let's figure out an efficient way to solve this problem. The core trick here is to turn each string into a standardized key—so any two strings with exactly the same words (no matter the order) will share the same key. Then we just group by this key and count how many times each group shows up. This approach is way faster than the n-gram/join methods you tried, perfect for large datasets.
Python (Pandas) Solution
This works seamlessly with pandas and scales well even for millions of rows:
- First, create a function to generate the standardized key: split the string into words, sort them, and join them back into a fixed string.
import pandas as pd def standardize_string(s): # Optional: add .lower() if you want case-insensitive matching (e.g., "Dog" = "dog") # Optional: use re.sub(r'[^\w\s]', '', s) first to strip punctuation words = s.split() sorted_words = sorted(words) return ' '.join(sorted_words)
- Apply this function to your text column, then group and count:
# Assume your dataframe is named `df` and the text column is `sentences` df['standardized_key'] = df['sentences'].apply(standardize_string) # Get the count per group, plus example strings for clarity result = df.groupby('standardized_key').agg( count=('sentences', 'count'), example_strings=('sentences', lambda x: ', '.join(x.unique())) ).reset_index()
R Solution
If you're working with R and dplyr, the same logic applies:
- Define the standardization function:
library(dplyr) library(stringr) standardize_string <- function(s) { # Optional: add str_to_lower() for case-insensitive matching # Optional: add str_remove_all(s, "[^\\w\\s]") to strip punctuation words <- str_split(s, " ")[[1]] sorted_words <- sort(words) paste(sorted_words, collapse = " ") }
- Process your dataframe:
# Assume your dataframe is `df` with text column `sentences` df <- df %>% mutate(standardized_key = sapply(sentences, standardize_string)) %>% group_by(standardized_key) %>% summarize( count = n(), example_strings = str_c(unique(sentences), collapse = ", ") ) %>% ungroup()
Why This Works (and Is Efficient)
- This method avoids expensive pairwise comparisons or join operations (which are O(n²) and will crash with large data). Instead, each string is processed in O(k log k) time (where k is the number of words in the string)—super fast even for huge datasets.
- It correctly handles cases where words are repeated: e.g., "dog dog is" and "is dog dog" will get the same key, while "dog is good" and "dog are good" get different keys (since their word sets don't match).
内容的提问来源于stack exchange,提问作者ayush varshney
相关产品推荐
相关产品推荐

