You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用R语言统计相似字符串出现次数的优化方案求助

Got it, let's figure out an efficient way to solve this problem. The core trick here is to turn each string into a standardized key—so any two strings with exactly the same words (no matter the order) will share the same key. Then we just group by this key and count how many times each group shows up. This approach is way faster than the n-gram/join methods you tried, perfect for large datasets.

Python (Pandas) Solution

This works seamlessly with pandas and scales well even for millions of rows:

  1. First, create a function to generate the standardized key: split the string into words, sort them, and join them back into a fixed string.
import pandas as pd

def standardize_string(s):
    # Optional: add .lower() if you want case-insensitive matching (e.g., "Dog" = "dog")
    # Optional: use re.sub(r'[^\w\s]', '', s) first to strip punctuation
    words = s.split()
    sorted_words = sorted(words)
    return ' '.join(sorted_words)
  1. Apply this function to your text column, then group and count:
# Assume your dataframe is named `df` and the text column is `sentences`
df['standardized_key'] = df['sentences'].apply(standardize_string)

# Get the count per group, plus example strings for clarity
result = df.groupby('standardized_key').agg(
    count=('sentences', 'count'),
    example_strings=('sentences', lambda x: ', '.join(x.unique()))
).reset_index()

R Solution

If you're working with R and dplyr, the same logic applies:

  1. Define the standardization function:
library(dplyr)
library(stringr)

standardize_string <- function(s) {
  # Optional: add str_to_lower() for case-insensitive matching
  # Optional: add str_remove_all(s, "[^\\w\\s]") to strip punctuation
  words <- str_split(s, " ")[[1]]
  sorted_words <- sort(words)
  paste(sorted_words, collapse = " ")
}
  1. Process your dataframe:
# Assume your dataframe is `df` with text column `sentences`
df <- df %>%
  mutate(standardized_key = sapply(sentences, standardize_string)) %>%
  group_by(standardized_key) %>%
  summarize(
    count = n(),
    example_strings = str_c(unique(sentences), collapse = ", ")
  ) %>%
  ungroup()

Why This Works (and Is Efficient)

  • This method avoids expensive pairwise comparisons or join operations (which are O(n²) and will crash with large data). Instead, each string is processed in O(k log k) time (where k is the number of words in the string)—super fast even for huge datasets.
  • It correctly handles cases where words are repeated: e.g., "dog dog is" and "is dog dog" will get the same key, while "dog is good" and "dog are good" get different keys (since their word sets don't match).

内容的提问来源于stack exchange,提问作者ayush varshney

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 07:42:52