You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言文本相似度计算:排除相同条目并优化性能

Hey Adam! Let's work through your two requirements to get a faster, more accurate name similarity comparison.

First: Excluding Same-Position Matches

Your original code compares every name in names1 to every name in names2, including the exact same row (e.g., row 1 of names1 vs row 1 of names2). We need to exclude those diagonal matches to meet your first rule.

Second: Speeding Up Large Datasets

The sapply loop with RecordLinkage::levenshteinSim is slow for big data because it processes each pair one at a time. Instead, we'll use vectorized functions from the stringdist package (which you already have loaded) to compute all pairwise similarities in one go—this is way more efficient, especially with thousands of names.


Full Optimized Code

# Load required packages
library(stringdist)
library(dplyr)
library(scales)

# Your original data
id1 <- 1:8
names1 <- c("Prabhudev Ramanujam","Deepak Subramaniam","Sangamer Mahapatra","SriramKishore Sharma", "Deepak Subramaniam","SriramKishore Sharma","Deepak Subramaniam","Sangamer Mahapatra")
id2 <- c(1,2,3,4,11,13,9,10)
names2 <- c("Prabhudev Ramanujam","Deepak Subramaniam","Sangamer Mahapatra","SriramKishore Sharma", "Deepak Subramaniam","Sangamer Mahapatra","SriramKishore Sharma","Deepak Subramaniam")
Name_Data <- data.frame(id1, names1, id2, names2, stringsAsFactors = FALSE)

# Step 1: Compute all pairwise Levenshtein distances (vectorized, fast)
dist_matrix <- stringdistmatrix(names1, names2, method = "lv")

# Step 2: Convert distances to similarity scores (1 - distance / max length of the two strings)
sim_matrix <- 1 - dist_matrix / outer(nchar(names1), nchar(names2), FUN = max)

# Step 3: Exclude same-position matches (set diagonal values to NA)
diag(sim_matrix) <- NA

# Step 4: Reshape into a clean data frame (all valid pairs + percentages)
final_results <- as.data.frame.table(sim_matrix, stringsAsFactors = FALSE) %>%
  rename(names1_row = Var1, names2_row = Var2, similarity = Freq) %>%
  # Map row indices back to original IDs and names
  left_join(Name_Data %>% mutate(names1_row = as.character(row_number())), by = "names1_row") %>%
  left_join(Name_Data %>% mutate(names2_row = as.character(row_number())) %>% select(names2_row, id2, names2), by = "names2_row") %>%
  # Remove the excluded same-position pairs
  filter(!is.na(similarity)) %>%
  # Convert similarity to percentage format
  mutate(similarity_percent = percent(similarity)) %>%
  # Reorder columns for readability
  select(id1, names1, id2, names2, similarity, similarity_percent)

# Optional: If you want a summary per names1 entry (e.g., average similarity)
summary_per_name <- final_results %>%
  group_by(id1, names1) %>%
  summarize(avg_similarity = mean(similarity),
            avg_similarity_percent = percent(avg_similarity),
            .groups = "drop")

What Makes This Better?

  • Blazing Fast: stringdistmatrix uses optimized C code to compute all pairs at once, which is drastically faster than looping with sapply for large datasets (think 10x+ speedup with 10k+ names).
  • Exact Exclusion: We explicitly zero out the diagonal of the similarity matrix to exclude same-row matches, so you never compare row i of names1 to row i of names2.
  • Lean Dependencies: We ditch RecordLinkage entirely since stringdist can handle the distance/similarity calculation natively—no extra packages needed beyond what you already loaded.

内容的提问来源于stack exchange,提问作者Adam Shaw

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 07:28:02