R语言文本相似度计算:排除相同条目并优化性能
Hey Adam! Let's work through your two requirements to get a faster, more accurate name similarity comparison.
First: Excluding Same-Position Matches
Your original code compares every name in names1 to every name in names2, including the exact same row (e.g., row 1 of names1 vs row 1 of names2). We need to exclude those diagonal matches to meet your first rule.
Second: Speeding Up Large Datasets
The sapply loop with RecordLinkage::levenshteinSim is slow for big data because it processes each pair one at a time. Instead, we'll use vectorized functions from the stringdist package (which you already have loaded) to compute all pairwise similarities in one go—this is way more efficient, especially with thousands of names.
Full Optimized Code
# Load required packages library(stringdist) library(dplyr) library(scales) # Your original data id1 <- 1:8 names1 <- c("Prabhudev Ramanujam","Deepak Subramaniam","Sangamer Mahapatra","SriramKishore Sharma", "Deepak Subramaniam","SriramKishore Sharma","Deepak Subramaniam","Sangamer Mahapatra") id2 <- c(1,2,3,4,11,13,9,10) names2 <- c("Prabhudev Ramanujam","Deepak Subramaniam","Sangamer Mahapatra","SriramKishore Sharma", "Deepak Subramaniam","Sangamer Mahapatra","SriramKishore Sharma","Deepak Subramaniam") Name_Data <- data.frame(id1, names1, id2, names2, stringsAsFactors = FALSE) # Step 1: Compute all pairwise Levenshtein distances (vectorized, fast) dist_matrix <- stringdistmatrix(names1, names2, method = "lv") # Step 2: Convert distances to similarity scores (1 - distance / max length of the two strings) sim_matrix <- 1 - dist_matrix / outer(nchar(names1), nchar(names2), FUN = max) # Step 3: Exclude same-position matches (set diagonal values to NA) diag(sim_matrix) <- NA # Step 4: Reshape into a clean data frame (all valid pairs + percentages) final_results <- as.data.frame.table(sim_matrix, stringsAsFactors = FALSE) %>% rename(names1_row = Var1, names2_row = Var2, similarity = Freq) %>% # Map row indices back to original IDs and names left_join(Name_Data %>% mutate(names1_row = as.character(row_number())), by = "names1_row") %>% left_join(Name_Data %>% mutate(names2_row = as.character(row_number())) %>% select(names2_row, id2, names2), by = "names2_row") %>% # Remove the excluded same-position pairs filter(!is.na(similarity)) %>% # Convert similarity to percentage format mutate(similarity_percent = percent(similarity)) %>% # Reorder columns for readability select(id1, names1, id2, names2, similarity, similarity_percent) # Optional: If you want a summary per names1 entry (e.g., average similarity) summary_per_name <- final_results %>% group_by(id1, names1) %>% summarize(avg_similarity = mean(similarity), avg_similarity_percent = percent(avg_similarity), .groups = "drop")
What Makes This Better?
- Blazing Fast:
stringdistmatrixuses optimized C code to compute all pairs at once, which is drastically faster than looping withsapplyfor large datasets (think 10x+ speedup with 10k+ names). - Exact Exclusion: We explicitly zero out the diagonal of the similarity matrix to exclude same-row matches, so you never compare row i of
names1to row i ofnames2. - Lean Dependencies: We ditch
RecordLinkageentirely sincestringdistcan handle the distance/similarity calculation natively—no extra packages needed beyond what you already loaded.
内容的提问来源于stack exchange,提问作者Adam Shaw
相关产品推荐
相关产品推荐

