R语言循环实现序列匹配:筛选相似度≥0.8的描述
Hey there! No need to apologize for asking basic questions—we all start somewhere 😊 Let's tackle your problem step by step.
First, let's clarify what we need to do: calculate the similarity between your new description (new_cr) and every description in your dataframe, then pull out all entries where the similarity score is ≥0.8 (adjust if your score range is 0-100 instead of 0-1) as a list.
Method 1: Base R (no explicit loops)
We can use sapply to iterate over each description in your dataframe, compute the similarity score, then filter based on your threshold:
library(fuzzywuzzyR) # Example value for new_cr (replace with your actual input) new_cr <- "grant system admin rights" # Calculate similarity scores for all descriptions similarity_scores <- sapply(df$description, function(desc) { matcher <- SequenceMatcher$new(string1 = desc, string2 = new_cr) matcher$ratio() }) # Extract descriptions with score ≥0.8 and convert to a list similar_descriptions <- as.list(df$description[similarity_scores >= 0.8])
Method 2: Using purrr (cleaner & efficient for larger datasets)
If you're open to using the purrr package (part of the tidyverse), it provides a more readable way to handle iteration, which is especially useful if your dataframe is large:
library(fuzzywuzzyR) library(purrr) new_cr <- "grant system admin rights" # Compute similarity scores with map_dbl (returns a numeric vector) similarity_scores <- map_dbl(df$description, ~{ SequenceMatcher$new(string1 = .x, string2 = new_cr)$ratio() }) # Filter and convert to list similar_descriptions <- df$description[similarity_scores >= 0.8] %>% as.list()
Important Note:
Double-check the range of the ratio() output! If SequenceMatcher$ratio() returns values between 0-100 (percentages) instead of 0-1, you'll need to adjust your threshold to 80 instead of 0.8.
Both methods avoid manual for loops and are optimized for R's vectorized operations, making them more efficient than writing a loop yourself.
内容的提问来源于stack exchange,提问作者user2906657

