使用R语言stringdist匹配店铺名称:如何识别ADIDAS STORE与Adidas为同一店铺?
Got it, let's tackle this matching problem—your goal of linking "ADIDAS STORE" to "Adidas" makes total sense, and the issue right now is that raw string distance calculations are getting thrown off by case differences, extra words like "STORE", and string length mismatches. Here are two solid approaches to fix this:
1. Preprocess Names First (Most Reliable)
The best way to get accurate matches is to standardize your shop names first—strip out common suffixes, normalize case, and clean up whitespace. This removes the noise that's messing with the distance calculation.
Here's a reusable preprocessing function and how to use it:
library(stringdist) # Create a preprocessing function to standardize shop names preprocess_shop <- function(name) { # Convert all text to lowercase name <- tolower(name) # Remove common suffixes like " store" (adjust this if you have other common terms) name <- gsub("\\s+store$", "", name) # Trim any extra leading/trailing whitespace name <- trimws(name) return(name) } # Your original data a <- c("Adidas", "macys","apple store", "VANS Store New York", "new wave","adsasds") target <- "ADIDAS STORE" # Preprocess all names a_clean <- sapply(a, preprocess_shop) target_clean <- preprocess_shop(target) # Calculate Jaro-Winkler distance on cleaned names distances <- stringdist(target_clean, a_clean, method = "jw") # Create and sort results results <- data.frame( original_name = a, cleaned_name = a_clean, jw_distance = distances ) %>% arrange(jw_distance) print(results)
What this does:
- The cleaned target becomes
"adidas"and the cleaned "Adidas" is also"adidas", so their Jaro-Winkler distance drops to 0—a perfect match at the top of your results. - You can expand the
gsubline to remove other common terms like "shop", "boutique", etc., if needed for your dataset.
2. Use Partial Matching (Quick Fix)
If you don't want to preprocess, you can use the partial = TRUE argument in stringdist to allow partial matches. This tells the function to find the best overlap between the shorter string and parts of the longer string, which works great when one name has extra words like "STORE".
library(stringdist) a <- c("Adidas", "macys","apple store", "VANS Store New York", "new wave","adsasds") target <- "ADIDAS STORE" # Calculate partial Jaro-Winkler distance (case-normalized first) distances <- stringdist(tolower(target), tolower(a), method = "jw", partial = TRUE) results <- data.frame( original_name = a, jw_distance = distances ) %>% arrange(jw_distance) print(results)
What this does:
- By normalizing case and using partial matching, the function prioritizes the overlap between "adidas" (from the target) and "adidas" (from your list), pushing that match to the top with a much lower distance.
Both methods will fix your issue—preprocessing is better for long-term use with messy data, while partial matching is a quick win for one-off tasks.
内容的提问来源于stack exchange,提问作者 DarkSide

