在R语言中批量重命名列变量及基于字符匹配实现品牌分组的技术咨询
Hey there! Let's tackle your two R questions one by one—they’re both super common (and tricky!) data wrangling tasks, so I’ve got you covered.
First, let’s cover marking rows where a column’s value starts with your target text. The stringr package (part of the tidyverse) makes this straightforward with str_starts(). Here’s a quick example using a sample data frame:
# Load required packages library(dplyr) library(stringr) # Sample data df <- tibble(product = c("Apple iPhone", "Samsung Galaxy", "apple Watch", "Google Pixel", "Xiaomi Mi")) # Add a flag column for values starting with "Apple" (case-insensitive) df <- df %>% mutate(is_apple_brand = str_starts(product, fixed("Apple", ignore_case = TRUE)))
If you want to rename those matching values instead of just flagging them, use if_else() with the same str_starts() check:
# Rename all values starting with "Apple" (case-insensitive) to "Apple Product" df <- df %>% mutate(product = if_else( str_starts(product, fixed("Apple", ignore_case = TRUE)), "Apple Product", product ))
Pro tip: Use fixed() with ignore_case = TRUE to avoid case sensitivity issues (like missing "apple Watch" if you only check for "Apple").
Spelling and whitespace issues are the worst for grouping—luckily, we can build a function that standardizes text first, then matches brands using flexible rules. Here’s a step-by-step solution:
Step 1: Build a Custom Matching Function
This function will first clean the product names (fix spacing, lowercase, remove junk characters) then use pattern matching to assign brands. We’ll prioritize exact matches first, then broader keyword matches, and add a fallback for unknowns.
library(dplyr) library(stringr) match_brand <- function(product_name) { # Step 1: Standardize the text to reduce matching errors cleaned_name <- product_name %>% str_to_lower() %>% # Convert to lowercase str_squish() %>% # Remove extra spaces (e.g., " Samsung " → "samsung") str_remove_all("[^a-z0-9 ]") # Remove special characters (e.g., "Samsung!" → "samsung") # Step 2: Match brands using flexible rules (order matters—put specific rules first!) case_when( # Match values starting with the brand name str_detect(cleaned_name, "^apple") ~ "Apple", str_detect(cleaned_name, "^samsung") ~ "Samsung", str_detect(cleaned_name, "^google") ~ "Google", # Match values containing brand-specific keywords (for misspellings or partial names) str_detect(cleaned_name, "iphone|ipad") ~ "Apple", str_detect(cleaned_name, "galaxy|note") ~ "Samsung", str_detect(cleaned_name, "pixel|nexus") ~ "Google", # Fallback for unmatchable entries TRUE ~ "Unknown" ) }
Step 2: Use the Function on Your Data
Apply it to your product name column with mutate():
# Sample messy data (with spacing/spelling errors) messy_df <- tibble(product_name = c(" AppleIphone ", "samsunggalaxy", "Google pixel", "IPad Pro", "random phone")) # Add the brand column messy_df <- messy_df %>% mutate(brand = match_brand(product_name))
Bonus: Fuzzy Matching for Severe Spelling Errors
If you have really bad typos (like "Samsang" instead of "Samsung"), use fuzzy matching with the stringdist package to allow small edit distances:
library(stringdist) fuzzy_match_brand <- function(product_name, max_edit_distance = 2) { cleaned_name <- str_to_lower(str_squish(product_name)) brand_list <- c("Apple", "Samsung", "Google") # Calculate edit distance between cleaned name and each brand distances <- stringdist(cleaned_name, str_to_lower(brand_list), method = "lv") # Return the closest brand if it's within the threshold if (min(distances) <= max_edit_distance) { brand_list[which.min(distances)] } else { "Unknown" } } # Test it with a typo fuzzy_match_brand("Samsang Galaxy") # Returns "Samsung"
内容的提问来源于stack exchange,提问作者Urb_Ink

