You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R语言中批量重命名列变量及基于字符匹配实现品牌分组的技术咨询

Hey there! Let's tackle your two R questions one by one—they’re both super common (and tricky!) data wrangling tasks, so I’ve got you covered.

1. Marking or Renaming Values in a Column That Start With a Specific Character/Word

First, let’s cover marking rows where a column’s value starts with your target text. The stringr package (part of the tidyverse) makes this straightforward with str_starts(). Here’s a quick example using a sample data frame:

# Load required packages
library(dplyr)
library(stringr)

# Sample data
df <- tibble(product = c("Apple iPhone", "Samsung Galaxy", "apple Watch", "Google Pixel", "Xiaomi Mi"))

# Add a flag column for values starting with "Apple" (case-insensitive)
df <- df %>%
  mutate(is_apple_brand = str_starts(product, fixed("Apple", ignore_case = TRUE)))

If you want to rename those matching values instead of just flagging them, use if_else() with the same str_starts() check:

# Rename all values starting with "Apple" (case-insensitive) to "Apple Product"
df <- df %>%
  mutate(product = if_else(
    str_starts(product, fixed("Apple", ignore_case = TRUE)),
    "Apple Product",
    product
  ))

Pro tip: Use fixed() with ignore_case = TRUE to avoid case sensitivity issues (like missing "apple Watch" if you only check for "Apple").

2. Creating a Brand-Matching Function to Handle Spelling/Spacing Errors

Spelling and whitespace issues are the worst for grouping—luckily, we can build a function that standardizes text first, then matches brands using flexible rules. Here’s a step-by-step solution:

Step 1: Build a Custom Matching Function

This function will first clean the product names (fix spacing, lowercase, remove junk characters) then use pattern matching to assign brands. We’ll prioritize exact matches first, then broader keyword matches, and add a fallback for unknowns.

library(dplyr)
library(stringr)

match_brand <- function(product_name) {
  # Step 1: Standardize the text to reduce matching errors
  cleaned_name <- product_name %>%
    str_to_lower() %>%          # Convert to lowercase
    str_squish() %>%            # Remove extra spaces (e.g., "  Samsung  " → "samsung")
    str_remove_all("[^a-z0-9 ]") # Remove special characters (e.g., "Samsung!" → "samsung")
  
  # Step 2: Match brands using flexible rules (order matters—put specific rules first!)
  case_when(
    # Match values starting with the brand name
    str_detect(cleaned_name, "^apple") ~ "Apple",
    str_detect(cleaned_name, "^samsung") ~ "Samsung",
    str_detect(cleaned_name, "^google") ~ "Google",
    
    # Match values containing brand-specific keywords (for misspellings or partial names)
    str_detect(cleaned_name, "iphone|ipad") ~ "Apple",
    str_detect(cleaned_name, "galaxy|note") ~ "Samsung",
    str_detect(cleaned_name, "pixel|nexus") ~ "Google",
    
    # Fallback for unmatchable entries
    TRUE ~ "Unknown"
  )
}

Step 2: Use the Function on Your Data

Apply it to your product name column with mutate():

# Sample messy data (with spacing/spelling errors)
messy_df <- tibble(product_name = c("  AppleIphone  ", "samsunggalaxy", "Google pixel", "IPad Pro", "random phone"))

# Add the brand column
messy_df <- messy_df %>%
  mutate(brand = match_brand(product_name))

Bonus: Fuzzy Matching for Severe Spelling Errors

If you have really bad typos (like "Samsang" instead of "Samsung"), use fuzzy matching with the stringdist package to allow small edit distances:

library(stringdist)

fuzzy_match_brand <- function(product_name, max_edit_distance = 2) {
  cleaned_name <- str_to_lower(str_squish(product_name))
  brand_list <- c("Apple", "Samsung", "Google")
  
  # Calculate edit distance between cleaned name and each brand
  distances <- stringdist(cleaned_name, str_to_lower(brand_list), method = "lv")
  
  # Return the closest brand if it's within the threshold
  if (min(distances) <= max_edit_distance) {
    brand_list[which.min(distances)]
  } else {
    "Unknown"
  }
}

# Test it with a typo
fuzzy_match_brand("Samsang Galaxy") # Returns "Samsung"

内容的提问来源于stack exchange,提问作者Urb_Ink

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 04:07:44