如何用stringr正则表达式从零件描述提取重复词生成产品类型
Let’s walk through this step by step—you’ve got 50k+ parts with unstructured descriptions, no predefined product types, and need to group them using repeating keywords. We’ll use text processing with stringr and tidyverse tools to make this work.
Step 1: Preprocess Messy Description Text
First, we need to normalize inconsistent formatting (like camelCase terms, mixed case) so we can extract meaningful keywords. For example, "WaterBotHold" gets split into readable words to align with other mentions of "water".
library(tidyverse) library(stringr) # Your sample dataset df <- tribble( ~PartNo, ~Description, ~ProductType, "A000443", "Water Bottle", "", "A000445", "Contain Water", "", "A000448", "WaterBotHold", "" ) # Clean descriptions: split camelCase, lowercase, extract words df_processed <- df %>% mutate( # Add space before uppercase letters to split camelCase clean_desc = str_replace_all(Description, "(?<=[a-z])(?=[A-Z])", " "), # Convert to lowercase and pull out all word characters keywords = str_extract_all(tolower(clean_desc), "\\w+") )
Step 2: Identify High-Frequency Keywords
Next, we’ll count keyword occurrences to spot terms that define product groups (like "water" in your sample). For 50k parts, this will highlight the most common themes across descriptions.
# Flatten keywords to count their frequency keyword_counts <- df_processed %>% unnest(keywords) %>% count(keywords, sort = TRUE) # Filter for meaningful keywords (adjust threshold for your full dataset) # For 50k parts, use a higher threshold like n >= 10 to filter noise top_keywords <- keyword_counts %>% filter(n >= 1) # Sample output: # keywords | n # ---------|--- # water | 3 # bottle | 1 # contain | 1 # bot | 1 # hold | 1
Step 3: Map Keywords to Consistent Product Types
Create a lookup table to group related keywords into standardized product types. For your sample, all water-related terms get mapped to "Water Container".
# Build a keyword-to-product-type mapping (expand this for your full dataset) product_type_lookup <- tibble( keyword = c("water", "bottle", "contain", "bot", "hold"), ProductType = rep("Water Container", 5) ) # Match keywords back to the original data and fill the ProductType column df_final <- df_processed %>% unnest(keywords) %>% left_join(product_type_lookup, by = "keywords") %>% group_by(PartNo) %>% summarise( Description = first(Description), # Pick the first matching product type (adjust if handling conflicts) ProductType = first(ProductType) )
Step 4: Scaling to 50k+ Parts
For large datasets, optimize with these tips:
- Fuzzy matching: Use the
stringdistpackage to handle misspelled keywords (e.g., "water" vs "watter"). - Auto-clustering: Use word embeddings or k-means clustering to auto-group similar descriptions if manual lookup is too time-consuming.
- Vectorized operations: Stick to
dplyr’s vectorized functions instead of loops to handle 50k rows efficiently.
Final Sample Output
Running the code on your sample gives this cleaned, grouped data:
| PartNo | Description | ProductType |
|---|---|---|
| A000443 | Water Bottle | Water Container |
| A000445 | Contain Water | Water Container |
| A000448 | WaterBotHold | Water Container |
内容的提问来源于stack exchange,提问作者stackinator

