基于stringr正则表达式的未知产品类型零件编号分组咨询
Hey there! As someone who’s tackled messy inventory data before, I totally get how frustrating it is to group parts when descriptions are unstandardized. Let’s break this down step by step, starting with your sample data and building up to a solution that works for your 50k+ rows.
1. First: Spot Core Keywords in Descriptions
The key here is to identify recurring terms that define each product group. In your sample, we can see two clear clusters: water-related items and snack/protein bars. Since descriptions are inconsistent (like "WaterBotHold" or "Hershey_Bar"), we’ll use regex to match variations of these core terms.
2. Build a Regex-Based Matching System
We’ll create a mapping of product types to their associated keywords, then write a simple function to check each description against these keywords. This leverages the stringr tools you’re learning from R for Data Science and is beginner-friendly.
Here’s the code tailored to your sample data:
library(tidyverse) library(stringr) # Your original data frame df <- tribble( ~PartNo, ~Description, ~ProductType, "A000443", "Water Bottle", "", "A000445", "Contain Water", "", "A000448", "WaterBotHold", "", "HRZ55", "Hershey_Bar", "", "RRB55", "Candy Energy", "", "QMU55", "Bar Protein", "" ) # Define keyword groups (adjust these based on your full dataset!) product_keywords <- list( "Water Container" = c("water", "bottl", "contain"), # Match water/bottle variations "Snack & Protein Bar" = c("bar", "candy", "protein", "hershey") ) # Function to assign product type based on keyword matches assign_product_type <- function(desc) { # Convert to lowercase to avoid case sensitivity issues desc_lower <- str_to_lower(desc) # Loop through our keyword groups to find a match for (type_name in names(product_keywords)) { if (any(str_detect(desc_lower, product_keywords[[type_name]]))) { return(type_name) } } # If no matches found, label as unclassified for later review return("Unclassified") } # Apply the function to fill the ProductType column df <- df %>% mutate(ProductType = map_chr(Description, assign_product_type)) # View the final result df
3. Handle Edge Cases & Scale to 50k Rows
- Case Insensitivity: Converting descriptions to lowercase ensures "Water", "water", and "WATER" all match the same keywords.
- Unclassified Items: The function labels non-matching entries as "Unclassified"—you can review these later to add more keywords to your groups.
- Efficiency: This approach is fast even for large datasets because
str_detectis vectorized, andmap_chrefficiently applies the function across the column.
4. Advanced: Use Text Analysis to Find Hidden Keywords
If you’re unsure what keywords to use for your full dataset, the tidytext package can help you split descriptions into individual words and find the most frequent terms. This reveals patterns you might miss manually:
library(tidytext) # Split descriptions into words and count frequencies desc_word_counts <- df %>% unnest_tokens(word, Description) %>% count(word, sort = TRUE) # View the most common words to inform your keyword groups desc_word_counts
内容的提问来源于stack exchange,提问作者stackinator

