You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于stringr正则表达式的未知产品类型零件编号分组咨询

Hey there! As someone who’s tackled messy inventory data before, I totally get how frustrating it is to group parts when descriptions are unstandardized. Let’s break this down step by step, starting with your sample data and building up to a solution that works for your 50k+ rows.

Solution Approach & Step-by-Step Implementation

1. First: Spot Core Keywords in Descriptions

The key here is to identify recurring terms that define each product group. In your sample, we can see two clear clusters: water-related items and snack/protein bars. Since descriptions are inconsistent (like "WaterBotHold" or "Hershey_Bar"), we’ll use regex to match variations of these core terms.

2. Build a Regex-Based Matching System

We’ll create a mapping of product types to their associated keywords, then write a simple function to check each description against these keywords. This leverages the stringr tools you’re learning from R for Data Science and is beginner-friendly.

Here’s the code tailored to your sample data:

library(tidyverse)
library(stringr)

# Your original data frame
df <- tribble(
  ~PartNo, ~Description, ~ProductType,
  "A000443", "Water Bottle", "",
  "A000445", "Contain Water", "",
  "A000448", "WaterBotHold", "",
  "HRZ55", "Hershey_Bar", "",
  "RRB55", "Candy Energy", "",
  "QMU55", "Bar Protein", ""
)

# Define keyword groups (adjust these based on your full dataset!)
product_keywords <- list(
  "Water Container" = c("water", "bottl", "contain"), # Match water/bottle variations
  "Snack & Protein Bar" = c("bar", "candy", "protein", "hershey")
)

# Function to assign product type based on keyword matches
assign_product_type <- function(desc) {
  # Convert to lowercase to avoid case sensitivity issues
  desc_lower <- str_to_lower(desc)
  
  # Loop through our keyword groups to find a match
  for (type_name in names(product_keywords)) {
    if (any(str_detect(desc_lower, product_keywords[[type_name]]))) {
      return(type_name)
    }
  }
  
  # If no matches found, label as unclassified for later review
  return("Unclassified")
}

# Apply the function to fill the ProductType column
df <- df %>%
  mutate(ProductType = map_chr(Description, assign_product_type))

# View the final result
df

3. Handle Edge Cases & Scale to 50k Rows

  • Case Insensitivity: Converting descriptions to lowercase ensures "Water", "water", and "WATER" all match the same keywords.
  • Unclassified Items: The function labels non-matching entries as "Unclassified"—you can review these later to add more keywords to your groups.
  • Efficiency: This approach is fast even for large datasets because str_detect is vectorized, and map_chr efficiently applies the function across the column.

4. Advanced: Use Text Analysis to Find Hidden Keywords

If you’re unsure what keywords to use for your full dataset, the tidytext package can help you split descriptions into individual words and find the most frequent terms. This reveals patterns you might miss manually:

library(tidytext)

# Split descriptions into words and count frequencies
desc_word_counts <- df %>%
  unnest_tokens(word, Description) %>%
  count(word, sort = TRUE)

# View the most common words to inform your keyword groups
desc_word_counts

内容的提问来源于stack exchange,提问作者stackinator

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 07:08:24