You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用stringr正则表达式从零件描述提取重复词生成产品类型

Solution for Grouping Parts by Product Type Using Stringr & Tidyverse

Let’s walk through this step by step—you’ve got 50k+ parts with unstructured descriptions, no predefined product types, and need to group them using repeating keywords. We’ll use text processing with stringr and tidyverse tools to make this work.

Step 1: Preprocess Messy Description Text

First, we need to normalize inconsistent formatting (like camelCase terms, mixed case) so we can extract meaningful keywords. For example, "WaterBotHold" gets split into readable words to align with other mentions of "water".

library(tidyverse)
library(stringr)

# Your sample dataset
df <- tribble(
  ~PartNo, ~Description, ~ProductType,
  "A000443", "Water Bottle", "",
  "A000445", "Contain Water", "",
  "A000448", "WaterBotHold", ""
)

# Clean descriptions: split camelCase, lowercase, extract words
df_processed <- df %>%
  mutate(
    # Add space before uppercase letters to split camelCase
    clean_desc = str_replace_all(Description, "(?<=[a-z])(?=[A-Z])", " "),
    # Convert to lowercase and pull out all word characters
    keywords = str_extract_all(tolower(clean_desc), "\\w+")
  )

Step 2: Identify High-Frequency Keywords

Next, we’ll count keyword occurrences to spot terms that define product groups (like "water" in your sample). For 50k parts, this will highlight the most common themes across descriptions.

# Flatten keywords to count their frequency
keyword_counts <- df_processed %>%
  unnest(keywords) %>%
  count(keywords, sort = TRUE)

# Filter for meaningful keywords (adjust threshold for your full dataset)
# For 50k parts, use a higher threshold like n >= 10 to filter noise
top_keywords <- keyword_counts %>%
  filter(n >= 1)

# Sample output:
# keywords | n
# ---------|---
# water    | 3
# bottle   | 1
# contain  | 1
# bot      | 1
# hold     | 1

Step 3: Map Keywords to Consistent Product Types

Create a lookup table to group related keywords into standardized product types. For your sample, all water-related terms get mapped to "Water Container".

# Build a keyword-to-product-type mapping (expand this for your full dataset)
product_type_lookup <- tibble(
  keyword = c("water", "bottle", "contain", "bot", "hold"),
  ProductType = rep("Water Container", 5)
)

# Match keywords back to the original data and fill the ProductType column
df_final <- df_processed %>%
  unnest(keywords) %>%
  left_join(product_type_lookup, by = "keywords") %>%
  group_by(PartNo) %>%
  summarise(
    Description = first(Description),
    # Pick the first matching product type (adjust if handling conflicts)
    ProductType = first(ProductType)
  )

Step 4: Scaling to 50k+ Parts

For large datasets, optimize with these tips:

  • Fuzzy matching: Use the stringdist package to handle misspelled keywords (e.g., "water" vs "watter").
  • Auto-clustering: Use word embeddings or k-means clustering to auto-group similar descriptions if manual lookup is too time-consuming.
  • Vectorized operations: Stick to dplyr’s vectorized functions instead of loops to handle 50k rows efficiently.

Final Sample Output

Running the code on your sample gives this cleaned, grouped data:

PartNoDescriptionProductType
A000443Water BottleWater Container
A000445Contain WaterWater Container
A000448WaterBotHoldWater Container

内容的提问来源于stack exchange,提问作者stackinator

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 07:08:09