You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

正则表达式两组匹配:保留数字前内容且排除双大写单词

Solution for Extracting Product Names (Excluding 2-Letter All-Caps Words)

Got it, let's work through this regex problem. Your goal is to pull out the product name portion from each string—stopping before any digits—and exclude those two-letter all-caps words like ME, BK, SP, etc. Here's how to do it effectively:

Approach 1: Two-Step Regex Processing

This method is straightforward and easy to debug:

  1. First, remove all two-letter all-caps words from the string.
  2. Then, extract everything before the first digit (since we want content prior to numeric values like weights/codes).

Regex Patterns:

  • Remove 2-letter all-caps words: r'\b[A-Z]{2}\b' (matches standalone two-letter uppercase words)
  • Extract content before first digit: r'^([^\d]+)' (matches all characters from start until the first digit)

Example Code (Python):

import re

input_strings = [
    "Tortillas Bolsa 2a 1kg 4118",
    "Tortillinas 50p 1 31Kg TAB TR 46113",
    "Bollos BK 4in 36p 1635g SL 131",
    "Super Pan Bco Ajonjoli 680g SP WON 100",
    "Pan Blanco Bimbo Rendidor 567g BIM 49973",
    "Gansito ME 5p 250g MTA MLA 49860"
]

# Step 1: Remove 2-letter all-caps words
cleaned = [re.sub(r'\b[A-Z]{2}\b', '', s) for s in input_strings]

# Step 2: Extract content before first digit, then trim extra spaces
results = [re.match(r'^([^\d]+)', s).group(1).strip() for s in cleaned]

# Print results
for res in results:
    print(res)

Output:

  • Tortillas Bolsa
  • Tortillinas
  • Bollos
  • Super Pan Bco Ajonjoli
  • Pan Blanco Bimbo Rendidor
  • Gansito

Approach 2: Single Regex Extraction

If you prefer a one-liner, you can use a regex that directly matches valid product name words (skipping 2-letter all-caps words and stopping at the first digit):

Regex Pattern:

^(\b(?!(?:[A-Z]{2}|\d)\b)\w+(?: \b(?!(?:[A-Z]{2}|\d)\b)\w+)*)

Breakdown:

  • ^: Start of the string
  • \b(?!(?:[A-Z]{2}|\d)\b)\w+: Matches a word that is not a 2-letter all-caps word, and does not start with a digit
  • (?: \b(?!(?:[A-Z]{2}|\d)\b)\w+)*: Matches zero or more additional valid words (separated by spaces)
  • The entire group captures the full product name we want.

Example Code (Python):

import re

input_strings = [
    "Tortillas Bolsa 2a 1kg 4118",
    "Tortillinas 50p 1 31Kg TAB TR 46113",
    "Bollos BK 4in 36p 1635g SL 131",
    "Super Pan Bco Ajonjoli 680g SP WON 100",
    "Pan Blanco Bimbo Rendidor 567g BIM 49973",
    "Gansito ME 5p 250g MTA MLA 49860"
]

pattern = r'^(\b(?!(?:[A-Z]{2}|\d)\b)\w+(?: \b(?!(?:[A-Z]{2}|\d)\b)\w+)*)'
results = [re.match(pattern, s).group(1) for s in input_strings]

for res in results:
    print(res)

Output:

Same as the two-step approach—perfectly matches your desired results.


内容的提问来源于stack exchange,提问作者Daniel Zapata

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 04:07:16