正则表达式两组匹配:保留数字前内容且排除双大写单词
Solution for Extracting Product Names (Excluding 2-Letter All-Caps Words)
Got it, let's work through this regex problem. Your goal is to pull out the product name portion from each string—stopping before any digits—and exclude those two-letter all-caps words like ME, BK, SP, etc. Here's how to do it effectively:
Approach 1: Two-Step Regex Processing
This method is straightforward and easy to debug:
- First, remove all two-letter all-caps words from the string.
- Then, extract everything before the first digit (since we want content prior to numeric values like weights/codes).
Regex Patterns:
- Remove 2-letter all-caps words:
r'\b[A-Z]{2}\b'(matches standalone two-letter uppercase words) - Extract content before first digit:
r'^([^\d]+)'(matches all characters from start until the first digit)
Example Code (Python):
import re input_strings = [ "Tortillas Bolsa 2a 1kg 4118", "Tortillinas 50p 1 31Kg TAB TR 46113", "Bollos BK 4in 36p 1635g SL 131", "Super Pan Bco Ajonjoli 680g SP WON 100", "Pan Blanco Bimbo Rendidor 567g BIM 49973", "Gansito ME 5p 250g MTA MLA 49860" ] # Step 1: Remove 2-letter all-caps words cleaned = [re.sub(r'\b[A-Z]{2}\b', '', s) for s in input_strings] # Step 2: Extract content before first digit, then trim extra spaces results = [re.match(r'^([^\d]+)', s).group(1).strip() for s in cleaned] # Print results for res in results: print(res)
Output:
- Tortillas Bolsa
- Tortillinas
- Bollos
- Super Pan Bco Ajonjoli
- Pan Blanco Bimbo Rendidor
- Gansito
Approach 2: Single Regex Extraction
If you prefer a one-liner, you can use a regex that directly matches valid product name words (skipping 2-letter all-caps words and stopping at the first digit):
Regex Pattern:
^(\b(?!(?:[A-Z]{2}|\d)\b)\w+(?: \b(?!(?:[A-Z]{2}|\d)\b)\w+)*)
Breakdown:
^: Start of the string\b(?!(?:[A-Z]{2}|\d)\b)\w+: Matches a word that is not a 2-letter all-caps word, and does not start with a digit(?: \b(?!(?:[A-Z]{2}|\d)\b)\w+)*: Matches zero or more additional valid words (separated by spaces)- The entire group captures the full product name we want.
Example Code (Python):
import re input_strings = [ "Tortillas Bolsa 2a 1kg 4118", "Tortillinas 50p 1 31Kg TAB TR 46113", "Bollos BK 4in 36p 1635g SL 131", "Super Pan Bco Ajonjoli 680g SP WON 100", "Pan Blanco Bimbo Rendidor 567g BIM 49973", "Gansito ME 5p 250g MTA MLA 49860" ] pattern = r'^(\b(?!(?:[A-Z]{2}|\d)\b)\w+(?: \b(?!(?:[A-Z]{2}|\d)\b)\w+)*)' results = [re.match(pattern, s).group(1) for s in input_strings] for res in results: print(res)
Output:
Same as the two-step approach—perfectly matches your desired results.
内容的提问来源于stack exchange,提问作者Daniel Zapata
相关产品推荐
相关产品推荐

