正则表达式递进式匹配多段结果的技术求助
Got it, let's tackle this problem. The issue with your current regex is that it's designed to match the entire longest valid sequence (like 6211-10/20/30/40), but you want all progressive prefixes of that sequence—from the first segment up to the full string.
Here's how to solve this, with two approaches depending on whether you want to use regex alone or combine it with a little code (the latter is often more straightforward):
Approach 1: Combine Regex with Simple String Splitting (Recommended)
First, use a tightened-up version of your regex to capture the full valid sequence. Then split that sequence into segments and generate all possible progressive prefixes.
Step 1: Capture the Full Sequence
Use this regex to grab the complete matching string:
\b\d{4}[\.\-/\\ _]{0,2}\d{2,4}(?:/\d{2,4})+
\bensures we don't match partial numbers[\.\-/\\ _]{0,2}allows 0-2 separators between the initial 4-digit segment and the next one(?:/\d{2,4})+matches one or more trailing/{2-4 digits}segments (non-capturing group to avoid extra noise)
Step 2: Generate Progressive Prefixes
Once you have the full match (e.g., 6211-10/20/30/40), split it by / and build each prefix. Here's an example in JavaScript:
const inputStr = "... chemical tank 6211-10/20/30/40 and other equipment ..."; const fullMatch = inputStr.match(/\b\d{4}[\.\-/\\ _]{0,2}\d{2,4}(?:/\d{2,4})+/)[0]; const segments = fullMatch.split('/'); const progressiveMatches = segments.reduce((acc, _, index) => { acc.push(segments.slice(0, index + 1).join('/')); return acc; }, []); // Result: ["6211-10", "6211-10/20", "6211-10/20/30", "6211-10/20/30/40"]
And here's the equivalent in Python:
import re input_str = "... chemical tank 6211-10/20/30/40 and other equipment ..." full_match = re.search(r'\b\d{4}[\.\-/\\ _]{0,2}\d{2,4}(?:/\d{2,4})+', input_str).group() segments = full_match.split('/') progressive_matches = ['/'.join(segments[:i+1]) for i in range(len(segments))] // Result: ['6211-10', '6211-10/20', '6211-10/20/30', '6211-10/20/30/40']
Approach 2: Pure Regex (For Engines That Support Global Matching + Lookaheads)
If you want to use regex alone to capture all progressive matches, you can leverage positive lookaheads to "lock in" the remaining segments while matching each prefix. Here's the regex:
\b\d{4}[\.\-/\\ _]{0,2}\d{2,4}(?:/\d{2,4})*?(?=(?:/\d{2,4})*$)
*?makes the trailing segment match non-greedy (so it starts with the shortest possible prefix)(?=(?:/\d{2,4})*$)is a positive lookahead that ensures the rest of the string (from the current match end) is zero or more valid trailing segments (this forces the regex to match every possible prefix that leads to the end of the full sequence)
Run this with the global match flag (g in JavaScript, re.findall in Python). For example, in JavaScript:
const inputStr = "... chemical tank 6211-10/20/30/40 and other equipment ..."; const regex = /\b\d{4}[\.\-/\\ _]{0,2}\d{2,4}(?:/\d{2,4})*?(?=(?:/\d{2,4})*$)/g; const progressiveMatches = [...inputStr.matchAll(regex)].map(m => m[0]); // Result: ["6211-10", "6211-10/20", "6211-10/20/30", "6211-10/20/30/40"]
Key Notes
- Both approaches work for any number of trailing segments (whether it's 2 like
6311-22/42or 7 like6158-47/84/85/86/87/88/89) - The string splitting approach is usually easier to read and maintain, especially if you need to adjust how prefixes are generated later
内容的提问来源于stack exchange,提问作者Yann Gueguen

