正则表达式优化求助:提取竖线分隔文本中的数字
Hey there! Your current regex works, but as you noticed, hardcoding all that fixed text makes it brittle—if any part of the preceding fields (like the date, link, or tooling info) changes, the regex breaks. Let's fix that with some more flexible patterns:
Option 1: Target specific positions without hardcoding
Since your target numbers are in specific columns (5th, 6th, 7th, 8th, then 10th, 11th, 12th when splitting by |), we can skip the first 4 columns using non-greedy matching, then capture the numbers, skip the 9th column, and capture the rest. This way, we don't care what's in those preceding columns—we just care about their structure.
Here's the optimized regex:
^\|\s*.+?\s*\|\s*.+?\s*\|\s*.+?\s*\|\s*.+?\s*\|\s*(\d+)\s*\|\s*(\d+)\s*\|\s*(\d+)\s*\|\s*(\d+)\s*\|\s*.+?\s*\|\s*(\d+)\s*\|\s*(\d+)\s*\|\s*(\d+)\s*\|$
Breakdown of improvements:
\s*matches any number of spaces (including zero) around the|—so if spacing varies, the regex still works..+?is a non-greedy match that grabs everything up to the next|—perfect for skipping entire columns without hardcoding their content.- We only focus on capturing the
(\d+)groups where your numbers live.
Option 2: Capture all numbers between pipes (regardless of position)
If you just need every number that's wrapped between | characters (and don't care which column they're in), use a simpler global regex:
\|\s*(\d+)\s*\|
Run this in global match mode (e.g., re.findall() in Python) and it will extract all the numbers between pipes, ignoring non-numeric columns entirely.
Bonus: Use named capture groups for readability
If you want to make the regex easier to maintain (so you remember which number is which), use named capture groups (supported in most modern regex engines):
^\|\s*.+?\s*\|\s*.+?\s*\|\s*.+?\s*\|\s*.+?\s*\|\s*(?P<metric1>\d+)\s*\|\s*(?P<metric2>\d+)\s*\|\s*(?P<metric3>\d+)\s*\|\s*(?P<metric4>\d+)\s*\|\s*.+?\s*\|\s*(?P<metric5>\d+)\s*\|\s*(?P<metric6>\d+)\s*\|\s*(?P<metric7>\d+)\s*\|$
Now when you extract matches, you can reference them by name (like match['metric1']) instead of just index numbers.
Key Takeaways
- Avoid hardcoding static content (dates, links, etc.)—it makes your regex fragile to changes.
- Use non-greedy matching (
.*?/.+?) to skip entire columns safely. - Account for variable spacing with
\s*instead of fixed spaces. - Choose the approach that fits your actual need: position-specific capture, or all numeric columns.
内容的提问来源于stack exchange,提问作者Oualid EL HASNAOUI

