如何使用正则表达式从包含多种数字格式的字符串中提取纯数字?
Got it, let's tackle this problem step by step. You've already nailed extracting those space-and-comma formatted numbers like 4 500,5, but the tricky part is grabbing standalone integers like 12 and 34 without picking up fragments from other formatted numbers (like the 4, 500, or 5 from 4 500,5).
Here's a regex pattern that handles both cases cleanly, using negative lookarounds to exclude partial number fragments:
import re text = ''' 4 500,5 12% 1,63% 568768,74832 days in between 34 cars in a row''' # Combined pattern to match all target number types number_pattern = re.compile(r'(?:\d+ )?\d+,\d+|\b(?!\d*,\d)(?<!\d,\d*)\d+\b') extracted_numbers = number_pattern.findall(text) print(extracted_numbers) # Output: ['4 500,5', '12', '568768,74832', '34']
Let's break down how this works:
(?:\d+ )?\d+,\d+: This is your original pattern (wrapped in a non-capturing group) that matches numbers with optional leading spaces and a comma decimal separator (covers4 500,5and568768,74832).\b(?!\d*,\d)(?<!\d,\d*)\d+\b: This part targets standalone integers:\bensures we're matching whole words (so we don't get partial digits from longer numbers).(?!\d*,\d): Negative lookahead to exclude any number that's followed by a comma and more digits (prevents matching500from4 500,5).(?<!\d,\d*): Negative lookbehind to exclude any number that's preceded by digits and a comma (prevents matching5from4 500,5or63from1,63%).
The key here is that we're prioritizing the more complex formatted numbers first (the left side of the |), so the regex doesn't split those into smaller integer fragments before matching them as whole units.
If you ever need to include numbers like 1,63 from the 1,63% string, you can adjust the pattern to add another case: (?:\d+ )?\d+,\d+%?|\b(?!\d*,\d)(?<!\d,\d*)\d+\b — but based on your question, this extra case probably isn't needed.
内容的提问来源于stack exchange,提问作者Rufat

