使用正则表达式提取带千分位和逗号分隔的EUR后缀多规格金额
Hey there! I see you're struggling with extracting those EUR amounts that use dots as thousand separators and commas as decimal points, and your current approach of stacking multiple re.finditer calls is leading to duplicate results and unnecessary complexity. Let's fix that with a single, clean regex that covers all your cases.
The Solution Regex
First, here's a regex that will match all the amount formats you've shown:
import re # Your sample text input_text = "……金额 35.550,35 EUR ……某内容 7,2 EUR ……花费 98.556.732,93 EUR……" # The regex pattern amount_pattern = r'(\d+(?:\.\d+)*,\d+)\sEUR' # Extract all matches amounts = re.findall(amount_pattern, input_text) print(amounts) # Output: ['35.550,35', '7,2', '98.556.732,93']
How This Regex Works
Let's break down the pattern to understand why it works for your cases:
(\d+(?:\.\d+)*,\d+): The capture group that grabs the actual amount\d+: Matches the initial sequence of digits (like35,7, or98)(?:\.\d+)*: A non-capturing group that matches zero or more instances of a dot followed by digits (this handles the thousand separators, like.550in35.550or.556.732in98.556.732),\d+: Matches the decimal separator (comma) followed by one or more digits (the decimal part, like,35,,2, or,93)
\sEUR: Ensures we only match amounts that are directly followed by a space and the fixedEURsuffix, so we don't accidentally grab unrelated numbers.
Handling Integer-Only Amounts (Optional)
If you might also have amounts without a decimal part (like 100 EUR), just adjust the regex to make the decimal part optional:
input_text_with_integer = "……金额 35.550,35 EUR ……某内容 7,2 EUR ……花费 98.556.732,93 EUR……还有100 EUR……" amount_pattern_with_integers = r'(\d+(?:\.\d+)*(?:,\d+)?)\sEUR' amounts_with_integers = re.findall(amount_pattern_with_integers, input_text_with_integer) print(amounts_with_integers) # Output: ['35.550,35', '7,2', '98.556.732,93', '100']
Why Your Original Approach Was Problematic
Your initial method of using multiple separate regexes (matching \d+, \d+\.\d+, etc.) leads to duplicates because shorter patterns will match parts of longer valid amounts. For example, \d+ would match 35, 550, and 35.550 from 35.550,35 EUR—creating unnecessary duplicate entries. Using one targeted regex avoids this entirely.
内容的提问来源于stack exchange,提问作者DavyJones

