求助:Regex匹配字符串中关联dorm的指定数字异常问题
Your original regex falls short for two key reasons:
- You used the plural
dormsinstead of the singulardorm, so it misses matches where the singular form is used (like your first input string). - It only targets single digits immediately before the keyword, but doesn't account for number lists separated by commas or
e(Portuguese for "and") that lead up todormordormitórios.
Corrected Regex
Here's a regex that captures all relevant numbers correctly:
\d+(?=(?:\s*(?:,|e)\s*\d+)*\s+(?:dorm|dormitórios)\b)
Breakdown of the Regex
Let's unpack how this works:
\d+: Matches one or more digits (the actual number we want to extract).(?=...): A positive lookahead that ensures the number is part of a sequence leading to our target keywords.(?:\s*(?:,|e)\s*\d+)*: Matches zero or more additional number entries separated by commas ore, with optional whitespace around the separators. The?:makes this a non-capturing group so it doesn't interfere with our desired matches.\s+(?:dorm|dormitórios)\b: Matches the whitespace before either keyword, followed by a word boundary (\b) to avoid partial matches (e.g., if a longer word contained "dorm"). Again,?:keeps this as a non-capturing group.
How to Use It
If you're using Python, here's a quick implementation to get your desired output:
import re regex = r"\d+(?=(?:\s*(?:,|e)\s*\d+)*\s+(?:dorm|dormitórios)\b)" input_strings = [ "lori ipsum 1, 2 e 3 dorm kietjiojwoijdej 162, 131 e 107 m²", "lori fsdfsd ipsum 2 e 3 dormitórios fsrfsrfrfrfkietjiojwoijdej 162, 131 e 107 m²", "lori ipsum dfs 3 dorm kidfsrfrfrffretjiojwoijdej 62, 13 e 10 m²", "lori ipsum 1 dormitórios kietjiojwoijdej 16, 31 e 107 m²" ] for text in input_strings: matches = re.findall(regex, text) formatted_output = ", ".join([f"[{i}] => {num}" for i, num in enumerate(matches)]) print(f"[{formatted_output}]")
Expected Output
Running this code will produce exactly the output you requested:
[[0] => 1, [1] => 2, [2] => 3] [[0] => 2, [1] => 3] [[0] => 3] [[0] => 1]
This regex works for all your test cases and handles the list structures with commas and e correctly.
内容的提问来源于stack exchange,提问作者MagicHat
相关产品推荐
相关产品推荐

