You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求助:Regex匹配字符串中关联dorm的指定数字异常问题

Your original regex falls short for two key reasons:

  1. You used the plural dorms instead of the singular dorm, so it misses matches where the singular form is used (like your first input string).
  2. It only targets single digits immediately before the keyword, but doesn't account for number lists separated by commas or e (Portuguese for "and") that lead up to dorm or dormitórios.

Corrected Regex

Here's a regex that captures all relevant numbers correctly:

\d+(?=(?:\s*(?:,|e)\s*\d+)*\s+(?:dorm|dormitórios)\b)

Breakdown of the Regex

Let's unpack how this works:

  • \d+: Matches one or more digits (the actual number we want to extract).
  • (?=...): A positive lookahead that ensures the number is part of a sequence leading to our target keywords.
    • (?:\s*(?:,|e)\s*\d+)*: Matches zero or more additional number entries separated by commas or e, with optional whitespace around the separators. The ?: makes this a non-capturing group so it doesn't interfere with our desired matches.
    • \s+(?:dorm|dormitórios)\b: Matches the whitespace before either keyword, followed by a word boundary (\b) to avoid partial matches (e.g., if a longer word contained "dorm"). Again, ?: keeps this as a non-capturing group.

How to Use It

If you're using Python, here's a quick implementation to get your desired output:

import re

regex = r"\d+(?=(?:\s*(?:,|e)\s*\d+)*\s+(?:dorm|dormitórios)\b)"
input_strings = [
    "lori ipsum 1, 2 e 3 dorm kietjiojwoijdej 162, 131 e 107 m²",
    "lori fsdfsd ipsum 2 e 3 dormitórios fsrfsrfrfrfkietjiojwoijdej 162, 131 e 107 m²",
    "lori ipsum dfs 3 dorm kidfsrfrfrffretjiojwoijdej 62, 13 e 10 m²",
    "lori ipsum 1 dormitórios kietjiojwoijdej 16, 31 e 107 m²"
]

for text in input_strings:
    matches = re.findall(regex, text)
    formatted_output = ", ".join([f"[{i}] => {num}" for i, num in enumerate(matches)])
    print(f"[{formatted_output}]")

Expected Output

Running this code will produce exactly the output you requested:

[[0] => 1, [1] => 2, [2] => 3]
[[0] => 2, [1] => 3]
[[0] => 3]
[[0] => 1]

This regex works for all your test cases and handles the list structures with commas and e correctly.

内容的提问来源于stack exchange,提问作者MagicHat

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 10:36:22