You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

正则表达式为何返回多个匹配结果?Python re.findall代码求助

Why Your Regex Returns Duplicate Results & How to Fix It for Name Matching

The Immediate Duplicate Result Issue

Your regex r"((\w+\s*){1,3} from)" uses nested capturing groups, which is why re.findall returns duplicate-looking results. Here's the breakdown:

  • re.findall returns every captured group from each match. Your regex has two capturing groups: the outer (...) wrapping the whole pattern, and the inner (\w+\s*) for repeated word chunks.
  • Since the inner group repeats {1,3} times, only the final iteration of that group gets stored. So each result is a tuple containing the full matched string (from the outer group) and the last word chunk captured by the inner group.

Fix for Duplicates

Convert the inner group to a non-capturing group using (?:...) to exclude it from the results. This way, you only get the full match you care about:

import re

# Option 1: Keep outer capturing group (returns list of matched strings)
regex = r"((?:\w+\s*){1,3} from)"
test_str = "text text this is Alex Smith from text text"
print(re.findall(regex, test_str))  # Output: ['Alex Smith from']

# Option 2: Remove outer group (returns full matches directly)
regex = r"(?:\w+\s*){1,3} from"
print(re.findall(regex, test_str))  # Output: ['Alex Smith from']

Matching 2/3-Word Names (Your Actual Goal)

Now let's tackle your real objective: capturing names like "Alex Smith" (2 words) or "Mr. Alex Smith" (3 words) in contexts like after "this is" or before "from".

Key Considerations for Name Matching

  • Names typically start with capital letters, and titles like "Mr." include a trailing dot.
  • We need to target names in specific contexts to avoid matching random word sequences.

Regex Patterns for Common Scenarios

1. Capture Names Before " from "

This pattern grabs 2-3 word names immediately preceding " from ":

regex = r"(\b(?:[A-Z][a-z]*\.?\s*){2,3})\s+from"
test_str = "this is Alex Smith from Japan"
matches = re.findall(regex, test_str)
print(matches)  # Output: ['Alex Smith']

test_str2 = "this is Mr. Alex Smith from Canada"
matches2 = re.findall(regex, test_str2)
print(matches2)  # Output: ['Mr. Alex Smith']

2. Capture Names After "this is "

This pattern captures 2-3 word names right after "this is ", stopping before " from " or the end of the string:

regex = r"this is (\b(?:[A-Z][a-z]*\.?\s*){2,3})(?=\s+from|$)"
test_str = "this is Alex Smith from Japan"
matches = re.findall(regex, test_str)
print(matches)  # Output: ['Alex Smith']

test_str2 = "hello this is Mr. Alex Smith"
matches2 = re.findall(regex, test_str2)
print(matches2)  # Output: ['Mr. Alex Smith']

Pattern Breakdown

  • [A-Z][a-z]*: Matches a capitalized word (e.g., "Alex", "Smith").
  • \.?: Optional dot to account for titles like "Mr.".
  • (?:...): Non-capturing group to bundle the word pattern without adding extra captures.
  • {2,3}: Ensures we match exactly 2 or 3 word components.
  • (?=\s+from|$): Positive lookahead to stop the match before " from " or at the end of the string, preventing extra words from being captured.

内容的提问来源于stack exchange,提问作者Droid-Bird

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 06:18:36