正则表达式为何返回多个匹配结果?Python re.findall代码求助
Why Your Regex Returns Duplicate Results & How to Fix It for Name Matching
The Immediate Duplicate Result Issue
Your regex r"((\w+\s*){1,3} from)" uses nested capturing groups, which is why re.findall returns duplicate-looking results. Here's the breakdown:
re.findallreturns every captured group from each match. Your regex has two capturing groups: the outer(...)wrapping the whole pattern, and the inner(\w+\s*)for repeated word chunks.- Since the inner group repeats
{1,3}times, only the final iteration of that group gets stored. So each result is a tuple containing the full matched string (from the outer group) and the last word chunk captured by the inner group.
Fix for Duplicates
Convert the inner group to a non-capturing group using (?:...) to exclude it from the results. This way, you only get the full match you care about:
import re # Option 1: Keep outer capturing group (returns list of matched strings) regex = r"((?:\w+\s*){1,3} from)" test_str = "text text this is Alex Smith from text text" print(re.findall(regex, test_str)) # Output: ['Alex Smith from'] # Option 2: Remove outer group (returns full matches directly) regex = r"(?:\w+\s*){1,3} from" print(re.findall(regex, test_str)) # Output: ['Alex Smith from']
Matching 2/3-Word Names (Your Actual Goal)
Now let's tackle your real objective: capturing names like "Alex Smith" (2 words) or "Mr. Alex Smith" (3 words) in contexts like after "this is" or before "from".
Key Considerations for Name Matching
- Names typically start with capital letters, and titles like "Mr." include a trailing dot.
- We need to target names in specific contexts to avoid matching random word sequences.
Regex Patterns for Common Scenarios
1. Capture Names Before " from "
This pattern grabs 2-3 word names immediately preceding " from ":
regex = r"(\b(?:[A-Z][a-z]*\.?\s*){2,3})\s+from" test_str = "this is Alex Smith from Japan" matches = re.findall(regex, test_str) print(matches) # Output: ['Alex Smith'] test_str2 = "this is Mr. Alex Smith from Canada" matches2 = re.findall(regex, test_str2) print(matches2) # Output: ['Mr. Alex Smith']
2. Capture Names After "this is "
This pattern captures 2-3 word names right after "this is ", stopping before " from " or the end of the string:
regex = r"this is (\b(?:[A-Z][a-z]*\.?\s*){2,3})(?=\s+from|$)" test_str = "this is Alex Smith from Japan" matches = re.findall(regex, test_str) print(matches) # Output: ['Alex Smith'] test_str2 = "hello this is Mr. Alex Smith" matches2 = re.findall(regex, test_str2) print(matches2) # Output: ['Mr. Alex Smith']
Pattern Breakdown
[A-Z][a-z]*: Matches a capitalized word (e.g., "Alex", "Smith").\.?: Optional dot to account for titles like "Mr.".(?:...): Non-capturing group to bundle the word pattern without adding extra captures.{2,3}: Ensures we match exactly 2 or 3 word components.(?=\s+from|$): Positive lookahead to stop the match before " from " or at the end of the string, preventing extra words from being captured.
内容的提问来源于stack exchange,提问作者Droid-Bird
相关产品推荐
相关产品推荐

