正则表达式捕获特定大写开头连续单词的问题排查
Fixing Your Name Extraction Regex Issue
Let's break down why your original approach failed and fix it to meet your exact requirements.
Problem with Your Original Code
The main issues were:
- The unnecessary custom delimiter (
3hsdfhdg11sdf) paired with an extra literal<in your regex, which made it impossible to match any text. - The boundary handling for the name wasn't correctly targeting sentence-ending punctuation or the end of the input string.
Corrected Solution
We can simplify this by directly targeting valid boundaries after the name (your specified punctuation or end of string) without needing a hacky custom delimiter. Here's the revised code:
import re def test_extraction_func(input_text): word = "" # Regex pattern breakdown: # 1. Case-insensitive trigger phrase to identify name context # 2. Capture group for the multi-word capitalized name (supports accented characters) # 3. Optional valid boundary punctuation OR end of string trigger_phrase = r"(?i:no\s*identifiques\s*como\s*(?:a\s*un|un|)\s*nombre\s*(?:a\s*el\s*nombre\s*(?:de|)|al\s*nombre|a)\s+)" name_capture = r"([A-ZÁÉÍÓÚÜÑ][a-záéíóúüñ]+(?:\s*[A-ZÁÉÍÓÚÜÑ][a-záéíóúüñ]+)*)" valid_boundaries = r"\s*(?:\.\s*\n|;\s*|,\s*|\.\s*|$)" full_pattern = trigger_phrase + name_capture + valid_boundaries match = re.search(full_pattern, input_text) if match: word = match.group(1).strip() print(repr(word)) else: print(repr(word)) # Test all your examples print("Example 1:") test_extraction_func("SAFJHDFH no identifiques como nombre a María del Carmen asjdjhs") # Output: "" print("Example 2:") test_extraction_func("no identifiques como nombre a María Carmen") # Output: "María Carmen" print("Example 3:") test_extraction_func("sagasghas no identifiques como a un nombre a María Carmen Azul, k9kfjfjfd") # Output: "María Carmen Azul" print("Example 4:") test_extraction_func("sagasghas no identifiques como a un nombre a María Carmen Azul; Aun que no estoy realmente segura de ello") # Output: "María Carmen Azul" print("Example 5:") test_extraction_func("no identifiques como nombre a María hghdshgsd") # Output: ""
Key Changes Explained
- Removed the custom delimiter: We use
$to match the end of the string, combined with your required punctuation (. \n,.\n,.,,,;) in a non-capturing group. This ensures the name is followed only by valid boundaries. - Simplified trigger phrase: Kept it case-insensitive to handle variations in input capitalization, while maintaining flexibility to match all your specified trigger phrase formats.
- Cleaner capture logic: The code directly checks for a valid match and strips any extra whitespace from the captured name for clean output.
Test Results
Running the code produces exactly what you expected:
- Example 1: Empty string (name is followed by random lowercase text, not a valid boundary)
- Example 2: Captures
"María Carmen"(ends at the end of the text) - Example 3: Captures
"María Carmen Azul"(followed by a comma) - Example 4: Captures
"María Carmen Azul"(followed by a semicolon) - Example 5: Empty string (name is followed by random lowercase text)
内容的提问来源于stack exchange,提问作者Matt095
相关产品推荐
相关产品推荐

