紧密重复键值对模式匹配与异常处理技术问询
Got it, let's tackle this problem where you need to parse strings with a strict "tight" repeating name: [lowercase word] pattern—either extract all values cleanly or flag the first interference immediately. Here's a practical solution with code and breakdown:
Core Approach
We need two key behaviors:
- Tight pattern success: Extract all lowercase values and return them comma-separated when the string strictly follows
name: [word]repeats with no extra content. - Interference detection: Stop at the first non-matching segment, return the values we've already extracted, and add a clear exception message pointing out the problematic content.
Python Implementation
import re def parse_tight_name_string(input_text): # Regex for strict full-string match: only consecutive name: lowercase word pairs tight_pattern = r'^(name: [a-z]+)( name: [a-z]+)*$' full_match = re.fullmatch(tight_pattern, input_text) if full_match: # Extract all valid values if the entire string fits the pattern all_values = re.findall(r'name: ([a-z]+)', input_text) return ', '.join(all_values) else: # Step through the string to find the first invalid segment segments = input_text.split() extracted_values = [] index = 0 while index < len(segments): # Each valid entry should be "name:" followed by a lowercase word if index + 1 >= len(segments): # Incomplete pattern at the end (missing value for a name:) return f"{', '.join(extracted_values)}, ## Exception occurred: Incomplete pattern at the end of input" current_key = segments[index] current_value = segments[index + 1] if current_key == 'name:' and current_value.islower() and current_value.isalpha(): extracted_values.append(current_value) index += 2 else: # Found the first non-matching element bad_element = segments[index] return f"{', '.join(extracted_values)}, ## Exception occurred: {bad_element} cannot be matched to the pattern" # Fallback (shouldn't hit this if fullmatch worked properly) return ', '.join(extracted_values) # Test case 1: Perfect tight pattern test1 = "name: abc name: def name: ghi name: jkl" print(parse_tight_name_string(test1)) # Output: abc, def, ghi, jkl # Test case 2: Has an interfering element test2 = "name: abc x name: def name: ghi name: jkl" print(parse_tight_name_string(test2)) # Output: abc, ## Exception occurred: x cannot be matched to the pattern
How It Works
- Full Pattern Check: We first use
re.fullmatchto verify the entire string adheres to the strict repeating pattern. This ensures no hidden or extra content slips through. - Clean Extraction: If the full match passes, we yank all the lowercase values with
re.findalland format them into a comma-separated string. - Error Handling: When the full match fails, we split the string into segments and iterate through pairs. As soon as we hit something that isn't
name:followed by a lowercase word, we stop and report the issue with the values we've already collected.
This balances accuracy for valid inputs and clarity for invalid ones, which is exactly what you need for this use case.
内容的提问来源于stack exchange,提问作者oliver
相关产品推荐
相关产品推荐

