You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

紧密重复键值对模式匹配与异常处理技术问询

Got it, let's tackle this problem where you need to parse strings with a strict "tight" repeating name: [lowercase word] pattern—either extract all values cleanly or flag the first interference immediately. Here's a practical solution with code and breakdown:

Core Approach

We need two key behaviors:

  • Tight pattern success: Extract all lowercase values and return them comma-separated when the string strictly follows name: [word] repeats with no extra content.
  • Interference detection: Stop at the first non-matching segment, return the values we've already extracted, and add a clear exception message pointing out the problematic content.

Python Implementation

import re

def parse_tight_name_string(input_text):
    # Regex for strict full-string match: only consecutive name: lowercase word pairs
    tight_pattern = r'^(name: [a-z]+)( name: [a-z]+)*$'
    full_match = re.fullmatch(tight_pattern, input_text)

    if full_match:
        # Extract all valid values if the entire string fits the pattern
        all_values = re.findall(r'name: ([a-z]+)', input_text)
        return ', '.join(all_values)
    else:
        # Step through the string to find the first invalid segment
        segments = input_text.split()
        extracted_values = []
        index = 0
        while index < len(segments):
            # Each valid entry should be "name:" followed by a lowercase word
            if index + 1 >= len(segments):
                # Incomplete pattern at the end (missing value for a name:)
                return f"{', '.join(extracted_values)}, ## Exception occurred: Incomplete pattern at the end of input"
            
            current_key = segments[index]
            current_value = segments[index + 1]
            
            if current_key == 'name:' and current_value.islower() and current_value.isalpha():
                extracted_values.append(current_value)
                index += 2
            else:
                # Found the first non-matching element
                bad_element = segments[index]
                return f"{', '.join(extracted_values)}, ## Exception occurred: {bad_element} cannot be matched to the pattern"
        
        # Fallback (shouldn't hit this if fullmatch worked properly)
        return ', '.join(extracted_values)

# Test case 1: Perfect tight pattern
test1 = "name: abc name: def name: ghi name: jkl"
print(parse_tight_name_string(test1))  # Output: abc, def, ghi, jkl

# Test case 2: Has an interfering element
test2 = "name: abc x name: def name: ghi name: jkl"
print(parse_tight_name_string(test2))  # Output: abc, ## Exception occurred: x cannot be matched to the pattern

How It Works

  1. Full Pattern Check: We first use re.fullmatch to verify the entire string adheres to the strict repeating pattern. This ensures no hidden or extra content slips through.
  2. Clean Extraction: If the full match passes, we yank all the lowercase values with re.findall and format them into a comma-separated string.
  3. Error Handling: When the full match fails, we split the string into segments and iterate through pairs. As soon as we hit something that isn't name: followed by a lowercase word, we stop and report the issue with the values we've already collected.

This balances accuracy for valid inputs and clarity for invalid ones, which is exactly what you need for this use case.

内容的提问来源于stack exchange,提问作者oliver

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 09:50:57