You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

正则表达式提取需求:过滤{}及内部内容,匹配指定规则记录

Solution to Extract Text After First = and Exclude {...} Blocks

Let's break down how to fix this properly—your current regex =.*(^[{*}]) isn't targeting the right pattern for what you need.

Why Your Current Regex Fails

That regex has two critical flaws:

  • .* greedily matches everything from the = to the end of the string, so you're not isolating the parts outside {} at all.
  • The character class (^[{*}]) is miswritten: the ^ inside a class negates it, but placing * inside doesn't do what you think. Even if you fixed it to [^{}], this would only capture a single character, not all valid text.

Correct Approach

We need two core steps:

  1. Isolate all text after the first = (since we only care about content following that first equal sign).
  2. Remove every instance of {...} (including the curly braces themselves) from that isolated text.

Here are actionable solutions in common languages:

Python Example

import re

def extract_clean_text(input_str):
    # Step 1: Grab everything after the first '='
    if '=' not in input_str:
        return ""
    after_equals = input_str.split('=', 1)[1]
    
    # Step 2: Strip out all { ... } blocks (non-greedy match to avoid spanning multiple blocks)
    cleaned_text = re.sub(r'\{.*?\}', '', after_equals)
    
    # Optional: Trim extra leading/trailing whitespace if needed
    return cleaned_text.strip()

Testing with Your Samples

Let's run this against your test cases:

  • Input: ERR_KEY_VALUE=Error logging id = {0,number,integer}; patient email = {1};
    Output: Error logging id = ; patient email = ;
  • Input: STATUS_INVITE_SENT=Invite sent
    Output: Invite sent
  • Input: WEBUSER_DUPLICATE=An account with username ''{0}'' already exists
    Output: An account with username '' '' already exists

Alternative Single-Regex Approach (Advanced)

If you prefer a one-liner (works in engines supporting lookarounds like Python, JavaScript, PCRE):

import re

def extract_clean_text(input_str):
    # Match all valid segments after the first '=', skipping { ... } blocks
    matches = re.findall(r'(?<=^[^=]=)(?:[^{}]|(?<=\}))+', input_str)
    return ''.join(matches).strip()

Key Notes

  • Use .*? (non-greedy quantifier) when matching {...} to ensure we only capture one block at a time—this prevents accidentally matching from the first { to the last } in the string.
  • The split-on-first-= method is simpler and more readable than regex for isolating the initial content, but both approaches get the job done.

内容的提问来源于stack exchange,提问作者Andy Vincent

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 03:53:12