You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用正则表达式替换指定标签间的多个CDATA标记

Got it, let's work through this problem together. Your initial regex is pointing in the right direction, but it has a few quirks that we need to fix, and we'll adjust it to remove those CDATA markers inside your <tag> blocks reliably—even if we need a tiny bit of callback logic (which is standard for this kind of task, and counts as a single regex-driven workflow).

First, let's fix your tag-matching regex. The issues with your current one are:

  • (.)* is greedy, so it'll match from the very first <tag> to the last </tag> in your file, not the closest closing tag.
  • The (?!&lt;/tag&gt;) lookahead is misplaced here—it doesn't stop the greedy match, and it's not doing what you think it does.
  • . doesn't match newlines, so if your tag content spans multiple lines, the regex will break.

Here's the corrected regex to match the entire <tag> block (adjust based on whether you're working with escaped XML or raw XML):

For escaped XML (where < = &lt; and > = &gt;):

(&lt;tag&gt;)([\s\S]*?)(&lt;/tag&gt;)

For raw XML:

(<tag>)([\s\S]*?)(</tag>)
  • [\s\S]*? matches every character (including newlines) in non-greedy mode, so it stops at the first </tag>.
  • The three capture groups let us keep the opening tag, the inner content, and the closing tag separate.

Now, to remove the CDATA markers (<!CDATA[ and ]]>) inside the inner content, we'll use a substitution callback. Most programming languages and regex-aware tools support this—it lets us take the matched tag block, clean up the CDATA markers in the inner part, then put it all back together.

Example in Python:

import re

def clean_cdata_in_tag(match):
    # Grab the parts of our matched tag block
    opening_tag = match.group(1)
    inner_content = match.group(2)
    closing_tag = match.group(3)
    
    # Strip out CDATA start/end markers from the inner content
    # Use this line for escaped XML:
    cleaned_content = re.sub(r'&lt;!CDATA\[|]]&gt;', '', inner_content)
    # Use this line instead for raw XML:
    # cleaned_content = re.sub(r'<!CDATA\[|]]>', '', inner_content)
    
    # Reassemble the tag with cleaned content
    return f"{opening_tag}{cleaned_content}{closing_tag}"

# Test it out with messy XML
messy_xml = """
&lt;tag&gt;Sample text &lt;!CDATA[with CDATA stuff]]&gt; more text &lt;!CDATA[another CDATA block]]&gt; done&lt;/tag&gt;
&lt;other-tag&gt;This content should stay as-is&lt;/other-tag&gt;
"""

cleaned_xml = re.sub(r'(&lt;tag&gt;)([\s\S]*?)(&lt;/tag&gt;)', clean_cdata_in_tag, messy_xml)
print(cleaned_xml)

This will output:

&lt;tag&gt;Sample text with CDATA stuff more text another CDATA block done&lt;/tag&gt;
&lt;other-tag&gt;This content should stay as-is&lt;/other-tag&gt;

If you're set on a "pure regex" solution without a callback (like in basic text editors), it's tricky—you'd have to run a regex multiple times to catch all CDATA blocks, since most basic regex engines can't do nested global replacements in one pass. For example, in VS Code, you could use this find/repair regex for escaped XML:

  • Find: (&lt;tag&gt;.*?)&lt;!CDATA\[([\s\S]*?)]]&gt;(.*?&lt;/tag&gt;)
  • Replace: $1$2$3
    But you'd need to hit "Replace All" repeatedly until no more matches show up. The callback method is way more reliable for handling multiple CDATA blocks in one go.

To wrap up:

  1. Use the corrected non-greedy regex to capture entire <tag> blocks
  2. Use a substitution callback to clean CDATA markers in the inner content
  3. This works in almost all languages/tools and handles all edge cases like multi-line content or multiple CDATA blocks.

内容的提问来源于stack exchange,提问作者Lestat86

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 07:30:33