You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用正则表达式匹配除<em>TEST</em>外的所有<字符?

解决方案

核心思路

先标记需要保留的合法<em>标签(或其转义实体),再删除所有剩余的孤立<字符或不完整标签片段,最后恢复保留的内容。

情况1:字符串中是原始HTML标签(未转义,如<em>TEST</em>)

采用三步法处理,示例代码(Python):

import re

input_str = "这是<单独的<字符,还有<abc片段,<em>TEST</em>需要保留,另一个<xyz"

# 1. 用占位符暂存所有合法<em>标签内容
temp_pattern = re.compile(r'(<em>.*?</em>)')
placeholders = []
def replace_with_placeholder(match):
    placeholders.append(match.group(1))
    return f'__PLACEHOLDER_{len(placeholders)-1}__'

temp_str = temp_pattern.sub(replace_with_placeholder, input_str)

# 2. 删除所有无效的<开头片段
cleaned_str = re.sub(r'<(?!em>)[^>]*', '', temp_str)

# 3. 恢复占位符对应的标签内容
for i, content in enumerate(placeholders):
    cleaned_str = cleaned_str.replace(f'__PLACEHOLDER_{i}__', content)

print(cleaned_str)
# 输出:这是字符,还有片段,<em>TEST</em>需要保留,另一个

情况2:字符串中是转义后的HTML实体(如&lt;em&gt;TEST&lt;/em&gt;)

调整正则匹配规则,同样用占位法处理,示例代码(Python):

import re

input_str = "这是&lt;单独的&lt;字符,还有&lt;abc片段,&lt;em&gt;TEST&lt;/em&gt;需要保留,另一个&lt;xyz"

# 1. 暂存合法的转义<em>标签
temp_pattern = re.compile(r'(&lt;em&gt;.*?&lt;/em&gt;)')
placeholders = []
def replace_with_placeholder(match):
    placeholders.append(match.group(1))
    return f'__PLACEHOLDER_{len(placeholders)-1}__'

temp_str = temp_pattern.sub(replace_with_placeholder, input_str)

# 2. 删除无效的转义<或不完整片段
cleaned_str = re.sub(r'&lt;(?!em&gt;)[^;]*', '', temp_str)

# 3. 恢复占位符内容
for i, content in enumerate(placeholders):
    cleaned_str = cleaned_str.replace(f'__PLACEHOLDER_{i}__', content)

print(cleaned_str)
# 输出:这是字符,还有片段,&lt;em&gt;TEST&lt;/em&gt;需要保留,另一个

正则规则说明

  • 匹配合法标签:(<em>.*?</em>) 或 (&lt;em&gt;.*?&lt;/em&gt;),使用非贪婪模式.*?避免误匹配多个标签的内容。
  • 删除无效片段:<(?!em>)[^>]* 匹配以<开头但后续不是em>的内容,直到遇到>或字符串结束;转义版本&lt;(?!em&gt;)[^;]* 匹配以&lt;开头但后续不是em&gt;的转义片段。

内容的提问来源于stack exchange,提问作者Keannylen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 07:55:16