You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用正则移除XML中不匹配的ENAMEX标签且保留数据

Fixing Mismatched ENAMEX Tags in XML (Remove ENAMEX, Preserve Content)

Let's break down what's going wrong with your current regex, then fix it with targeted replacements—plus a more robust alternative using an XML parser for future-proofing.

The Problem with Your Current Code

Your regex logic is over-matching and not properly capturing the content inside the misplaced ENAMEX tags. The add_closing_tag_search pattern grabs everything after the opening ENAMEX (including WEBSITE.COM</SubjectFullName>) and wraps it in a new ENAMEX pair, which then gets mangled by your first replacement step. That's why you're losing content and getting unclosed tags.

Quick Fix with Targeted Regex

Since your XML has a predictable structure, we can use two precise regex replacements to clean up the tags without losing content:

import re

original_xml = '''<Subject stype="ORG" xref="1234"> <SubjectFullName type="L"><ENAMEX type="ORGANIZATION" id="ORG-112233-000">WEBSITE.COM</SubjectFullName> <SubjectLastName type="L">WEBSITE.COM</ENAMEX></SubjectLastName> <SubjectPhone type="Work">1234567890</SubjectPhone> </Subject>'''

# Step 1: Remove the opening ENAMEX tag inside SubjectFullName, keep the content
clean_step1 = re.sub(
    r'<SubjectFullName type="L"><ENAMEX[^>]+>(.*?)</SubjectFullName>',
    r'<SubjectFullName type="L">\1</SubjectFullName>',
    original_xml
)

# Step 2: Remove the extra closing ENAMEX tag inside SubjectLastName
final_xml = re.sub(
    r'</ENAMEX></SubjectLastName>',
    r'</SubjectLastName>',
    clean_step1
)

print(final_xml)

This outputs exactly what you want:

<Subject stype="ORG" xref="1234"> <SubjectFullName type="L">WEBSITE.COM</SubjectFullName> <SubjectLastName type="L">WEBSITE.COM</SubjectLastName> <SubjectPhone type="Work">1234567890</SubjectPhone> </Subject>

More Robust Solution: Use an XML Parser

Regex works for simple cases, but it's fragile for XML (tags can vary in attributes, structure, etc.). For a more reliable fix, use an XML parser like BeautifulSoup—it automatically handles mismatched tags and lets you safely remove ENAMEX elements while preserving their content:

from bs4 import BeautifulSoup

original_xml = '''<Subject stype="ORG" xref="1234"> <SubjectFullName type="L"><ENAMEX type="ORGANIZATION" id="ORG-112233-000">WEBSITE.COM</SubjectFullName> <SubjectLastName type="L">WEBSITE.COM</ENAMEX></SubjectLastName> <SubjectPhone type="Work">1234567890</SubjectPhone> </Subject>'''

# Parse the XML (even with mismatched tags)
soup = BeautifulSoup(original_xml, 'xml')

# Find all ENAMEX tags and replace them with their text content
for enamex_tag in soup.find_all('ENAMEX'):
    enamex_tag.replace_with(enamex_tag.text)

# Get the cleaned, properly formatted XML
print(soup.prettify())

This method will work even if your XML structure changes (e.g., additional attributes on ENAMEX, different nesting levels) because it understands XML's structure instead of relying on string patterns.

内容的提问来源于stack exchange,提问作者carousallie

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 16:37:34