如何用正则移除XML中不匹配的ENAMEX标签且保留数据
Let's break down what's going wrong with your current regex, then fix it with targeted replacements—plus a more robust alternative using an XML parser for future-proofing.
The Problem with Your Current Code
Your regex logic is over-matching and not properly capturing the content inside the misplaced ENAMEX tags. The add_closing_tag_search pattern grabs everything after the opening ENAMEX (including WEBSITE.COM</SubjectFullName>) and wraps it in a new ENAMEX pair, which then gets mangled by your first replacement step. That's why you're losing content and getting unclosed tags.
Quick Fix with Targeted Regex
Since your XML has a predictable structure, we can use two precise regex replacements to clean up the tags without losing content:
import re original_xml = '''<Subject stype="ORG" xref="1234"> <SubjectFullName type="L"><ENAMEX type="ORGANIZATION" id="ORG-112233-000">WEBSITE.COM</SubjectFullName> <SubjectLastName type="L">WEBSITE.COM</ENAMEX></SubjectLastName> <SubjectPhone type="Work">1234567890</SubjectPhone> </Subject>''' # Step 1: Remove the opening ENAMEX tag inside SubjectFullName, keep the content clean_step1 = re.sub( r'<SubjectFullName type="L"><ENAMEX[^>]+>(.*?)</SubjectFullName>', r'<SubjectFullName type="L">\1</SubjectFullName>', original_xml ) # Step 2: Remove the extra closing ENAMEX tag inside SubjectLastName final_xml = re.sub( r'</ENAMEX></SubjectLastName>', r'</SubjectLastName>', clean_step1 ) print(final_xml)
This outputs exactly what you want:
<Subject stype="ORG" xref="1234"> <SubjectFullName type="L">WEBSITE.COM</SubjectFullName> <SubjectLastName type="L">WEBSITE.COM</SubjectLastName> <SubjectPhone type="Work">1234567890</SubjectPhone> </Subject>
More Robust Solution: Use an XML Parser
Regex works for simple cases, but it's fragile for XML (tags can vary in attributes, structure, etc.). For a more reliable fix, use an XML parser like BeautifulSoup—it automatically handles mismatched tags and lets you safely remove ENAMEX elements while preserving their content:
from bs4 import BeautifulSoup original_xml = '''<Subject stype="ORG" xref="1234"> <SubjectFullName type="L"><ENAMEX type="ORGANIZATION" id="ORG-112233-000">WEBSITE.COM</SubjectFullName> <SubjectLastName type="L">WEBSITE.COM</ENAMEX></SubjectLastName> <SubjectPhone type="Work">1234567890</SubjectPhone> </Subject>''' # Parse the XML (even with mismatched tags) soup = BeautifulSoup(original_xml, 'xml') # Find all ENAMEX tags and replace them with their text content for enamex_tag in soup.find_all('ENAMEX'): enamex_tag.replace_with(enamex_tag.text) # Get the cleaned, properly formatted XML print(soup.prettify())
This method will work even if your XML structure changes (e.g., additional attributes on ENAMEX, different nesting levels) because it understands XML's structure instead of relying on string patterns.
内容的提问来源于stack exchange,提问作者carousallie

