如何将XML文件中的非法< >字符转义为<和>
To fix your malformed XML and escape all invalid </> characters into </>, the most reliable approach is to use a tolerant XML parser that can handle invalid input and serialize it back to valid XML. Here's how to do it in common languages, plus a regex workaround for simple cases:
Python Solution (Recommended)
Using lxml with recovery mode enabled will automatically parse the malformed XML and escape invalid characters during serialization:
from lxml import etree # Your invalid XML string invalid_xml = "<tag1> <tag2>This is<>a<AA>text</tag2> <ABC>0123-</xyz>-89</ABC> </tag1>" # Create a parser that recovers from errors parser = etree.XMLParser(recover=True) root = etree.fromstring(invalid_xml, parser=parser) # Serialize back to valid XML with proper escaping valid_xml = etree.tostring(root, encoding='unicode', pretty_print=False) print(valid_xml)
This will output exactly your desired result:
<tag1> <tag2>This is<>a<AA>text</tag2> <ABC>0123-</xyz>-89</ABC> </tag1>
Java Solution
Use JDOM2 with a SAX parser configured to recover from fatal errors:
import org.jdom2.Document; import org.jdom2.input.SAXBuilder; import org.jdom2.output.Format; import org.jdom2.output.XMLOutputter; import java.io.StringReader; public class FixInvalidXml { public static void main(String[] args) throws Exception { String invalidXml = "<tag1> <tag2>This is<>a<AA>text</tag2> <ABC>0123-</xyz>-89</ABC> </tag1>"; // Enable error recovery in the SAX parser SAXBuilder builder = new SAXBuilder(); builder.setFeature("http://apache.org/xml/features/continue-after-fatal-error", true); Document doc = builder.build(new StringReader(invalidXml)); // Output valid XML with escaped characters XMLOutputter outputter = new XMLOutputter(Format.getRawFormat()); String validXml = outputter.outputString(doc); System.out.println(validXml); } }
Regex Workaround (For Simple, Known Tags)
If you can't use a parser and know all valid tags in advance, you can use this two-step regex method (note: this isn't foolproof for complex nested XML):
import re invalid_xml = "<tag1> <tag2>This is<>a<AA>text</tag2> <ABC>0123-</xyz>-89</ABC> </tag1>" # Step 1: Escape all < and > characters temp = invalid_xml.replace('<', '<').replace('>', '>') # Step 2: Unescape only the valid tags valid_tags = ['tag1', 'tag2', 'ABC'] for tag in valid_tags: temp = re.sub(f'<{tag}>', f'<{tag}>', temp) temp = re.sub(f'</{tag}>', f'</{tag}>', temp) valid_xml = temp print(valid_xml)
Key Explanation
The problem with your input is that some </> characters don't form properly nested, valid XML tags. A tolerant parser will interpret these invalid characters as plain text content, and when you serialize the parsed document, it will automatically escape them to </> to ensure the XML is valid. This method handles edge cases like unclosed or mismatched tags that regex can't reliably address.
内容的提问来源于stack exchange,提问作者Tedxxxx

