You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将XML文件中的非法< >字符转义为&lt;和&gt;

How to Escape Illegal < and > Characters in Malformed XML

To fix your malformed XML and escape all invalid </> characters into &lt;/&gt;, the most reliable approach is to use a tolerant XML parser that can handle invalid input and serialize it back to valid XML. Here's how to do it in common languages, plus a regex workaround for simple cases:

Using lxml with recovery mode enabled will automatically parse the malformed XML and escape invalid characters during serialization:

from lxml import etree

# Your invalid XML string
invalid_xml = "<tag1> <tag2>This is<>a<AA>text</tag2> <ABC>0123-</xyz>-89</ABC> </tag1>"

# Create a parser that recovers from errors
parser = etree.XMLParser(recover=True)
root = etree.fromstring(invalid_xml, parser=parser)

# Serialize back to valid XML with proper escaping
valid_xml = etree.tostring(root, encoding='unicode', pretty_print=False)

print(valid_xml)

This will output exactly your desired result:

<tag1> <tag2>This is&lt;&gt;a&lt;AA&gt;text</tag2> <ABC>0123-&lt;/xyz&gt;-89</ABC> </tag1>

Java Solution

Use JDOM2 with a SAX parser configured to recover from fatal errors:

import org.jdom2.Document;
import org.jdom2.input.SAXBuilder;
import org.jdom2.output.Format;
import org.jdom2.output.XMLOutputter;

import java.io.StringReader;

public class FixInvalidXml {
    public static void main(String[] args) throws Exception {
        String invalidXml = "<tag1> <tag2>This is<>a<AA>text</tag2> <ABC>0123-</xyz>-89</ABC> </tag1>";
        
        // Enable error recovery in the SAX parser
        SAXBuilder builder = new SAXBuilder();
        builder.setFeature("http://apache.org/xml/features/continue-after-fatal-error", true);
        
        Document doc = builder.build(new StringReader(invalidXml));
        
        // Output valid XML with escaped characters
        XMLOutputter outputter = new XMLOutputter(Format.getRawFormat());
        String validXml = outputter.outputString(doc);
        
        System.out.println(validXml);
    }
}

Regex Workaround (For Simple, Known Tags)

If you can't use a parser and know all valid tags in advance, you can use this two-step regex method (note: this isn't foolproof for complex nested XML):

import re

invalid_xml = "<tag1> <tag2>This is<>a<AA>text</tag2> <ABC>0123-</xyz>-89</ABC> </tag1>"

# Step 1: Escape all < and > characters
temp = invalid_xml.replace('<', '&lt;').replace('>', '&gt;')

# Step 2: Unescape only the valid tags
valid_tags = ['tag1', 'tag2', 'ABC']
for tag in valid_tags:
    temp = re.sub(f'&lt;{tag}&gt;', f'<{tag}>', temp)
    temp = re.sub(f'&lt;/{tag}&gt;', f'</{tag}>', temp)

valid_xml = temp
print(valid_xml)

Key Explanation

The problem with your input is that some </> characters don't form properly nested, valid XML tags. A tolerant parser will interpret these invalid characters as plain text content, and when you serialize the parsed document, it will automatically escape them to &lt;/&gt; to ensure the XML is valid. This method handles edge cases like unclosed or mismatched tags that regex can't reliably address.

内容的提问来源于stack exchange,提问作者Tedxxxx

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 03:54:34