如何使用XSLT去除合并XML文件中的重复XML声明与根元素?
Absolutely! XSLT can absolutely solve this problem—you just need to work around the fact that your merged file isn’t well-formed XML (thanks to those duplicate declarations and root elements). Here’s a practical, self-contained solution using XSLT 3.0 (which has robust text-processing tools to handle non-well-formed input):
How it works
We’ll break the task into four key steps:
- Load the merged file as raw text (since we can’t parse it directly as XML).
- Strip out all duplicate XML declarations (
<?xml ...?>). - Split the cleaned text into individual document fragments.
- Wrap all those fragments in a single new root element.
XSLT Code Example
<xsl:stylesheet version="3.0" xmlns:xsl="http://www.w3.org/1999/XSL/Transform"> <xsl:output method="xml" indent="yes" encoding="UTF-8"/> <!-- Main entry point (runs when no initial template is specified) --> <xsl:template name="xsl:initial-template"> <!-- Load your merged input file --> <xsl:variable name="raw-merged-text" select="unparsed-text('your-merged-file.xml')"/> <!-- Remove all XML declarations from the text --> <xsl:variable name="text-without-decls" select="replace($raw-merged-text, '<\?xml[^>]+?\?>', '', 's')"/> <!-- Split the text into distinct document fragments (adjust regex if your roots have unique names) --> <xsl:variable name="document-fragments" select="tokenize($text-without-decls, '<([^/][^>]+)>')[normalize-space(.) != '']"/> <!-- Wrap everything in a new root element --> <merged-documents> <!-- Reconstruct each original document with its root element --> <xsl:for-each select="tokenize($text-without-decls, '<([^/][^>]+)>')[position() mod 2 = 0]"> <xsl:variable name="root-tag" select="."/> <xsl:variable name="fragment-content" select="$document-fragments[position() = current()/position() + 1]"/> <xsl:value-of select="concat('<', $root-tag, '>', $fragment-content, '</', $root-tag, '>')" disable-output-escaping="yes"/> </xsl:for-each> </merged-documents> </xsl:template> </xsl:stylesheet>
Key Notes
- XSLT 3.0 Support: This uses
unparsed-text()and advanced regex, which are part of XSLT 3.0. Tools like Saxon HE/PE/EE support this natively. - Regex Tweaks: If your original XML files all use the same root element (e.g.,
<record>), replace the generic split regex with something like'<record>'for better reliability. - Disable Output Escaping: We use this to ensure the reconstructed XML elements are rendered as actual tags, not escaped text like
<doc1>.
Example Input/Output
Input (merged file):
<?xml version="1.0" encoding="UTF-8"?> <customer> <id>1</id> <name>Alice</name> </customer> <?xml version="1.0" encoding="UTF-8"?> <customer> <id>2</id> <name>Bob</name> </customer>
Output:
<?xml version="1.0" encoding="UTF-8"?> <merged-documents> <customer> <id>1</id> <name>Alice</name> </customer> <customer> <id>2</id> <name>Bob</name> </customer> </merged-documents>
If you’re stuck with XSLT 1.0 (which lacks text-processing features), you’ll need to pre-process the file first (e.g., with a simple Python script or sed command) to remove XML declarations and wrap everything in a temporary root, then use basic XSLT to clean up if needed.
内容的提问来源于stack exchange,提问作者Brendon S

