XSLT字符串拆分合并:移除指定PrimeLineNo的OrderLine片段
Alright, let's tackle this problem step by step. The goal is to remove the entire escaped <OrderLine> element with PrimeLineNo="1" from the Inputxml attribute of your XML nodes—even when there are multiple OrderLine entries present.
Instead of relying on error-prone string manipulation or regex (which can break if the XML structure varies slightly), we'll treat the escaped XML in the Inputxml attribute as actual XML data. Here's the plan:
- Decode the escaped string: Convert the
Inputxmlattribute's escaped XML content into a proper XML node set. - Filter out the target element: Remove any
OrderLineelement wherePrimeLineNo="1". - Re-encode the remaining XML: Turn the filtered node set back into an escaped string to replace the original
Inputxmlattribute value.
We'll use XSLT 3.0 here (supported by processors like Saxon-HE) because it has built-in functions for parsing and serializing XML, making the solution clean and reliable:
<xsl:stylesheet version="3.0" xmlns:xsl="http://www.w3.org/1999/XSL/Transform"> <xsl:output method="xml" indent="yes"/> <!-- Handle elements with the Inputxml attribute --> <xsl:template match="*[@Inputxml]"> <xsl:copy> <!-- Copy all original attributes except Inputxml --> <xsl:copy-of select="@* except @Inputxml"/> <!-- Rebuild the Inputxml attribute with filtered content --> <xsl:attribute name="Inputxml"> <!-- Wrap escaped content in a root node to parse it properly --> <xsl:variable name="decoded-content" select="parse-xml(concat('<root>', @Inputxml, '</root>'))"/> <!-- Serialize only the OrderLine elements that don't have PrimeLineNo=1 --> <xsl:value-of select="serialize($decoded-content/root/*[not(self::OrderLine[@PrimeLineNo='1'])])"/> </xsl:attribute> <!-- Copy any child nodes of the original element --> <xsl:apply-templates/> </xsl:copy> </xsl:template> <!-- Default template: copy all other nodes and attributes as-is --> <xsl:template match="@*|node()"> <xsl:copy> <xsl:apply-templates select="@*|node()"/> </xsl:copy> </xsl:template> </xsl:stylesheet>
parse-xml(): This function takes the escaped string fromInputxml, wraps it in a temporary<root>element (since XML requires a single root node), and converts it into a usable XML node set.serialize(): Converts the filtered XML nodes back into an escaped string, which we use to update theInputxmlattribute.@* except @Inputxml: Ensures all other attributes of the original element (likeAttrib1andAttrib2in your test case) are preserved exactly as they are.- The filter
[not(self::OrderLine[@PrimeLineNo='1'])]: Targets onlyOrderLineelements with the specific attribute value we want to remove, leaving all other content (like the<Line>element in your first example) untouched.
Let's test this with your sample XML:
Test Input
<root> <input Inputxml="<Order..<Line PPlineNO="1"/> >" /> <Test Attrib1="b" Attrib2="C" Inputxml=" <OrderLine OrderedQty="1" PrimeLineNo="1" ShipNode="ABC" > </OrderLine > <OrderLine OrderedQty="1" PrimeLineNo="2" ShipNode="ABC" > </OrderLine >" /> </root>
Test Output
<root> <input Inputxml="<Order..<Line PPlineNO="1"/> >"/> <Test Attrib1="b" Attrib2="C" Inputxml="<OrderLine OrderedQty="1" PrimeLineNo="2" ShipNode="ABC"/>"/> </root>
Notice that the first <input> element's content is unchanged (since there's no OrderLine with PrimeLineNo="1"), and the <Test> element now only has the OrderLine with PrimeLineNo="2".
内容的提问来源于stack exchange,提问作者Tina Dsouza

