如何使用XSLT检索XML文档并按sourceCode分组去重过滤区域
Hey there! Let's work through this XSLT transformation problem you've got. You already nailed loading the external XML document with the document() function, so we can focus on the grouping, deduplication, and filtering logic to get your desired output.
XSLT 2.0 makes this kind of grouping and deduplication straightforward with the <xsl:for-each-group> element. Here's the complete code:
<?xml version="1.0" encoding="UTF-8"?> <xsl:stylesheet version="2.0" xmlns:xsl="http://www.w3.org/1999/XSL/Transform"> <xsl:output method="xml" indent="yes" encoding="UTF-8"/> <!-- Load your external mapping document (you already had this right!) --> <xsl:variable name="lookupDoc" select="document('../common/lookup/mapping.xml')"/> <xsl:template match="/"> <codes> <!-- First: Filter out MAIL entries, then group by sourceCode --> <xsl:for-each-group select="$lookupDoc/codes/code[not(@area = 'MAIL')]" group-by="@sourceCode"> <code> <!-- Convert sourceCode attribute to an element --> <sourceCode><xsl:value-of select="current-grouping-key()"/></sourceCode> <areas> <!-- Deduplicate areas within the current sourceCode group --> <xsl:for-each-group select="current-group()" group-by="@area"> <area><xsl:value-of select="current-grouping-key()"/></area> </xsl:for-each-group> </areas> </code> </xsl:for-each-group> </codes> </xsl:template> </xsl:stylesheet>
1. Filtering First
We start by excluding any <code> elements where the area attribute is "MAIL" directly in our selection:
select="$lookupDoc/codes/code[not(@area = 'MAIL')]"
This is more efficient than processing those entries and filtering them out later—we just skip them entirely from the get-go.
2. Grouping by sourceCode
The outer <xsl:for-each-group> clusters all remaining <code> elements by their sourceCode attribute. The current-grouping-key() gives us the value of the sourceCode for the current group, which we output as a proper <sourceCode> element (instead of keeping it as an attribute).
3. Deduplicating area Values
Inside each sourceCode group, we run a second <xsl:for-each-group> that groups by the area attribute. This automatically removes duplicates—each unique area gets output once as an <area> element inside the <areas> wrapper.
If you're working with an older processor that only supports XSLT 1.0, we can use the Muenchian grouping technique to achieve the same result. Here's how:
<?xml version="1.0" encoding="UTF-8"?> <xsl:stylesheet version="1.0" xmlns:xsl="http://www.w3.org/1999/XSL/Transform"> <xsl:output method="xml" indent="yes" encoding="UTF-8"/> <xsl:variable name="lookupDoc" select="document('../common/lookup/mapping.xml')"/> <!-- Key for grouping code elements by sourceCode (excludes MAIL) --> <xsl:key name="sourceCodeGroup" match="$lookupDoc/codes/code[not(@area = 'MAIL')]" use="@sourceCode"/> <!-- Key for deduplicating areas: combine sourceCode + area to ensure uniqueness per group --> <xsl:key name="uniqueArea" match="$lookupDoc/codes/code[not(@area = 'MAIL')]" use="concat(@sourceCode, '|', @area)"/> <xsl:template match="/"> <codes> <!-- Iterate over unique sourceCode values --> <xsl:for-each select="$lookupDoc/codes/code[not(@area = 'MAIL')][count(. | key('sourceCodeGroup', @sourceCode)[1]) = 1]"> <code> <sourceCode><xsl:value-of select="@sourceCode"/></sourceCode> <areas> <!-- Iterate over unique areas for this sourceCode --> <xsl:for-each select="key('sourceCodeGroup', @sourceCode)[count(. | key('uniqueArea', concat(@sourceCode, '|', @area))[1]) = 1]"> <area><xsl:value-of select="@area"/></area> </xsl:for-each> </areas> </code> </xsl:for-each> </codes> </xsl:template> </xsl:stylesheet>
XSLT 1.0 Explanation
- We define two keys: one to group by
sourceCode(excluding MAIL entries), and another to group by the combination ofsourceCodeandarea(this ensures we only get unique areas per sourceCode group). - The
count(. | key(...)[1]) = 1check is a trick to select only the first occurrence of each unique group—this gives us the distinct values we need without duplicates.
Both solutions will produce exactly the output you're looking for: filtered MAIL entries, grouped by sourceCode, deduplicated areas, and the restructured XML format.
内容的提问来源于stack exchange,提问作者xray1986

