使用Saxon流处理XSLT移除重复对象:大文件处理异常求助
解决大型XML Feed去重无输出问题的替代方案
问题分析
你当前用<xsl:for-each-group>按id分组去重的方案,在小型XML中正常,但处理200万条数据时无输出,核心原因是流模式下的group-by需要缓存所有分组键和对应条目,200万条数据会耗尽内存,导致转换异常终止。
替代方案:使用XSLT 3.0累加器(xsl:accumulator)
无需分组,通过累加器跟踪已处理的id,流处理时仅处理首次出现的条目,内存占用极低,适合超大型Feed。
修正后的完整XSLT代码
<?xml version="1.0"?> <xsl:stylesheet xmlns:xsl="http://www.w3.org/1999/XSL/Transform" version="3.0"> <xsl:output method="text" encoding="UTF-8" indent="no" omit-xml-declaration="yes"/> <xsl:mode streamable="yes" use-accumulators="seen-ids"/> <xsl:strip-space elements="*"/> <!-- 累加器:跟踪已处理的id集合 --> <xsl:accumulator name="seen-ids" as="xs:string*" initial-value="()" streamable="yes"> <xsl:accumulator-rule match="item/id" select="if (not(. = $value)) then ($value, string(.)) else $value"/> </xsl:accumulator> <xsl:variable name="dq" select="'"'"/> <xsl:variable name="qcq" select="'","'"/> <xsl:variable name="lf" select="' '"/> <!-- Static defaults --> <xsl:variable name="data_source" select="'HHTestMedia'"/> <!-- 处理items节点下的每个item --> <xsl:template match="items"> <xsl:apply-templates select="item"/> </xsl:template> <!-- 仅处理首次出现的item --> <xsl:template match="item[string(id) not(= accumulator-before('seen-ids'))]"> <xsl:variable name="currentItem" select="copy-of(.)"/> <xsl:variable name="p_title" select="substring(replace($currentItem/title,'"',''), 1, 128)"/> <xsl:variable name="p_city" select="substring-before($currentItem/location, ',')"/> <xsl:variable name="p_state_code" select="substring(normalize-space(substring-after($currentItem/location, ',')), 1, 2)"/> <xsl:variable name="p_postal_code" select="$currentItem/postcode"/> <xsl:variable name="p_country_code" select="$currentItem/country"/> <xsl:variable name="p_company_name" select="replace(substring($currentItem/company, 1, 64), '"', '')"/> <xsl:value-of select="concat($dq, $p_title, $qcq, $p_city, $qcq, $p_state_code, $qcq, $p_postal_code, $qcq, $p_country_code, $qcq, $p_company_name, $dq, $lf)"/> </xsl:template> <!-- 跳过重复的item --> <xsl:template match="item"/> </xsl:stylesheet>
方案说明
- 累加器
seen-ids:以字符串集合形式存储已处理的id,每次遇到新id就添加到集合中,流处理时增量更新,无需缓存全量数据。 - 模板匹配逻辑:
- 匹配
items节点,遍历所有item; - 仅处理
id不在累加器历史中的item(即首次出现的条目); - 重复的
item被空模板匹配,直接跳过。
- 匹配
- 内存优化:仅保留已处理的
id集合,200万条不同id的情况下,内存占用远低于分组方案(分组需缓存所有条目)。
关键注意事项
- 确保Saxon EE 10.5启用了流处理(已在
<xsl:mode>中设置streamable="yes"); copy-of(.)仅在需要处理当前item的子节点时使用,避免不必要的节点缓存;- 若
id是CDATA格式,string(id)会自动提取CDATA中的文本内容,无需额外处理。
内容的提问来源于stack exchange,提问作者Gopinath
相关产品推荐
相关产品推荐

