You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Saxon流处理XSLT移除重复对象:大文件处理异常求助

解决大型XML Feed去重无输出问题的替代方案

问题分析

你当前用<xsl:for-each-group>按id分组去重的方案,在小型XML中正常,但处理200万条数据时无输出,核心原因是流模式下的group-by需要缓存所有分组键和对应条目,200万条数据会耗尽内存,导致转换异常终止。

替代方案:使用XSLT 3.0累加器(xsl:accumulator)

无需分组,通过累加器跟踪已处理的id,流处理时仅处理首次出现的条目,内存占用极低,适合超大型Feed。

修正后的完整XSLT代码

<?xml version="1.0"?>
<xsl:stylesheet
    xmlns:xsl="http://www.w3.org/1999/XSL/Transform" version="3.0">
    <xsl:output method="text" encoding="UTF-8" indent="no" omit-xml-declaration="yes"/>
    <xsl:mode streamable="yes" use-accumulators="seen-ids"/>
    <xsl:strip-space elements="*"/>

    <!-- 累加器:跟踪已处理的id集合 -->
    <xsl:accumulator name="seen-ids" as="xs:string*" initial-value="()" streamable="yes">
        <xsl:accumulator-rule match="item/id" select="if (not(. = $value)) then ($value, string(.)) else $value"/>
    </xsl:accumulator>

    <xsl:variable name="dq" select="'&quot;'"/>
    <xsl:variable name="qcq" select="'&quot;,&quot;'"/>
    <xsl:variable name="lf" select="'&#10;'"/>
    <!-- Static defaults -->
    <xsl:variable name="data_source" select="'HHTestMedia'"/>   

    <!-- 处理items节点下的每个item -->
    <xsl:template match="items">
        <xsl:apply-templates select="item"/>
    </xsl:template>

    <!-- 仅处理首次出现的item -->
    <xsl:template match="item[string(id) not(= accumulator-before('seen-ids'))]">
        <xsl:variable name="currentItem" select="copy-of(.)"/>                    
        <xsl:variable name="p_title" select="substring(replace($currentItem/title,'&quot;',''), 1, 128)"/>
        <xsl:variable name="p_city" select="substring-before($currentItem/location, ',')"/>
        <xsl:variable name="p_state_code" select="substring(normalize-space(substring-after($currentItem/location, ',')), 1, 2)"/>
        <xsl:variable name="p_postal_code" select="$currentItem/postcode"/>
        <xsl:variable name="p_country_code" select="$currentItem/country"/>
        <xsl:variable name="p_company_name" select="replace(substring($currentItem/company, 1, 64), '&quot;', '')"/>                                       

        <xsl:value-of select="concat($dq, $p_title, $qcq, $p_city, $qcq, $p_state_code, $qcq, $p_postal_code, $qcq, $p_country_code, $qcq, $p_company_name, $dq, $lf)"/>
    </xsl:template>

    <!-- 跳过重复的item -->
    <xsl:template match="item"/>
</xsl:stylesheet>

方案说明

  1. 累加器seen-ids:以字符串集合形式存储已处理的id,每次遇到新id就添加到集合中,流处理时增量更新,无需缓存全量数据。
  2. 模板匹配逻辑:
    • 匹配items节点,遍历所有item;
    • 仅处理id不在累加器历史中的item(即首次出现的条目);
    • 重复的item被空模板匹配,直接跳过。
  3. 内存优化:仅保留已处理的id集合,200万条不同id的情况下,内存占用远低于分组方案(分组需缓存所有条目)。

关键注意事项

  • 确保Saxon EE 10.5启用了流处理(已在<xsl:mode>中设置streamable="yes");
  • copy-of(.)仅在需要处理当前item的子节点时使用,避免不必要的节点缓存;
  • 若id是CDATA格式,string(id)会自动提取CDATA中的文本内容,无需额外处理。

内容的提问来源于stack exchange,提问作者Gopinath

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.23 01:32:03