如何用XSLT 2.0实现食谱配料单序列与数据库序列的精准匹配?
高效配料序列匹配的XSLT优化方案
核心问题分析
原方案遍历全量3500条数据库记录并调用contains匹配,存在两个关键缺陷:
- 全量遍历效率极低;
contains无法实现精确配料匹配(例如误将"salted"识别为匹配"salt"),且无法统计匹配的配料数量。
以下是针对XSLT 2.0+的高效解决方案:
方案1:利用索引键快速筛选候选节点
通过xsl:key建立配料到对应Ingredients节点的索引,直接定位所有包含至少一个源配料的候选节点,避免全量遍历:
步骤1:定义索引键
<!-- 建立单个配料到包含该配料的Ingredients节点的映射 --> <xsl:key name="ingredient-map" match="Ingredients" use="tokenize(@list, ',\s*')"/>
注:
,\s*用于处理逗号后的空格,避免因空格差异导致的匹配失败。
步骤2:计算匹配度并筛选最优结果
<!-- 先将源配料拆分为去重的序列 --> <xsl:variable name="source-ingredients" select="distinct-values(tokenize(@list, ',\s*'))"/> <!-- 通过索引快速获取所有相关候选节点(自动去重) --> <xsl:variable name="candidates" select="key('ingredient-map', $source-ingredients)"/> <!-- 按匹配配料数量降序排序,取排名第一的节点 --> <xsl:for-each select="$candidates"> <xsl:sort select="count(tokenize(@list, ',\s*')[. = $source-ingredients])" order="descending"/> <xsl:if test="position() = 1"> <!-- 输出匹配度最高的配料序列 --> <xsl:copy-of select="."/> </xsl:if> </xsl:for-each>
方案2:预处理数据库配料集合(适合多批次匹配)
若需多次执行匹配操作,提前将数据库的配料序列转换为标准化集合,避免重复执行tokenize:
步骤1:预处理数据库数据
<!-- 生成包含标准化配料集合的中间节点 --> <xsl:variable name="processed-db"> <xsl:for-each select="$data/Ingredients"> <IngredientEntry original-list="{@list}" ingredients="{tokenize(@list, ',\s*')}"/> </xsl:for-each> </xsl:variable>
步骤2:快速匹配最优结果
<xsl:variable name="source-set" select="tokenize(@list, ',\s*')"/> <xsl:for-each select="$processed-db/IngredientEntry"> <!-- 计算当前条目与源配料的匹配数量 --> <xsl:variable name="match-count" select="count(tokenize(@ingredients, ' ')[. = $source-set])"/> <xsl:sort select="$match-count" order="descending"/> <xsl:if test="position() = 1"> <Ingredients list="{@original-list}"/> </xsl:if> </xsl:for-each>
方案3:用XPath集合操作直接计算最优结果
无需循环,通过XPath内置函数直接计算最大匹配数并筛选对应节点:
<xsl:variable name="source-ingredients" select="tokenize(@list, ',\s*')"/> <!-- 计算所有数据库节点的最大匹配数量 --> <xsl:variable name="max-match-count" select="max($data/Ingredients/count(tokenize(@list, ',\s*')[. = $source-ingredients]))"/> <!-- 输出所有匹配数量等于最大值的节点(若只需一个则加[1]) --> <xsl:copy-of select="$data/Ingredients[count(tokenize(@list, ',\s*')[. = $source-ingredients]) = $max-match-count]"/>
额外优化建议
- 大小写兼容:若需忽略大小写匹配,将
. = $source-ingredients替换为lower-case(.) = lower-case($source-ingredients); - 去重处理:对源配料和数据库配料执行
distinct-values,避免重复配料影响匹配计数; - XSLT 3.0优化:若使用XSLT 3.0,可利用
xsl:accumulator或流处理进一步提升大数据集的处理效率。
内容的提问来源于stack exchange,提问作者Michael Friedman
相关产品推荐
相关产品推荐

