You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何仅用XPath 3.1找出出现频次最高的属性值?

问题:找出戏剧XML中台词最多的角色(XPath 3.1实现)

XML示例片段

<play>
    <speech>
        <spkr name="CAL">CAL.</spkr>Et tu mecastor salve, Lysistrata. Sed quid conturbata es?
        exporge frontem, carissima: non enim te decent contracta supercilia.</speech>
    <speech>
        <spkr name="LYS">LYS.</spkr>Sed, ô Calonice, uritur mihi cor, et valde me piget sexus
        nostri, quoniam viri existimant<endnote orig="transcriber" n="1"/> nos esse nequam.</speech>
    <speech>
        <spkr name="CAL">CAL.</spkr>Quippe tales pol sumus.</speech>
</play>

需求说明

需要在oXygen中使用XPath 3.1找出该戏剧里台词最多的角色(示例中应返回CAL)。此前尝试过distinct-values()、count()、max()的组合,也研究过saxon:highest(),但没能实现“去重角色名→统计每个角色的台词次数→取次数最多的角色”的需求。已找到XQuery的循环排序方案,但希望得到简洁优雅的XPath方案,方便后续迁移到XSLT中使用。

XPath 3.1解决方案

利用XPath 3.1的分组和映射功能,可以写出简洁的表达式:

方案1:使用group by分组统计

(for $speaker in //spkr/@name
 group by $name := $speaker
 return map {
     "name": $name,
     "count": count($speaker)
 })[count = max(current()/count)]?name

方案2:先去重再统计

let $all-speakers := //spkr/@name
let $speaker-stats := 
    for $name in distinct-values($all-speakers)
    return map {
        "name": $name,
        "count": count($all-speakers[. = $name])
    }
return $speaker-stats[count = max($speaker-stats?count)]?name

说明

  • 两个方案都会返回所有台词数并列最多的角色;如果只需要返回第一个结果,可在末尾添加[1],比如return $speaker-stats[count = max($speaker-stats?count)]?name[1]
  • 表达式基于XPath 3.1标准,oXygen默认使用的Saxon处理器完全支持该语法,可直接在XPath编辑器或XSLT中使用

内容的提问来源于stack exchange,提问作者haggis78

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.14 06:57:24