You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过Vespa统计字符串或数组中关键词的出现次数?

当前你用matchCount得到的结果始终为1,是因为matchCount的作用是统计匹配查询条件的字段/数组元素的数量,而非字段内关键词的出现次数:

  • 对于Name字符串字段:只要字段整体匹配了matches正则,matchCount(Name)就返回1,和字段内关键词出现次数无关;
  • 对于NameArray数组字段:你的查询是针对整个数组字段的正则匹配,而非数组单个元素,因此matchCount(NameArray)也返回1。
一、统计完整关键词(如"cat",大小写不敏感)的出现次数

1. 针对数组字段NameArray

修改Schema配置

给NameArray添加小写归一化,确保大小写不敏感匹配:

field NameArray type array<string> {
    indexing: summary | attribute
    normalizer: lowercase # 索引时将元素转为小写
}

查询方式

使用contains匹配数组元素,而非matches,同时指定排名配置searchByName1:

{
    "hits": 150,
    "ranking": {
        "profile": "searchByName1"
    },
    "offset": 0,
    "yql": "select * from search where contains(NameArray, 'cat')"
}

此时matchCount(NameArray)会返回数组中匹配"cat"(大小写不敏感)的元素数量,示例文档会返回3(两个Cat+一个сat,注意若с是西里尔字母则无法匹配,需统一为拉丁字母)。

2. 针对字符串字段Name

修改Schema配置

将Name字段类型改为text,配置分词、归一化,让Vespa能识别字段内的独立词汇:

field Name type text {
    indexing: summary | index # 需要index才能统计词频
    index: enable-bm25
    normalizer: lowercase # 实现大小写不敏感
    tokenizer: whitespace # 按空格分词,可根据需求替换为其他分词器
}

使用termFrequency统计词频

修改排名配置,用termFrequency获取关键词出现次数:

rank-profile searchByName {
    first-phase {
        expression: termFrequency(Name, "cat")
    }
}

查询方式

直接查询匹配的文档:

{
    "hits": 150,
    "ranking": {
        "profile": "searchByName"
    },
    "offset": 0,
    "yql": "select * from search where contains(Name, 'cat')"
}

此时返回的relevance就是"cat"在Name字段中的出现次数(示例文档会返回3)。

二、统计子串(如"ca")的出现次数

子串统计无法通过常规分词实现,可选择以下两种方式:

1. 索引层方案(性能优,适合大数据量)

给Name字段配置n-gram分词器,提前索引所有2字符的子串:

field Name type text {
    indexing: summary | index
    index: enable-bm25
    normalizer: lowercase
    tokenizer: n-gram(2,2) # 生成所有连续2字符的子串
}

然后用termFrequency(Name, "ca")统计子串出现次数,排名配置和查询方式同完整词统计。

2. 查询层方案(无需改Schema,适合小数据量)

使用Vespa的脚本字段,在查询时通过JavaScript计算子串出现次数:

{
    "hits": 150,
    "offset": 0,
    "yql": "select *, script(\"function countSubstr(s, sub) { let cnt=0, pos=0; const lowerS = s.toLowerCase(), lowerSub = sub.toLowerCase(); while ((pos = lowerS.indexOf(lowerSub, pos)) !== -1) { cnt++; pos += lowerSub.length; } return cnt; }\", Name, \"ca\") as ca_count from search where Name matches '(?i).*ca.*'"
}

查询结果中会新增ca_count字段,值为"ca"的出现次数。

内容的提问来源于stack exchange,提问作者igor tarashchuk

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.10 09:47:04