You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从tsvector获取源文档中词位的实际起始位置?

PostgreSQL tsvector词位位置映射到源文档的问题

示例场景

执行以下SQL生成tsvector并展开:

select
    *
from
    unnest(to_tsvector('english', 'something wide this more wider and wider social-economy wide somethings'))

得到的词位与位置对应表:

lexemepositions
economi10
social9
social-economi8
someth1,12
wide2,11
wider5,7

PostgreSQL文档提到“位置通常表示源词在文档中的位置”,但这里的“位置”并非源文档中的起始符号索引,而是tsvector处理后词位的顺序索引,这让使用者产生困惑。

需求目标

需要在客户端将tsvector中的词位映射回源文档,实现类似PostgreSQL的高亮效果(不使用ts_headline),最终效果如下:

something wide this more wider and wider social-economy wide somethings

尝试方案与问题

最初在C#中尝试按空格和连字符拆分源文档为令牌,通过string.StartsWith匹配词位,代码如下:

var source = "something wide this more wider and wider social-economy wide somethings";
source
    .Split(new[] { ' ', '-' }, StringSplitOptions.RemoveEmptyEntries | StringSplitOptions.TrimEntries)
    .Select((w, i) => new
    {
        Word = w,
        Pos = i + 1,
    })
    .OrderBy(w => w.Word)
;

拆分后得到的词与位置表:

lexemepositions
and6
economy9
more4
social8
something1
somethings11
this3
wide2
wide10
wider5
wider7

但因为tsvector的词干提取(比如lexemeeconomi对应源词economy)、同义词等处理逻辑,导致词位与拆分后的源词无法直接匹配,位置对应关系也不吻合。现在需要解决的核心问题是:如何获取tsvector词位在源文档中的实际符号位置?

内容的提问来源于stack exchange,提问作者Kasbolat Kumakhov

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.10 21:40:18