You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

LanceDB全文搜索为何无法在Top-K结果中匹配精确文本?

LanceDB全文搜索(FTS)结果排序异常问题

尝试使用lancedb执行全文搜索时遇到异常结果:精确匹配的文本无法在小Top-K(如k<500)结果中优先返回,必须设置足够大的limit才能找到。

最小复现代码

# 数据生成
import lancedb 
import polars as pl 
from string import ascii_lowercase

words = [] # 生成所有双词组合:'aa aa', 'aa ab', ..., 'zz zz'
for c1 in ascii_lowercase:
    for c2 in ascii_lowercase:
        for c3 in ascii_lowercase:
            for c4 in ascii_lowercase:
                words.append(f'{c1}{c2} {c3}{c4}')

df = pl.DataFrame({'text':words}).with_row_index('row_id')

# 创建本地数据库并写入数据
client = lancedb.connect('local-lancedb')
client.create_table('for_testing', df)

# 打开表并创建索引
tb = client.open_table('for_testing')
tb.create_fts_index('text', remove_stop_words=False, with_position=True, replace=True, stem=False)
tb.create_scalar_index('row_id')

现象说明

  1. 精确匹配SQL查询正常:
tb.search().where('''text = 'aa az' ''').to_polars()

返回结果:

shape: (1, 2)
┌────────┬───────┐
│ row_id ┆ text  │
│ ---    ┆ ---   │
│ u32    ┆ str   │
╞════════╪═══════╡
│ 456975 ┆ zz zz │
└────────┴───────┘
  1. FTS查询需大limit才能找到精确匹配:
    当设置limit=2000时,能在结果中找到精确匹配的'zz zz',且它的_score最高:
tb.search('zz zz', fts_columns='text', query_type='fts').limit(2_000).to_polars()

返回结果片段:

shape: (1_351, 3)
┌────────┬───────┬──────────┐
│ row_id ┆ text  ┆ _score   │
│ ---    ┆ ---   ┆ ---      │
│ u32    ┆ str   ┆ f32      │
╞════════╪═══════╪══════════╡
│ 456975 ┆ zz zz ┆ 8.0072   │
│ 17575  ┆ az zz ┆ 5.823418 │
│ 18927  ┆ bb zz ┆ 5.823418 │
│ …      ┆ …     ┆ …        │
└────────┴───────┴──────────┘
  1. 小limit下精确匹配结果丢失:
    当limit=1000时,结果中完全找不到'zz zz':
# 丢失正确结果
tb.search('zz zz', fts_columns='text', query_type='fts').limit(1_000).to_polars()

返回结果片段:

shape: (1_000, 3)
┌────────┬───────┬──────────┐
│ row_id ┆ text  ┆ _score   │
│ ---    ┆ ---   ┆ ---      │
│ u32    ┆ str   ┆ f32      │
╞════════╪═══════╪══════════╡
│ 17575  ┆ az zz ┆ 5.823418 │
│ 18251  ┆ ba zz ┆ 5.823418 │
│ …      ┆ …     ┆ …        │
│ 16899  ┆ ay zz ┆ 5.823418 │
└────────┴───────┴──────────┘
  1. 其他精确文本的类似问题:
    比如查询'ab cd'时,必须设置limit>=298才能找到精确匹配,小limit下过滤后结果为空:
# 仅当limit >=298时才能找到结果
tb.search('ab cd', fts_columns='text', query_type='fts').limit(100).to_polars().filter(text='ab cd')

返回结果:

shape: (0, 3)
┌────────┬──────┬────────┐
│ row_id ┆ text ┆ _score │
│ ---    ┆ ---  ┆ ---    │
│ u32    ┆ str  ┆ f32    │
╞════════╪══════╪════════╡
└────────┴──────┴────────┘

索引配置说明

已使用以下参数创建FTS索引:

tb.create_fts_index('text', remove_stop_words=False, with_position=True, replace=True, stem=False)

疑问:为何FTS无法将'zz zz'这类精确文本查询结果优先排在Top-K(如k<500)结果中?


内容的提问来源于stack exchange,提问作者MKWL

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.12 09:24:52