LanceDB全文搜索为何无法在Top-K结果中匹配精确文本?
LanceDB全文搜索(FTS)结果排序异常问题
尝试使用lancedb执行全文搜索时遇到异常结果:精确匹配的文本无法在小Top-K(如k<500)结果中优先返回,必须设置足够大的limit才能找到。
最小复现代码
# 数据生成 import lancedb import polars as pl from string import ascii_lowercase words = [] # 生成所有双词组合:'aa aa', 'aa ab', ..., 'zz zz' for c1 in ascii_lowercase: for c2 in ascii_lowercase: for c3 in ascii_lowercase: for c4 in ascii_lowercase: words.append(f'{c1}{c2} {c3}{c4}') df = pl.DataFrame({'text':words}).with_row_index('row_id') # 创建本地数据库并写入数据 client = lancedb.connect('local-lancedb') client.create_table('for_testing', df) # 打开表并创建索引 tb = client.open_table('for_testing') tb.create_fts_index('text', remove_stop_words=False, with_position=True, replace=True, stem=False) tb.create_scalar_index('row_id')
现象说明
- 精确匹配SQL查询正常:
tb.search().where('''text = 'aa az' ''').to_polars()
返回结果:
shape: (1, 2) ┌────────┬───────┐ │ row_id ┆ text │ │ --- ┆ --- │ │ u32 ┆ str │ ╞════════╪═══════╡ │ 456975 ┆ zz zz │ └────────┴───────┘
- FTS查询需大limit才能找到精确匹配:
当设置limit=2000时,能在结果中找到精确匹配的'zz zz',且它的_score最高:
tb.search('zz zz', fts_columns='text', query_type='fts').limit(2_000).to_polars()
返回结果片段:
shape: (1_351, 3) ┌────────┬───────┬──────────┐ │ row_id ┆ text ┆ _score │ │ --- ┆ --- ┆ --- │ │ u32 ┆ str ┆ f32 │ ╞════════╪═══════╪══════════╡ │ 456975 ┆ zz zz ┆ 8.0072 │ │ 17575 ┆ az zz ┆ 5.823418 │ │ 18927 ┆ bb zz ┆ 5.823418 │ │ … ┆ … ┆ … │ └────────┴───────┴──────────┘
- 小limit下精确匹配结果丢失:
当limit=1000时,结果中完全找不到'zz zz':
# 丢失正确结果 tb.search('zz zz', fts_columns='text', query_type='fts').limit(1_000).to_polars()
返回结果片段:
shape: (1_000, 3) ┌────────┬───────┬──────────┐ │ row_id ┆ text ┆ _score │ │ --- ┆ --- ┆ --- │ │ u32 ┆ str ┆ f32 │ ╞════════╪═══════╪══════════╡ │ 17575 ┆ az zz ┆ 5.823418 │ │ 18251 ┆ ba zz ┆ 5.823418 │ │ … ┆ … ┆ … │ │ 16899 ┆ ay zz ┆ 5.823418 │ └────────┴───────┴──────────┘
- 其他精确文本的类似问题:
比如查询'ab cd'时,必须设置limit>=298才能找到精确匹配,小limit下过滤后结果为空:
# 仅当limit >=298时才能找到结果 tb.search('ab cd', fts_columns='text', query_type='fts').limit(100).to_polars().filter(text='ab cd')
返回结果:
shape: (0, 3) ┌────────┬──────┬────────┐ │ row_id ┆ text ┆ _score │ │ --- ┆ --- ┆ --- │ │ u32 ┆ str ┆ f32 │ ╞════════╪══════╪════════╡ └────────┴──────┴────────┘
索引配置说明
已使用以下参数创建FTS索引:
tb.create_fts_index('text', remove_stop_words=False, with_position=True, replace=True, stem=False)
疑问:为何FTS无法将'zz zz'这类精确文本查询结果优先排在Top-K(如k<500)结果中?
内容的提问来源于stack exchange,提问作者MKWL
相关产品推荐
相关产品推荐

