Typesense按category前缀过滤失效问题求助
在Perl项目中使用Typesense(依赖Search::Typesense和Search::Typesense::Collection模块),创建的集合schema里category字段类型为string且开启了facet。精确过滤filter_by => 'category:=animations/2d/animals/birds'能正常工作,但使用前缀过滤filter_by => 'category:[prefix="animations/2d"]'时得不到预期结果,请问问题出在哪里?
相关schema代码:
my $collection_schema = { name => 'images-test', fields => [ { name => 'title', type => 'string', index => \1, facet => \0 }, { name => 'description', type => 'string', index => \1, facet => \0 }, { name => 'category', type => 'string', index => \1, facet => \1 }, { name => 'keywords', type => 'string', index => \1, facet => \1 }, { name => 'download_count', type => 'int32', index => \1, sort => \1 }, { name => 'has_dxf', type => 'bool', index => \1 }, { name => 'has_eps', type => 'bool', index => \1 }, { name => 'has_fla', type => 'bool', index => \1 }, { name => 'has_gif', type => 'bool', index => \1 }, { name => 'has_jpg', type => 'bool', index => \1 }, { name => 'has_pdf', type => 'bool', index => \1 }, { name => 'has_png', type => 'bool', index => \1 }, { name => 'has_svg', type => 'bool', index => \1 }, { name => 'has_swf', type => 'bool', index => \1 }, { name => 'has_wmf', type => 'bool', index => \1 }, { name => 'add_date', type => 'int64', index => \1, facet => \1, "sort" => \1 }, { name => 'colors', type => 'string[]', index => \1, facet => \1 }, { name => "black_white", type => "bool", index => \1 }, ], default_sorting_field => 'download_count' };
示例category值:animations/2d/animals/fish and amphibians(中文:动画/2D/动物/鱼类与两栖类)
核心原因
Typesense的前缀过滤[prefix="..."]基于字段的分词结果匹配,普通string类型的category字段会被默认分词器按/、空格等分隔符拆分成多个独立词(比如animations、2d、animals)。当你用prefix="animations/2d"时,分词结果中不存在以该字符串开头的词,自然匹配不到结果。
开启facet仅用于分组统计,不会改变字段的分词和索引逻辑。
解决方法
方法1:开启infix索引并使用范围过滤
修改schema中category字段的定义,添加infix: \1,让Typesense生成包含所有子字符串的索引:
{ name => 'category', type => 'string', index => \1, facet => \1, infix => \1 },
之后用范围过滤实现前缀匹配效果:
filter_by => 'category:>=animations/2d AND category:<animations/2e'
利用字符串字典序,animations/2e是animations/2d之后的第一个前缀边界,可匹配所有以animations/2d开头的category值。
方法2:设置空的token_separators
将category字段的分词分隔符设为空数组,让整个路径被当作单个词索引:
{ name => 'category', type => 'string', index => \1, facet => \1, token_separators => [] },
此时原前缀过滤语法category:[prefix="animations/2d"]即可正常工作,因为整个路径会被作为单个词进行前缀匹配。
方法3:拆分层级字段
如果分类是多层级结构,可拆分为category_level1、category_level2等独立字段,便于层级过滤,同时避免路径字符串的匹配问题。
内容的提问来源于stack exchange,提问作者Andrew Newby

