You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何解决Hibernate Search+Lucene无法搜索this、not等停用词的问题?

解决Hibernate Search无法搜索"this""not"等停用词的问题

问题根源

你遇到的情况是Lucene默认英文分词器(如StandardAnalyzer)会自动过滤停用词——像"this""is""not"这类高频、无明确业务语义的词汇会被跳过,不加入索引,自然无法通过这些词搜索到结果。

解决方案

1. 自定义分词器,控制停用词过滤

这是最彻底的解决方式,需要修改索引配置并重新索引数据:

方式一:完全禁用停用词过滤

在实体类的搜索字段上指定自定义分词器:

@FullTextField(analyzer = "no-stopwords-analyzer")
private String productName;

然后创建分词器配置类(适配Hibernate Search 6.x,对应Spring Boot 2.7.x的版本):

import org.hibernate.search.backend.lucene.analysis.LuceneAnalysisConfigurationContext;
import org.hibernate.search.backend.lucene.analysis.LuceneAnalysisConfigurer;

public class CustomAnalysisConfigurer implements LuceneAnalysisConfigurer {
    @Override
    public void configure(LuceneAnalysisConfigurationContext context) {
        // 定义无停用词的分词器:仅做标准分词、小写转换,跳过停用词过滤
        context.analyzer("no-stopwords-analyzer").custom()
            .tokenizer("standard")
            .tokenFilter("lowercase");
            // 若需要词干提取可添加.tokenFilter("porterstem"),不需要则省略
    }
}

在application.properties中指定配置类:

hibernate.search.backend.analysis.configurer=com.yourpackage.CustomAnalysisConfigurer

方式二:使用自定义停用词列表

如果你仍想过滤部分停用词,但保留"this""not",可以创建自定义停用词文件custom-stopwords.txt(放在src/main/resources下),只写入需要过滤的词:

a
an
the

然后修改分词器配置:

@Override
public void configure(LuceneAnalysisConfigurationContext context) {
    context.analyzer("custom-stopwords-analyzer").custom()
        .tokenizer("standard")
        .tokenFilter("lowercase")
        .tokenFilter("stop")
            .param("words", "custom-stopwords.txt")
            .param("ignoreCase", "true");
}

2. 重新索引已有数据

修改分词器配置后,必须对现有数据重新索引,确保旧数据使用新规则生成索引:

SearchSession searchSession = Search.session(entityManager);
searchSession.massIndexer(Product.class)
    .startAndWait();

3. 查询时的临时适配(不推荐)

如果暂时无法重新索引,可在查询时指定无停用词的分词器处理查询语句,但仅当索引中存在对应词汇时有效——如果旧索引已经过滤了停用词,这种方式依然搜不到结果:

SearchSession searchSession = Search.session(entityManager);
List<Product> results = searchSession.search(Product.class)
    .where(f -> f.match()
        .field("productName")
        .matching("This Is Not Milk")
        .analyzer("no-stopwords-analyzer"))
    .fetchHits(20);

注意事项

  • 完全禁用停用词会增大索引体积、降低搜索效率,建议优先选择自定义停用词列表的方案,只过滤无业务价值的词汇。
  • 若你使用的是Hibernate Search 5.x(对应Hibernate 5的旧版本),需用@AnalyzerDef注解定义分词器:
    @AnalyzerDef(name = "no-stopwords-analyzer",
        tokenizer = @TokenizerDef(factory = StandardTokenizerFactory.class),
        filters = {
            @TokenFilterDef(factory = LowerCaseFilterFactory.class)
            // 不要添加StopFilterFactory
        })
    
    字段上标注@Analyzer(definition = "no-stopwords-analyzer")即可。

内容的提问来源于stack exchange,提问作者user1034461

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 00:25:16