You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Spring Data JPA实现多字段词形兼容的模糊查询

解决方案

方案1:利用数据库全文索引(以MySQL为例)

这种方式依赖数据库原生的全文搜索能力,能高效实现词根/词干匹配,同时支持多字段(question和answer)的联合查询。

步骤1:给数据表添加全文索引

先给questions表的目标字段创建全文索引:

ALTER TABLE questions ADD FULLTEXT INDEX idx_fulltext_qa (question, answer);

步骤2:在Repository中添加自定义查询方法

使用@Query注解编写原生SQL,调用MySQL的MATCH AGAINST功能:

package com.example.kr2db.repository;

import com.example.kr2db.model.Question;
import org.springframework.data.jpa.repository.JpaRepository;
import org.springframework.data.jpa.repository.Query;
import org.springframework.data.repository.query.Param;

import java.util.List;

public interface QuestionRepository extends JpaRepository<Question, Integer> {
    @Query(value = "SELECT * FROM questions WHERE MATCH(question, answer) AGAINST(:keyword IN BOOLEAN MODE)", nativeQuery = true)
    List<Question> findByKeywordInQuestionOrAnswer(@Param("keyword") String keyword);
}

说明:IN BOOLEAN MODE支持灵活匹配规则,MySQL默认的英文分词器会自动处理词干转换,输入rats时能匹配到包含rat的记录。

方案2:使用Hibernate Search(跨数据库全文搜索框架)

如果需要跨数据库兼容,或者更强大的分词、词干分析能力,可以采用Hibernate Search。

步骤1:添加依赖

在pom.xml中引入适配Spring Boot版本的Hibernate Search依赖:

<dependency>
    <groupId>org.hibernate.search</groupId>
    <artifactId>hibernate-search-mapper-orm</artifactId>
    <version>6.2.0.Final</version>
</dependency>
<dependency>
    <groupId>org.hibernate.search</groupId>
    <artifactId>hibernate-search-backend-lucene</artifactId>
    <version>6.2.0.Final</version>
</dependency>

步骤2:给实体类添加全文索引注解

修改Question实体,标记需要被索引的字段:

package com.example.kr2db.model;

import javax.persistence.Entity;
import javax.persistence.GeneratedValue;
import javax.persistence.Id;
import org.hibernate.search.mapper.pojo.mapping.definition.annotation.FullTextField;
import org.hibernate.search.mapper.pojo.mapping.definition.annotation.Indexed;

@Entity
@Indexed // 标记该实体需生成全文索引
public class Question {
    @Id
    @GeneratedValue
    private int id;

    @FullTextField(analyzer = "english") // 使用英文分词器,自动处理词干
    private String question;

    @FullTextField(analyzer = "english")
    private String answer;

    // 构造函数、getter/setter(必须添加getter,Hibernate Search需要访问字段内容)
    public Question() {
    }

    public Question(String question, String answer) {
        this.question = question;
        this.answer = answer;
    }

    public int getId() {
        return id;
    }

    public String getQuestion() {
        return question;
    }

    public void setQuestion(String question) {
        this.question = question;
    }

    public String getAnswer() {
        return answer;
    }

    public void setAnswer(String answer) {
        this.answer = answer;
    }
}

步骤3:在Repository中实现搜索方法

通过Spring Data JPA注入的EntityManager调用Hibernate Search的API:

package com.example.kr2db.repository;

import com.example.kr2db.model.Question;
import org.springframework.data.jpa.repository.JpaRepository;
import org.hibernate.search.mapper.orm.Search;
import org.hibernate.search.mapper.orm.session.SearchSession;
import javax.persistence.EntityManager;
import java.util.List;

public interface QuestionRepository extends JpaRepository<Question, Integer> {

    default List<Question> searchByKeyword(String keyword) {
        SearchSession searchSession = Search.session(entityManager());
        return searchSession.search(Question.class)
                .where(f -> f.match()
                        .fields("question", "answer")
                        .matching(keyword)
                        .analyzer("english"))
                .fetchHits(20); // 可根据需求调整返回结果数量
    }

    EntityManager entityManager(); // Spring Data JPA会自动注入EntityManager实例
}

说明:Hibernate Search的英文分词器会自动将rats解析为词根rat,从而匹配包含该词根的记录。

方案3:简易词干处理+模糊查询(适合小型应用)

如果不想依赖数据库特性或额外框架,可以手动提取关键词的词干,再执行模糊查询,实现成本低但准确性略逊于前两种方案。

步骤1:添加词干分析依赖

引入Apache Lucene的分词工具:

<dependency>
    <groupId>org.apache.lucene</groupId>
    <artifactId>lucene-analyzers-common</artifactId>
    <version>9.7.0</version>
</dependency>

步骤2:在Repository中实现词干处理+查询逻辑

package com.example.kr2db.repository;

import com.example.kr2db.model.Question;
import org.springframework.data.jpa.repository.JpaRepository;
import org.springframework.data.jpa.repository.Query;
import org.springframework.data.repository.query.Param;
import org.apache.lucene.analysis.en.EnglishAnalyzer;
import org.apache.lucene.analysis.tokenattributes.CharTermAttribute;
import java.io.IOException;
import java.io.StringReader;
import java.util.List;

public interface QuestionRepository extends JpaRepository<Question, Integer> {

    @Query("SELECT q FROM Question q WHERE LOWER(q.question) LIKE CONCAT('%', :stem, '%') OR LOWER(q.answer) LIKE CONCAT('%', :stem, '%')")
    List<Question> findByStemInQuestionOrAnswer(@Param("stem") String stem);

    default List<Question> searchByKeyword(String keyword) throws IOException {
        // 提取关键词的词干
        EnglishAnalyzer analyzer = new EnglishAnalyzer();
        String stem = null;
        try (var tokenStream = analyzer.tokenStream(null, new StringReader(keyword))) {
            CharTermAttribute termAttr = tokenStream.addAttribute(CharTermAttribute.class);
            tokenStream.reset();
            if (tokenStream.incrementToken()) {
                stem = termAttr.toString();
            }
            tokenStream.end();
        }
        return stem == null ? List.of() : findByStemInQuestionOrAnswer(stem);
    }
}

说明:先将输入的关键词转换为词根(比如rats→rat),再通过模糊查询匹配包含该词根的字段。

内容的提问来源于stack exchange,提问作者Stasis2

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.15 04:25:31