You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Solr中实现文本片段模糊搜索及NGramFilter应用咨询

Great question! Let's break down your problem and walk through the best solutions tailored to your use case.

First, let's understand why your current setup isn't working: your text_ci field type uses KeywordTokenizerFactory, which treats the entire description value as a single token. When you run a fuzzy search like acomodation~2, Solr calculates the edit distance against the entire string (e.g., "accomodation in paris") instead of just the "accomodation" part. The edit distance between your query and the full string is way larger than 2, so no matches come up.

Now, your idea of using NGramFilterFactory is a solid approach for partial/fuzzy matching, but there's another even better fit for your fixed field format. Let's cover both options:


Option 1: Word-level fuzzy matching (best for your strict format)

Since your description follows a consistent pattern ("Accomodation in {city}"), splitting the text into individual words and applying fuzzy matching to those words is more efficient and precise. Here's how to set it up:

  1. Update your field type to use StandardTokenizerFactory (which splits text into words) instead of KeywordTokenizerFactory:
<fieldType name="text_ci_word" class="solr.TextField" omitNorms="true" sortMissingLast="true">
  <analyzer type="index">
    <tokenizer class="solr.StandardTokenizerFactory"/>
    <filter class="solr.LowerCaseFilterFactory"/>
  </analyzer>
  <analyzer type="query">
    <tokenizer class="solr.StandardTokenizerFactory"/>
    <filter class="solr.LowerCaseFilterFactory"/>
  </analyzer>
</fieldType>
  1. Update your description field to use this new type:
<field name="description" type="text_ci_word" indexed="true" multiValued="false" required="true"/>
  1. Reindex your data, then run your fuzzy search like this:
description:acomodation~2

This works because StandardTokenizer splits "Accomodation in Paris" into three tokens: accomodation, in, paris. The fuzzy search now calculates the edit distance against just the accomodation token—since "acomodation" is only one character off from "accomodation" (missing a 'c'), the edit distance is 1 (well within your ~2 limit), so the match will return.

Pros: Smaller index size, faster searches, more precise matches.
Cons: Only works for full-word fuzzy matches (won't help if you want to search partial fragments like "modation").


Option 2: NGramFilter for partial/fuzzy matching (flexible for any fragment)

If you need to support searches for any substring (e.g., "modation" or "in par") along with spelling errors, NGramFilterFactory is the way to go. This filter splits the text into overlapping character sequences (n-grams) at index time, so even partial or misspelled queries can match.

Here's how to implement it:

  1. Create a new field type with NGramFilterFactory:
<fieldType name="text_ci_ngram" class="solr.TextField" omitNorms="true" sortMissingLast="true">
  <analyzer type="index">
    <tokenizer class="solr.KeywordTokenizerFactory"/> <!-- Keep entire string as one token first -->
    <filter class="solr.LowerCaseFilterFactory"/>
    <!-- Split into 3-15 character n-grams (adjust sizes based on your needs) -->
    <filter class="solr.NGramFilterFactory" minGramSize="3" maxGramSize="15"/>
  </analyzer>
  <analyzer type="query">
    <tokenizer class="solr.KeywordTokenizerFactory"/>
    <filter class="solr.LowerCaseFilterFactory"/>
    <!-- Apply n-gram to query too, so partial queries match -->
    <filter class="solr.NGramFilterFactory" minGramSize="3" maxGramSize="15"/>
  </analyzer>
</fieldType>
  1. Update your description field to use this type and reindex.

  2. Now, a query like description:acomodation will match documents with "Accomodation in {city}" because the n-grams of your query will overlap with the n-grams from the indexed text. You can even drop the ~2 fuzzy modifier if you want, since the n-grams already handle partial matches—though you can still use it for extra spelling tolerance if needed.

Pros: Supports partial substring matches and spelling errors, highly flexible.
Cons: Larger index size (since we store many n-grams), slightly slower searches (a trade-off for flexibility).


Which should you choose?

  • Go with Option 1 if your searches are primarily focused on the "Accomodation" word (or other full words in the field) with spelling errors—it's simpler and more efficient.
  • Go with Option 2 if you need to allow users to search any part of the description string (e.g., searching for "par" to find Paris accommodations).

内容的提问来源于stack exchange,提问作者Petar Yakov

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 07:17:44