You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Solr WordDelimiterFilterFactory搜索字母数字组合及排序异常问题

Solr Relevance Issue: "2 STAGE" Query Returns "4 STAGE" First

Alright, let's diagnose why your Solr search for "2 STAGE" is prioritizing the 4-stage document over the correct 2-stage one. The issue boils down to how your text field's analyzer is processing tokens, and how relevance scoring is calculating matches.

The Root Cause

Looking at your text field type's analyzer chain, the culprit is likely the WordDelimiterFilterFactory configuration:

<filter class="solr.WordDelimiterFilterFactory" 
        protected="protwords.txt" 
        generateWordParts="1" 
        generateNumberParts="1" 
        catenateWords="1" 
        catenateNumbers="0" 
        catenateAll="1" 
        splitOnCaseChange="0" 
        preserveOriginal="1"/>

The catenateAll="1" setting tells Solr to combine all tokens (words and numbers) from the title into a single massive token. For example, "BASECOAT PROCESS 4 STAGE - 90 LINE" gets turned into tokens like basecoatprocess4stage90line alongside individual tokens like 4, stage, etc.

When you search for "2 STAGE", Solr looks for the tokens 2 and stage. The 4-stage document ends up with more overlapping token matches (thanks to that big concatenated token), which inflates its relevance score and pushes it to the top—even though it doesn't have the exact "2 STAGE" phrase.

Fixes to Get the Right Result

1. Adjust the WordDelimiterFilter Settings

First, disable catenateAll since it's creating overly broad tokens that skew relevance. You almost certainly don't need this unless you specifically want to match concatenated phrases (like "basecoatprocess"). Update that line to:

catenateAll="0"

If you don't need combined word tokens (e.g., basecoatprocess), you can also set catenateWords="0"—this keeps tokens more granular and helps with precise matching.

2. Prioritize Phrase Matches

To make exact phrase matches like "2 STAGE" rank higher, you have two options:

  • Encourage quoted queries: Tell users to search with "2 STAGE" (quotes) to force a phrase match. This will prioritize documents where 2 and stage appear adjacent to each other.
  • Default to phrase boosting: Use the edismax query parser with the pf (phrase fields) parameter to automatically boost exact phrase matches. Add this to your request handler configuration:
    <requestHandler name="/select" class="solr.SearchHandler">
      <lst name="defaults">
        <str name="defType">edismax</str>
        <str name="qf">title^2</str> <!-- Give title field more weight overall -->
        <str name="pf">title^5</str> <!-- Boost exact phrase matches in title heavily -->
      </lst>
    </requestHandler>
    

3. Test Tokenization with Solr Admin

Always use the Analysis tab in Solr Admin to verify how your fields are tokenizing. Paste your title text and the query text to see exactly what tokens are generated at index and query time. This will confirm that your WordDelimiterFilter changes are working as expected.

4. Add an Exact Match Boost

If you want to guarantee the 2-stage document ranks first, add a boost query (bq) to your search request that gives extra weight to documents with the exact "2 STAGE" phrase in the title:

bq=title:"2 STAGE"^10

This adds a significant relevance boost to the correct document, ensuring it stays at the top.

Final Steps

After updating your schema, reindex all your documents (since analyzer changes only affect new or reindexed content), then test the "2 STAGE" query again. The 2-stage document should now appear as the top result.

内容的提问来源于stack exchange,提问作者mohit_rocks

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 03:45:20