Solr WordDelimiterFilterFactory搜索字母数字组合及排序异常问题
Alright, let's diagnose why your Solr search for "2 STAGE" is prioritizing the 4-stage document over the correct 2-stage one. The issue boils down to how your text field's analyzer is processing tokens, and how relevance scoring is calculating matches.
The Root Cause
Looking at your text field type's analyzer chain, the culprit is likely the WordDelimiterFilterFactory configuration:
<filter class="solr.WordDelimiterFilterFactory" protected="protwords.txt" generateWordParts="1" generateNumberParts="1" catenateWords="1" catenateNumbers="0" catenateAll="1" splitOnCaseChange="0" preserveOriginal="1"/>
The catenateAll="1" setting tells Solr to combine all tokens (words and numbers) from the title into a single massive token. For example, "BASECOAT PROCESS 4 STAGE - 90 LINE" gets turned into tokens like basecoatprocess4stage90line alongside individual tokens like 4, stage, etc.
When you search for "2 STAGE", Solr looks for the tokens 2 and stage. The 4-stage document ends up with more overlapping token matches (thanks to that big concatenated token), which inflates its relevance score and pushes it to the top—even though it doesn't have the exact "2 STAGE" phrase.
Fixes to Get the Right Result
1. Adjust the WordDelimiterFilter Settings
First, disable catenateAll since it's creating overly broad tokens that skew relevance. You almost certainly don't need this unless you specifically want to match concatenated phrases (like "basecoatprocess"). Update that line to:
catenateAll="0"
If you don't need combined word tokens (e.g., basecoatprocess), you can also set catenateWords="0"—this keeps tokens more granular and helps with precise matching.
2. Prioritize Phrase Matches
To make exact phrase matches like "2 STAGE" rank higher, you have two options:
- Encourage quoted queries: Tell users to search with
"2 STAGE"(quotes) to force a phrase match. This will prioritize documents where2andstageappear adjacent to each other. - Default to phrase boosting: Use the
edismaxquery parser with thepf(phrase fields) parameter to automatically boost exact phrase matches. Add this to your request handler configuration:<requestHandler name="/select" class="solr.SearchHandler"> <lst name="defaults"> <str name="defType">edismax</str> <str name="qf">title^2</str> <!-- Give title field more weight overall --> <str name="pf">title^5</str> <!-- Boost exact phrase matches in title heavily --> </lst> </requestHandler>
3. Test Tokenization with Solr Admin
Always use the Analysis tab in Solr Admin to verify how your fields are tokenizing. Paste your title text and the query text to see exactly what tokens are generated at index and query time. This will confirm that your WordDelimiterFilter changes are working as expected.
4. Add an Exact Match Boost
If you want to guarantee the 2-stage document ranks first, add a boost query (bq) to your search request that gives extra weight to documents with the exact "2 STAGE" phrase in the title:
bq=title:"2 STAGE"^10
This adds a significant relevance boost to the correct document, ensuring it stays at the top.
Final Steps
After updating your schema, reindex all your documents (since analyzer changes only affect new or reindexed content), then test the "2 STAGE" query again. The 2-stage document should now appear as the top result.
内容的提问来源于stack exchange,提问作者mohit_rocks

