Elasticsearch自动查询功能实现求助:结果排序异常问题
Let's build a solution that meets all your requirements, starting with a proper index setup and a tailored query to ensure correct matching and ranking.
Step 1: Index Setup with Custom Analyzer
First, we'll create an index with a custom analyzer that handles synonym normalization (for BTech variants), word delimiters (to split dots/spaces), and ngrams (for partial matches like "Tec").
PUT /auto_lookup_index { "settings": { "analysis": { "filter": { // Normalize all BTech variants to a single term "btech_synonyms": { "type": "synonym", "synonyms": ["btech, b.tech, b tech, b-tech, BTech"] }, // Split terms on dots and preserve original forms "custom_word_delimiter": { "type": "word_delimiter", "preserve_original": true, "catenate_words": true, "stem_english_possessive": true }, // Generate ngrams for partial match support "custom_ngram": { "type": "ngram", "min_gram": 3, "max_gram": 15 } }, "analyzer": { "auto_lookup_analyzer": { "type": "custom", "tokenizer": "standard", "filter": [ "custom_word_delimiter", "lowercase", "btech_synonyms", "custom_ngram" ] } } } }, "mappings": { "properties": { "title": { "type": "text", "analyzer": "auto_lookup_analyzer", "search_analyzer": "auto_lookup_analyzer" } } } }
Step 2: Index Sample Documents
Let's add your sample data to test the query:
POST /auto_lookup_index/_doc/1 { "title": "BTech in cse" } POST /auto_lookup_index/_doc/2 { "title": "b.tech in computer science" } POST /auto_lookup_index/_doc/3 { "title": "b tech in computer" } POST /auto_lookup_index/_doc/4 { "title": "Technological Advance" } POST /auto_lookup_index/_doc/5 { "title": "Artificial Technology" } POST /auto_lookup_index/_doc/6 { "title": "b.e. in mechanical engineering" }
Step 3: The Query
This query uses function_score to prioritize BTech-related documents, ensuring they rank above unrelated matches like "b.e." when querying BTech variants. It also handles partial matches for terms like "Tec".
Replace {{USER_INPUT}} with your actual query term (e.g., "BTech", "B.Tech", "Tec"):
GET /auto_lookup_index/_search { "query": { "function_score": { // Base query matches all relevant documents "query": { "match": { "title": "{{USER_INPUT}}" } }, // Boost BTech-related documents by 5x to ensure they rank higher "functions": [ { "filter": { "match": { "title": "btech" } }, "boost_factor": 5 } ], // Multiply the base score with the boost factor for relevant docs "boost_mode": "multiply", // Return only top matches (adjust size as needed) "size": 10 } }, // Optional: Highlight matches for clarity "highlight": { "fields": { "title": {} } } }
How It Works
- Synonym Normalization: All BTech variants ("BTech", "B.Tech", "b tech") are converted to "btech" during indexing and searching, ensuring consistent matching.
- Word Delimiter: Splits terms like "B.Tech" into "B" and "Tech" while preserving the original, so both partial and full matches work.
- Ngrams: Generates partial terms (e.g., "tec" from "btech" or "technological") to support queries like "Tec".
- Score Boosting: BTech-related documents get a 5x score boost, so they will always rank above unrelated matches like "b.e." when querying any BTech variant.
Testing the Requirements
- Input "BTech": Returns the three BTech documents (ranked top).
- Input "B.Tech": Same results as above, thanks to synonym normalization.
- Input "Tec": Returns all five relevant documents (tech-related + BTech), with BTech docs prioritized (adjust boost factor if you want them mixed).
- Input "b.tech": BTech documents are ranked first, with "b.e." appearing much lower (or not at all if size is set to 5).
内容的提问来源于stack exchange,提问作者pankaj

