You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Elasticsearch自动查询功能实现求助:结果排序异常问题

Elasticsearch Query for Auto-Lookup with Prioritized BTech Results

Let's build a solution that meets all your requirements, starting with a proper index setup and a tailored query to ensure correct matching and ranking.

Step 1: Index Setup with Custom Analyzer

First, we'll create an index with a custom analyzer that handles synonym normalization (for BTech variants), word delimiters (to split dots/spaces), and ngrams (for partial matches like "Tec").

PUT /auto_lookup_index
{
  "settings": {
    "analysis": {
      "filter": {
        // Normalize all BTech variants to a single term
        "btech_synonyms": {
          "type": "synonym",
          "synonyms": ["btech, b.tech, b tech, b-tech, BTech"]
        },
        // Split terms on dots and preserve original forms
        "custom_word_delimiter": {
          "type": "word_delimiter",
          "preserve_original": true,
          "catenate_words": true,
          "stem_english_possessive": true
        },
        // Generate ngrams for partial match support
        "custom_ngram": {
          "type": "ngram",
          "min_gram": 3,
          "max_gram": 15
        }
      },
      "analyzer": {
        "auto_lookup_analyzer": {
          "type": "custom",
          "tokenizer": "standard",
          "filter": [
            "custom_word_delimiter",
            "lowercase",
            "btech_synonyms",
            "custom_ngram"
          ]
        }
      }
    }
  },
  "mappings": {
    "properties": {
      "title": {
        "type": "text",
        "analyzer": "auto_lookup_analyzer",
        "search_analyzer": "auto_lookup_analyzer"
      }
    }
  }
}

Step 2: Index Sample Documents

Let's add your sample data to test the query:

POST /auto_lookup_index/_doc/1
{ "title": "BTech in cse" }

POST /auto_lookup_index/_doc/2
{ "title": "b.tech in computer science" }

POST /auto_lookup_index/_doc/3
{ "title": "b tech in computer" }

POST /auto_lookup_index/_doc/4
{ "title": "Technological Advance" }

POST /auto_lookup_index/_doc/5
{ "title": "Artificial Technology" }

POST /auto_lookup_index/_doc/6
{ "title": "b.e. in mechanical engineering" }

Step 3: The Query

This query uses function_score to prioritize BTech-related documents, ensuring they rank above unrelated matches like "b.e." when querying BTech variants. It also handles partial matches for terms like "Tec".

Replace {{USER_INPUT}} with your actual query term (e.g., "BTech", "B.Tech", "Tec"):

GET /auto_lookup_index/_search
{
  "query": {
    "function_score": {
      // Base query matches all relevant documents
      "query": {
        "match": {
          "title": "{{USER_INPUT}}"
        }
      },
      // Boost BTech-related documents by 5x to ensure they rank higher
      "functions": [
        {
          "filter": {
            "match": {
              "title": "btech"
            }
          },
          "boost_factor": 5
        }
      ],
      // Multiply the base score with the boost factor for relevant docs
      "boost_mode": "multiply",
      // Return only top matches (adjust size as needed)
      "size": 10
    }
  },
  // Optional: Highlight matches for clarity
  "highlight": {
    "fields": {
      "title": {}
    }
  }
}

How It Works

  1. Synonym Normalization: All BTech variants ("BTech", "B.Tech", "b tech") are converted to "btech" during indexing and searching, ensuring consistent matching.
  2. Word Delimiter: Splits terms like "B.Tech" into "B" and "Tech" while preserving the original, so both partial and full matches work.
  3. Ngrams: Generates partial terms (e.g., "tec" from "btech" or "technological") to support queries like "Tec".
  4. Score Boosting: BTech-related documents get a 5x score boost, so they will always rank above unrelated matches like "b.e." when querying any BTech variant.

Testing the Requirements

  • Input "BTech": Returns the three BTech documents (ranked top).
  • Input "B.Tech": Same results as above, thanks to synonym normalization.
  • Input "Tec": Returns all five relevant documents (tech-related + BTech), with BTech docs prioritized (adjust boost factor if you want them mixed).
  • Input "b.tech": BTech documents are ranked first, with "b.e." appearing much lower (or not at all if size is set to 5).

内容的提问来源于stack exchange,提问作者pankaj

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 07:45:42