You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Elasticsearch构建精准查询:动态获取内容的优化咨询

问题描述

我正在构建一个基于Elasticsearch的应用,流程为检索文档并拆分为片段,当前使用查询:query = "what is your pricing?",随后将文本与该查询传入以下函数:

def construct_enhanced_query(user_query):
    keywords = extract_keywords(user_query)
    should_clauses = [{"match": {"content": {"query": keyword, "boost": 2}}} for keyword in keywords]

    query = {
        "query": {
            "bool": {
                "should": should_clauses,
                "minimum_should_match": 1
            }
        }
    }
    return query

其中extract_keywords()函数实现如下:

def extract_keywords(text):
    doc = nlp(text)
    keywords = set([chunk.text for chunk in doc.noun_chunks] + [ent.text for ent in doc.ents])
    return keywords

现有文档片段包含:

  • 短导航片段:ABOUT PRICING CONTACT Learn more
  • 含定价详情的长片段:PRICING PLANS Private Lessons $150 PER STUDENT 30 minute lessons 1 instructor and 1 student (1:1) Book Now Semi-Private $130 /PER STUDENT 30 minute lessons 1 instructor and 2 students (1:2)

期望检索到$150 PER STUDENT及其相关信息,但当前仅返回含"PRICING"的短片段。问题由当前查询的关键词搜索及权重设置导致,因用户查询不固定无法硬编码,现寻求动态获取相关信息的优化建议。

优化建议

1. 改进关键词提取逻辑

当前extract_keywords仅提取名词短语和实体,对"what is your pricing?"只会得到pricing单个关键词,匹配范围过窄。可做以下调整:

  • 扩展实词范围:保留查询中的所有实词(名词、动词、形容词),避免仅依赖名词短语,比如从查询中提取pricing之外,补充与"询问详情"相关的语义词;
  • 动态同义词扩展:用WordNet或Elasticsearch同义词过滤器,给核心关键词自动扩展同义词,比如pricing扩展为["pricing", "price", "cost", "rate"],扩大匹配覆盖面。

2. 调整Elasticsearch查询结构

短片段因包含pricing且长度短,TF/IDF评分更高,需优化查询结构提升长详情片段的权重:

  • 加入短语匹配子句:针对核心关键词pricing添加match_phrase,优先匹配上下文连贯的片段(如长片段的PRICING PLANS):
    {
      "query": {
        "bool": {
          "should": [
            {"match": {"content": {"query": "pricing", "boost": 2}}},
            {"match_phrase": {"content": {"query": "pricing", "slop": 10, "boost": 3}}}
          ],
          "minimum_should_match": 1
        }
      }
    }
    
  • 用function_score按长度加权:给包含关键词的长片段额外加分,抵消短片段的评分优势:
    {
      "query": {
        "function_score": {
          "query": {
            "bool": {
              "should": [{"match": {"content": {"query": keyword, "boost": 2}}} for keyword in keywords]
            }
          },
          "functions": [
            {
              "script_score": {
                "script": {
                  "source": "Math.log(doc['content'].value.length() + 1)"
                }
              }
            }
          ],
          "boost_mode": "multiply"
        }
      }
    }
    

3. 利用上下文关联查询

  • 使用more_like_this关联同主题片段:先执行原查询拿到初始匹配的短片段,再用该片段内容构建more_like_this查询,获取同主题的长详情片段:
    def construct_mlt_query(initial_hits):
        mlt_query = {
            "query": {
                "more_like_this": {
                    "fields": ["content"],
                    "like": [hit['_source']['content'] for hit in initial_hits],
                    "min_term_freq": 1,
                    "max_query_terms": 10
                }
            }
        }
        return mlt_query
    

4. 优化片段索引策略

  • 标记片段类型:拆分片段时,自动识别导航型短片段和详情型长片段,添加type字段(如type: "navigation"/type: "detail"),查询时给detail类型加权:
    {
      "query": {
        "bool": {
          "should": [
            {"match": {"content": {"query": "pricing", "boost": 2}}},
            {"term": {"type": {"value": "detail", "boost": 3}}}
          ],
          "minimum_should_match": 1
        }
      }
    }
    
  • 单独索引价格字段:用正则或NLP实体识别提取片段中的价格信息,存入price字段,查询时优先匹配含该字段的片段:
    {
      "query": {
        "bool": {
          "should": [
            {"match": {"content": {"query": "pricing", "boost": 2}}},
            {"exists": {"field": "price", "boost": 3}}
          ]
        }
      }
    }
    

5. 动态调整关键词权重

针对不同类型的关键词设置差异化boost,提升核心词的权重:

def extract_keywords_with_weights(text):
    doc = nlp(text)
    keywords = []
    for chunk in doc.noun_chunks:
        # 核心名词短语权重设为3
        keywords.append((chunk.text, 3))
    for ent in doc.ents:
        # 实体权重设为2
        keywords.append((ent.text, 2))
    # 去重并保留最高权重
    unique_keywords = {}
    for kw, weight in keywords:
        if kw not in unique_keywords or weight > unique_keywords[kw]:
            unique_keywords[kw] = weight
    return unique_keywords.items()

def construct_enhanced_query(user_query):
    keyword_weights = extract_keywords_with_weights(user_query)
    should_clauses = [{"match": {"content": {"query": keyword, "boost": weight}}} for keyword, weight in keyword_weights]

    query = {
        "query": {
            "bool": {
                "should": should_clauses,
                "minimum_should_match": 1
            }
        }
    }
    return query

内容的提问来源于stack exchange,提问作者bobthebuilder

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.28 17:04:56