You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何动态将字符串转换为Elasticsearch查询结构?代码问题修复

问题:动态构建Elasticsearch查询体不符合预期需求

我需要将字符串'doc and specia not foo'动态传入search_body,规则如下:

  • 初始词放入must数组
  • and后的词放入should数组
  • not后的词放入must_not数组

但当前代码输出不符合预期,以下是我的代码、实际输出及期望输出:

我的代码

search_string = 'doc and specia not foo'
search_terms = search_string.split()
search_body = {
    "query": {
        "bool": {
            "must": [],
            "should": [],
            "must_not": []
        }
    }
}

# Process each search term
for i, term in enumerate(search_terms):
    if term.lower() == "and" or term.lower() == "or":
        continue  # Skip the "and" and "or" keywords

    # Get the next search term if it exists
    if i < len(search_terms) - 1:
        next_term = search_terms[i + 1]

        if next_term.lower() == "not":
            # Handle "not" operator
            search_body['query']['bool']['must_not'].append({"match": {"your_field_name": term}})
        else:
            # Handle other terms
            operator = "must"
            if term.lower() == "or":
                operator = "should"

            search_body['query']['bool'][operator].append({"match": {"your_field_name": next_term}})
        continue

    # Handle the last search term
    search_body['query']['bool']['must'].append({"match": {"your_field_name": term}})

print(search_body)

实际输出

{
  'query': {
    'bool': {
      'must': [
        {
          'match': {
            'your_field_name': 'and'
          }
        },
        {
          'match': {
            'your_field_name': 'foo'
          }
        }
      ],
      'should': [
        
      ],
      'must_not': [
        {
          'match': {
            'your_field_name': 'specia'
          }
        }
      ]
    }
  }
}

期望输出

{
    "query": {
        "bool": {
            "must": [
                {"match": {"your_field_name": "doc"}}
            ],
            "should": [
                {"match": {"your_field_name": "specia"}}
            ],
            "must_not": [
                {"match": {"your_field_name": "foo"}}
            ]
        }
    }
}

修正方案

原代码逻辑存在多处错误:

  1. 错误地将not前面的词放入must_not,实际应该是not后面的词才是需要排除的内容
  2. 处理and的逻辑混乱,错误地将后续词的判断逻辑搞反
  3. 循环中对关键词的跳过和后续词的处理逻辑冲突,导致无关内容被加入must数组

修正后的代码通过跟踪当前操作符的方式处理,逻辑更清晰:

search_string = 'doc and specia not foo'
search_terms = search_string.split()
search_body = {
    "query": {
        "bool": {
            "must": [],
            "should": [],
            "must_not": []
        }
    }
}

current_operator = "must"  # 默认初始操作符是must

i = 0
while i < len(search_terms):
    term = search_terms[i].lower()
    if term == "and":
        current_operator = "should"
        i += 1
        continue
    elif term == "not":
        current_operator = "must_not"
        i += 1
        continue
    # 处理当前词,放入对应操作符的数组
    search_body['query']['bool'][current_operator].append({
        "match": {"your_field_name": search_terms[i]}
    })
    # 重置默认操作符为must,避免后续词继承之前的should/must_not
    current_operator = "must"
    i += 1

print(search_body)

代码说明

  • 用while循环替代for循环,更灵活地控制索引跳转,避免重复处理关键词
  • 通过current_operator变量跟踪当前应该使用的操作符,初始为must
  • 遇到and时,将操作符设为should,跳过关键词后处理下一个词
  • 遇到not时,将操作符设为must_not,跳过关键词后处理下一个词
  • 每个词处理完成后,重置操作符为must,确保后续默认词回到must数组

运行修正后的代码即可得到期望的输出结果。

内容的提问来源于stack exchange,提问作者Mak Hus

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.16 16:37:27