You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Elasticsearch输入数组过滤数组字段:需mentions元素全包含于输入数组

Filter Elasticsearch Documents Where All Array Elements Match Input List

Hey there, I get what you're trying to do here—you need to only keep documents where every single element in the mentions array is present in your input list. The standard terms query or basic bool/must won't work here because those just check if at least one element matches, not all.

Here's the exact solution you need:

Use a Script Query to Validate Full Array Inclusion

The most straightforward way to enforce this "all elements must be in input" rule is with a script query that checks every element in the document's mentions array against your input list.

Here's the query tailored to your sample input (["X", "Y", "K", "L"]):

{
  "query": {
    "bool": {
      "filter": {
        "script": {
          "script": {
            "source": "def allowedTerms = params.allowed; doc['mentions'].values.stream().allMatch(term -> allowedTerms.contains(term))",
            "params": {
              "allowed": ["X", "Y", "K", "L"]
            }
          }
        }
      }
    }
  }
}

How This Works

Let's break down the script logic step by step:

  • doc['mentions'].values grabs the array of mentions values from the current document
  • stream().allMatch(...) runs a validation check against every element in that array
  • allowedTerms.contains(term) verifies each mentions element is present in your input list

If every element passes that check, the document is included in the results.

Why Your Initial Approaches Didn't Work

  • A plain terms query on mentions returns any document with at least one matching element—so it would incorrectly include the yyyy document (since it has K and L, which are in your input, even though M isn't)
  • A bool/must with terms does the same thing—it's designed for partial matches, not full array inclusion

Performance Note

If you're dealing with an extremely large dataset, script queries can be slightly slower than pure DSL queries. But for most use cases, this will work perfectly fine. If you need maximum speed, you could precompute a field like all_mentions_allowed at index time, but that adds extra complexity to your data pipeline.

Testing this with your sample data:

  • {"name": "xxxx", "mentions": ["X", "Y"]} → all elements are allowed, so it's included
  • {"name": "yyyy", "mentions": ["K", "L", "M"]} → M isn't allowed, so it's excluded
  • {"name": "zzz", "mentions": ["X", "L"]} → all elements are allowed, so it's included

Which gives you exactly the expected output you listed.


内容的提问来源于stack exchange,提问作者umit.kas

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 10:39:05