You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Elasticsearch中返回命中的查询词

How to Return Only Matched Query Terms Instead of Full Array in Elasticsearch

Got it, let's tackle this problem: you want to get only the exact student IDs from your query that matched each document, instead of the entire studentIds array stored in the document. Your initial aggregation approach works but feels clunky because you're repeating query parameters in the script—here are two cleaner solutions:

Option 1: Use script_fields for Per-Document Matched Terms

This approach directly computes the intersection between your query's student IDs and each document's studentIds array, returning only the matches for every hit.

GET /es_episode/_search
{
  "_source": false, // Disable full source since we only need matched terms
  "query": {
    "terms": { "studentIds": [300, 400] }
  },
  "script_fields": {
    "matched_studentIds": {
      "script": {
        "source": """
          def queryIds = params.queryIds;
          def docIds = doc['studentIds'];
          return queryIds.stream().filter(id -> docIds.contains(id)).collect(Collectors.toList());
        """,
        "params": {
          "queryIds": [300, 400] // Pass your query IDs here once
        }
      }
    }
  }
}

What this does:

  • The script takes your query ID list as a parameter (no hardcoding in the script logic)
  • It compares each query ID against the document's studentIds values
  • Returns only the IDs that exist in both the query and the document

Example Response:

"hits" : {
  "total" : { "value" : 1, "relation" : "eq" },
  "max_score" : 1.0,
  "hits" : [
    {
      "_index" : "es_episode",
      "_type" : "episode",
      "_id" : "2",
      "_score" : 1.0,
      "fields" : {
        "matched_studentIds" : [ 300 ]
      }
    }
  ]
}

Option 2: Simplified Aggregation with include Parameter

If you need aggregated stats (like which query IDs matched how many episodes), you can skip the bucket_selector script entirely by using the include parameter in the terms aggregation. This filters the aggregation buckets to only your query IDs upfront.

GET /es_episode/_search
{
  "_source": false,
  "query": {
    "terms": { "studentIds": [300, 400] }
  },
  "aggs": {
    "matched_studentIds": {
      "terms": {
        "field": "studentIds",
        "include": [300, 400], // Only aggregate IDs from your query
        "size": 10
      },
      "aggs": {
        "related_episodes": {
          "terms": { "field": "episodeId" }
        }
      }
    }
  }
}

Why this is better than your original aggregation:

  • No need for a script to filter buckets—Elasticsearch handles the filtering directly
  • You only define your query IDs once (in both the query and aggregation include), which is easier to maintain
  • The response will only include buckets for IDs that actually matched (so 300 appears, 400 doesn't since no documents have it)

Which Option to Choose?

  • Use Option 1 if you need to see exactly which query IDs matched each individual document
  • Use Option 2 if you need aggregated data (like count of episodes per matched student ID)

Note on Script Permissions:

If you get script-related errors, make sure your Elasticsearch cluster allows inline scripts. You can enable this in elasticsearch.yml:

script.allowed_types: inline

内容的提问来源于stack exchange,提问作者yyyyyyinstance

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 13:02:49