如何在Elasticsearch中返回命中的查询词
Got it, let's tackle this problem: you want to get only the exact student IDs from your query that matched each document, instead of the entire studentIds array stored in the document. Your initial aggregation approach works but feels clunky because you're repeating query parameters in the script—here are two cleaner solutions:
Option 1: Use script_fields for Per-Document Matched Terms
This approach directly computes the intersection between your query's student IDs and each document's studentIds array, returning only the matches for every hit.
GET /es_episode/_search { "_source": false, // Disable full source since we only need matched terms "query": { "terms": { "studentIds": [300, 400] } }, "script_fields": { "matched_studentIds": { "script": { "source": """ def queryIds = params.queryIds; def docIds = doc['studentIds']; return queryIds.stream().filter(id -> docIds.contains(id)).collect(Collectors.toList()); """, "params": { "queryIds": [300, 400] // Pass your query IDs here once } } } } }
What this does:
- The script takes your query ID list as a parameter (no hardcoding in the script logic)
- It compares each query ID against the document's
studentIdsvalues - Returns only the IDs that exist in both the query and the document
Example Response:
"hits" : { "total" : { "value" : 1, "relation" : "eq" }, "max_score" : 1.0, "hits" : [ { "_index" : "es_episode", "_type" : "episode", "_id" : "2", "_score" : 1.0, "fields" : { "matched_studentIds" : [ 300 ] } } ] }
Option 2: Simplified Aggregation with include Parameter
If you need aggregated stats (like which query IDs matched how many episodes), you can skip the bucket_selector script entirely by using the include parameter in the terms aggregation. This filters the aggregation buckets to only your query IDs upfront.
GET /es_episode/_search { "_source": false, "query": { "terms": { "studentIds": [300, 400] } }, "aggs": { "matched_studentIds": { "terms": { "field": "studentIds", "include": [300, 400], // Only aggregate IDs from your query "size": 10 }, "aggs": { "related_episodes": { "terms": { "field": "episodeId" } } } } } }
Why this is better than your original aggregation:
- No need for a script to filter buckets—Elasticsearch handles the filtering directly
- You only define your query IDs once (in both the query and aggregation
include), which is easier to maintain - The response will only include buckets for IDs that actually matched (so 300 appears, 400 doesn't since no documents have it)
Which Option to Choose?
- Use Option 1 if you need to see exactly which query IDs matched each individual document
- Use Option 2 if you need aggregated data (like count of episodes per matched student ID)
Note on Script Permissions:
If you get script-related errors, make sure your Elasticsearch cluster allows inline scripts. You can enable this in elasticsearch.yml:
script.allowed_types: inline
内容的提问来源于stack exchange,提问作者yyyyyyinstance

