Elasticsearch按指定项查询下一项及出现次数的方案咨询
Hey there! Let's work through this problem together. The goal is to take an input item name, find the next item in each document (based on the order field being exactly one higher than the target item's order), then count how many times each next item appears across all documents. I'll cover both nested and non-nested mapping scenarios since you mentioned itemArray could be either.
First: Mapping Setup
Nested Mapping (Recommended)
Nested types are the safer choice here because they preserve the relationship between each name and its corresponding order (Elasticsearch flattens regular arrays by default, which can lead to accidental cross-matches if you're not careful). Here's how to define the mapping:
PUT /items_index { "mappings": { "properties": { "id": { "type": "integer" }, "itemArray": { "type": "nested", "properties": { "name": { "type": "keyword" }, // Use keyword for exact matches "order": { "type": "integer" } } } } } }
Non-Nested Mapping
If you need to use a non-nested array, the mapping is simpler, but note that the order of elements in the name and order arrays will be preserved (so we can match them by index):
PUT /items_index { "mappings": { "properties": { "id": { "type": "integer" }, "itemArray": { "properties": { "name": { "type": "keyword" }, "order": { "type": "integer" } } } } } }
Query to Find Next Item & Count Occurrences
We'll use a scripted_metric aggregation because it lets us handle the logic of finding the target item's order, locating the next item, and counting results all in one place. This works for both nested and non-nested mappings (the script adjusts automatically based on the array structure).
Replace {{TARGET_ITEM_NAME}} with your input item name (e.g., "X"):
POST /items_index/_search { "size": 0, // We don't need the raw documents, just the aggregation results "query": { // First, filter out documents that don't contain the target item "bool": { "should": [ { "nested": { "path": "itemArray", "query": { "term": { "itemArray.name": "{{TARGET_ITEM_NAME}}" } } } }, { "term": { "itemArray.name": "{{TARGET_ITEM_NAME}}" } } ], "minimum_should_match": 1 } }, "aggs": { "next_item_counts": { "scripted_metric": { "init_script": "state.counts = [:];", // Initialize an empty map to hold counts "map_script": """ // Step 1: Find the target item's order def targetOrder = -1; def names = doc['itemArray.name']; def orders = doc['itemArray.order']; for (int i = 0; i < names.length; i++) { if (names[i].value == params.targetName) { targetOrder = orders[i].value; break; // Assume each item name appears once per document } } // Skip if target item wasn't found (though our query should filter these out) if (targetOrder == -1) return; // Step 2: Find the item with order = targetOrder + 1 def nextOrder = targetOrder + 1; for (int i = 0; i < orders.length; i++) { if (orders[i].value == nextOrder) { def nextItemName = names[i].value; // Update the count for this next item state.counts[nextItemName] = (state.counts[nextItemName] ?: 0) + 1; break; // Assume each order is unique per document } } """, "params": { "targetName": "{{TARGET_ITEM_NAME}}" }, // Pass in the target item name "combine_script": "return state.counts;", // Combine results from each shard "reduce_script": """ // Merge counts across all shards def finalCounts = [:]; for (shardCounts in states) { for (entry in shardCounts.entrySet()) { finalCounts[entry.getKey()] = (finalCounts[entry.getKey()] ?: 0) + entry.getValue(); } } return finalCounts; """ } } } }
How This Works
- Query Filter: We first filter out any documents that don't contain the target item (using a bool query that works for both nested and non-nested mappings).
- Init Script: Sets up an empty map to track counts of next items.
- Map Script: Runs per document:
- Finds the
ordervalue of the target item by looping through thenamearray. - Looks for the item with an
ordervalue exactly one higher than the target's. - Updates the count map with the name of that next item.
- Finds the
- Combine & Reduce Scripts: Merge counts from all shards into a final result set.
Example Output (For Target Item "X")
Using your sample data, the response will look like this:
{ "aggregations": { "next_item_counts": { "value": { "Y": 1, "Z": 1 } } } }
This tells us that "Y" was the next item once, and "Z" was the next item once across your three documents.
内容的提问来源于stack exchange,提问作者gokhansari

