如何使用Elasticsearch提取所有文档中嵌套字段color1的全部值
Hey there! Let's figure out how to get all the color1 values mapped to each website in your Elasticsearch documents. The reason your match_all with size:0 didn't work is that size:0 just tells ES not to return raw document hits—you need to use aggregations to group and extract the data you want. Here are a couple of tailored solutions depending on your exact needs:
1. Group by Website and Extract color1 Values (Ideal for Aggregated Results)
If you want to see each unique website paired with its associated color1 values, use a combination of terms aggregation (to group websites) and either top_hits or nested aggregation (to handle the color array):
Option A: For Single Color Entry Per Website
If each website's document only has one color object in the color array, this simple query works perfectly:
GET /your_index_name/_search { "size": 0, // Skip returning raw docs, focus only on aggregations "aggs": { "grouped_websites": { "terms": { "field": "website.keyword", // Use .keyword to avoid text field tokenization issues "size": 10000 // Adjust this number to cover all your unique websites }, "aggs": { "extract_color1": { "top_hits": { "size": 1, "_source": { "includes": ["color.color1"] } } } } } } }
Option B: For Multiple Color Entries Per Website
If a website has multiple documents or multiple color objects in the array, use a nested aggregation to capture all unique color1 values for each website:
GET /your_index_name/_search { "size": 0, "aggs": { "grouped_websites": { "terms": { "field": "website.keyword", "size": 10000 }, "aggs": { "color_nested": { "nested": { "path": "color" // Target the nested color array to iterate through each object }, "aggs": { "all_color1_values": { "terms": { "field": "color.color1.keyword", "size": 10000 } } } } } } } }
2. Get All Website-color1 Pairs Directly (No Aggregation)
If you just want a flat list of every document's website and corresponding color1, use _source filtering to only return the fields you care about, and set a large enough size to fetch all documents (use scroll/search_after if you have more than 10k docs):
GET /your_index_name/_search { "size": 10000, // Adjust based on your total document count "_source": { "includes": ["website", "color.color1"] } }
Just replace your_index_name with the actual name of your Elasticsearch index, and tweak the size values to make sure you capture all your data.
内容的提问来源于stack exchange,提问作者Victoire

