AWS Elasticsearch V6.4多字段分组及分页实现技术问询
Hey there! Let's work through the best way to group your jobs data by userId, companyId, and title while adding reliable pagination support. First, let's touch on why your current approach with the samejobs field might not be the most ideal: maintaining a pre-calculated field like samejobs requires ongoing sync as your data changes, which can lead to inconsistencies over time, and it doesn't give you the flexibility to adjust grouping logic on the fly.
Instead, Elasticsearch's Composite Aggregation is the perfect fit here—it's built specifically for paginating multi-dimensional aggregations, works efficiently with large datasets, and avoids the memory limitations of standard terms aggregations. Here's how to implement this properly:
Step 1: Core Requirements Breakdown
We need to:
- Apply all your existing filters (active job status, expired date range, geo-distance, search phrases) first to narrow down the relevant dataset.
- Group the filtered documents using keyword-type fields (since text fields can't be used for aggregation):
userId.keywordfor exact user matchescompanyId.keywordfor exact company matchestitle.normalize(your lowercase-normalized field) for case-insensitive title grouping
- Get the count of documents per group and paginate through these groups seamlessly.
Step 2: Composite Aggregation Query Examples
First Page Query
This query runs all your filters, then returns the first 100 grouped results with their document counts:
GET jobs/_search { "size": 0, // We only need aggregation results, not individual documents "query": { "bool": { "must": [ {"term": {"jobStatus.keyword": "active"}}, {"term": {"status.keyword": "active"}}, {"range": {"dateExpired": {"gt": 1568007661898}}}, { "multi_match": { "query": "sales executive", "type": "phrase_prefix", "fields": ["title", "industryName", "primarySkills.value.keyword", "secondarySkills"] } }, { "bool": { "should": [ { "multi_match": { "query": "Rajkot", "type": "phrase_prefix", "fields": ["area", "cityName"] } } ], "minimum_should_match": 1 // Ensure at least one location condition is met } } ], "filter": { "geo_distance": { "distance": "15000km", "location": {"lat": 23.0066678, "lon": 72.5017585} } } } }, "aggs": { "grouped_jobs": { "composite": { "size": 100, // Number of groups per page "sources": [ {"user_id": {"terms": {"field": "userId.keyword"}}}, {"company_id": {"terms": {"field": "companyId.keyword"}}}, {"job_title": {"terms": {"field": "title.normalize"}}} // Use normalized title for consistent case-insensitive grouping ], "order": {"doc_count": "desc"} // Optional: sort groups by document count (descending) } } } }
Pagination (Next Page)
To fetch the next page of groups, use the after parameter with the last group's key from your previous result:
GET jobs/_search { "size": 0, "query": { /* Reuse the same filter query from above */ }, "aggs": { "grouped_jobs": { "composite": { "size": 100, "sources": [ /* Same sources as the first page */ ], "order": {"doc_count": "desc"}, "after": { "user_id": "user_123", "company_id": "company_456", "job_title": "sales executive" } // Replace this with the "key" value of the last group from your first page response } } } }
Step 3: Important Notes
- Stick to Keyword/Normalized Fields: Never use text fields for aggregation—they're analyzed and won't give you exact group matches. Your
title.normalizefield is perfect here because it standardizes case, ensuring "Sales Executive" and "sales executive" are grouped together. - Filter Before Aggregating: Applying your filters first reduces the dataset size, making the aggregation faster and more efficient.
- Composite Aggregation Advantages: Unlike standard terms aggregations, composite doesn't load all groups into memory, so it works reliably even with massive datasets. The
afterparameter ensures stable, gap-free pagination. - Customize Sorting: Adjust the
orderparameter if you want to sort groups by a different field (e.g.,user_idascending instead of document count).
Why This Beats Your Current Approach
- No more maintaining the
samejobsfield—aggregations calculate groups in real-time, so your results are always accurate as data changes. - Flexible grouping logic: you can tweak the grouped fields or sorting without updating existing data.
- Proper pagination that scales: even if you have thousands of unique groups, this method won't hit memory limits or return inconsistent pages.
内容的提问来源于stack exchange,提问作者Nitin Muchhadiya

