You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Elasticsearch按email分组并按score排序的Java实现求助

Solution to Implement Elasticsearch Full-Text Search, Group by Email, and Sort by Score in Java

I'll walk you through replicating your Postman request (plus the email grouping and score sorting functionality you mentioned) using the Elasticsearch Rest High Level Client (v7.x—adjust the version to match your Elasticsearch cluster).

Step 1: Add Dependencies

First, make sure you have the required dependencies in your pom.xml (Maven) or equivalent for Gradle:

<!-- Maven -->
<dependency>
    <groupId>org.elasticsearch.client</groupId>
    <artifactId>elasticsearch-rest-high-level-client</artifactId>
    <version>7.17.9</version> <!-- Match your ES cluster version -->
</dependency>
<dependency>
    <groupId>org.elasticsearch</groupId>
    <artifactId>elasticsearch</artifactId>
    <version>7.17.9</version>
</dependency>

Step 2: Full Java Implementation

Here's the complete code that covers all your requirements:

  • Full-text search for the keyword "Innovation"
  • Highlighting matches (matching your Postman settings exactly)
  • Grouping results by the email field
  • Sorting groups by the sum of _score (since you mentioned the score summation is now working correctly)
import org.elasticsearch.action.search.SearchRequest;
import org.elasticsearch.action.search.SearchResponse;
import org.elasticsearch.client.RequestOptions;
import org.elasticsearch.client.RestHighLevelClient;
import org.elasticsearch.common.unit.TimeValue;
import org.elasticsearch.index.query.BoolQueryBuilder;
import org.elasticsearch.index.query.QueryBuilders;
import org.elasticsearch.search.SearchHit;
import org.elasticsearch.search.aggregations.AggregationBuilders;
import org.elasticsearch.search.aggregations.bucket.terms.Terms;
import org.elasticsearch.search.aggregations.metrics.sum.Sum;
import org.elasticsearch.search.builder.SearchSourceBuilder;
import org.elasticsearch.search.fetch.subphase.highlight.HighlightBuilder;
import org.elasticsearch.search.sort.SortOrder;

import java.io.IOException;
import java.util.concurrent.TimeUnit;

public class ElasticsearchInnovationSearch {

    public static void main(String[] args) throws IOException {
        // Initialize the client (update host/port to match your cluster)
        try (RestHighLevelClient client = new RestHighLevelClient(
                org.elasticsearch.client.RestClient.builder(
                        new org.apache.http.HttpHost("localhost", 9200, "http")))) {

            // 1. Target your specific index
            SearchRequest searchRequest = new SearchRequest("your-index-name"); // Replace with your index name
            SearchSourceBuilder sourceBuilder = new SearchSourceBuilder();

            // 2. Set pagination (matches your Postman from/size)
            sourceBuilder.from(0);
            sourceBuilder.size(10);
            sourceBuilder.timeout(new TimeValue(60, TimeUnit.SECONDS));

            // 3. Build the full-text query for "Innovation"
            BoolQueryBuilder boolQuery = QueryBuilders.boolQuery();
            boolQuery.must(QueryBuilders.queryStringQuery("Innovation"));
            sourceBuilder.query(boolQuery);

            // 4. Configure highlighting (mirrors your Postman settings)
            HighlightBuilder highlightBuilder = new HighlightBuilder();
            highlightBuilder.requireFieldMatch(false);
            highlightBuilder.preTags("<b>");
            highlightBuilder.postTags("</b>...");
            // Uncomment below to target specific fields, or leave as-is for all fields
            // highlightBuilder.field("content");
            sourceBuilder.highlighter(highlightBuilder);

            // 5. Add aggregation: Group by email, sum scores, sort by total score descending
            sourceBuilder.aggregation(
                    AggregationBuilders.terms("group_by_email")
                            .field("email.keyword") // Use .keyword for exact email matches
                            .order(Terms.Order.aggregation("sum_score", false)) // Sort groups by total score descending
                            .subAggregation(
                                    AggregationBuilders.sum("sum_score")
                                            .field("_score")
                            )
            );

            searchRequest.source(sourceBuilder);

            // 6. Execute the search
            SearchResponse searchResponse = client.search(searchRequest, RequestOptions.DEFAULT);

            // 7. Process raw search results
            System.out.println("Raw Search Hits:");
            for (SearchHit hit : searchResponse.getHits().getHits()) {
                System.out.println("Document ID: " + hit.getId());
                System.out.println("Source Data: " + hit.getSourceAsString());
                System.out.println("Highlighted Matches: " + hit.getHighlightFields());
                System.out.println("------------------------");
            }

            // 8. Process grouped results (by email)
            System.out.println("\nGrouped Results by Email:");
            Terms emailGroups = searchResponse.getAggregations().get("group_by_email");
            for (Terms.Bucket bucket : emailGroups.getBuckets()) {
                String email = bucket.getKeyAsString();
                long documentCount = bucket.getDocCount();
                Sum totalScore = bucket.getAggregations().get("sum_score");

                System.out.println("Email: " + email);
                System.out.println("Total Matching Documents: " + documentCount);
                System.out.println("Total Score for Group: " + totalScore.getValue());
                System.out.println("------------------------");
            }

        } catch (IOException e) {
            e.printStackTrace();
        }
    }
}

Key Notes to Adjust for Your Setup:

  • Index Name: Replace "your-index-name" with the actual name of your Elasticsearch index.
  • Email Field: We use email.keyword for aggregation to ensure exact matches—this assumes your email field is mapped as a text field with a keyword sub-field (standard for email fields).
  • Highlighting: If you only want to highlight specific fields (e.g., content or title), replace the wildcard behavior with highlightBuilder.field("your-target-field").
  • Client Configuration: Update the host/port if your Elasticsearch cluster isn't running on localhost:9200.
  • Sorting: Change false to true in Terms.Order.aggregation("sum_score", false) if you want ascending sort by total score.

Quick Troubleshooting:

  • If you're using Elasticsearch 8.x, the client API has changed (use co.elastic.clients.elasticsearch.ElasticsearchClient instead). Let me know if you need the 8.x adapted version.
  • Double-check that your email field is mapped correctly in the index—if you get aggregation errors, verify the field exists and has the keyword sub-field.

内容的提问来源于stack exchange,提问作者Jeff

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 09:51:51