You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Elasticsearch中基于字段去重并按其他字段分组(Nest实现)

Solution for Your Elasticsearch + Nest Query Problem

Hey there! Let's work through this problem step by step. You need to first get unique documents based on complexProperty1.A, then group those unique docs by the combination of complexProperty1.D and complexProperty1.E—and we'll use Nest to pull this off.

Step 1: Define Your POCO Classes

First, map your document structure to C# classes so Nest can handle serialization/deserialization smoothly:

public class ComplexProperty1
{
    public string A { get; set; }
    public string B { get; set; }
    public bool D { get; set; }
    public string E { get; set; }
    public List<string> F { get; set; }
}

public class ComplexProperty2
{
    public string X { get; set; }
    public List<string> Y { get; set; }
    public string Z { get; set; }
}

public class Document
{
    public ComplexProperty1 ComplexProperty1 { get; set; }
    public ComplexProperty2 ComplexProperty2 { get; set; }
}

Step 2: Build the Query with Nest

We'll leverage two core Elasticsearch features here: collapse (to fetch unique docs by A) and script-based terms aggregation (to group by D + E combinations). Here's the full query code:

// Initialize Nest client with your cluster URL
var settings = new ConnectionSettings(new Uri("http://localhost:9200"))
    .DefaultIndex("your-index-name"); // Replace with your actual index name

var client = new ElasticClient(settings);

// Execute the search request
var response = client.Search<Document>(s => s
    .Size(0) // Skip top-level raw hits—we only care about aggregation results
    .Collapse(c => c
        .Field(f => f.ComplexProperty1.A) // Collapse to unique values of `complexProperty1.A`
        .InnerHits(i => i.Size(1)) // Keep exactly 1 document per unique `A` value
    )
    .Aggregations(a => a
        .Terms("group_by_d_e", t => t
            // Combine D and E into a single group key (e.g., "true|case") using a script
            .Script(s => s.Source("doc['complexProperty1.D'].value + '|' + doc['complexProperty1.E'].value"))
            // Fetch all documents in each group with TopHits
            .Aggregations(subAgg => subAgg
                .TopHits("docs_in_group", th => th
                    .Size(1000) // Adjust this based on the max docs per group you need
                    .Sort(sort => sort.Ascending("_score"))
                )
            )
        )
    )
);

Step 3: Process the Response

Once you have the query response, extract and iterate over the grouped documents like this:

if (response.IsValid)
{
    var deGroups = response.Aggregations.Terms("group_by_d_e");
    
    foreach (var group in deGroups.Buckets)
    {
        Console.WriteLine($"Group Key: {group.Key}");
        var groupDocuments = group.Aggregations.TopHits("docs_in_group").Documents<Document>();
        
        foreach (var doc in groupDocuments)
        {
            Console.WriteLine($"  - A: {doc.ComplexProperty1.A}, D: {doc.ComplexProperty1.D}, E: {doc.ComplexProperty1.E}");
        }
    }
}
else
{
    // Handle query errors—use debug info to troubleshoot
    Console.WriteLine($"Query failed: {response.DebugInformation}");
}

Key Considerations

  • Mapping Validation: Ensure complexProperty1.E is mapped as a keyword type (not text) to avoid tokenization breaking grouping. D should be mapped as boolean for accurate script evaluation.
  • Version Support: The collapse feature requires Elasticsearch 6.3 or later—confirm your cluster meets this requirement.
  • Performance Tuning: If working with very large datasets, add filters to narrow down results before aggregation, and adjust Size parameters to avoid unnecessary memory usage.

内容的提问来源于stack exchange,提问作者Vista

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 10:52:58