You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Elasticsearch技术咨询:双时间段事件统计与doc_count直方图实现

Can Elasticsearch handle these two requirements?

Absolutely! Both of your needs—counting events that occurred in both 2016 and 2017, and generating a histogram from the aggregated doc_count field—are fully achievable in Elasticsearch. Let’s break down how to implement each one:

1. Count events present in both 2016 and 2017

To find keys that exist in both time periods and count them, we’ll use nested aggregations with a bucket selector to filter only keys that have data in both years. Here’s the query:

GET /your_index_name/_search
{
  "size": 0, // We don't need individual documents, just aggregation results
  "aggs": {
    "all_keys": {
      "terms": {
        "field": "Key",
        "size": 10000 // Adjust based on your total unique keys; use a high enough value to capture all
      },
      "aggs": {
        // Check if the key has data in 2016
        "has_2016_data": {
          "filter": {
            "range": {
              "@timestamp": { // Replace with your actual date field name
                "gte": "2016-01-01T00:00:00Z",
                "lte": "2016-12-31T23:59:59Z"
              }
            }
          }
        },
        // Check if the key has data in 2017
        "has_2017_data": {
          "filter": {
            "range": {
              "@timestamp": {
                "gte": "2017-01-01T00:00:00Z",
                "lte": "2017-12-31T23:59:59Z"
              }
            }
          }
        },
        // Only keep keys that have data in both years
        "keep_both_years": {
          "bucket_selector": {
            "buckets_path": {
              "count_2016": "has_2016_data._count",
              "count_2017": "has_2017_data._count"
            },
            "script": "params.count_2016 > 0 && params.count_2017 > 0"
          }
        }
      }
    },
    // Final count of keys present in both years
    "total_common_events": {
      "bucket_count": {
        "buckets_path": "all_keys"
      }
    }
  }
}

How this works:

  • We first group all documents by the Key field.
  • For each key, we check if there are any documents in 2016 and 2017 using filter aggregations.
  • The bucket_selector removes keys that only exist in one year.
  • Finally, bucket_count gives us the total number of keys (events) that occurred in both 2016 and 2017.

2. Generate a histogram from the aggregated doc_count field

If you want to create a histogram showing the distribution of doc_count values (e.g., how many keys have between 1-10 documents, 11-20 documents, etc.), you can extend the previous query with a histogram aggregation. Here’s how:

GET /your_index_name/_search
{
  "size": 0,
  "aggs": {
    "all_keys": {
      "terms": {
        "field": "Key",
        "size": 10000
      },
      "aggs": {
        "has_2016_data": {
          "filter": {
            "range": {
              "@timestamp": {
                "gte": "2016-01-01T00:00:00Z",
                "lte": "2016-12-31T23:59:59Z"
              }
            }
          }
        },
        "has_2017_data": {
          "filter": {
            "range": {
              "@timestamp": {
                "gte": "2017-01-01T00:00:00Z",
                "lte": "2017-12-31T23:59:59Z"
              }
            }
          }
        },
        "keep_both_years": {
          "bucket_selector": {
            "buckets_path": {
              "count_2016": "has_2016_data._count",
              "count_2017": "has_2017_data._count"
            },
            "script": "params.count_2016 > 0 && params.count_2017 > 0"
          }
        }
      }
    },
    "total_common_events": {
      "bucket_count": {
        "buckets_path": "all_keys"
      }
    },
    // Histogram of doc_count values for common keys
    "doc_count_distribution": {
      "histogram": {
        "field": "_count", // This refers to the total documents per key
        "interval": 10, // Adjust the interval to fit your data (e.g., 5, 20, 100)
        "min_doc_count": 1 // Only include intervals with at least one key
      }
    }
  }
}

How this works:

  • The histogram aggregation takes the _count value (total documents per key from the terms aggregation) and groups them into bins based on the interval you set.
  • The result will show how many keys fall into each document count range—perfect for building a histogram visualization.

内容的提问来源于stack exchange,提问作者Chris

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 04:05:42