如何对Elasticsearch5.3查询请求做统计?含热门搜索词统计咨询
Hey Allen! Great question—you absolutely can calculate last month's top search keywords using Elasticsearch 5.3, no extra tools are mandatory (though there are optional add-ons if you want to expand functionality later). Let me walk you through the practical ways to do this:
First, I assume you're storing user search requests (including the keyword, timestamp, etc.) in an Elasticsearch index (let's call it user_searches), where each document has fields like search_keyword (the user's input) and search_timestamp (when the search happened).
Core Query
Use a terms aggregation combined with a time range filter to get the top keywords from last month:
{ "size": 0, // Skip returning individual documents—we only care about the aggregation results "query": { "range": { "search_timestamp": { "gte": "now-1M/M", // Start of last month (midnight) "lt": "now/M" // Start of current month (midnight) } } }, "aggs": { "top_search_keywords": { "terms": { "field": "search_keyword.keyword", // Use the keyword sub-field to avoid tokenization issues "size": 10, // Adjust this to get top N keywords (e.g., 20 for top 20) "order": { "_count": "desc" // Sort by search frequency, highest first } } } } }
Critical Notes
- If your
search_keywordis atextfield, you can't aggregate on it directly (since it gets tokenized). You need to either:- Add a
keywordsub-field in your index mapping (e.g.,search_keyword.keyword) and reindex your data, or - Set the field type to
keywordupfront if you don't need full-text search on the keyword itself.
- Add a
- Elasticsearch 5.3 disables
fielddataby default fortextfields, so using thekeywordsub-field is non-negotiable here—otherwise you'll hit a runtime error.
If your application is querying Elasticsearch directly for user searches, you can add a step to write each search's keyword and timestamp to the user_searches index (using ES's Bulk API or your client library of choice). This way, you can run the aggregation query above at any time to get up-to-date top keywords.
You can handle the core use case with ES 5.3 alone, but these tools can help with more complex needs:
- Synonym grouping: If you want to treat similar keywords as one (e.g., "iPhone" and "Apple phone"), use Elasticsearch's synonym filter during indexing/querying, or pre-process data with Logstash before it hits ES.
- Visualization: For dashboards or monthly reports, use Kibana 5.3 (make sure the version matches your ES instance)—it integrates seamlessly with ES and lets you build charts, tables, and dashboards with zero code.
If you have high search volumes, consider using time-based sharding for your user_searches index (e.g., one shard per month). This will make querying last month's data much faster, as ES only needs to scan the relevant shard instead of the entire index.
内容的提问来源于stack exchange,提问作者Allen_Tsang

