Elasticsearch子文档问答排列统计方案及性能咨询
针对父子文档结构的用户问答组合统计方案
首先明确你的场景:你用Elasticsearch的父子文档存储用户与问答记录(user为父文档,answer为子文档),想要统计所有用户的问答组合(比如同时拥有{Q:5,A:6}和{Q:6,A:6}的用户数量),但之前尝试的嵌套聚合因过滤逻辑排除子文档而失效;同时你担心把所有问答嵌入父文档会因问题数量多(2000个)、用户量大(数百万)导致性能问题。
先整理你的现有Schema与数据导入命令:
现有Schema与数据导入
# 创建带父子关系的索引 curl -X PUT "localhost:9200/test_example" -H 'Content-Type: application/json' -d' { "mappings": { "answer" : { "_parent" : { "type" : "user" } } } } ' # 批量导入用户父文档 curl -X PUT "localhost:9200/test_example/user/_bulk?pretty" -H 'Content-Type: application/json' -d' { "index": {"_id": 1} } { "user": 1, "foo": 1} { "index": {"_id": 2} } { "user": 2, "foo": 2} { "index": {"_id": 3} } { "user": 3, "foo": 1} { "index": {"_id": 4} } { "user": 4, "foo": 1, "answers": {"5": [6,3], "6":[6], "7":[6]}, "potential_approach": true} ' # 批量导入问答子文档 curl -X PUT "localhost:9200/test_example/answer/_bulk?pretty" -H 'Content-Type: application/json' -d' { "index": { "parent": 1}} {"question":5, "answer":6, "negative": false} { "index": { "parent": 1}} {"question":5, "answer":3, "negative": false} { "index": { "parent": 1}} {"question":6, "answer":6, "negative": false} { "index": { "parent": 1}} {"question":7, "answer":6, "negative": false} { "index": { "parent": 2}} {"question":5, "answer":6, "negative": false} { "index": { "parent": 2}} {"question":5, "answer":3, "negative": false} { "index": { "parent": 2}} {"question":6, "answer":1, "negative": false} { "index": { "parent": 2}} {"question":7, "answer":2, "negative": false} { "index": { "parent": 3}} {"question":5, "answer":6, "negative": false} { "index": { "parent": 3}} {"question":5, "answer":2, "negative": false} { "index": { "parent": 3}} {"question":6, "answer":4, "negative": false} { "index": { "parent": 3}} {"question":7, "answer":3, "negative": false} '
你之前尝试的聚合查询因在子文档层面嵌套过滤,导致上层过滤排除了后续需要的子文档,无法得到正确的组合统计:
curl -X POST "localhost:9200/test_example/user/_search?pretty&size=1" -H 'Content-Type: application/json' -d' { "query": {}, "aggs": { "permutations": { "children": { "type": "answer" }, "aggs": { "filtered": { "filter": { "bool": { "should": [], "must_not": [], "must": [ { "term": { "question": 5 } } ] } }, "aggs": { "5": { "terms": { "field": "answer", "size": 500 }, "aggs": { "filtered": { "filter": { "bool": { "should": [], "must_not": [], "must": [ { "term": { "question": 6 } } ] } }, "aggs": { "6": { "terms": { "field": "answer", "size": 500 } } } } } } } } } } } } '
接下来给你两种可行的解决方案:
方案一:基于现有父子文档结构的正确聚合方式
核心思路是先定位符合单个问答条件的用户,再叠加过滤统计组合用户数,或通过composite聚合遍历所有问答对并关联父文档计数:
方式1:统计指定问答组合的用户数
如果你需要统计特定组合(比如同时有{Q:5,A:6}和{Q:6,A:6}的用户),可以用filters聚合配合多个has_child查询:
curl -X POST "localhost:9200/test_example/user/_search?pretty&size=0" -H 'Content-Type: application/json' -d' { "aggs": { "target_combinations": { "filters": { "filters": { "Q5A6_plus_Q6A6": { "bool": { "must": [ { "has_child": { "type": "answer", "query": { "bool": { "must": [ {"term": {"question": 5}}, {"term": {"answer": 6}} ] } } } }, { "has_child": { "type": "answer", "query": { "bool": { "must": [ {"term": {"question": 6}}, {"term": {"answer": 6}} ] } } } } ] } }, "Q5A3_plus_Q6A1": { "bool": { "must": [ { "has_child": { "type": "answer", "query": { "bool": { "must": [ {"term": {"question": 5}}, {"term": {"answer": 3}} ] } } } }, { "has_child": { "type": "answer", "query": { "bool": { "must": [ {"term": {"question": 6}}, {"term": {"answer": 1}} ] } } } } ] } } } }, "aggs": { "unique_users": { "cardinality": { "field": "user" } } } } } } '
方式2:遍历所有问答组合并统计用户数
如果需要遍历所有可能的问答组合,用composite聚合在子文档层面先分组,再关联父文档计数(避免内存溢出):
curl -X POST "localhost:9200/test_example/answer/_search?pretty&size=0" -H 'Content-Type: application/json' -d' { "aggs": { "all_question_answer_pairs": { "composite": { "size": 10000, "sources": [ {"question": {"terms": {"field": "question"}}}, {"answer": {"terms": {"field": "answer"}}} ] }, "aggs": { "user_count": { "cardinality": { "field": "_parent" } } } } } } '
这个查询会返回每个(question, answer)对对应的用户数量,你可以通过客户端逻辑进一步组合多个问答对的用户交集。
方案二:验证嵌入父文档方案的可行性
如果你想把问答存入父文档,绝对不推荐用动态键的结构(比如{"5": [6,3]}),因为会导致mapping膨胀,推荐改成嵌套数组结构:
{ "user": 1, "foo": 1, "answers": [ {"question":5, "answer":6}, {"question":5, "answer":3}, {"question":6, "answer":6}, {"question":7, "answer":6} ] }
然后把answers设置为嵌套类型:
curl -X PUT "localhost:9200/test_example_nested" -H 'Content-Type: application/json' -d' { "mappings": { "user": { "properties": { "answers": { "type": "nested", "properties": { "question": {"type": "keyword"}, "answer": {"type": "keyword"} } } } } } } '
此时可以用nested聚合统计组合:
curl -X POST "localhost:9200/test_example_nested/user/_search?pretty&size=0" -H 'Content-Type: application/json' -d' { "aggs": { "q5_answers": { "nested": { "path": "answers" }, "aggs": { "filter_q5": { "filter": {"term": {"answers.question": 5}}, "aggs": { "answer_terms": { "terms": {"field": "answers.answer", "size": 500}, "aggs": { "back_to_user": { "reverse_nested": {}, "aggs": { "q6_answers": { "nested": { "path": "answers" }, "aggs": { "filter_q6": { "filter": {"term": {"answers.question": 6}}, "aggs": { "answer_terms": { "terms": {"field": "answers.answer", "size": 500}, "aggs": { "user_count": { "cardinality": {"field": "user"} } } } } } } } } } } } } } } } } } '
性能验证建议
对于数百万用户、每个用户500-700条问答的场景,你可以:
- 创建测试索引,导入10万条模拟用户数据
- 运行上述聚合查询,监控Elasticsearch节点的CPU、内存、响应时间
- 优化点:
- 给
answers.question和answers.answer设置keyword类型(不需要分词) - 优先用
composite聚合替代多层嵌套terms,减少内存占用 - 调整分片数(比如按用户数设置5-10个分片)
- 给
总结
- 如果不想修改现有数据结构,方案一是最优选择,通过
has_child过滤和composite聚合可以高效统计组合用户数 - 如果考虑嵌入父文档,一定要用嵌套数组结构而非动态键,通过小批量测试验证性能;Elasticsearch在配置合理的情况下(足够内存、合适分片数),完全可以处理你的规模
内容的提问来源于stack exchange,提问作者Ergo
相关产品推荐
相关产品推荐

