Gremlin多场景高效计数查询实现及多选择场景扩展咨询
多选择场景Gremlin查询实现方案

上图定义了本次查询所用的图Schema,全图总节点数650万,总边数300万,以下为适配多选择场景的查询实现:
通用计数查询(支持多路径组合、公共节点自动去重)
你之前的单路径硬编码写法扩展性较差,改用union包裹所有选中的路径组合即可适配多选择场景,公共节点因为共用聚合键会自动去重,不会重复计数。
适配场景1(Node2→relation1→Node1、Node1→relation5→Node3组合)的查询代码如下:
g.V().union( // 第一条路径:Node2 -> relation1 -> Node1 has('property1', 'Node2'). filter(outE().has('property1', 'relation1')). dedup().aggregate('Node2').by(constant(1)). outE().has('property1', 'relation1'). dedup().aggregate('relation1').by(constant(1)). inV().dedup().aggregate('Node1').by(constant(1)), // 第二条路径:Node1 -> relation5 -> Node3 has('property1', 'Node1'). filter(outE().has('property1', 'relation5')). dedup().aggregate('Node1').by(constant(1)). outE().has('property1', 'relation5'). dedup().aggregate('relation5').by(constant(1)). inV().dedup().aggregate('Node3').by(constant(1)) ).limit(1). // 不需要遍历所有路径结果,仅需聚合结果即可 project('Node2', 'relation1', 'Node1', 'Node3', 'relation5'). by(coalesce(select('Node2').unfold().sum(), constant(0))). by(coalesce(select('relation1').unfold().sum(), constant(0))). by(coalesce(select('Node1').unfold().sum(), constant(0))). by(coalesce(select('Node3').unfold().sum(), constant(0))). by(coalesce(select('relation5').unfold().sum(), constant(0)))
这里用
coalesce是为了处理某类节点/关系没有匹配结果时返回0,避免空指针异常,不需要可直接去掉。
输出结果和要求格式一致:
==>[Node2:7003,relation1:200166,Node1:22000,Node3:167,relation5:11000]
全图范围选择适配
如果是全图统计指定节点、关系类型的数量,不需要走路径匹配,直接单独统计性能更高,示例代码如下:
g.V().union( has('property1', 'Node2').count().aggregate('Node2'), has('property1', 'Node1').count().aggregate('Node1'), has('property1', 'Node3').count().aggregate('Node3') ).E().union( has('property1', 'relation1').count().aggregate('relation1'), has('property1', 'relation5').count().aggregate('relation5') ).limit(1). project('Node2', 'relation1', 'Node1', 'Node3', 'relation5'). by(select('Node2')). by(select('relation1')). by(select('Node1')). by(select('Node3')). by(select('relation5'))
返回实际数据的查询方案
如果需要返回选中内容的实际数据而非计数,将aggregate的计数逻辑替换为收集元素本身,最后去重即可,示例如下:
g.V().union( // 第一条路径 has('property1', 'Node2').filter(outE().has('property1', 'relation1')).dedup().aggregate('Node2'). outE().has('property1', 'relation1').dedup().aggregate('relation1'). inV().dedup().aggregate('Node1'), // 第二条路径 has('property1', 'Node1').filter(outE().has('property1', 'relation5')).dedup().aggregate('Node1'). outE().has('property1', 'relation5').dedup().aggregate('relation5'). inV().dedup().aggregate('Node3') ).limit(1). project('Node2', 'relation1', 'Node1', 'Node3', 'relation5'). by(select('Node2').dedup().fold()). by(select('relation1').dedup().fold()). by(select('Node1').dedup().fold()). by(select('Node3').dedup().fold()). by(select('relation5').dedup().fold())
该查询返回的每个键对应的数据都是去重后的节点/边列表,可按需取出属性展示。
内容的提问来源于stack exchange,提问作者Phoenix
相关产品推荐
相关产品推荐

