优化含repeat与times的Neptune查询,解决执行耗时过长问题
Neptune查询性能优化方案
原始查询语句
g.V().hasLabel('User') .has('user_id', 1004) .repeat(both('USES_UPI','USES_ACCOUNT','USES_HARDWARE_ID','USES_GAID','HAS_COOKIES').simplePath().dedup()) .times(3) .hasLabel('Gaid') .dedup() .count()
查询性能分析结果
原始遍历计划
[GraphStep(vertex,[]), HasStep([~label.eq(User), user_id.eq(159017810)]), RepeatStep([VertexStep(BOTH,[USES_UPI, USES_ACCOUNT, USES_HARDWARE_ID, USES_GAID, HAS_COOKIES],vertex), PathFilterStep(simple,null,null), DedupGlobalStep(null,null), RepeatEndStep],until(loops(3)),emit(false)), HasStep([~label.eq(Gaid)]), DedupGlobalStep(null,null), CountGlobalStep]
优化后遍历计划
Neptune steps: [ NeptuneGraphQueryStep(Vertex) { JoinGroupNode { PatternNode[(?1, <user_id>, ?9, ?) . project distinct ?1 . ContainsFilter(?9 in (159017810^^<INT>, 159017810^^<LONG>, 1.59017808E8^^<FLOAT>, 1.5901781E8^^<DOUBLE>)) .], {estimatedCardinality=1, expectedTotalOutput=1, indexTime=0, joinTime=0, numSearches=1, actualTotalOutput=1} PatternNode[(?1, <~label>, ?2=<User>, <~>) . project ask .], {estimatedCardinality=1327714, expectedTotalOutput=679, actualTotalOutput=379683, indexTime=0, joinTime=0, numSearches=1} RepeatNode { Repeat { JoinGroupNode { UnionNode { PatternNode[(?3, ?6, ?4, ?7) . project ?3,?4 . IsEdgeIdFilter(?7) . ContainsFilter(?6 in (<USES_UPI>, <USES_ACCOUNT>, <USES_HARDWARE_ID>, <USES_GAID>, <HAS_COOKIES>)) .], {cacheJoin=true, estimatedCardinality=1923966, indexTime=354, joinTime=26212, numSearches=379683} PatternNode[(?4, ?6, ?3, ?7) . project ?3,?4 . IsEdgeIdFilter(?7) . ContainsFilter(?6 in (<USES_UPI>, <USES_ACCOUNT>, <USES_HARDWARE_ID>, <USES_GAID>, <HAS_COOKIES>)) .], {cacheJoin=true, estimatedCardinality=1923966, indexTime=356, joinTime=19633, numSearches=379683} }, annotations={estimatedCardinality=3847932} SimplePathFilter(?1, ?4)) . } } LoopsCondition { LoopsFilter(?3,eq(3)) } }, annotations={emitFirst=false, untilFirst=false, repeatMode=BFS, dedup=true} }, annotations={path=[Vertex(?1):GraphStep, Repeat[̶V̶e̶r̶t̶e̶x̶(̶?̶3̶)̶:̶G̶r̶a̶p̶h̶S̶t̶e̶p̶, Vertex(?4):VertexStep, ̶V̶e̶r̶t̶e̶x̶(̶?̶8̶)̶:̶V̶e̶r̶t̶e̶x̶S̶t̶e̶p̶]], joinStats=true, optimizationTime=2, maxVarId=10, executionTime=107826} }, NeptuneTraverserConverterStep ] + not converted into Neptune steps: NeptuneHasStep([~label.eq(Gaid)]), Neptune steps: [ NeptuneMemoryTrackerStep ] + not converted into Neptune steps: DedupGlobalStep(null,null),CountGlobalStep, WARNING: >> [NeptuneHasStep([~label.eq(Gaid)]), DedupGlobalStep(null,null)] << (or one of the children for each step) is not supported natively yet
运行时指标
Query Execution: 107828.341 ms
遍历指标
Step Count Traversers Time (ms) % Dur ------------------------------------------------------------------------------------------------------------- NeptuneGraphQueryStep(Vertex) 743022 743022 52892.767 49.06 NeptuneTraverserConverterStep 743022 743022 8142.297 7.55 NeptuneHasStep([~label.eq(Gaid)]) 4138 4138 46695.267 43.32 DedupGlobalStep(null,null) 4138 4138 48.927 0.05 CountGlobalStep 1 1 22.215 0.02 >TOTAL - - 107801.475 -
重复遍历指标
Iteration Visited Output Until Emit Next ------------------------------------------------------ 0 1 0 0 0 1 1 3 0 0 0 3 2 379679 0 0 0 379679 3 743022 743022 743022 0 0 ------------------------------------------------------ 1122705 743022 743022 0 379683
其他指标
Predicates ========== # of predicates: 26 Results ======= Count: 1 Index Operations ================ Query execution: # of statement index ops: 1,502,390 # of unique statement index ops: 1,502,390 Duplication ratio: 1.0 # of terms materialized: 162
优化建议
从性能数据来看,NeptuneHasStep([~label.eq(Gaid)])耗时占比超43%,且无法被Neptune原生支持;同时重复遍历的第2、3轮处理了大量无效节点,导致整体耗时过长。可从以下方向优化:
- 提前过滤目标节点,减少无效遍历
原查询在3轮遍历完成后才过滤Gaid节点,导致大量非目标节点被传输处理。修改查询,在遍历过程中一旦匹配到Gaid节点就提前终止该分支:
g.V().hasLabel('User').has('user_id', 1004) .repeat(both('USES_UPI','USES_ACCOUNT','USES_HARDWARE_ID','USES_GAID','HAS_COOKIES') .simplePath() .dedup() .not(hasLabel('Gaid'))) .times(3) .both('USES_UPI','USES_ACCOUNT','USES_HARDWARE_ID','USES_GAID','HAS_COOKIES') .hasLabel('Gaid') .dedup() .count()
- 缩小去重范围,降低全局去重开销
原查询中repeat内部的dedup()是全局去重,性能损耗大。可改为基于当前节点的局部去重,或仅按节点ID去重:
g.V().hasLabel('User').has('user_id', 1004) .repeat(both('USES_UPI','USES_ACCOUNT','USES_HARDWARE_ID','USES_GAID','HAS_COOKIES') .simplePath() .dedup(local)) .times(3) .hasLabel('Gaid') .dedup() .count()
- 优化索引配置
- 确保
User标签的user_id属性已创建复合索引,加速初始用户节点定位; - 为
Gaid标签创建标签索引,减少客户端过滤Gaid节点的耗时。
- 调整遍历模式为DFS
当前repeat使用BFS模式,会占用大量内存存储中间节点。若对遍历顺序无要求,改为DFS模式可降低内存开销:
g.V().hasLabel('User').has('user_id', 1004) .repeat(both('USES_UPI','USES_ACCOUNT','USES_HARDWARE_ID','USES_GAID','HAS_COOKIES') .simplePath() .dedup()) .times(3) .hasLabel('Gaid') .dedup() .count() .option('neptune.repeatMode', 'DFS')
- 移除冗余去重步骤
若repeat内部的去重已能保证节点唯一性,可移除结尾的dedup(),减少重复操作。
内容的提问来源于stack exchange,提问作者Ankit Sharma
相关产品推荐
相关产品推荐

