如何优化Gremlin的project().by()查询,减少子遍历重复执行?
优化Gremlin校园家庭统计查询:减少重复子女遍历
可以通过缓存每个Parent的子女集合,让所有统计项复用同一集合来避免重复遍历,以下是两种可行的改写方案:
方案一:使用sideEffect缓存子女集合
这种方式最直观,先通过sideEffect一次性获取并缓存每个家长的子女列表,后续统计均基于该缓存集合计算:
g.V().hasLabel('Parent'). sideEffect(out('has_child').fold().as('children')). project('Parent', 'boys', 'girls', 'STEM_students', 'sport_participants'). by('name'). by(select('children').unfold().has('gender', 'male').count()). by(select('children').unfold().has('gender', 'female').count()). by(select('children').unfold().out('enrolled_in').has('type', 'STEM').count()). by(select('children').unfold().out('participates_in').hasLabel('Sport').count())
方案二:使用match预加载子女数据
通过match步骤预定义家长与子女的关联关系,同样实现一次遍历复用:
g.V().hasLabel('Parent').as('p'). match( __.as('p').out('has_child').fold().as('children') ). project('Parent', 'boys', 'girls', 'STEM_students', 'sport_participants'). by(select('p').values('name')). by(select('children').unfold().has('gender', 'male').count()). by(select('children').unfold().has('gender', 'female').count()). by(select('children').unfold().out('enrolled_in').has('type', 'STEM').count()). by(select('children').unfold().out('participates_in').hasLabel('Sport').count())
核心优势
两种方案均保证每个Parent仅执行一次子女查找遍历,后续统计直接复用缓存的子女集合,既保留了原查询中project能包含无符合条件子女家长的特性,又大幅提升了查询效率。
内容的提问来源于stack exchange,提问作者Tom Kelly
相关产品推荐
相关产品推荐

