生物知识图谱Gremlin变长路径查询性能优化问询
问题背景
我正在为生物相关知识图谱开展Gremlin性能基准测试,需要编写与以下Neo4j/Cypher等价的Gremlin查询:
MATCH path = (gene:Gene) - [:enc] -> (prot:Protein) - [:h_s_s|ortho|xref*0..2] - (prot1:Protein) - [:is_a|ac_by] - (enz:Enzyme) - [:ac_by|in_by] -> (cmp:Comp) - [:cs_by|pd_by] -> (trn:Transport) - [:part_of*0..3] -> (pwy:Path) RETURN [ n in nodes(path) | n.iri ] as nodeIris, rand() AS rnd ORDER BY rnd LIMIT 100
查询逻辑说明
- 蛋白质可通过0-2步
h_s_s、ortho、xref关系关联其他蛋白质 - 通路可通过0-3步
part_of关系关联其他通路 - 需要捕获所有符合最大长度限制的链
我编写了如下等价Gremlin查询(标签名称已调整以支持多标签):
g.V().hasLabel ( 'Concept:Gene:Resource' ) .out ( 'enc' ).hasLabel ( 'Concept:Protein:Resource' ) .emit () .repeat ( both ( 'h_s_s', 'ortho', 'xref' ).simplePath().hasLabel ( 'Concept:Protein:Resource' ) ) .times ( 2 ) .both ( 'is_a', 'ac_by' ).hasLabel ( 'Concept:Enzyme:Resource' ) .out ( 'ac_by', 'in_by' ).hasLabel ( 'Comp:Concept:Resource' ) .out ( 'cs_by', 'pd_by' ).hasLabel ( 'Concept:Resource:Transport' ) .emit () .repeat ( both ( 'part_of' ).simplePath().hasLabel ( 'Concept:Path:Resource' ) ) .times ( 3 ) .sample ( 100 ) .path ().by ( 'iri' )
该查询可正常运行,但速度极慢(耗时10-20秒)。请问emit()/repeat()/times()是否为实现该逻辑的最高效方式?我曾考虑用显式变长路径的union实现,但该方式不够直观易写。
解决方案与优化建议
1. 关于emit()/repeat()/times()的效率
emit()+repeat()+times()是Gremlin中实现变长路径的标准方式,本身不存在效率缺陷,你的查询慢主要是步骤顺序和过滤时机导致的冗余计算:
- 当前
simplePath()放在repeat()内部每一步,会对每个遍历到的顶点都执行路径唯一性检查,开销较高 - 标签过滤步骤的位置可以优化,减少不必要的顶点遍历
2. 关键优化点
(1)调整过滤步骤顺序
把hasLabel()前置到关系遍历之后,减少需要执行simplePath()检查的顶点数量:
g.V().hasLabel('Concept:Gene:Resource') .out('enc').hasLabel('Concept:Protein:Resource') .emit() // 先过滤标签,再检查路径唯一性,减少计算量 .repeat(both('h_s_s', 'ortho', 'xref').hasLabel('Concept:Protein:Resource').simplePath()) .times(2) .both('is_a', 'ac_by').hasLabel('Concept:Enzyme:Resource') .out('ac_by', 'in_by').hasLabel('Comp:Concept:Resource') .out('cs_by', 'pd_by').hasLabel('Concept:Resource:Transport') .emit() .repeat(both('part_of').hasLabel('Concept:Path:Resource').simplePath()) .times(3) .sample(100) .path().by('iri')
(2)添加索引加速查询
确保所有用到的标签(如Concept:Gene:Resource、Concept:Protein:Resource)都创建标签索引,iri属性创建属性索引,能大幅减少初始顶点查找和属性提取的时间。
(3)调整sample()的执行时机
当前sample()在所有路径生成后执行,若符合条件的路径数量极大,可尝试在生成通路顶点后立即采样,再回溯路径(需保证逻辑一致性):
g.V().hasLabel('Concept:Gene:Resource') // ... 前面的步骤不变 .emit() .repeat(both('part_of').hasLabel('Concept:Path:Resource').simplePath()) .times(3) .sample(100) // 先采样顶点,再获取路径 .path().by('iri')
(4)替代方案:显式路径的union写法
如果你的图数据库对固定长度路径优化更好,可以尝试用union()分情况处理0到N步的路径,虽然代码更长,但避免了repeat()的迭代开销:
g.V().hasLabel('Concept:Gene:Resource') .out('enc').hasLabel('Concept:Protein:Resource') // 处理蛋白质的0-2步关联 .union( __.identity(), // 0步 __.both('h_s_s', 'ortho', 'xref').hasLabel('Concept:Protein:Resource'), // 1步 __.both('h_s_s', 'ortho', 'xref').hasLabel('Concept:Protein:Resource') .both('h_s_s', 'ortho', 'xref').hasLabel('Concept:Protein:Resource').simplePath() // 2步 ) .both('is_a', 'ac_by').hasLabel('Concept:Enzyme:Resource') .out('ac_by', 'in_by').hasLabel('Comp:Concept:Resource') .out('cs_by', 'pd_by').hasLabel('Concept:Resource:Transport') // 处理通路的0-3步关联 .union( __.identity(), // 0步 __.both('part_of').hasLabel('Concept:Path:Resource'), //1步 __.both('part_of').hasLabel('Concept:Path:Resource') .both('part_of').hasLabel('Concept:Path:Resource').simplePath(), //2步 __.both('part_of').hasLabel('Concept:Path:Resource') .both('part_of').hasLabel('Concept:Path:Resource') .both('part_of').hasLabel('Concept:Path:Resource').simplePath() //3步 ) .sample(100) .path().by('iri')
3. 总结
emit()/repeat()/times()是实现变长路径最直观的方式,优先通过调整过滤顺序、添加索引来优化性能;若仍不满足需求,再考虑显式union的写法。
内容的提问来源于stack exchange,提问作者zakmck
相关产品推荐
相关产品推荐

