Gremlin查询问题:关联边属性与独立顶点属性获取指定数据
图数据库Gremlin查询优化问题
图结构
v1: Protein{prefName: 'QP1'} -- r1: part_of{evidence: 'ns:testdb'} --> v2: Protcmplx{prefName: 'P12 Complex'} ev: EvidenceType{ iri = "ns:testdb", label = "Test Database" }
需求
编写Gremlin查询,获取part_of关系实例,返回以下内容:
- 来源Protein顶点的
prefName - 目标Protcmplx顶点的
prefName - 边的
evidence对应的EvidenceType顶点的label
问题查询及症状
尝试的查询语句:
g.V().hasLabel( containing('Protein') ).as('p') .outE().hasLabel( 'is_part_of' ).as('pr') .inV().hasLabel( containing('Protcmplx') ).as('cpx') .V().hasLabel( containing('EvidenceType') ).as('ev') .has( 'iri', eq( select('pr').by('evidence') ) ) .select( 'p', 'cpx', 'ev', 'pr' ) .by('prefName') .by('prefName') .by('label') .by('evidence') .limit(100)
该查询在仅数千个节点和边的场景下耗时极长,最终无结果返回,但确认数据确实存在。推测问题出在has( 'iri', ... )的属性匹配逻辑上,不知道如何正确将边属性与未关联的顶点属性做匹配(采用这种建模方式是因为LPG模型不支持超边)。
问题原因及修正方案
核心问题
原查询中途调用.V()会重新遍历所有顶点,完全断开了之前的Protein->part_of->Protcmplx路径,导致后续的匹配变成了笛卡尔积式的全量关联,性能直接爆炸。另外还有一个细节错误:图中边的标签是part_of,但查询里写的是is_part_of,这也会导致匹配不到边。
正确查询写法
推荐使用match语句进行结构化查询,逻辑更清晰,性能更优:
g.match( __.as('p').hasLabel(containing('Protein')), __.as('p').outE('part_of').as('pr'), __.as('pr').inV().hasLabel(containing('Protcmplx')).as('cpx'), __.as('ev').hasLabel(containing('EvidenceType')).has('iri', select('pr').by('evidence')) ) .select('p', 'cpx', 'ev', 'pr') .by('prefName') .by('prefName') .by('label') .by('evidence') .limit(100)
或者用链式遍历的写法,在遍历到边之后,基于边的evidence属性直接查找对应的EvidenceType顶点:
g.V().hasLabel(containing('Protein')).as('p') .outE('part_of').as('pr') .inV().hasLabel(containing('Protcmplx')).as('cpx') // 基于边的evidence属性精准查找对应证据顶点 .V().hasLabel(containing('EvidenceType')) .has('iri', select('pr').by('evidence')) .as('ev') .select('p', 'cpx', 'ev', 'pr') .by('prefName') .by('prefName') .by('label') .by('evidence') .limit(100)
这两种写法都会保持路径的关联性,避免全量笛卡尔积匹配,同时修正了边标签的错误,能够快速匹配到目标数据。
内容的提问来源于stack exchange,提问作者zakmck
相关产品推荐
相关产品推荐

