AWS Neptune多可用区集群部署出现解压缩异常问题
AWS Neptune多AZ集群Gremlin查询抛出Decompression Exception问题
图结构插入语句
以下是构建测试图的Gremlin语句:
g.addV('ORG').as('1'). property(single, 'orgId', 'f5c').addV('COMP'). as('2'). property(single, 'compId', 2112896). property(single, 'owner', 'def').addV('COMP'). as('3'). property(single, 'compId', 2100198). property(single, 'owner', 'def').addV('COMP'). as('4'). property(single, 'compId', 4007384). property(single, 'owner', 'def').addV('COMP'). as('5'). property(single, 'compId', 2106827). property(single, 'owner', 'abc').addV('COMP'). as('6'). property(single, 'compId', 2106829). property(single, 'owner', 'abc').addV('COMP'). as('7'). property(single, 'compId', 2104080). property(single, 'owner', 'abc').addV('COMP'). as('8'). property(single, 'compId', 2110851). property(single, 'owner', 'abc').addV('ORG'). as('9'). property(single, 'orgId', 28932).addE('edge'). from('1').to('5').property('sub', '951a'). addE('edge').from('1').to('2'). property(single, 'sub', 5779).addE('edge').from('2'). to('3').property(single, 'sub', 5779).addE('edge'). from('3').to('4').property(single, 'sub', 5779). addE('edge').from('5').to('6'). property(single, 'sub', '951a').addE('edge'). from('6').to('7').property(single, 'sub', 951a').addE('edge'). from('7').to('8'). property(single, 'sub', '951a').addE('edge'). from('4').to('9').property(single, 'sub', 1234). addE('edge').from('8').to('9'). property(single, 'sub', 465474)
查询语句
用于获取路径的Gremlin查询:
g.V().has("orgId", "f5c") .repeat(bothE().otherV().simplePath()) .until(hasLabel("ORG")).path().by(valueMap())
异常现象
- 该查询在Neptune单实例环境运行正常
- 在多可用区(Multi-AZ)集群环境(Neptune引擎版本1.2.1.0,数据量较大)执行时,抛出Decompression Exception异常,异常栈信息如下:
{...}
慢查询日志
查询产生的慢日志相关信息:
"queryStats": { "query": "g.V().has(\"orgId\",\"abc\").repeat(__.bothE(\"service\").otherV().simplePath()).until(__.has(\"orgId\",P.neq(\"\"))).path().by(__.valueMap())", "queryFingerprint": "g.V().has(string0,string1).repeat(__.bothE(string2).otherV().simplePath()).until(__.has(string0,P.neq(string3))).path().by(__.valueMap())", "queryLanguage": "Gremlin" }
使用的依赖
项目中使用的Maven依赖:
<dependency> <groupId>com.amazonaws</groupId> <artifactId>amazon-neptune-sigv4-signer</artifactId> <version>2.4.0</version> </dependency> <dependency> <groupId>org.apache.tinkerpop</groupId> <artifactId>gremlin-driver</artifactId> <version>3.6.2</version> </dependency> <dependency> <groupId>org.apache.tinkerpop</groupId> <artifactId>gremlin-core</artifactId> <version>3.6.2</version> </dependency>
集群连接代码
使用SigV4签名的Neptune集群连接代码:
public Cluster cluster() { return Cluster.build(neptuneUrl).port(portNo).enableSsl(true).handshakeInterceptor(r -> { try { NeptuneNettyHttpSigV4Signer sigV4Signer = new NeptuneNettyHttpSigV4Signer(neptuneRegion, new DefaultAWSCredentialsProviderChain()); sigV4Signer.signRequest(r); } catch (NeptuneSigV4SignerException e) { log.error("** Error while signing the request ** " + e.getMessage()); } return r; }).maxConnectionPoolSize(10).minConnectionPoolSize(5).maxInProcessPerConnection(3).minSimultaneousUsagePerConnection(1) .maxSimultaneousUsagePerConnection(3) .create(); }
排查与解决方案
1. 限制查询结果规模
Decompression Exception通常因查询返回数据量过大,超出驱动解压处理能力导致,可:
- 在查询末尾添加
limit()限制结果数量,比如.limit(100) - 优化
valueMap(),仅返回所需属性,比如.by(valueMap('orgId', 'compId')),减少单条结果的数据体积
2. 优化查询性能
原查询在大数据量下性能低下,易导致响应数据膨胀:
- 明确边方向:业务允许时,将
bothE()改为outE('edge')或inE('edge'),避免遍历所有类型边 - 添加遍历深度限制:在
repeat()中加入times()限制最大深度,比如.repeat(...).times(6),防止无限制遍历 - 过滤冗余数据:在
repeat()中添加边标签过滤,比如bothE('edge'),仅遍历目标类型边 - 替换
simplePath():业务允许时,改用dedup()等更轻量的去重方式,降低遍历开销
3. 调整驱动配置
检查Gremlin驱动的压缩与缓冲区配置:
- 显式指定序列化器,避免压缩兼容性问题:
.serializer(Serializers.GRAPHBINARY_V1) - 调整驱动最大内容长度:添加
.maxContentLength(10485760)(10MB),适配大响应数据 - 对齐压缩设置:确保驱动与Neptune集群压缩配置一致,Neptune默认启用gzip,可在驱动中显式配置
.enableSsl(true).compress(true)
4. 升级Neptune引擎版本
当前使用的1.2.1.0版本可能存在已知的解压或查询性能问题,查看AWS Neptune版本发布日志,升级至较新稳定版本(如1.2.2.0及以上)验证是否修复问题。
5. 检查SigV4签名拦截器
确保签名过程未修改HTTP请求的压缩相关头(如Accept-Encoding),避免干扰Neptune的压缩响应处理。
内容的提问来源于stack exchange,提问作者Parvesh Kumar
相关产品推荐
相关产品推荐

