You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

AWS ECS Service Connect 间歇性10秒超时问题求助(直接IP调用稳定)

AWS ECS Service Connect 间歇性10秒超时问题求助(直接IP调用稳定)

大家好,我现在遇到一个非常棘手的AWS ECS Service Connect问题——服务间调用总是间歇性超时,折腾了好久排除了不少可能,但现在卡壳了,迹象都指向Service Connect代理本身有问题,想请教下大家有没有遇到过类似情况或者有排查思路?

基础设施概述

  • AWS ECS Fargate环境,采用标准客户端-服务端架构
  • 网络模式为awsvpc,部署在双栈(IPv4/IPv6)VPC的私有子网中
  • 使用ECS Service Connect默认的.local命名空间
  • 用AWS CDK拆分栈,通过SSM参数存储管理运行时配置

我观察到的现象

当客户端服务(service-a)调用下游服务(service-b)时,请求经常会卡10秒左右然后失败,具体情况如下:

1. 负载测试结果

负载测试把问题暴露得非常明显:

  • 通过Service Connect路由时,失败率约12%,99分位延迟高达14000ms
  • 但当service-a直接调用service-b的私有IP时,失败率为0%,99分位延迟只有120ms(原本有测试截图,暂时无法展示)

2. 手动curl测试

为了排除并发负载的影响,我在service-a的容器里单独执行curl请求,结果同样出现间歇性10秒超时:

# 快速的IPv4调用
$ curl -4 -s -w '→ IP=%{remote_ip} total=%{time_total}s\n' -o /dev/null http://service-b.local:5000/api/v1/b/test
→ IP=127.255.0.2 total=0.007104s

# 超时的IPv6调用
$ curl -6 -s -w '→ IP=%{remote_ip} total=%{time_total}s\n' -o /dev/null http://service-b.local:5000/api/v1/b/test
→ IP=2600:f0f0::2 total=9.999628s

# 另一个超时的IPv6调用
$ curl -6 -s -w '→ IP=%{remote_ip} total=%{time_total}s\n' -o /dev/null http://service-b.local:5000/api/v1/b/test
→ IP=2600:f0f0::2 total=9.994741s

# 快速的IPv6调用
$ curl -6 -s -w '→ IP=%{remote_ip} total=%{time_total}s\n' -o /dev/null http://service-b.local:5000/api/v1/b/test
→ IP=2600:f0f0::2 total=0.005802s

# 另一个超时的IPv6调用
$ curl -6 -s -w '→ IP=%{remote_ip} total=%{time_total}s\n' -o /dev/null http://service-b.local:5000/api/v1/b/test
→ IP=2600:f0f0::2 total=10.001393s

# 超时的IPv4调用
$ curl -4 -s -w '→ IP=%{remote_ip} total=%{time_total}s\n' -o /dev/null http://service-b.local:5000/api/v1/b/test
→ IP=127.255.0.2 total=9.999234s

# 快速的IPv4调用
$ curl -4 -s -w '→ IP=%{remote_ip} total=%{time_total}s\n' -o /dev/null http://service-b.local:5000/api/v1/b/test
→ IP=127.255.0.2 total=0.002342s

已经排除的可能性

我已经逐一排查并排除了以下因素:

  • 目标服务(service-b)本身健康问题:直接IP调用完全正常,说明service-b没有问题
  • HTTP客户端库问题:用Python的httpx和aiohttp都能复现同样的间歇性超时,排除了客户端库的锅
  • 基础网络连通性问题:很多请求能成功,说明安全组、NACL、VPC路由配置没有根本性错误
  • 仅IPv6问题:curl测试显示IPv4和IPv6连接都会出现间歇性超时
  • 仅高负载问题:单个孤立请求也会失败,和并发流量无关

求助

现在实在想不到还有什么排查方向了,有没有大佬能指点下问题可能出在哪?或者有没有什么我没考虑到的排查步骤?万分感谢!

内容来源于stack exchange

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.07 09:53:10