AKS扩缩容异常:4 Pod比1 Pod的Requests per second更低的原因
AKS扩缩容测试异常:多Pod反而性能下降的原因分析
我在使用Azure Kubernetes Service(AKS)进行服务扩缩容测试时,遇到了反常现象:4个Pod的性能表现反而不如1个Pod,具体Apache Benchmark测试结果如下:
1个Pod时的测试结果
Server Software: nginx/1.16.0 Server Hostname: demo Server Port: 80 Document Path: / Document Length: 612 bytes Concurrency Level: 500 Time taken for tests: 1.712 seconds Complete requests: 10000 Failed requests: 0 Total transferred: 8450000 bytes HTML transferred: 6120000 bytes Requests per second: 5842.36 [#/sec] (mean) Time per request: 85.582 [ms] (mean) Time per request: 0.171 [ms] (mean, across all concurrent requests) Transfer rate: 4821.09 [Kbytes/sec] received Connection Times (ms) min mean[+/-sd] median max Connect: 2 36 6.1 36 73 Processing: 9 47 9.1 46 73 Waiting: 1 35 8.1 33 55 Total: 48 84 8.2 83 121 Percentage of the requests served within a certain time (ms) 50% 83 66% 86 75% 88 80% 90 90% 94 95% 97 98% 101 99% 103 100% 121 (longest request)
4个Pod时的测试结果
Server Software: nginx/1.16.0 Server Hostname: demo Server Port: 80 Document Path: / Document Length: 612 bytes Concurrency Level: 500 Time taken for tests: 2.537 seconds Complete requests: 10000 Failed requests: 0 Total transferred: 8450000 bytes HTML transferred: 6120000 bytes Requests per second: 3941.67 [#/sec] (mean) Time per request: 126.850 [ms] (mean) Time per request: 0.254 [ms] (mean, across all concurrent requests) Transfer rate: 3252.65 [Kbytes/sec] received Connection Times (ms) min mean[+/-sd] median max Connect: 3 51 13.4 52 136 Processing: 19 73 24.3 66 171 Waiting: 1 55 23.0 49 148 Total: 59 125 25.2 121 219 Percentage of the requests served within a certain time (ms) 50% 121 66% 130 75% 137 80% 141 90% 152 95% 179 98% 195 99% 207 100% 219 (longest request)
问题
为什么4个Pod时的Requests per second反而比1个Pod更低,且Time per request有所增加?
可能的原因分析
- 负载均衡调度开销:AKS中的负载均衡器(或Ingress Controller)在分发请求到多个Pod时,会额外增加调度、会话管理或健康检查的开销。如果负载均衡策略存在问题(比如会话粘性导致请求集中到部分Pod,或者调度算法延迟过高),会拖慢整体响应速度。
- 节点资源瓶颈:如果4个Pod都部署在同一个节点上,节点的CPU、内存、网络带宽可能已达饱和状态,多个Pod之间争抢资源,导致单个Pod的处理能力下降,最终整体吞吐量反而降低。可以通过
kubectl top nodes命令检查节点资源使用率。 - Ingress Controller性能限制:若使用Ingress路由请求,Ingress Controller本身可能成为性能瓶颈。比如单个Ingress Pod无法处理高并发请求,导致请求排队,即使后端Pod数量增加,整体吞吐量也无法提升。
- 测试环境干扰:两次测试时集群状态可能存在差异,比如节点上运行了其他负载、网络出现波动,或者Azure底层资源调度(如节点迁移)影响了性能。建议多次测试取平均值,排除偶然因素。
- Pod预热不充分:如果4个Pod刚启动就进行测试,可能未完成预热(如缓存加载、运行时初始化),导致初始处理性能较低。可以等待Pod稳定运行一段时间后再执行测试。
内容的提问来源于stack exchange,提问作者Duc Vo
相关产品推荐
相关产品推荐

