为何AWS EC2上nvidia-docker的TensorFlow Serving推理耗时翻倍?
Great question—this performance gap is actually pretty common when comparing these two GPUs, and it boils down to a few key hardware and environment differences:
Architectural Generational Gap
The GTX 1080 uses NVIDIA's Pascal architecture (GP104), while the Tesla K80 is based on the older Kepler architecture (GK210). Pascal is a significant leap forward:- The GTX 1080 delivers ~8.2 TFLOPS of single-precision floating-point performance, whereas a single K80 GPU (since K80 is a dual-core card, and EC2 instances often expose one core by default) only hits ~4.35 TFLOPS. That's nearly a 2x difference right there.
- Pascal also includes improved instruction sets and core efficiency, making it better at handling modern ML workloads like TensorFlow's optimized operators.
Memory Bandwidth Limitations
The GTX 1080 uses GDDR5X memory with a bandwidth of 320 GB/s, while the K80's GDDR5 only offers 240 GB/s. For many deep learning inference tasks—especially with large models or batch sizes—memory bandwidth is a major bottleneck. The lower bandwidth on K80 can slow down data transfer between GPU memory and cores, directly increasing inference time.EC2 Instance Power/Performance Modes
Tesla K80s on EC2 are often configured to run in power-saving modes by default to reduce costs. This can throttle the GPU's clock speed, limiting its peak performance. In contrast, your local GTX 1080 is likely running at full power with no throttling, maximizing its output. You can verify this by runningnvidia-smion both systems to compare GPU clock speeds and utilization rates.Software Optimization Prioritization
TensorFlow and TensorFlow Serving prioritize optimizations for newer GPU architectures. Pascal has been around longer and is more widely used in consumer and professional ML setups, so it benefits from more refined operator implementations and compiler optimizations. Kepler, being an older architecture, doesn't get the same level of ongoing optimization love, leading to less efficient execution of inference workloads.
To confirm these hypotheses, you could:
- Run
nvidia-smi dmonon both systems to monitor GPU clock speeds, memory usage, and utilization during inference. - Run a simple CUDA benchmark (like a custom TensorFlow-based compute test) to compare raw compute and memory performance between the two GPUs.
内容的提问来源于stack exchange,提问作者saurav agarwal

