OpenFOAM在GCE与AWS EC2上的性能差异求助:EC2更优原因排查
Background & Observations
From your benchmarks running OpenFOAM on Intel Skylake-based instances:
- EC2 C5 (Frankfurt) delivers 30% faster execution speeds compared to GCE (Netherlands)
- EC2 also ends up being cheaper overall thanks to shorter runtime (full performance metrics are in the attached image)
The critical clue here is the Open MPI warning that pops up on GCE but not on EC2:
"A high-performance Open MPI point-to-point messaging module was unable to find any relevant network interfaces. Another transport will be used instead, although this may result in lower performance."
This almost certainly explains the performance difference—Open MPI is falling back to a slower network transport on GCE, while it’s leveraging a high-speed option on EC2.
Troubleshooting Steps to Fix the GCE Network Transport Issue
Let’s walk through actionable steps to diagnose and resolve this:
1. Check Available Network Interfaces
First, confirm what network interfaces your GCE instance has access to:
ip addr show
GCE’s primary high-bandwidth interface is usually ens4 (or nic0 on older instances). Make sure this interface is active and has the correct IP configuration.
2. Inspect Open MPI’s Transport Detection
Run this command to see which byte transfer layers (BTLs) Open MPI can detect:
ompi_info --param btl all --level 9
On a properly configured instance, you should see tcp (at minimum) or ofi (if RDMA is supported) listed as active high-performance transports. If these are missing, Open MPI will default to slower options like vader (shared memory only) or sm.
3. Verify GCE Network Configuration
- Premium Network Tier: Ensure your GCE instance uses the premium tier—standard tier has bandwidth caps that can bottleneck MPI traffic.
- SR-IOV Support: Check if your Skylake-based GCE instance type supports SR-IOV (this enables direct hardware access for networking, which MPI relies on for top performance). You can confirm this in GCE’s instance documentation.
- Firewall Rules: Make sure firewall rules allow MPI traffic. Open MPI uses dynamic ports by default, so either open the necessary port ranges or restrict MPI to a specific range with
--mca btl_tcp_port_range <start>:<end>when launching.
4. Force Open MPI to Use a High-Performance Transport
Try explicitly telling mpirun to use the TCP transport (reliable and high-speed) on GCE:
mpirun --mca btl tcp,self -np <number_of_processes> <your_openfoam_command>
If your instance supports RDMA, you can test the OFI transport instead:
mpirun --mca btl ofi,self -np <number_of_processes> <your_openfoam_command>
Run your benchmark again—if the warning disappears and performance improves, you’ve found the fix.
5. Match Open MPI & OpenFOAM Versions
Double-check that both instances are running the same versions of Open MPI and OpenFOAM. Version mismatches can lead to inconsistent transport detection behavior. Verify with:
mpirun --version foamVersion
内容的提问来源于stack exchange,提问作者Raj Chaudhary

