Eigen Tensor计算Z-Scores代码比Numpy慢500倍,求问题原因
Hey there! Let's walk through the key issues with your current C++ code and how to get it matching (or even beating!) NumPy's performance.
1. You're Probably Not Compiling With Optimization Flags
This is the most common culprit when Eigen code lags behind optimized libraries like NumPy. NumPy is built with aggressive optimizations (-O3, -march=native, vectorization enabled) out of the box, but default C++ compiler settings use -O0 (no optimizations) which leaves massive performance on the table.
Fix: Compile your code with these flags:
g++ -O3 -march=native -std=c++17 your_file.cpp -o your_program
-O3: Enables all high-level optimizations for speed-march=native: Optimizes for your CPU's specific instruction set (like AVX2, AVX-512)-std=c++17: Ensures modern C++ support for Eigen's full feature set
2. Redundant Computations & Unnecessary Memory Allocations
Your code calculates scaled_a - scaled_b twice (once for means, once for submean), forcing Eigen to redo expensive arithmetic and allocate new memory for the result each time—this is a huge waste of cycles.
Worse, your standard deviation calculation has redundant and mathematically incorrect steps:
auto stds = submean.mean(dims).reshape(good_dims).abs().square().mean(dims).reshape(good_dims).sqrt();
Standard deviation requires squaring deviations first, then taking the mean, then the square root. Your current code computes the mean of deviations first, which is wrong and adds extra slow operations.
3. Suboptimal Memory Layout
Eigen's Tensor defaults to column-major storage (Fortran-style), while NumPy uses row-major (C-style) by default. If your data access patterns don't match the storage layout, you'll get poor cache utilization, which cripples performance.
Fix: Explicitly use row-major layout for your tensors to match NumPy:
Eigen::Tensor<float, 2, Eigen::RowMajor> a(2, 5); Eigen::Tensor<float, 2, Eigen::RowMajor> b(2, 5);
4. Revised, Optimized Code
Here's a fixed version of your code that addresses all the above issues:
#include <Eigen/Dense> #include <Eigen/Tensor> #include <iostream> #include <cstdio> #include <time.h> using namespace Eigen; // Helper to get nanosecond timestamp inline uint64_t get_nanos() { struct timespec ts; clock_gettime(CLOCK_MONOTONIC, &ts); return (uint64_t)ts.tv_sec * 1000000000L + ts.tv_nsec; } int main(int argc, char** argv) { int scale = atoi(argv[1]); Eigen::array<int, 2> bbcast({scale, 1}); long startTime = get_nanos(); // Use RowMajor to match NumPy's memory layout for better cache performance Eigen::Tensor<float, 2, Eigen::RowMajor> a(2, 5); a.setRandom(); Eigen::Tensor<float, 2, Eigen::RowMajor> b(2, 5); b.setRandom(); // Broadcast once, compute difference once and reuse everywhere auto scaled_a = a.broadcast(bbcast); auto scaled_b = b.broadcast(bbcast); auto diff = scaled_a - scaled_b; // Store the core difference once // Calculate means with minimal overhead Eigen::array<int, 1> reduce_dims({0}); auto means = diff.mean(reduce_dims); // Broadcast means to match diff's shape without extra reshaping Eigen::array<int, 2> mean_bcast({scaled_a.dimension(0), 1}); auto means_broadcasted = means.broadcast(mean_bcast); std::cout << means << std::endl; printf("Calculated means, took %f seconds\n",(float)(get_nanos() - startTime) / 1000000000L); // Calculate standard deviation correctly and efficiently auto squared_deviations = (diff - means_broadcasted).square(); auto stds = squared_deviations.mean(reduce_dims).sqrt(); std::cout << stds << std::endl; printf("Calculated std, took %f seconds\n",(float)(get_nanos() - startTime) / 1000000000L); // Calculate final z-scores (uncomment if you need them) // auto z_scores = (diff - means_broadcasted) / stds.broadcast(mean_bcast); // std::cout << z_scores << std::endl; return 0; }
Key Improvements in This Code:
- Single calculation of
diff: Avoids redundant arithmetic and memory allocations - Correct standard deviation calculation: Fixes the mathematical error and removes unnecessary operations
- RowMajor layout: Improves cache efficiency for row-wise access patterns
- Minimized reshaping: Uses broadcast directly on reduced tensors instead of extra reshape steps
- Lazy evaluation: Eigen's Tensor uses expression templates where possible, so intermediate operations don't allocate memory until necessary
Final Notes
After compiling with the optimized flags, you should see performance that's on par with NumPy—if not faster! Eigen's Tensor module is just as capable as NumPy for numerical computations, but it requires careful use of optimizations and avoiding redundant operations to reach its full potential.
内容的提问来源于stack exchange,提问作者Moshe Kravchik

