如何处理参数数量未知的用户函数以开展高性能性能测试?
Absolutely! You can pull this off with templates and variadic arguments—no heavy std::function overhead, and full support for optimizations like inlining. Let’s break down the practical approaches:
Solution 1: Variadic Template Test Function
The simplest approach is to write a generic template that accepts any function (or callable) and its arguments directly. Since everything is resolved at compile time, the compiler can fully inline the function call and eliminate any unnecessary overhead.
#include <chrono> #include <iostream> // Generic performance test template - works with any function signature template <typename Func, typename... Args> void perf_test(Func func, Args&&... args) { // Start high-resolution timer const auto start = std::chrono::high_resolution_clock::now(); // Execute the function with perfect-forwarded arguments func(std::forward<Args>(args)...); // Calculate and print duration const auto end = std::chrono::high_resolution_clock::now(); const auto duration = std::chrono::duration_cast<std::chrono::nanoseconds>(end - start); std::cout << "Execution time: " << duration.count() << " nanoseconds\n"; } // Your example user functions void user_function1(double x[], double y[]) { // Dummy work for demonstration for (int i = 0; i < 1000; ++i) { y[i] = x[i] * 2.0; } } void user_function2(double x[], double y[], int z, double A[]) { // Dummy work for demonstration for (int i = 0; i < z; ++i) { y[i] = x[i] * A[i]; } } int main() { double x[1000], y[1000], A[1000]; const int z = 1000; // Test function 1 with its arguments perf_test(user_function1, x, y); // Test function 2 with its arguments perf_test(user_function2, x, y, z, A); return 0; }
How This Optimizes Well
- The template generates a specialized version of
perf_testfor each unique function signature. - The compiler knows exactly which function is being called at compile time, so it can fully inline
user_function1/user_function2into the test code. - Perfect forwarding (
std::forward) ensures no unnecessary copies of your arguments.
Solution 2: Pre-Bound Commands (Matching Your Original Idea)
If you need to wrap functions with pre-bound arguments (like your Command struct), you can use a template-based wrapper instead of std::function. This keeps all type information intact for the optimizer.
#include <chrono> #include <iostream> // Template-based Command wrapper - holds any callable template <typename Func> struct Command { Func func; // Overload operator() to accept whatever arguments the callable needs template <typename... Args> void operator()(Args&&... args) const { func(std::forward<Args>(args)...); } }; // Helper to create Command instances without explicit template arguments template <typename Func> Command<Func> make_command(Func func) { return {std::move(func)}; } // Reuse the same perf_test template from Solution 1 template <typename Func, typename... Args> void perf_test(Func func, Args&&... args) { const auto start = std::chrono::high_resolution_clock::now(); func(std::forward<Args>(args)...); const auto end = std::chrono::high_resolution_clock::now(); const auto duration = std::chrono::duration_cast<std::chrono::nanoseconds>(end - start); std::cout << "Execution time: " << duration.count() << " nanoseconds\n"; } // Your user functions (same as before) void user_function1(double x[], double y[]) { for (int i = 0; i < 1000; ++i) { y[i] = x[i] * 2.0; } } void user_function2(double x[], double y[], int z, double A[]) { for (int i = 0; i < z; ++i) { y[i] = x[i] * A[i]; } } int main() { double x[1000], y[1000], A[1000]; const int z = 1000; // Pre-bind arguments to user_function1 using a lambda auto cmd1 = make_command([&]() { user_function1(x, y); }); perf_test(cmd1); // No arguments needed here - lambda handles it // Pre-bind arguments to user_function2 auto cmd2 = make_command([&]() { user_function2(x, y, z, A); }); perf_test(cmd2); return 0; }
Why This Works Better Than std::bind
- The lambda captures variables by reference (or value, if you prefer) and the template wrapper preserves all type details.
- The compiler can inline the entire chain:
perf_testcalls the lambda, which calls the user function—no runtime indirection. - Unlike
std::function, there’s no type erasure, so no hidden virtual calls or heap allocations.
Why std::function Is a Bad Fit Here
std::function uses type erasure to wrap any callable, but this comes with overhead:
- It typically uses a virtual function call to invoke the wrapped function, which blocks inlining.
- There may be small heap allocations (for larger callables) or extra indirection that skews performance measurements.
- Since the target function’s type is hidden at compile time, the optimizer can’t do much to optimize the call.
Pro Tips for Accurate Performance Testing
- Enable Compiler Optimizations: Always compile with
-O2or-O3(GCC/Clang) or/O2(MSVC). Without optimizations, inlining won’t happen, and your results will be meaningless. - Run Multiple Iterations: For fast functions, run them thousands/millions of times to get a measurable duration. Modify
perf_testto loop over the function call:template <typename Func, typename... Args> void perf_test(Func func, Args&&... args, size_t iterations = 1'000'000) { const auto start = std::chrono::high_resolution_clock::now(); for (size_t i = 0; i < iterations; ++i) { func(std::forward<Args>(args)...); } const auto end = std::chrono::high_resolution_clock::now(); const auto total = std::chrono::duration_cast<std::chrono::nanoseconds>(end - start); std::cout << "Total: " << total.count() << " ns | Avg per call: " << total.count() / iterations << " ns\n"; } - Avoid Side Effects: Make sure your test functions do real work that the compiler can’t optimize away. For example, write results to a
volatilevariable if needed.
内容的提问来源于stack exchange,提问作者Georg RP

