Visual Studio中能否零开销安全处理std::variant?
能否在Visual Studio中让std::variant实现零开销且安全的容器逻辑?
我需要处理这样的场景:比较两个std::variant<int*, float*, double*>的大小,逻辑实现如下:
#include <variant> extern bool bar(),foo(); bool lessForMyVariants(const std::variant<int*, float*, double *> x, const std::variant<int*, float*, double *> y) { if (x.index()!=y.index()) { return x.index()<y.index(); } else { switch(x.index()) { case 0: if (x.index()==0 && y.index()==0) return *std::get<0>(x)<*std::get<0>(y); break; case 1: if (x.index()==1 && y.index()==1) return *std::get<1>(x)<*std::get<1>(y); break; case 2: if (x.index()==2 && y.index()==2) return *std::get<2>(x)<*std::get<2>(y); break; default: return foo(); } } return bar(); }
由于在switch的case分支中,x.index()和y.index()的值已经被确保一致且正确,理论上std::get调用不会触发异常,也不会走到bar()或foo()的分支。GCC可以完成这个优化,但Visual Studio无法实现相同效果。
注意事项
- 我知道代码里的if检查是冗余的,但这是触发GCC优化的必要条件,如果能去掉当然更好。
- 和《Unsafe,
noexceptand no-overhead way of accessingstd::variant》这篇内容不同,我不关注不安全的访问方式。 - 欢迎提出
std::variant的替代方案,只要能解决这个问题就行。 foo和bar的调用只是用来验证编译器输出的,没有实际业务逻辑。
性能测试:switch vs std::visit
很多人认为std::visit比switch更快,但我的测试结果正好相反。测试代码如下:
// ConsoleApplication1.cpp : This file contains the 'main' function. Program execution begins and ends there. // #include <variant> #include <vector> #include <iostream> #include <algorithm> #include <chrono> typedef std::variant<int*, float*, double*> V; extern bool bar() { throw 2; } extern bool foo() { throw 3; } bool lessForMyVariantsSwitch(const std::variant<int*, float*, double*> x, const std::variant<int*, float*, double*> y) { if (x.index() != y.index()) { return x.index() < y.index(); } else { switch (x.index()) { case 0: if (x.index() == 0 && y.index() == 0) return *std::get<0>(x) < *std::get<0>(y); break; case 1: if (x.index() == 1 && y.index() == 1) return *std::get<1>(x) < *std::get<1>(y); break; case 2: if (x.index() == 2 && y.index() == 2) return *std::get<2>(x) < *std::get<2>(y); break; default: return foo(); } } return bar(); } bool lessForMyVariantsSwitchConstSimple(const std::variant<int*, float*, double*>&x, const std::variant<int*, float*, double*>&y) { if (x.index() != y.index()) { return x.index() < y.index(); } else { switch (x.index()) { case 0: return *std::get<0>(x) < *std::get<0>(y); break; case 1: return *std::get<1>(x) < *std::get<1>(y); break; case 2: return *std::get<2>(x) < *std::get<2>(y); break; default: return foo(); } } return bar(); } bool lessForMyVariantsSwitchConst(const std::variant<int*, float*, double*>&x, const std::variant<int*, float*, double*>&y) { if (x.index() != y.index()) { return x.index() < y.index(); } else { switch (x.index()) { case 0: if (x.index() == 0 && y.index() == 0) return *std::get<0>(x) < *std::get<0>(y); break; case 1: if (x.index() == 1 && y.index() == 1) return *std::get<1>(x) < *std::get<1>(y); break; case 2: if (x.index() == 2 && y.index() == 2) return *std::get<2>(x) < *std::get<2>(y); break; default: return foo(); } } return bar(); } // helper type for the visitor #4 template<class... Ts> struct overloaded : Ts... { using Ts::operator()...; }; // explicit deduction guide (not needed as of C++20) template<class... Ts> overloaded(Ts...)->overloaded<Ts...>; bool lessForMyVariantsVisit(const std::variant<int*, float*, double*> x, const std::variant<int*, float*, double*> y) { return std::visit(overloaded{ [] <typename T>(T * lhs, T * rhs) { return *lhs < *rhs; }, [&](auto,auto) { return x.index() < y.index(); } }, x, y); } bool lessForMyVariantsVisitConst(const std::variant<int*, float*, double*>&x, const std::variant<int*, float*, double*>&y) { return std::visit(overloaded{ [] <typename T>(T * lhs, T * rhs) { return *lhs < *rhs; }, [&](auto,auto) { return x.index() < y.index(); } }, x, y); } template <class P> size_t checkSort(std::vector<V> const& v, P p, const char* whichSort) { size_t z=0; std::vector<V> v2; auto t1 = std::chrono::high_resolution_clock::now(); constexpr int maxNum = 1000000; for (int j = 0; j < maxNum; ++j) { v2 = v; //std::ranges::sort(v2, p); std::sort(v2.begin(), v2.end(), p); z += v2[0].index(); } auto t2 = std::chrono::high_resolution_clock::now(); std::cout << whichSort <<" took " << std::chrono::duration_cast<std::chrono::nanoseconds>(t2 - t1).count()*1.0 / maxNum << " nanoseconds per sort\n"; return z; } int main() { int varr[4] = { 6,7,1,10 }; float farr[4] = { 5.0f, 1.2f, 4.5f, 2.2f }; double darr[4] = { 5.0, 1.2, 4.5, 2.2 }; std::vector<V> v; for (int i = 0; i < 4; ++i) { v.emplace_back(varr + i); v.emplace_back(farr + i); v.emplace_back(darr + i); } double z=0; z+=checkSort(v, lessForMyVariantsSwitch, "lessForMyVariantsSwitch"); z += checkSort(v, lessForMyVariantsVisit, "lessForMyVariantsVisit"); z += checkSort(v, lessForMyVariantsSwitchConst, "lessForMyVariantsSwitchConst const&"); z += checkSort(v, lessForMyVariantsVisitConst, "lessForMyVariantsVisitConst const&"); z += checkSort(v, lessForMyVariantsSwitchConstSimple, "lessForMyVariantsSwitchConstSimple const&"); std::cout << "dummy: " << z; return 0; }
在Visual Studio 2022的/O2优化下,switch版本的实现速度更快:
- lessForMyVariantsSwitch:单次排序耗时90.0676纳秒
- lessForMyVariantsVisit:单次排序耗时125.441纳秒
- lessForMyVariantsSwitchConst const&:单次排序耗时88.1962纳秒
- lessForMyVariantsVisitConst const&:单次排序耗时121.182纳秒
- lessForMyVariantsSwitchConstSimple const&:单次排序耗时96.642纳秒
在WSL-Ubuntu下使用g++ 11.3.0和clang++ 14.0.0测试也得到类似结果,唯一区别是clang++中lessForMyVariantsSwitchConstSimple的速度最快。
内容的提问来源于stack exchange,提问作者Hans Olsson
相关产品推荐
相关产品推荐

