如何将char*转换为std::u8string?(Windows环境专属方案)
将char*转换为std::u8string(Windows环境)
在Windows环境下,标准库、Boost及Eigen抛出的异常中,what()返回的char*通常采用系统默认ANSI代码页(比如GBK,取决于系统区域设置)编码,而非UTF-8。直接用std::u8string(error.what())构造会导致乱码,必须经过编码转换。
核心实现思路
Windows提供了MultiByteToWideChar和WideCharToMultiByte两个API,可完成「ANSI多字节 → UTF-16 → UTF-8」的转换,最终得到标准的std::u8string。
辅助函数实现
#include <string> #include <windows.h> std::u8string char_to_u8string(const char* mb_str) { if (!mb_str) return {}; // 第一步:将ANSI多字节字符串转为UTF-16(wchar_t) const int wide_buf_len = MultiByteToWideChar(CP_ACP, 0, mb_str, -1, nullptr, 0); if (wide_buf_len == 0) return {}; std::wstring wide_str(wide_buf_len, L'\0'); MultiByteToWideChar(CP_ACP, 0, mb_str, -1, wide_str.data(), wide_buf_len); // 第二步:将UTF-16转为UTF-8 const int u8_buf_len = WideCharToMultiByte(CP_UTF8, 0, wide_str.data(), -1, nullptr, 0, nullptr, nullptr); if (u8_buf_len == 0) return {}; std::u8string u8_str(u8_buf_len - 1, u8'\0'); // 减去自动添加的终止符长度 WideCharToMultiByte(CP_UTF8, 0, wide_str.data(), -1, reinterpret_cast<char*>(u8_str.data()), u8_buf_len, nullptr, nullptr); return u8_str; }
异常捕获场景的使用示例
#include <stdexcept> #include <boost/exception/diagnostic_information.hpp> #include <Eigen/Core> int main() { try { // 可能抛出目标异常的业务代码 throw std::runtime_error("测试异常信息"); } catch (const std::exception& e) { std::u8string u8_error = char_to_u8string(e.what()); // 处理UTF-8格式的错误信息 } catch (const boost::exception& e) { // Boost异常推荐用diagnostic_information获取详细信息 std::u8string u8_error = char_to_u8string(boost::diagnostic_information(e).c_str()); } catch (const Eigen::Exception& e) { std::u8string u8_error = char_to_u8string(e.what()); } return 0; }
关键说明
CP_ACP:代表Windows系统默认的ANSI代码页,精准匹配标准库/Boost/Eigen在Windows下输出的多字节错误信息编码。- 错误处理:示例中转换失败时返回空字符串,你可以根据业务需求调整为抛出异常或其他逻辑。
- 特殊情况:若确认某类异常的
what()已返回UTF-8编码(极少发生在Windows环境),可直接用std::u8string(mb_str)构造,但统一使用上述转换方法更稳妥。
内容的提问来源于stack exchange,提问作者user13583700
相关产品推荐
相关产品推荐

