含重音字符的UTF-8 std::string转UTF-16 std::wstring失败求助
解决Windows下std::string含重音字符转std::wstring的bad conversion异常
问题根源
你的代码中std::string s{"hello ê world"}在Windows中文环境下默认是GBK编码,但你使用的std::codecvt_utf8<wchar_t>是用于将UTF-8编码的字节序列转换为UTF-16(Windows下wchar_t对应UTF-16)的std::wstring,编码格式不匹配直接导致转换失败,抛出bad conversion异常。
解决方案
根据你的场景需求,提供两种可靠的解决方式:
方式一:确保源字符串为UTF-8编码
如果你的字符串是代码中的字面量,直接用u8前缀强制字符串为UTF-8编码,配合原有的转换器即可正常转换:
#include <iostream> #include <string> #include <locale> #include <codecvt> #pragma warning( disable : 4996 ) int main() { // u8前缀强制字面量为UTF-8编码 std::string s{u8"hello ê world"}; try { std::wstring ws = std::wstring_convert<std::codecvt_utf8<wchar_t>>().from_bytes(s); // 配置wcout的本地环境,避免宽字符输出乱码 std::wcout.imbue(std::locale("")); std::wcout << ws << "\n"; } catch (const std::exception& e) { std::cout << e.what() << "\n"; } }
方式二:处理系统默认编码(如GBK)的字符串
如果你的std::string来自外部输入(如文件、控制台),是Windows系统默认的ANSI编码(中文环境为GBK),推荐使用Windows原生APIMultiByteToWideChar来转换,兼容性和稳定性更好:
#include <iostream> #include <string> #include <windows.h> #pragma warning( disable : 4996 ) int main() { std::string s{"hello ê world"}; // 获取所需宽字符缓冲区的大小 int wstr_len = MultiByteToWideChar(CP_ACP, 0, s.c_str(), -1, nullptr, 0); if (wstr_len == 0) { std::cout << "转换失败,错误码: " << GetLastError() << "\n"; return 1; } std::wstring ws(wstr_len, L'\0'); // 执行转换 if (MultiByteToWideChar(CP_ACP, 0, s.c_str(), -1, ws.data(), wstr_len) == 0) { std::cout << "转换失败,错误码: " << GetLastError() << "\n"; return 1; } // 移除转换时自动添加的末尾空字符 ws.pop_back(); // 配置输出环境避免乱码 std::wcout.imbue(std::locale("")); std::wcout << ws << "\n"; return 0; }
补充说明
- 注意
std::wcout默认可能无法正确输出宽字符,需要调用std::wcout.imbue(std::locale(""))来适配系统本地编码。 std::wstring_convert和std::codecvt_utf8在C++17标准中被标记为弃用,若追求长期兼容性,优先选择Windows原生API或第三方编码库。
内容的提问来源于stack exchange,提问作者DailyLearner
相关产品推荐
相关产品推荐

