C++ std::string转std::wstring Windows正常Linux乱码解决方法
问题根因
std::wstring wsTmp(str.begin(), str.end());的逻辑是逐字符将char隐式转换为wchar_t,Windows下运行正常完全是巧合:
- Windows下
wchar_t占2字节,存储UTF-16编码,日常使用的窄字符串多为单字节ANSI代码页编码,字符取值范围0-255,零扩展到2字节后刚好匹配对应宽字符值 - Linux下
char默认为有符号类型,系统默认窄字符串编码为UTF-8:"pokémon"中的é在UTF-8下对应2个字节0xC3、0xA9,两个值作为有符号char均为负数,隐式转换为Linux下4字节长度的wchar_t(存储UTF-32编码)时会触发符号扩展,最终得到错误的\xffffffc3、\xffffffa9值。
可跨平台正常工作的转换方法
不要使用逐字符直接拷贝的错误逻辑,必须显式做编码转换,将对应编码的std::string转为平台wstring使用的宽字符编码:
方法1:C标准库实现(适配C11~C++20,跨平台通用)
C++17后std::wstring_convert和std::codecvt_utf8被标记为弃用,但目前GCC、Clang、MSVC所有主流编译器均保留支持,写跨平台工具代码稳定性足够:
#include <string> #include <locale> #include <codecvt> std::wstring utf8_to_wstring(const std::string& str) { std::wstring_convert<std::codecvt_utf8<wchar_t>> conv; return conv.from_bytes(str); }
调用时直接执行std::wstring wsTmp = utf8_to_wstring(str);即可,在Windows、Linux平台下都能正确转换UTF-8编码的"pokémon"为合法宽字符串。
方法2:POSIX原生接口实现(无标准库弃用接口顾虑)
如果不想使用被标记弃用的标准库接口,可直接用POSIX标准提供的多字节转宽字符接口,适配Linux等类Unix系统:
#include <string> #include <cstdlib> #include <clocale> std::wstring utf8_to_wstring(const std::string& str) { // 保存原有locale配置,转换完成后恢复,避免影响其他业务逻辑 const char* old_locale = std::setlocale(LC_ALL, nullptr); std::setlocale(LC_ALL, "en_US.UTF-8"); const char* src = str.c_str(); size_t len = std::mbstowcs(nullptr, src, 0); std::wstring res(len, L'\0'); std::mbstowcs(res.data(), src, len + 1); // 恢复原有locale配置 std::setlocale(LC_ALL, old_locale); return res; }
注意事项
- 逐字符直接拷贝的写法不存在编码转换逻辑,哪怕在Windows平台,只要
std::string存储的是UTF-8编码的多字节字符,一样会返回错误结果,不属于可通用的转换方案 - 转换逻辑必须和
std::string的实际编码匹配:如果窄字符串使用GBK、Latin-1等其他编码,要对应修改转换时的源编码参数,不能硬套UTF-8转换逻辑
内容的提问来源于stack exchange,提问作者Deep Learner
相关产品推荐
相关产品推荐

