You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

C++ std::string转std::wstring Windows正常Linux乱码解决方法

问题根因

std::wstring wsTmp(str.begin(), str.end());的逻辑是逐字符将char隐式转换为wchar_t,Windows下运行正常完全是巧合:

  • Windows下wchar_t占2字节,存储UTF-16编码,日常使用的窄字符串多为单字节ANSI代码页编码,字符取值范围0-255,零扩展到2字节后刚好匹配对应宽字符值
  • Linux下char默认为有符号类型,系统默认窄字符串编码为UTF-8:"pokémon"中的é在UTF-8下对应2个字节0xC3、0xA9,两个值作为有符号char均为负数,隐式转换为Linux下4字节长度的wchar_t(存储UTF-32编码)时会触发符号扩展,最终得到错误的\xffffffc3、\xffffffa9值。
可跨平台正常工作的转换方法

不要使用逐字符直接拷贝的错误逻辑,必须显式做编码转换,将对应编码的std::string转为平台wstring使用的宽字符编码:

方法1:C标准库实现(适配C11~C++20,跨平台通用)

C++17后std::wstring_convert和std::codecvt_utf8被标记为弃用,但目前GCC、Clang、MSVC所有主流编译器均保留支持,写跨平台工具代码稳定性足够:

#include <string>
#include <locale>
#include <codecvt>

std::wstring utf8_to_wstring(const std::string& str) {
    std::wstring_convert<std::codecvt_utf8<wchar_t>> conv;
    return conv.from_bytes(str);
}

调用时直接执行std::wstring wsTmp = utf8_to_wstring(str);即可,在Windows、Linux平台下都能正确转换UTF-8编码的"pokémon"为合法宽字符串。

方法2:POSIX原生接口实现(无标准库弃用接口顾虑)

如果不想使用被标记弃用的标准库接口,可直接用POSIX标准提供的多字节转宽字符接口,适配Linux等类Unix系统:

#include <string>
#include <cstdlib>
#include <clocale>

std::wstring utf8_to_wstring(const std::string& str) {
    // 保存原有locale配置,转换完成后恢复,避免影响其他业务逻辑
    const char* old_locale = std::setlocale(LC_ALL, nullptr);
    std::setlocale(LC_ALL, "en_US.UTF-8");
    const char* src = str.c_str();
    size_t len = std::mbstowcs(nullptr, src, 0);
    std::wstring res(len, L'\0');
    std::mbstowcs(res.data(), src, len + 1);
    // 恢复原有locale配置
    std::setlocale(LC_ALL, old_locale);
    return res;
}
注意事项
  • 逐字符直接拷贝的写法不存在编码转换逻辑,哪怕在Windows平台,只要std::string存储的是UTF-8编码的多字节字符,一样会返回错误结果,不属于可通用的转换方案
  • 转换逻辑必须和std::string的实际编码匹配:如果窄字符串使用GBK、Latin-1等其他编码,要对应修改转换时的源编码参数,不能硬套UTF-8转换逻辑

内容的提问来源于stack exchange,提问作者Deep Learner

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.29 02:33:08