You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Linux下UTF-8中文字符串显示异常问题求助

Linux下UTF-8中文显示异常的解决方法

问题根源

Linux环境中std::wcout的默认locale通常不兼容UTF-16,且终端普遍采用UTF-8编码,直接输出ICU库返回的UChar*(UTF-16格式)会因编码不匹配导致乱码。而Windows控制台默认支持UTF-16输出,因此能正常显示内容。

解决方案

方案1:将UnicodeString转回UTF-8后用std::cout输出

这是适配Linux终端UTF-8环境的最直接方式:

#include <unicode/unistr.h>
#include <iostream>

int main() {
    std::string utf8String = "1101\241@\245x\252d"; // "1101 台泥"
    icu::UnicodeString unicodeString = icu::UnicodeString::fromUTF8(utf8String.c_str());
    
    // 将UnicodeString转回UTF-8格式
    std::string outputUtf8;
    unicodeString.toUTF8String(outputUtf8);
    
    std::cout << "Decoded UTF-8 Data: " << outputUtf8 << std::endl;
    return 0;
}

方案2:配置std::wcout的locale以支持UTF-16输出

若必须使用std::wcout,需先设置匹配的locale,确保终端能识别UTF-16:

#include <unicode/unistr.h>
#include <iostream>
#include <locale>

int main() {
    // 设置locale为系统默认的UTF-8 locale(如"zh_CN.UTF-8"或"en_US.UTF-8")
    std::locale::global(std::locale(""));
    std::wcout.imbue(std::locale());
    
    std::string utf8String = "1101\241@\245x\252d"; // "1101 台泥"
    icu::UnicodeString unicodeString = icu::UnicodeString::fromUTF8(utf8String.c_str());
    const UChar* utf16Data = unicodeString.getBuffer();
    
    std::wcout << "Decoded UTF-16 Data: " << utf16Data << std::endl;
    return 0;
}

额外注意事项

  • 检查Linux终端的locale设置:执行locale命令确认当前locale为UTF-8格式,若未设置,可临时执行export LC_ALL=zh_CN.UTF-8(或对应地区的UTF-8 locale)。
  • 注意UnicodeString::getBuffer()返回的是内部缓冲区指针,不要在UnicodeString对象销毁后使用该指针。

内容的提问来源于stack exchange,提问作者Weimin Chan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.02 22:02:32