Linux下UTF-8中文字符串显示异常问题求助
Linux下UTF-8中文显示异常的解决方法
问题根源
Linux环境中std::wcout的默认locale通常不兼容UTF-16,且终端普遍采用UTF-8编码,直接输出ICU库返回的UChar*(UTF-16格式)会因编码不匹配导致乱码。而Windows控制台默认支持UTF-16输出,因此能正常显示内容。
解决方案
方案1:将UnicodeString转回UTF-8后用std::cout输出
这是适配Linux终端UTF-8环境的最直接方式:
#include <unicode/unistr.h> #include <iostream> int main() { std::string utf8String = "1101\241@\245x\252d"; // "1101 台泥" icu::UnicodeString unicodeString = icu::UnicodeString::fromUTF8(utf8String.c_str()); // 将UnicodeString转回UTF-8格式 std::string outputUtf8; unicodeString.toUTF8String(outputUtf8); std::cout << "Decoded UTF-8 Data: " << outputUtf8 << std::endl; return 0; }
方案2:配置std::wcout的locale以支持UTF-16输出
若必须使用std::wcout,需先设置匹配的locale,确保终端能识别UTF-16:
#include <unicode/unistr.h> #include <iostream> #include <locale> int main() { // 设置locale为系统默认的UTF-8 locale(如"zh_CN.UTF-8"或"en_US.UTF-8") std::locale::global(std::locale("")); std::wcout.imbue(std::locale()); std::string utf8String = "1101\241@\245x\252d"; // "1101 台泥" icu::UnicodeString unicodeString = icu::UnicodeString::fromUTF8(utf8String.c_str()); const UChar* utf16Data = unicodeString.getBuffer(); std::wcout << "Decoded UTF-16 Data: " << utf16Data << std::endl; return 0; }
额外注意事项
- 检查Linux终端的locale设置:执行
locale命令确认当前locale为UTF-8格式,若未设置,可临时执行export LC_ALL=zh_CN.UTF-8(或对应地区的UTF-8 locale)。 - 注意
UnicodeString::getBuffer()返回的是内部缓冲区指针,不要在UnicodeString对象销毁后使用该指针。
内容的提问来源于stack exchange,提问作者Weimin Chan
相关产品推荐
相关产品推荐

