如何在C++中将含Unicode转义序列的const char*转换为std::wstring
解决Unicode转义序列转std::wstring的问题
你需要分两种场景处理,取决于你的const char* s实际存储的内容:
场景1:字符串是编译器已解析的UTF-8编码
如果你的代码里const char* s = "\u0633\u0644\u0627\u0645"是在C++11及以上标准下编写的,编译器会自动把\uXXXX转义序列转换成对应的UTF-8字节序列(前提是编译环境使用UTF-8编码)。此时只需要把UTF-8编码的字符串转换成std::wstring即可。
实现代码(标准库方式)
#include <string> #include <codecvt> // 将UTF-8字符串转换为std::wstring std::wstring utf8_to_wstring(const std::string& utf8_str) { std::wstring_convert<std::codecvt_utf8<wchar_t>> converter; return converter.from_bytes(utf8_str); } int main() { const char* s = "\u0633\u0644\u0627\u0645"; // 编译器自动转为UTF-8编码的"سلام" std::wstring result = utf8_to_wstring(s); // result即为L"سلام" return 0; }
实现代码(Windows平台专用)
如果是Windows平台,也可以用系统API实现:
#include <string> #include <windows.h> std::wstring utf8_to_wstring(const char* utf8_str) { if (!utf8_str) return L""; // 计算所需宽字符长度 int wide_len = MultiByteToWideChar(CP_UTF8, 0, utf8_str, -1, nullptr, 0); std::wstring wide_str(wide_len - 1, L'\0'); // 执行转换 MultiByteToWideChar(CP_UTF8, 0, utf8_str, -1, &wide_str[0], wide_len); return wide_str; }
场景2:字符串包含字面的Unicode转义序列
如果你的const char* s存储的是字面的"\u0633\u0644\u0627\u0645"(即实际字符是\、u、0、6等,比如从文件或网络读取的原始转义文本),需要手动解析这些\uXXXX转义序列,转换成对应的Unicode码点,再构建std::wstring。
实现代码
#include <string> #include <cctype> #include <stdexcept> #include <cstdlib> // 解析包含\uXXXX转义序列的字符串为std::wstring std::wstring decode_unicode_escapes(const char* s) { std::wstring result; const char* p = s; while (*p != '\0') { // 识别\u转义序列 if (*p == '\\' && *(p + 1) == 'u') { p += 2; // 跳过"\u" // 检查后续是否有4个合法的十六进制字符 if (!isxdigit(*p) || !isxdigit(*(p+1)) || !isxdigit(*(p+2)) || !isxdigit(*(p+3))) { throw std::invalid_argument("无效的Unicode转义序列"); } // 提取4位十六进制字符 char hex_buf[5] = {0}; strncpy(hex_buf, p, 4); // 转换为Unicode码点 unsigned long codepoint = strtoul(hex_buf, nullptr, 16); // 处理BMP范围内的码点(0x0000-0xFFFF),如需支持非BMP可添加代理对转换逻辑 if (codepoint > 0xFFFF) { throw std::invalid_argument("暂不支持BMP以外的Unicode码点"); } result.push_back(static_cast<wchar_t>(codepoint)); p += 4; // 跳过4位十六进制字符 } else { // 普通字符直接转换为宽字符 result.push_back(static_cast<wchar_t>(*p)); p++; } } return result; } int main() { // 注意这里的反斜杠需要转义,否则编译器会提前解析 const char* s = "\\u0633\\u0644\\u0627\\u0645"; std::wstring result = decode_unicode_escapes(s); // result即为L"سلام" return 0; }
扩展说明
如果需要支持非BMP的Unicode码点(如emoji),需要将码点转换为UTF-16代理对(Windows下wchar_t是UTF-16),转换逻辑如下:
// 辅助函数:将Unicode码点转换为UTF-16代理对 void codepoint_to_utf16(unsigned long codepoint, std::wstring& out) { if (codepoint <= 0xFFFF) { out.push_back(static_cast<wchar_t>(codepoint)); } else { codepoint -= 0x10000; wchar_t high_surrogate = static_cast<wchar_t>((codepoint >> 10) + 0xD800); wchar_t low_surrogate = static_cast<wchar_t>((codepoint & 0x3FF) + 0xDC00); out.push_back(high_surrogate); out.push_back(low_surrogate); } }
将之前代码中result.push_back(static_cast<wchar_t>(codepoint));替换为调用这个函数即可。
内容的提问来源于stack exchange,提问作者Babak.Abad
相关产品推荐
相关产品推荐

