寻求C++实现非ASCII长连字符转ASCII连字符的解决方案
C++实现非ASCII长连字符替换为ASCII连字符
你提到的VI中显示为(▒~@~S)的非ASCII长连字符,大概率是UTF-8编码的**EN DASH(U+2013)或EM DASH(U+2014)**这类宽连字符,终端编码不匹配导致显示乱码。以下是几种实用的C++实现方案:
方案1:UTF-8字符串直接替换字节序列
如果你的文本是UTF-8编码,可直接匹配目标长连字符的UTF-8字节序列并替换:
#include <string> void replaceLongDashes(std::string& str) { // 替换EN DASH(UTF-8字节:0xE2 0x80 0x93) const std::string en_dash = "\xE2\x80\x93"; size_t pos = 0; while ((pos = str.find(en_dash, pos)) != std::string::npos) { str.replace(pos, en_dash.size(), "-"); pos += 1; // 移动到替换后位置,避免重复匹配 } // 替换EM DASH(UTF-8字节:0xE2 0x80 0x94) const std::string em_dash = "\xE2\x80\x94"; pos = 0; while ((pos = str.find(em_dash, pos)) != std::string::npos) { str.replace(pos, em_dash.size(), "-"); pos += 1; } }
方案2:宽字符(wstring)替换Unicode码点
如果程序使用宽字符字符串(std::wstring),可以直接基于Unicode码点匹配替换:
#include <string> #include <algorithm> void replaceLongDashes(std::wstring& wstr) { // 匹配EN DASH(U+2013)和EM DASH(U+2014),替换为ASCII连字符 std::replace_if(wstr.begin(), wstr.end(), [](wchar_t c) { return c == L'\u2013' || c == L'\u2014'; }, L'-'); }
方案3:兼容多种非ASCII连字符
如果需要覆盖更多类似连字符(比如FIGURE DASH U+2012、HORIZONTAL BAR U+2015),可以扩展匹配范围:
#include <string> #include <algorithm> void replaceAllNonAsciiDashes(std::wstring& wstr) { const wchar_t target_dashes[] = {L'\u2012', L'\u2013', L'\u2014', L'\u2015'}; for (wchar_t dash : target_dashes) { std::replace(wstr.begin(), wstr.end(), dash, L'-'); } }
确认目标字符的方法
如果不确定具体是哪种连字符,可以用以下代码打印UTF-8字符串中每个字符的Unicode码点:
#include <iostream> #include <string> #include <cstdint> #include <iomanip> void printUtf8Codepoints(const std::string& str) { size_t i = 0; while (i < str.size()) { uint8_t c = static_cast<uint8_t>(str[i]); uint32_t codepoint = 0; if ((c & 0x80) == 0) { // 单字节ASCII字符 codepoint = c; i++; } else if ((c & 0xE0) == 0xC0) { // 双字节UTF-8 codepoint = ((c & 0x1F) << 6) | (static_cast<uint8_t>(str[i+1]) & 0x3F); i += 2; } else if ((c & 0xF0) == 0xE0) { // 三字节UTF-8(对应U+2013/U+2014这类字符) codepoint = ((c & 0x0F) << 12) | ((static_cast<uint8_t>(str[i+1]) & 0x3F) << 6) | (static_cast<uint8_t>(str[i+2]) & 0x3F); i += 3; } else { // 四字节及以上UTF-8,此处简化跳过 i++; continue; } std::cout << "U+" << std::uppercase << std::hex << std::setw(4) << std::setfill('0') << codepoint << " "; } std::cout << std::endl; }
内容的提问来源于stack exchange,提问作者Kalyan Srujan
相关产品推荐
相关产品推荐

