You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

寻求C++实现非ASCII长连字符转ASCII连字符的解决方案

C++实现非ASCII长连字符替换为ASCII连字符

你提到的VI中显示为(▒~@~S)的非ASCII长连字符,大概率是UTF-8编码的**EN DASH(U+2013)或EM DASH(U+2014)**这类宽连字符,终端编码不匹配导致显示乱码。以下是几种实用的C++实现方案:

方案1:UTF-8字符串直接替换字节序列

如果你的文本是UTF-8编码,可直接匹配目标长连字符的UTF-8字节序列并替换:

#include <string>

void replaceLongDashes(std::string& str) {
    // 替换EN DASH(UTF-8字节:0xE2 0x80 0x93)
    const std::string en_dash = "\xE2\x80\x93";
    size_t pos = 0;
    while ((pos = str.find(en_dash, pos)) != std::string::npos) {
        str.replace(pos, en_dash.size(), "-");
        pos += 1; // 移动到替换后位置,避免重复匹配
    }

    // 替换EM DASH(UTF-8字节:0xE2 0x80 0x94)
    const std::string em_dash = "\xE2\x80\x94";
    pos = 0;
    while ((pos = str.find(em_dash, pos)) != std::string::npos) {
        str.replace(pos, em_dash.size(), "-");
        pos += 1;
    }
}

方案2:宽字符(wstring)替换Unicode码点

如果程序使用宽字符字符串(std::wstring),可以直接基于Unicode码点匹配替换:

#include <string>
#include <algorithm>

void replaceLongDashes(std::wstring& wstr) {
    // 匹配EN DASH(U+2013)和EM DASH(U+2014),替换为ASCII连字符
    std::replace_if(wstr.begin(), wstr.end(),
        [](wchar_t c) { return c == L'\u2013' || c == L'\u2014'; },
        L'-');
}

方案3:兼容多种非ASCII连字符

如果需要覆盖更多类似连字符(比如FIGURE DASH U+2012、HORIZONTAL BAR U+2015),可以扩展匹配范围:

#include <string>
#include <algorithm>

void replaceAllNonAsciiDashes(std::wstring& wstr) {
    const wchar_t target_dashes[] = {L'\u2012', L'\u2013', L'\u2014', L'\u2015'};
    for (wchar_t dash : target_dashes) {
        std::replace(wstr.begin(), wstr.end(), dash, L'-');
    }
}

确认目标字符的方法

如果不确定具体是哪种连字符,可以用以下代码打印UTF-8字符串中每个字符的Unicode码点:

#include <iostream>
#include <string>
#include <cstdint>
#include <iomanip>

void printUtf8Codepoints(const std::string& str) {
    size_t i = 0;
    while (i < str.size()) {
        uint8_t c = static_cast<uint8_t>(str[i]);
        uint32_t codepoint = 0;
        if ((c & 0x80) == 0) {
            // 单字节ASCII字符
            codepoint = c;
            i++;
        } else if ((c & 0xE0) == 0xC0) {
            // 双字节UTF-8
            codepoint = ((c & 0x1F) << 6) | (static_cast<uint8_t>(str[i+1]) & 0x3F);
            i += 2;
        } else if ((c & 0xF0) == 0xE0) {
            // 三字节UTF-8(对应U+2013/U+2014这类字符)
            codepoint = ((c & 0x0F) << 12) | ((static_cast<uint8_t>(str[i+1]) & 0x3F) << 6) 
                      | (static_cast<uint8_t>(str[i+2]) & 0x3F);
            i += 3;
        } else {
            // 四字节及以上UTF-8,此处简化跳过
            i++;
            continue;
        }
        std::cout << "U+" << std::uppercase << std::hex << std::setw(4) << std::setfill('0') 
                  << codepoint << " ";
    }
    std::cout << std::endl;
}

内容的提问来源于stack exchange,提问作者Kalyan Srujan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.17 22:07:54