You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

C++实现支持变音符号的字符串与十六进制互转及EEPROM适配

问题描述

我制作了一款EEPROM(类似简易U盘)编程器,正在编写程序读取TXT文件文本并转换为二进制/十六进制,供编程器将数据写入EEPROM。目前缺少字符串转十六进制的函数,尝试了以下代码,仅部分有效:

string Text = "This is a string.";
for(int i = 0; i < Text.size(); i++) {
    cout << uppercase << hex << (int)Text[i] << " ";
}

该代码输出:

54 68 69 73 20 69 73 20 61 20 73 74 72 69 6E 67 2E 

但输入含变音符号的文本:

Thïs ìs â stríng.

得到的输出为:

54 68 FFFFFFC3 FFFFFFAF 73 20 FFFFFFC3 FFFFFFAC 73 20 FFFFFFC3 FFFFFFA2 20 73 74 72 FFFFFFC3 FFFFFFAD 6E 67 2E 

此结果不符合预期,推测普通字符转ASCII,特殊字符转某种Unicode格式。希望实现全Unicode支持,且因EEPROM仅能存储2k字节,需尽可能节省空间。

最终目标

  • 实现字符串转十六进制的函数,支持变音符号且空间高效;
  • 实现十六进制转字符串的函数,同样支持变音符号。

若无法实现,可接受自定义格式,例如用"|e^"存储'ê',以"|"作为特殊字符标识。


解决方案

问题根源

你的代码里string在C++中默认是基于UTF-8编码的字节序列,变音符号(比如ï、ì)在UTF-8中是多字节字符(占2字节),而你直接把每个字节转成int输出时,因char是有符号类型,负数值会触发符号扩展,导致出现FFFFFFC3这类多余前缀。

方案1:UTF-8编码存储(空间高效,推荐)

UTF-8本身就是空间友好的编码:ASCII字符占1字节,欧洲语言变音符号占2字节,复杂Unicode字符占3-4字节,完全适配2k字节的EEPROM容量。只需修正输出逻辑避免符号扩展即可:

字符串转十六进制(UTF-8)

#include <string>
#include <iostream>
#include <cstdio>

std::string str_to_hex_utf8(const std::string& text) {
    std::string hex_str;
    for (unsigned char c : text) { // 用unsigned char避免符号扩展
        char buf[3];
        std::sprintf(buf, "%02X ", c);
        hex_str += buf;
    }
    // 移除末尾多余空格(可选)
    if (!hex_str.empty()) hex_str.pop_back();
    return hex_str;
}

// 使用示例
int main() {
    std::string text = "Thïs ìs â stríng.";
    std::cout << str_to_hex_utf8(text) << std::endl;
    return 0;
}

输出为:54 68 C3 AF 73 20 C3 AC 73 20 C3 A2 20 73 74 72 C3 AD 6E 67 2E,每个字节都是标准两位十六进制格式。

十六进制转字符串(UTF-8)

#include <string>
#include <sstream>
#include <algorithm>

std::string hex_to_str_utf8(const std::string& hex_str) {
    std::string text;
    std::istringstream hex_stream(hex_str);
    std::string byte_str;
    while (hex_stream >> byte_str) {
        // 把十六进制字符串转成无符号字节
        unsigned char byte = static_cast<unsigned char>(std::stoul(byte_str, nullptr, 16));
        text += byte;
    }
    return text;
}

// 使用示例
int main() {
    std::string hex = "54 68 C3 AF 73 20 C3 AC 73 20 C3 A2 20 73 74 72 C3 AD 6E 67 2E";
    std::cout << hex_to_str_utf8(hex) << std::endl;
    return 0;
}

该方案无需额外处理编码映射,直接利用UTF-8原生特性,既能支持全Unicode字符,又能最大化利用EEPROM存储空间。

方案2:自定义压缩格式(极端空间优化)

如果你的文本以欧洲语言为主,变音符号种类有限,可以用自定义单字节编码进一步压缩空间:

  • 0x00-0x7F范围保留ASCII原编码
  • 0x80-0xFF范围映射常用变音符号(比如ï=0x80,ì=0x81)

字符串转自定义十六进制

#include <string>
#include <unordered_map>
#include <cstdio>
#include <codecvt>

// 自定义字符映射表,可按需扩展
std::unordered_map<char32_t, unsigned char> char_to_code = {
    {U'ï', 0x80},
    {U'ì', 0x81},
    {U'â', 0x82},
    {U'í', 0x83}
};

std::string str_to_hex_custom(const std::u32string& text) {
    std::string hex_str;
    for (char32_t c : text) {
        if (c <= 0x7F) {
            char buf[3];
            std::sprintf(buf, "%02X ", static_cast<unsigned char>(c));
            hex_str += buf;
        } else {
            auto it = char_to_code.find(c);
            if (it != char_to_code.end()) {
                char buf[3];
                std::sprintf(buf, "%02X ", it->second);
                hex_str += buf;
            } else {
                // 未映射字符转为UTF-8多字节存储
                unsigned char bytes[4];
                int len = 0;
                if (c <= 0x7FF) {
                    bytes[0] = 0xC0 | ((c >> 6) & 0x1F);
                    bytes[1] = 0x80 | (c & 0x3F);
                    len = 2;
                } else if (c <= 0xFFFF) {
                    bytes[0] = 0xE0 | ((c >> 12) & 0x0F);
                    bytes[1] = 0x80 | ((c >> 6) & 0x3F);
                    bytes[2] = 0x80 | (c & 0x3F);
                    len = 3;
                } else {
                    bytes[0] = 0xF0 | ((c >> 18) & 0x07);
                    bytes[1] = 0x80 | ((c >> 12) & 0x3F);
                    bytes[2] = 0x80 | ((c >> 6) & 0x3F);
                    bytes[3] = 0x80 | (c & 0x3F);
                    len = 4;
                }
                for (int i = 0; i < len; i++) {
                    char buf[3];
                    std::sprintf(buf, "%02X ", bytes[i]);
                    hex_str += buf;
                }
            }
        }
    }
    if (!hex_str.empty()) hex_str.pop_back();
    return hex_str;
}

// 辅助函数:将普通string转成u32string(处理UTF-8输入)
std::u32string str_to_u32(const std::string& str) {
    std::wstring_convert<std::codecvt_utf8<char32_t>, char32_t> converter;
    return converter.from_bytes(str);
}

十六进制转自定义字符串

#include <string>
#include <unordered_map>
#include <sstream>
#include <codecvt>

std::unordered_map<unsigned char, char32_t> code_to_char = {
    {0x80, U'ï'},
    {0x81, U'ì'},
    {0x82, U'â'},
    {0x83, U'í'}
};

std::u32string hex_to_str_custom(const std::string& hex_str) {
    std::u32string text;
    std::istringstream hex_stream(hex_str);
    std::string byte_str;
    while (hex_stream >> byte_str) {
        unsigned char byte = static_cast<unsigned char>(std::stoul(byte_str, nullptr, 16));
        if (byte <= 0x7F) {
            text += static_cast<char32_t>(byte);
        } else {
            auto it = code_to_char.find(byte);
            if (it != code_to_char.end()) {
                text += it->second;
            } else {
                // 处理UTF-8多字节字符,需完整解析逻辑,此处简化为占位符
                text += U'?';
            }
        }
    }
    return text;
}

// 辅助函数:将u32string转成普通string(输出UTF-8)
std::string u32_to_str(const std::u32string& u32_str) {
    std::wstring_convert<std::codecvt_utf8<char32_t>, char32_t> converter;
    return converter.to_bytes(u32_str);
}

该方案能把常用变音符号压缩到1字节,比UTF-8更省空间,但需要维护映射表,扩展性不如UTF-8方案。


内容的提问来源于stack exchange,提问作者OADINC

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 16:20:31