You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

将char强制转换为USHORT时获取错误UTF-8值的技术求助

解决Linux下char转USHORT的非ASCII字符错误问题

嗨,我来帮你搞定这个移植中的字符转换问题!

首先得搞清楚问题出在哪:Linux里的char默认是有符号类型(signed char),当字符值超过127的时候(比如你例子里的°,在DOS常用的扩展ASCII里是176),会被当成负数存在char里。这时候直接强制转USHORT,会触发「符号扩展」——把高位全补1,结果就不是你想要的数值了。比如176作为signed char是-80,转成USHORT就变成了0xFFB0(也就是65464),而不是正确的0x00B0(176)。

下面给你几个针对性的解决方案:

方案1:先转unsigned char再转USHORT(针对单字节扩展ASCII)

如果你的旧代码用的是DOS的单字节扩展ASCII(比如CP437、CP850),那只需要先把char转成无符号的unsigned char,再转USHORT就能避免符号扩展:

char* test = "°...";
// 处理第一个字符
unsigned char unsigned_char_val = (unsigned char)test[0];
USHORT unicode_val = (USHORT)unsigned_char_val;
// 传入你的绘制函数
draw_glyph(unicode_val);

这样转换后,176就会被正确保留为0x00B0,不会变成错误的负数扩展值。

方案2:处理UTF-8编码的字符串(Linux默认编码)

如果你的Linux程序里的字符串是UTF-8编码的(这是Linux的默认情况),那直接转单字节就不行了——因为非ASCII的UTF-8字符是多字节存储的。这时候需要先把UTF-8字符串解析成对应的Unicode码点,再传入函数。这里给你一个简单的手动解析函数示例:

#include <stdint.h>

// 解析UTF-8字符,返回对应的Unicode码点,同时移动指针到下一个字符
uint32_t utf8_to_unicode(const char** str_ptr) {
    uint32_t code_point = (unsigned char)(*str_ptr)[0];
    
    if (code_point < 0x80) {
        // 单字节ASCII字符
        (*str_ptr)++;
        return code_point;
    } else if (code_point < 0xE0) {
        // 双字节UTF-8
        code_point = ((code_point & 0x1F) << 6) | ((unsigned char)(*str_ptr)[1] & 0x3F);
        (*str_ptr) += 2;
        return code_point;
    } else if (code_point < 0xF0) {
        // 三字节UTF-8
        code_point = ((code_point & 0x0F) << 12) | 
                     (((unsigned char)(*str_ptr)[1] & 0x3F) << 6) | 
                     ((unsigned char)(*str_ptr)[2] & 0x3F);
        (*str_ptr) += 3;
        return code_point;
    }
    // 四字节UTF-8(简化处理,实际可根据需求完善)
    code_point = ((code_point & 0x07) << 18) | 
                 (((unsigned char)(*str_ptr)[1] & 0x3F) << 12) | 
                 (((unsigned char)(*str_ptr)[2] & 0x3F) << 6) | 
                 ((unsigned char)(*str_ptr)[3] & 0x3F);
    (*str_ptr) += 4;
    return code_point;
}

// 使用示例
char* test = "°...";
const char* current_ptr = test;
while (*current_ptr) {
    uint32_t unicode = utf8_to_unicode(&current_ptr);
    // 如果你的点阵字体只覆盖Unicode BMP范围(U+0000到U+FFFF),直接转USHORT即可
    draw_glyph((USHORT)unicode);
}

方案3:编码转换(从DOS编码转Unicode)

如果你的旧代码依赖的是特定的DOS字符编码(比如CP437),而Linux下读取的字符串是这个编码的(比如从旧文件读取),那需要先把编码转换成Unicode。可以用标准库的iconv来做:

#include <iconv.h>
#include <stdlib.h>
#include <string.h>

// 编码转换函数:from_code是原编码,to_code是目标编码
char* convert_encoding(const char* input, const char* from_code, const char* to_code) {
    iconv_t conv_handle = iconv_open(to_code, from_code);
    if (conv_handle == (iconv_t)-1) {
        return NULL; // 转换失败
    }

    size_t input_len = strlen(input);
    size_t output_len = input_len * 4; // 预估足够的输出空间
    char* output = malloc(output_len);
    if (!output) {
        iconv_close(conv_handle);
        return NULL;
    }

    char* in_buf = (char*)input;
    char* out_buf = output;
    if (iconv(conv_handle, &in_buf, &input_len, &out_buf, &output_len) == (size_t)-1) {
        free(output);
        iconv_close(conv_handle);
        return NULL;
    }

    *out_buf = '\0'; // 添加字符串结束符
    iconv_close(conv_handle);
    return output;
}

// 使用示例:把CP437编码转成UTF-8
char* dos_cp437_str = "°..."; // 假设这个字符串是CP437编码
char* utf8_str = convert_encoding(dos_cp437_str, "CP437", "UTF-8");
if (utf8_str) {
    // 再用方案2的函数解析成Unicode码点
    const char* ptr = utf8_str;
    while (*ptr) {
        uint32_t unicode = utf8_to_unicode(&ptr);
        draw_glyph((USHORT)unicode);
    }
    free(utf8_str);
}

最后提醒你一下:先确认旧DOS代码用的是什么字符编码(比如CP437、CP850),这很关键——编码不匹配的话,转换肯定会出问题。

内容的提问来源于stack exchange,提问作者J.Panek

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 08:32:08