将char强制转换为USHORT时获取错误UTF-8值的技术求助
嗨,我来帮你搞定这个移植中的字符转换问题!
首先得搞清楚问题出在哪:Linux里的char默认是有符号类型(signed char),当字符值超过127的时候(比如你例子里的°,在DOS常用的扩展ASCII里是176),会被当成负数存在char里。这时候直接强制转USHORT,会触发「符号扩展」——把高位全补1,结果就不是你想要的数值了。比如176作为signed char是-80,转成USHORT就变成了0xFFB0(也就是65464),而不是正确的0x00B0(176)。
下面给你几个针对性的解决方案:
方案1:先转unsigned char再转USHORT(针对单字节扩展ASCII)
如果你的旧代码用的是DOS的单字节扩展ASCII(比如CP437、CP850),那只需要先把char转成无符号的unsigned char,再转USHORT就能避免符号扩展:
char* test = "°..."; // 处理第一个字符 unsigned char unsigned_char_val = (unsigned char)test[0]; USHORT unicode_val = (USHORT)unsigned_char_val; // 传入你的绘制函数 draw_glyph(unicode_val);
这样转换后,176就会被正确保留为0x00B0,不会变成错误的负数扩展值。
方案2:处理UTF-8编码的字符串(Linux默认编码)
如果你的Linux程序里的字符串是UTF-8编码的(这是Linux的默认情况),那直接转单字节就不行了——因为非ASCII的UTF-8字符是多字节存储的。这时候需要先把UTF-8字符串解析成对应的Unicode码点,再传入函数。这里给你一个简单的手动解析函数示例:
#include <stdint.h> // 解析UTF-8字符,返回对应的Unicode码点,同时移动指针到下一个字符 uint32_t utf8_to_unicode(const char** str_ptr) { uint32_t code_point = (unsigned char)(*str_ptr)[0]; if (code_point < 0x80) { // 单字节ASCII字符 (*str_ptr)++; return code_point; } else if (code_point < 0xE0) { // 双字节UTF-8 code_point = ((code_point & 0x1F) << 6) | ((unsigned char)(*str_ptr)[1] & 0x3F); (*str_ptr) += 2; return code_point; } else if (code_point < 0xF0) { // 三字节UTF-8 code_point = ((code_point & 0x0F) << 12) | (((unsigned char)(*str_ptr)[1] & 0x3F) << 6) | ((unsigned char)(*str_ptr)[2] & 0x3F); (*str_ptr) += 3; return code_point; } // 四字节UTF-8(简化处理,实际可根据需求完善) code_point = ((code_point & 0x07) << 18) | (((unsigned char)(*str_ptr)[1] & 0x3F) << 12) | (((unsigned char)(*str_ptr)[2] & 0x3F) << 6) | ((unsigned char)(*str_ptr)[3] & 0x3F); (*str_ptr) += 4; return code_point; } // 使用示例 char* test = "°..."; const char* current_ptr = test; while (*current_ptr) { uint32_t unicode = utf8_to_unicode(¤t_ptr); // 如果你的点阵字体只覆盖Unicode BMP范围(U+0000到U+FFFF),直接转USHORT即可 draw_glyph((USHORT)unicode); }
方案3:编码转换(从DOS编码转Unicode)
如果你的旧代码依赖的是特定的DOS字符编码(比如CP437),而Linux下读取的字符串是这个编码的(比如从旧文件读取),那需要先把编码转换成Unicode。可以用标准库的iconv来做:
#include <iconv.h> #include <stdlib.h> #include <string.h> // 编码转换函数:from_code是原编码,to_code是目标编码 char* convert_encoding(const char* input, const char* from_code, const char* to_code) { iconv_t conv_handle = iconv_open(to_code, from_code); if (conv_handle == (iconv_t)-1) { return NULL; // 转换失败 } size_t input_len = strlen(input); size_t output_len = input_len * 4; // 预估足够的输出空间 char* output = malloc(output_len); if (!output) { iconv_close(conv_handle); return NULL; } char* in_buf = (char*)input; char* out_buf = output; if (iconv(conv_handle, &in_buf, &input_len, &out_buf, &output_len) == (size_t)-1) { free(output); iconv_close(conv_handle); return NULL; } *out_buf = '\0'; // 添加字符串结束符 iconv_close(conv_handle); return output; } // 使用示例:把CP437编码转成UTF-8 char* dos_cp437_str = "°..."; // 假设这个字符串是CP437编码 char* utf8_str = convert_encoding(dos_cp437_str, "CP437", "UTF-8"); if (utf8_str) { // 再用方案2的函数解析成Unicode码点 const char* ptr = utf8_str; while (*ptr) { uint32_t unicode = utf8_to_unicode(&ptr); draw_glyph((USHORT)unicode); } free(utf8_str); }
最后提醒你一下:先确认旧DOS代码用的是什么字符编码(比如CP437、CP850),这很关键——编码不匹配的话,转换肯定会出问题。
内容的提问来源于stack exchange,提问作者J.Panek

