C语言中日文字符数组转换为等效十六进制值的实现方法
C语言日文字符数组转十六进制实现方案
实现原理
你给出的示例输入"アイコン"对应的输出EFBDB1EFBDB2EFBDBAEFBE9D,本质是该字符串以UTF-8编码存储时,每个字节的十进制值转换为两位大写十六进制后拼接的结果,所以实现逻辑分为两步:
- 确保待转换的日文字符串采用UTF-8编码(如果是其他编码如Shift-JIS需要先转码为UTF-8)
- 遍历字符串的每个字节,将每个字节值转为两位十六进制字符拼接即可
代码实现
场景1:源字符串已为UTF-8编码
如果你的编译环境默认用UTF-8存储字符串(多数Linux/macOS环境、Windows下指定/utf-8编译参数时满足),可以直接用以下代码:
#include <stdio.h> #include <string.h> // 字节转两位十六进制,结果写入out(至少3字节空间) void byte_to_hex(unsigned char byte, char *out) { const char *hex_chars = "0123456789ABCDEF"; out[0] = hex_chars[byte >> 4]; out[1] = hex_chars[byte & 0x0F]; out[2] = '\0'; } int main() { // 待转换的日文字符串 const char *jp_str = "アイコン"; // 输出缓冲区:每个字节对应2个十六进制字符,加结束符 char hex_out[strlen(jp_str) * 2 + 1]; int pos = 0; for (int i = 0; jp_str[i] != '\0'; i++) { char tmp[3]; byte_to_hex((unsigned char)jp_str[i], tmp); hex_out[pos++] = tmp[0]; hex_out[pos++] = tmp[1]; } hex_out[pos] = '\0'; printf("转换结果:%s\n", hex_out); return 0; }
运行后输出结果与你要求的EFBDB1EFBDB2EFBDBAEFBE9D完全一致。
场景2:源字符串为Shift-JIS等其他编码
如果你的源字符串是Shift-JIS编码,需要先转码为UTF-8,Windows下可以用内置的MultiByteToWideChar+WideCharToMultiByte函数实现,跨平台可以用libiconv库,以下是Windows内置API实现示例:
#include <windows.h> #include <stdio.h> #include <string.h> #include <stdlib.h> void byte_to_hex(unsigned char byte, char *out) { const char *hex_chars = "0123456789ABCDEF"; out[0] = hex_chars[byte >> 4]; out[1] = hex_chars[byte & 0x0F]; out[2] = '\0'; } int main() { // 假设源字符串为Shift-JIS编码的"アイコン" const char *sjis_str = "\xB1\xB2\xBA\xDD"; // 第一步:Shift-JIS转UTF-16 int wchar_len = MultiByteToWideChar(CP_ACP, 0, sjis_str, -1, NULL, 0); WCHAR *wstr = malloc(wchar_len * sizeof(WCHAR)); MultiByteToWideChar(CP_ACP, 0, sjis_str, -1, wstr, wchar_len); // 第二步:UTF-16转UTF-8 int utf8_len = WideCharToMultiByte(CP_UTF8, 0, wstr, -1, NULL, 0, NULL, NULL); char *utf8_str = malloc(utf8_len); WideCharToMultiByte(CP_UTF8, 0, wstr, -1, utf8_str, utf8_len, NULL, NULL); // 转十六进制逻辑 char hex_out[utf8_len * 2 + 1]; int pos = 0; for (int i = 0; utf8_str[i] != '\0'; i++) { char tmp[3]; byte_to_hex((unsigned char)utf8_str[i], tmp); hex_out[pos++] = tmp[0]; hex_out[pos++] = tmp[1]; } hex_out[pos] = '\0'; printf("转换结果:%s\n", hex_out); // 释放内存 free(wstr); free(utf8_str); return 0; }
注意事项
- 如果需要小写十六进制输出,只需要把
byte_to_hex函数里的hex_chars改为"0123456789abcdef"即可 - 处理长度不确定的字符串时建议动态分配输出缓冲区,避免栈溢出
内容的提问来源于stack exchange,提问作者Alok Ranjan Swain
相关产品推荐
相关产品推荐

