如何用C和Objective-C在'wb'模式下写入正确的Unicode字符十六进制值?
问题描述
我有一段包含Unicode字符的文本,示例代码如下:
NSString *fileName = @"Tên tình bạn dưới tình yêu.mp3"; const char *cStringFile = [fileName UTF8String];
我需要将该字符串以每个Unicode字符对应其码点的十六进制值写入文件,比如:
- 'ê'对应
EA - 'ạ'对应
1E A1 - 中文'最'对应
67 00(对应Unicode码点\u6700)
但我尝试的两种方法都没得到预期结果,实际写入的是UTF-8编码字节(比如'ê'写入C3AA,'ạ'写入E1BAA1),求问题原因和可行方案。
尝试的两种方法
方法1:转换为wchar_t后写入文件
// Determine the required size for the wchar_t string size_t input_length = strlen(cStringFile); size_t output_length = mbstowcs(NULL, stringText, input_length); // Allocate memory for the wchar_t string wchar_t *output = (wchar_t *)malloc((output_length + 1) * sizeof(wchar_t)); if (output == NULL) { printf("Memory allocation failed.\n"); return 1; } // Convert the C string to wchar_t string mbstowcs(output, cStringFile, input_length); output[output_length] = L'\0'; // Add null-termination unsigned long length = wcslen(output); // Loop through each character in the Unicode text for (int i = 0; i < length; i++) { // Write the Unicode character to the file fwprintf(fd, L"%lc", output[i]); } // Free the allocated memory free(output);
方法2:检查多字节字符后写入文件
NSString *fileName = @"Tên tình bạn dưới tình yêu.mp3"; const char *stringText = [fileName UTF8String]; unsigned long len = strlen(stringText); setlocale(LC_ALL, ""); for (char character = *stringText; character != '\0'; character = *++stringText) { if (!character) { continue; } putchar(character); int byteCount = numberOfBytesInChar((unsigned char)character); if (byteCount <= 1) { fprintf(fd, "%c", character); } else { for(int k = 0; k < byteCount; k++) { fprintf(fd, "%c", character); character = *++stringText; } } } int numberOfBytesInChar(unsigned char val) { if (val < 128) { return 1; } else if (val < 224) { return 2; } else if (val < 240) { return 3; } else { return 4; } }
问题原因分析
方法1的问题:
fwprintf(fd, L"%lc", output[i])的作用是将wchar_t字符按当前区域设置的编码(通常是UTF-8)写入文件,并非输出其Unicode码点的十六进制值。wchar_t仅存储宽字符,输出时会被转成多字节编码,而非直接输出码点。- 代码中存在变量名错误:
mbstowcs的第二个参数是stringText,但前面定义的UTF-8字符串变量是cStringFile,会导致转换逻辑异常。
方法2的问题:
- 整个逻辑只是在遍历并写入UTF-8的原始字节流,完全没有解析出Unicode码点。
numberOfBytesInChar仅用于判断UTF-8的字节长度,最终还是把这些字节原样写入文件,自然得到UTF-8编码结果。
- 整个逻辑只是在遍历并写入UTF-8的原始字节流,完全没有解析出Unicode码点。
可行解决方案
方案1:Objective-C原生API实现(推荐)
直接利用NSString的Unicode码点遍历API,无需手动处理编码转换,逻辑更可靠:
NSString *fileName = @"Tên tình bạn dưới tình yêu.mp3"; FILE *fd = fopen("output.txt", "w"); if (!fd) { NSLog(@"Failed to open file"); return; } // 遍历每个Unicode字符(含组合字符序列) [fileName enumerateSubstringsInRange:NSMakeRange(0, fileName.length) options:NSStringEnumerationByComposedCharacterSequences usingBlock:^(NSString * _Nullable substring, NSRange substringRange, NSRange enclosingRange, BOOL * _Nonnull stop) { // 获取当前字符的Unicode码点 uint32_t codePoint = [substring characterAtIndex:0]; // 选项1:写入十六进制字符串格式(如"EA 1EA1 6700 ") fprintf(fd, "%X ", codePoint); // 选项2:写入二进制字节(按小端字节序,符合你给出的中文'最'的示例) // if (codePoint <= 0xFF) { // fputc(codePoint, fd); // } else if (codePoint <= 0xFFFF) { // fputc(codePoint & 0xFF, fd); // 低字节 // fputc((codePoint >> 8) & 0xFF, fd); // 高字节 // } else { // // 处理4字节Unicode码点(按需调整字节序) // fputc(codePoint & 0xFF, fd); // fputc((codePoint >> 8) & 0xFF, fd); // fputc((codePoint >> 16) & 0xFF, fd); // fputc((codePoint >> 24) & 0xFF, fd); // } }]; fclose(fd);
说明:代码包含两种输出模式,可根据需求选择:
- 十六进制字符串:直接输出码点的十六进制文本,便于阅读
- 二进制字节:按小端字节序写入码点的原始字节,符合你给出的
\u6700输出67 00的要求
方案2:纯C语言解析UTF-8获取码点
如果必须用C语言处理UTF-8字符串,需要手动解析每个UTF-8序列对应的Unicode码点:
#include <stdio.h> #include <stdint.h> // 从UTF-8字符串中解析单个Unicode码点 uint32_t utf8_to_codepoint(const char **str) { uint8_t c = **str; uint32_t codepoint; int bytes; if (c < 0x80) { codepoint = c; bytes = 1; } else if (c < 0xE0) { codepoint = (c & 0x1F) << 6; codepoint |= (*(*str + 1) & 0x3F); bytes = 2; } else if (c < 0xF0) { codepoint = (c & 0x0F) << 12; codepoint |= ((*(*str + 1) & 0x3F) << 6); codepoint |= (*(*str + 2) & 0x3F); bytes = 3; } else { codepoint = (c & 0x07) << 18; codepoint |= ((*(*str + 1) & 0x3F) << 12); codepoint |= ((*(*str + 2) & 0x3F) << 6); codepoint |= (*(*str + 3) & 0x3F); bytes = 4; } *str += bytes; return codepoint; } int main() { const char *stringText = "Tên tình bạn dưới tình yêu.mp3"; FILE *fd = fopen("output.txt", "w"); if (!fd) { printf("Failed to open file\n"); return 1; } while (*stringText) { uint32_t cp = utf8_to_codepoint(&stringText); // 写入十六进制字符串 fprintf(fd, "%X ", cp); // 或者写入二进制字节(小端),同方案1的选项2代码 } fclose(fd); return 0; }
内容的提问来源于stack exchange,提问作者KamyFC
相关产品推荐
相关产品推荐

