You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用C和Objective-C在'wb'模式下写入正确的Unicode字符十六进制值?

问题描述

我有一段包含Unicode字符的文本,示例代码如下:

NSString *fileName = @"Tên tình bạn dưới tình yêu.mp3";
const char *cStringFile = [fileName UTF8String];

我需要将该字符串以每个Unicode字符对应其码点的十六进制值写入文件,比如:

  • 'ê'对应EA
  • 'ạ'对应1E A1
  • 中文'最'对应67 00(对应Unicode码点\u6700)

但我尝试的两种方法都没得到预期结果,实际写入的是UTF-8编码字节(比如'ê'写入C3AA,'ạ'写入E1BAA1),求问题原因和可行方案。


尝试的两种方法

方法1:转换为wchar_t后写入文件

// Determine the required size for the wchar_t string
size_t input_length = strlen(cStringFile);
size_t output_length = mbstowcs(NULL, stringText, input_length);

// Allocate memory for the wchar_t string
wchar_t *output = (wchar_t *)malloc((output_length + 1) * sizeof(wchar_t));
if (output == NULL) {
    printf("Memory allocation failed.\n");
    return 1;
}

// Convert the C string to wchar_t string
mbstowcs(output, cStringFile, input_length);
output[output_length] = L'\0'; // Add null-termination

unsigned long length = wcslen(output);
// Loop through each character in the Unicode text
for (int i = 0; i < length; i++) {
    // Write the Unicode character to the file
    fwprintf(fd, L"%lc", output[i]);
}

// Free the allocated memory
free(output);

方法2:检查多字节字符后写入文件

NSString *fileName = @"Tên tình bạn dưới tình yêu.mp3";
const char *stringText = [fileName UTF8String];
unsigned long len = strlen(stringText);
setlocale(LC_ALL, "");
for (char character = *stringText; character != '\0'; character = *++stringText)
{
    if (!character) {
        continue;
    }
    putchar(character);
    int byteCount = numberOfBytesInChar((unsigned char)character);
    if (byteCount <= 1) {
        fprintf(fd, "%c", character);
    } else {
        for(int k = 0; k < byteCount; k++)
        {
            fprintf(fd, "%c", character);
            character = *++stringText;
        }
    }
}

int numberOfBytesInChar(unsigned char val) {
  if (val < 128) {
     return 1;
  } else if (val < 224) {
     return 2;
  } else if (val < 240) {
     return 3;
  } else {
    return 4;
  }
}

问题原因分析
  1. 方法1的问题:

    • fwprintf(fd, L"%lc", output[i])的作用是将wchar_t字符按当前区域设置的编码(通常是UTF-8)写入文件,并非输出其Unicode码点的十六进制值。wchar_t仅存储宽字符,输出时会被转成多字节编码,而非直接输出码点。
    • 代码中存在变量名错误:mbstowcs的第二个参数是stringText,但前面定义的UTF-8字符串变量是cStringFile,会导致转换逻辑异常。
  2. 方法2的问题:

    • 整个逻辑只是在遍历并写入UTF-8的原始字节流,完全没有解析出Unicode码点。numberOfBytesInChar仅用于判断UTF-8的字节长度,最终还是把这些字节原样写入文件,自然得到UTF-8编码结果。

可行解决方案

方案1:Objective-C原生API实现(推荐)

直接利用NSString的Unicode码点遍历API,无需手动处理编码转换,逻辑更可靠:

NSString *fileName = @"Tên tình bạn dưới tình yêu.mp3";
FILE *fd = fopen("output.txt", "w");
if (!fd) {
    NSLog(@"Failed to open file");
    return;
}

// 遍历每个Unicode字符(含组合字符序列)
[fileName enumerateSubstringsInRange:NSMakeRange(0, fileName.length)
                           options:NSStringEnumerationByComposedCharacterSequences
                        usingBlock:^(NSString * _Nullable substring, NSRange substringRange, NSRange enclosingRange, BOOL * _Nonnull stop) {
    // 获取当前字符的Unicode码点
    uint32_t codePoint = [substring characterAtIndex:0];
    
    // 选项1:写入十六进制字符串格式(如"EA 1EA1 6700 ")
    fprintf(fd, "%X ", codePoint);
    
    // 选项2:写入二进制字节(按小端字节序,符合你给出的中文'最'的示例)
    // if (codePoint <= 0xFF) {
    //     fputc(codePoint, fd);
    // } else if (codePoint <= 0xFFFF) {
    //     fputc(codePoint & 0xFF, fd);       // 低字节
    //     fputc((codePoint >> 8) & 0xFF, fd); // 高字节
    // } else {
    //     // 处理4字节Unicode码点(按需调整字节序)
    //     fputc(codePoint & 0xFF, fd);
    //     fputc((codePoint >> 8) & 0xFF, fd);
    //     fputc((codePoint >> 16) & 0xFF, fd);
    //     fputc((codePoint >> 24) & 0xFF, fd);
    // }
}];

fclose(fd);

说明:代码包含两种输出模式,可根据需求选择:

  • 十六进制字符串:直接输出码点的十六进制文本,便于阅读
  • 二进制字节:按小端字节序写入码点的原始字节,符合你给出的\u6700输出67 00的要求

方案2:纯C语言解析UTF-8获取码点

如果必须用C语言处理UTF-8字符串,需要手动解析每个UTF-8序列对应的Unicode码点:

#include <stdio.h>
#include <stdint.h>

// 从UTF-8字符串中解析单个Unicode码点
uint32_t utf8_to_codepoint(const char **str) {
    uint8_t c = **str;
    uint32_t codepoint;
    int bytes;

    if (c < 0x80) {
        codepoint = c;
        bytes = 1;
    } else if (c < 0xE0) {
        codepoint = (c & 0x1F) << 6;
        codepoint |= (*(*str + 1) & 0x3F);
        bytes = 2;
    } else if (c < 0xF0) {
        codepoint = (c & 0x0F) << 12;
        codepoint |= ((*(*str + 1) & 0x3F) << 6);
        codepoint |= (*(*str + 2) & 0x3F);
        bytes = 3;
    } else {
        codepoint = (c & 0x07) << 18;
        codepoint |= ((*(*str + 1) & 0x3F) << 12);
        codepoint |= ((*(*str + 2) & 0x3F) << 6);
        codepoint |= (*(*str + 3) & 0x3F);
        bytes = 4;
    }

    *str += bytes;
    return codepoint;
}

int main() {
    const char *stringText = "Tên tình bạn dưới tình yêu.mp3";
    FILE *fd = fopen("output.txt", "w");
    if (!fd) {
        printf("Failed to open file\n");
        return 1;
    }

    while (*stringText) {
        uint32_t cp = utf8_to_codepoint(&stringText);
        // 写入十六进制字符串
        fprintf(fd, "%X ", cp);
        // 或者写入二进制字节(小端),同方案1的选项2代码
    }

    fclose(fd);
    return 0;
}

内容的提问来源于stack exchange,提问作者KamyFC

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.17 11:27:40