You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python与C编码差异:如何实现简洁的跨语言二进制文件读写?

解决方案:Python写入单字节编码文本到二进制文件,简化C读取

问题出在你用了array('u')存储字符串——这个类型对应UTF-32编码的Unicode字符,每个字符占4字节,所以"test"4个字符会占用16字节,导致C端需要跳过冗余字节才能读取有效内容。要实现4字节存储,只需把字符串编码为单字节编码格式(比如ASCII或Latin-1),直接写入二进制文件即可。

修改后的Python代码

方式1:直接编码为字节串写入

这是最简洁的实现,将字符串编码为ASCII字节串(每个字符1字节)后直接写入:

string = "test"
# 编码为ASCII字节串,"test"对应4个字节
byte_data = string.encode('ascii')

# 使用with语句自动管理文件关闭
with open('test_file.bin', 'wb') as output_file:
    output_file.write(byte_data)

方式2:用array模块的单字节类型存储

如果坚持使用array模块,选择'B'(无符号单字节)或'b'(有符号单字节)类型,先将字符串编码为字节串再转换为数组:

from array import array

string = "test"
# 编码为ASCII字节串,再转为无符号单字节数组
byte_array = array('B', string.encode('ascii'))

with open('test_file.bin', 'wb') as output_file:
    byte_array.tofile(output_file)

简化后的C读取代码

原C代码存在未初始化指针的问题(char *symbol;直接使用会触发未定义行为),以下是修复并简化后的版本:

#include <stdio.h>
#include <stdlib.h>

int main()
{
    FILE *file_pointer = fopen("test_file.bin", "rb");
    if (!file_pointer) {
        perror("Failed to open file");
        return 1;
    }

    // 分配5字节:4字节存"test" + 1字节存字符串终止符'\0'
    char *text = malloc(5 * sizeof(char));
    if (!text) {
        perror("Failed to allocate memory");
        fclose(file_pointer);
        return 1;
    }

    // 读取4个字节到text中
    size_t bytes_read = fread(text, sizeof(char), 4, file_pointer);
    // 手动添加字符串终止符,确保printf能正确识别字符串结尾
    text[bytes_read] = '\0';

    printf("%s\n", text);

    // 释放内存、关闭文件
    free(text);
    fclose(file_pointer);
    return 0;
}

补充说明

  • 如果字符串包含ASCII之外的字符(比如西欧语言字符),可以用encode('latin-1')替代encode('ascii')——Latin-1是单字节编码,覆盖更多字符,同样保证每个字符占1字节。
  • 如果需要处理中文等多字节字符,单字节编码无法满足,此时建议使用UTF-8编码(Python默认encode()就是UTF-8),但UTF-8是变长编码,C端读取时需要按UTF-8规则解析。

内容的提问来源于stack exchange,提问作者Evgueni Dinvay

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.12 01:43:31