You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

4字节char转unsigned long出现异常位移问题求助

问题根源:有符号char的符号扩展

你的代码问题出在默认char是有符号类型(多数编译器默认配置),当读取的字节最高位为1时(比如你例子里的0xD0,二进制11010000),会被解释为负数,参与运算时会触发符号扩展——把高位全部填充为1,直接打乱了字节拼接的逻辑,导致计算结果异常。

以你提到的目标数值184346为例,对应的四个字节是0x00、0x02、0xD0、0x1A:

  • bytes[2] = 0xD0作为有符号char时,实际值是-48
  • 执行(output << 8) + bytes[2]时,bytes[2]会被自动提升为int类型,变成0xFFFFFFD0
  • 之前的output是0x02 << 8 = 0x200,相加后得到0x200 + 0xFFFFFFD0 = 0x1FFFFFFD0,赋值给unsigned long后完全偏离预期,看起来像是高位的0x02被“替换”成了0x01,本质是符号扩展导致的高位溢出
修复方案

核心思路是避免符号扩展,所有字节运算前先转换为无符号类型:

方案1:直接用unsigned char存储字节

从根源上避免有符号类型的问题,把数组定义改为unsigned char:

unsigned long get_4b_size(FILE* file) {
  unsigned char bytes[4];
  fread(bytes, 1, 4, file);
  unsigned long output = bytes[0];
  for (int i = 1; i < 4; i++) {
    output = (output << 8) + bytes[i];
  }
  return output;
}

方案2:运算时强制转换

如果不想修改数组类型,在每个字节参与运算前强制转为unsigned char:

unsigned long get_4b_size(FILE* file) {
  char bytes[4];
  fread(bytes, 1, 4, file);
  unsigned long output = (unsigned char)bytes[0];
  for (int i = 1; i < 4; i++) {
    output = (output << 8) + (unsigned char)bytes[i];
  }
  return output;
}

方案3:修复直接位移的写法

针对你尝试过的直接位移求和写法,同样需要给每个字节加强制转换,同时建议用|代替+——逻辑上更符合字节拼接的语义,避免数值相加可能带来的意外进位:

unsigned long get_4b_size(FILE* file) {
  char bytes[4];
  fread(bytes, 1, 4, file);
  return ((unsigned char)bytes[0] << 24) | 
         ((unsigned char)bytes[1] << 16) | 
         ((unsigned char)bytes[2] << 8) | 
         (unsigned char)bytes[3];
}
额外注意事项
  • 字节序检查:你的代码默认是大端字节序(最高位字节先读取),如果目标文件是小端字节序,需要调整字节的拼接顺序。
  • 读取有效性校验:实际开发中要判断fread的返回值,确认是否成功读取了4个字节,避免文件读取失败导致的错误计算。

内容的提问来源于stack exchange,提问作者Marina Winter

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.03 03:12:34