二进制读取BMP结构体时丢失字节的原因排查
尝试读取.bmp文件头并输出到控制台,定义了两个结构体存储BMP头信息:
struct bmp_header{ uint16_t bmp_file_code{0x424D}; // BMP格式标识,固定为0x424D uint32_t bmp_file_size{0}; // 文件总大小 uint32_t application_data{0}; // 通常未使用 uint32_t pixel_offset{0}; // 像素数据起始偏移量 }; struct bmp_dib_header{ uint32_t dib_head_size; // DIB头大小(字节) uint32_t image_width; // 图像宽度(像素) uint32_t image_height; // 图像高度(像素) uint16_t color_planes; uint16_t bits_per_pixel; uint32_t compression_method; // 压缩方式代码 uint32_t image_size; // 图像数据大小(字节,宽度*高度*通道数) int32_t x_resolution; // 水平分辨率(像素/米) int32_t y_resolution; // 垂直分辨率(像素/米) uint32_t color_palette; // 使用的颜色数(默认0表示2^n种) uint32_t important_colors; // 重要颜色数(通常忽略) };
读取方法如下:
void read_header(std::ifstream &filestream) { auto end = filestream.tellg(); filestream.seekg(0, std::ios_base::beg); if (end - filestream.tellg() <= 0) throw std::runtime_error(strerror(errno)); // 读取文件头 filestream.read(reinterpret_cast<char*>(&bmp_head), sizeof(bmp_header)); // 读取DIB头 filestream.read(reinterpret_cast<char*>(&bmp_dib), sizeof(bmp_dib_header)); }
其中filestream以std::ios_base::binary和std::ios_base::in模式打开,输出函数为:
void debug_show_header_info() { std::cout << "BMP头数据:" << std::endl \ << "文件标识: " << std::hex << bmp_head.bmp_file_code << std::endl \ << "文件大小: " << std::hex << bmp_head.bmp_file_size << std::endl \ << "保留数据: " << std::hex << bmp_head.application_data << std::endl \ << "像素偏移: " << std::hex << bmp_head.pixel_offset << std::endl; }
十六进制编辑器中显示的数据正确,但程序输出的文件大小等字段值错误:
BMP头数据: 文件标识: 4d42 文件大小: c <<<< 实际应为0x528F0C 保留数据: 8a0000 像素偏移: 7c0000
尝试逐个读取值也得到相同结果,测试图片为filesampleshub.com的sample1.bmp。
核心问题:结构体对齐+字节序不匹配
结构体内存对齐问题
编译器默认会为结构体成员添加填充字节以保证内存对齐,但BMP文件头是紧凑连续存储的,无填充字节。比如bmp_header中,uint16_t类型的bmp_file_code占2字节,后续uint32_t类型的成员需要4字节对齐,编译器会自动在bmp_file_code后添加2字节填充,导致读取时把文件中原本属于bmp_file_size的前2字节当成填充字节跳过,最终读错数据。字节序问题
BMP文件采用小端字节序存储多字节数据,而主机字节序可能与小端不一致(比如部分大端架构)。例如文件标识0x424D(ASCII"BM")在文件中实际存储为4D 42,直接读取会被解析为0x4D42,和输出结果一致;文件大小0x528F0C在文件中存储为0C 8F 52 00,错误的对齐加上字节序问题会进一步放大解析错误。
解决方案
方法一:禁用结构体对齐
在结构体定义前添加编译器指令,强制紧凑存储,消除填充字节:
// GCC/Clang/VS通用指令 #pragma pack(push, 1) struct bmp_header{ uint16_t bmp_file_code{0x424D}; uint32_t bmp_file_size{0}; uint32_t application_data{0}; uint32_t pixel_offset{0}; }; struct bmp_dib_header{ uint32_t dib_head_size; uint32_t image_width; uint32_t image_height; uint16_t color_planes; uint16_t bits_per_pixel; uint32_t compression_method; uint32_t image_size; int32_t x_resolution; int32_t y_resolution; uint32_t color_palette; uint32_t important_colors; }; #pragma pack(pop)
方法二:手动处理字节序与字段读取
逐个读取每个字段,并手动转换为主机字节序(保证跨平台兼容性):
#include <cstdint> // 小端转主机字节序工具函数 uint16_t le_to_host16(uint16_t val) { #if __BYTE_ORDER__ == __ORDER_LITTLE_ENDIAN__ return val; #else return ((val >> 8) & 0xFF) | ((val << 8) & 0xFF00); #endif } uint32_t le_to_host32(uint32_t val) { #if __BYTE_ORDER__ == __ORDER_LITTLE_ENDIAN__ return val; #else return ((val >> 24) & 0xFF) | ((val >> 8) & 0xFF00) | ((val << 8) & 0xFF0000) | ((val << 24) & 0xFF000000); #endif } void read_header(std::ifstream &filestream) { auto end = filestream.tellg(); filestream.seekg(0, std::ios_base::beg); if (end - filestream.tellg() <= 0) throw std::runtime_error(strerror(errno)); // 逐个读取并转换文件头字段 filestream.read(reinterpret_cast<char*>(&bmp_head.bmp_file_code), sizeof(uint16_t)); bmp_head.bmp_file_code = le_to_host16(bmp_head.bmp_file_code); filestream.read(reinterpret_cast<char*>(&bmp_head.bmp_file_size), sizeof(uint32_t)); bmp_head.bmp_file_size = le_to_host32(bmp_head.bmp_file_size); filestream.read(reinterpret_cast<char*>(&bmp_head.application_data), sizeof(uint32_t)); bmp_head.application_data = le_to_host32(bmp_head.application_data); filestream.read(reinterpret_cast<char*>(&bmp_head.pixel_offset), sizeof(uint32_t)); bmp_head.pixel_offset = le_to_host32(bmp_head.pixel_offset); // DIB头字段同理,逐个读取并转换 filestream.read(reinterpret_cast<char*>(&bmp_dib.dib_head_size), sizeof(uint32_t)); bmp_dib.dib_head_size = le_to_host32(bmp_dib.dib_head_size); filestream.read(reinterpret_cast<char*>(&bmp_dib.image_width), sizeof(uint32_t)); bmp_dib.image_width = le_to_host32(bmp_dib.image_width); // 其余字段按相同方式处理... }
修改后,文件标识会正确解析为0x424D,文件大小也会显示为预期的0x528F0C。
内容的提问来源于stack exchange,提问作者Werder Foster

