如何在Python(NumPy)中无填充存储非字节对齐数据?
高效打包10/12位图像数据至HDF5的NumPy实现
核心思路
针对12位和10位数据分别设计向量化位打包逻辑,利用NumPy的视图操作与位运算实现无循环的高效转换,避免逐元素处理的性能损耗;同时存储原数组形状、位深等元数据,确保后续可完整恢复数据。
12位数据打包与恢复
打包逻辑(int16 → uint8)
2个12位数据共24位,刚好对应3个uint8字节。通过向量化位运算拆分重组,全程无Python循环:
import numpy as np import h5py def pack_12bit(arr): # 输入:shape=(H,W)的np.int16数组,仅低12位有效 arr = arr.astype(np.uint16) & 0x0FFF # 清除高位冗余填充 flat = arr.flatten() # 补0使长度为偶数,保证每2个元素一组 if len(flat) % 2 != 0: flat = np.pad(flat, (0, 1), mode='constant') pairs = flat.reshape(-1, 2) # 拆分重组为3个字节 byte0 = (pairs[:, 0] >> 4).astype(np.uint8) byte1 = ((pairs[:, 0] & 0x0F) << 4) | ((pairs[:, 1] >> 8) & 0x0F) byte2 = (pairs[:, 1] & 0xFF).astype(np.uint8) # 合并为一维uint8数组 packed = np.stack([byte0, byte1, byte2], axis=1).flatten() return packed, arr.shape # 示例:生成测试数据并写入HDF5 h, w = 1080, 1920 raw_12bit = np.random.randint(0, 4096, size=(h, w), dtype=np.int16) packed_data, orig_shape = pack_12bit(raw_12bit) with h5py.File('12bit_data.h5', 'w') as f: f.create_dataset('packed', data=packed_data, dtype=np.uint8) # 存储元数据用于恢复 f.attrs['original_shape'] = orig_shape f.attrs['bit_depth'] = 12
恢复逻辑(uint8 → int16)
从HDF5读取后反向拆分重组:
def unpack_12bit(packed_data, orig_shape): packed = packed_data.reshape(-1, 3) # 还原每组的两个12位值 val0 = (packed[:, 0].astype(np.uint16) << 4) | ((packed[:, 1] >> 4) & 0x0F) val1 = ((packed[:, 1] & 0x0F).astype(np.uint16) << 8) | packed[:, 2].astype(np.uint16) flat = np.stack([val0, val1], axis=1).flatten() # 截断到原数组长度(去掉补的0) flat = flat[:orig_shape[0] * orig_shape[1]] return flat.reshape(orig_shape).astype(np.int16) # 示例:读取并恢复数据 with h5py.File('12bit_data.h5', 'r') as f: packed = f['packed'][()] orig_shape = f.attrs['original_shape'] bit_depth = f.attrs['bit_depth'] restored_data = unpack_12bit(packed, orig_shape)
10位数据打包与恢复
打包逻辑(int16 → uint8)
每4个10位数据共40位,刚好对应5个uint8字节,通过向量化位运算实现无循环转换:
def pack_10bit(arr): arr = arr.astype(np.uint16) & 0x03FF # 保留低10位有效数据 flat = arr.flatten() # 补0使长度为4的倍数,保证每4个元素一组 pad_len = (4 - len(flat) % 4) % 4 if pad_len > 0: flat = np.pad(flat, (0, pad_len), mode='constant') quads = flat.reshape(-1, 4) # 拆分重组为5个字节 byte0 = (quads[:, 0] >> 2).astype(np.uint8) byte1 = ((quads[:, 0] & 0x3) << 6) | ((quads[:, 1] >> 4) & 0x3F) byte2 = ((quads[:, 1] & 0xF) << 4) | ((quads[:, 2] >> 6) & 0x0F) byte3 = ((quads[:, 2] & 0x3F) << 2) | ((quads[:, 3] >> 8) & 0x03) byte4 = (quads[:, 3] & 0xFF).astype(np.uint8) packed = np.stack([byte0, byte1, byte2, byte3, byte4], axis=1).flatten() return packed, arr.shape, pad_len # 示例:生成测试数据并写入HDF5 raw_10bit = np.random.randint(0, 1024, size=(h, w), dtype=np.int16) packed_data, orig_shape, pad_len = pack_10bit(raw_10bit) with h5py.File('10bit_data.h5', 'w') as f: f.create_dataset('packed', data=packed_data, dtype=np.uint8) f.attrs['original_shape'] = orig_shape f.attrs['bit_depth'] = 10 f.attrs['pad_length'] = pad_len # 记录补0长度用于恢复
恢复逻辑
def unpack_10bit(packed_data, orig_shape, pad_len): packed = packed_data.reshape(-1, 5) # 还原每组的4个10位值 val0 = (packed[:, 0].astype(np.uint16) << 2) | ((packed[:, 1] >> 6) & 0x3) val1 = ((packed[:, 1] & 0x3F).astype(np.uint16) << 4) | ((packed[:, 2] >> 4) & 0xF) val2 = ((packed[:, 2] & 0x0F).astype(np.uint16) << 6) | ((packed[:, 3] >> 2) & 0x3F) val3 = ((packed[:, 3] & 0x3).astype(np.uint16) << 8) | packed[:, 4].astype(np.uint16) flat = np.stack([val0, val1, val2, val3], axis=1).flatten() # 去掉补的0 if pad_len > 0: flat = flat[:-pad_len] return flat.reshape(orig_shape).astype(np.int16) # 示例:读取并恢复数据 with h5py.File('10bit_data.h5', 'r') as f: packed = f['packed'][()] orig_shape = f.attrs['original_shape'] bit_depth = f.attrs['bit_depth'] pad_len = f.attrs['pad_length'] restored_data = unpack_10bit(packed, orig_shape, pad_len)
性能优化说明
- 全程采用NumPy向量化运算,无Python循环,可充分利用CPU向量指令,适配数百MB/s的视频流处理需求
- 操作仅涉及位运算与数组形状重塑,内存占用可控,避免额外类型转换开销
- 直接对接h5py的数据集写入,无需中间数据格式转换
内容的提问来源于stack exchange,提问作者Daniel Konečný
相关产品推荐
相关产品推荐

