You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python(NumPy)中无填充存储非字节对齐数据?

高效打包10/12位图像数据至HDF5的NumPy实现

核心思路

针对12位和10位数据分别设计向量化位打包逻辑,利用NumPy的视图操作与位运算实现无循环的高效转换,避免逐元素处理的性能损耗;同时存储原数组形状、位深等元数据,确保后续可完整恢复数据。


12位数据打包与恢复

打包逻辑(int16 → uint8)

2个12位数据共24位,刚好对应3个uint8字节。通过向量化位运算拆分重组,全程无Python循环:

import numpy as np
import h5py

def pack_12bit(arr):
    # 输入:shape=(H,W)的np.int16数组,仅低12位有效
    arr = arr.astype(np.uint16) & 0x0FFF  # 清除高位冗余填充
    flat = arr.flatten()
    # 补0使长度为偶数,保证每2个元素一组
    if len(flat) % 2 != 0:
        flat = np.pad(flat, (0, 1), mode='constant')
    pairs = flat.reshape(-1, 2)
    # 拆分重组为3个字节
    byte0 = (pairs[:, 0] >> 4).astype(np.uint8)
    byte1 = ((pairs[:, 0] & 0x0F) << 4) | ((pairs[:, 1] >> 8) & 0x0F)
    byte2 = (pairs[:, 1] & 0xFF).astype(np.uint8)
    # 合并为一维uint8数组
    packed = np.stack([byte0, byte1, byte2], axis=1).flatten()
    return packed, arr.shape

# 示例:生成测试数据并写入HDF5
h, w = 1080, 1920
raw_12bit = np.random.randint(0, 4096, size=(h, w), dtype=np.int16)
packed_data, orig_shape = pack_12bit(raw_12bit)

with h5py.File('12bit_data.h5', 'w') as f:
    f.create_dataset('packed', data=packed_data, dtype=np.uint8)
    # 存储元数据用于恢复
    f.attrs['original_shape'] = orig_shape
    f.attrs['bit_depth'] = 12

恢复逻辑(uint8 → int16)

从HDF5读取后反向拆分重组:

def unpack_12bit(packed_data, orig_shape):
    packed = packed_data.reshape(-1, 3)
    # 还原每组的两个12位值
    val0 = (packed[:, 0].astype(np.uint16) << 4) | ((packed[:, 1] >> 4) & 0x0F)
    val1 = ((packed[:, 1] & 0x0F).astype(np.uint16) << 8) | packed[:, 2].astype(np.uint16)
    flat = np.stack([val0, val1], axis=1).flatten()
    # 截断到原数组长度(去掉补的0)
    flat = flat[:orig_shape[0] * orig_shape[1]]
    return flat.reshape(orig_shape).astype(np.int16)

# 示例:读取并恢复数据
with h5py.File('12bit_data.h5', 'r') as f:
    packed = f['packed'][()]
    orig_shape = f.attrs['original_shape']
    bit_depth = f.attrs['bit_depth']

restored_data = unpack_12bit(packed, orig_shape)

10位数据打包与恢复

打包逻辑(int16 → uint8)

每4个10位数据共40位,刚好对应5个uint8字节,通过向量化位运算实现无循环转换:

def pack_10bit(arr):
    arr = arr.astype(np.uint16) & 0x03FF  # 保留低10位有效数据
    flat = arr.flatten()
    # 补0使长度为4的倍数,保证每4个元素一组
    pad_len = (4 - len(flat) % 4) % 4
    if pad_len > 0:
        flat = np.pad(flat, (0, pad_len), mode='constant')
    quads = flat.reshape(-1, 4)
    # 拆分重组为5个字节
    byte0 = (quads[:, 0] >> 2).astype(np.uint8)
    byte1 = ((quads[:, 0] & 0x3) << 6) | ((quads[:, 1] >> 4) & 0x3F)
    byte2 = ((quads[:, 1] & 0xF) << 4) | ((quads[:, 2] >> 6) & 0x0F)
    byte3 = ((quads[:, 2] & 0x3F) << 2) | ((quads[:, 3] >> 8) & 0x03)
    byte4 = (quads[:, 3] & 0xFF).astype(np.uint8)
    packed = np.stack([byte0, byte1, byte2, byte3, byte4], axis=1).flatten()
    return packed, arr.shape, pad_len

# 示例:生成测试数据并写入HDF5
raw_10bit = np.random.randint(0, 1024, size=(h, w), dtype=np.int16)
packed_data, orig_shape, pad_len = pack_10bit(raw_10bit)

with h5py.File('10bit_data.h5', 'w') as f:
    f.create_dataset('packed', data=packed_data, dtype=np.uint8)
    f.attrs['original_shape'] = orig_shape
    f.attrs['bit_depth'] = 10
    f.attrs['pad_length'] = pad_len  # 记录补0长度用于恢复

恢复逻辑

def unpack_10bit(packed_data, orig_shape, pad_len):
    packed = packed_data.reshape(-1, 5)
    # 还原每组的4个10位值
    val0 = (packed[:, 0].astype(np.uint16) << 2) | ((packed[:, 1] >> 6) & 0x3)
    val1 = ((packed[:, 1] & 0x3F).astype(np.uint16) << 4) | ((packed[:, 2] >> 4) & 0xF)
    val2 = ((packed[:, 2] & 0x0F).astype(np.uint16) << 6) | ((packed[:, 3] >> 2) & 0x3F)
    val3 = ((packed[:, 3] & 0x3).astype(np.uint16) << 8) | packed[:, 4].astype(np.uint16)
    flat = np.stack([val0, val1, val2, val3], axis=1).flatten()
    # 去掉补的0
    if pad_len > 0:
        flat = flat[:-pad_len]
    return flat.reshape(orig_shape).astype(np.int16)

# 示例:读取并恢复数据
with h5py.File('10bit_data.h5', 'r') as f:
    packed = f['packed'][()]
    orig_shape = f.attrs['original_shape']
    bit_depth = f.attrs['bit_depth']
    pad_len = f.attrs['pad_length']

restored_data = unpack_10bit(packed, orig_shape, pad_len)

性能优化说明

  • 全程采用NumPy向量化运算,无Python循环,可充分利用CPU向量指令,适配数百MB/s的视频流处理需求
  • 操作仅涉及位运算与数组形状重塑,内存占用可控,避免额外类型转换开销
  • 直接对接h5py的数据集写入,无需中间数据格式转换

内容的提问来源于stack exchange,提问作者Daniel Konečný

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.08 19:04:50