基于起始/结束标记拆分存储UDP数据包的NumPy数组
按自定义首尾标记拆分UDP数据包NumPy数组
核心思路
通过NumPy的向量化操作定位所有起始(0xAAAA)和结束(0xD6D6)标记的位置,再根据标记对提取单个数据包,替代只能等分的np.array_split。
实现方案
情况1:数组为uint8类型(单个元素对应1字节)
如果二进制文件以单字节读取(每个元素范围0x00-0xFF),0xAAAA和0xD6D6是连续两个字节的标记,代码如下:
import numpy as np # 读取二进制文件为uint8数组(假设已完成此步骤,此处仅作示例) data = np.fromfile("udp_packets.bin", dtype=np.uint8) # 定位所有起始标记(0xAA, 0xAA)的起始索引 start_indices = np.where((data[:-1] == 0xAA) & (data[1:] == 0xAA))[0] # 定位所有结束标记(0xD6, 0xD6)的起始索引 end_indices = np.where((data[:-1] == 0xD6) & (data[1:] == 0xD6))[0] # 基础校验:确保标记数量匹配且顺序正确 if len(start_indices) != len(end_indices): raise ValueError("起始/结束标记数量不匹配,存在损坏数据包") for start, end in zip(start_indices, end_indices): if start >= end: raise ValueError("出现起始标记在结束标记之后的异常数据") # 提取每个完整数据包 packets = [] for start, end in zip(start_indices, end_indices): # 从起始标记第一个字节,到结束标记第二个字节(切片左闭右开,故取end+2) packet = data[start : end + 2] packets.append(packet)
情况2:数组为uint16类型(单个元素对应2字节)
如果读取时直接按16位解析(每个元素对应两个字节),标记可直接匹配单个元素,代码更简洁:
import numpy as np data = np.fromfile("udp_packets.bin", dtype=np.uint16) start_indices = np.where(data == 0xAAAA)[0] end_indices = np.where(data == 0xD6D6)[0] # 基础校验 if len(start_indices) != len(end_indices): raise ValueError("标记数量不匹配") for start, end in zip(start_indices, end_indices): if start >= end: raise ValueError("标记顺序异常") # 提取数据包 packets = [data[start : end + 1] for start, end in zip(start_indices, end_indices)]
补充说明
- 若数据包之间存在冗余字节,上述代码会自动跳过,仅提取首尾标记包裹的有效数据段。
- 若存在标记嵌套或错位,需根据实际数据格式调整校验逻辑,比如匹配每个起始标记对应的最近结束标记。
内容的提问来源于stack exchange,提问作者Prajjalak Chattopadhyay
相关产品推荐
相关产品推荐

