求助:如何用Python基于指定HEX值拆分文件并保存
文件拆分:提取包含指定4字节HEX标识符的部分
核心思路
你需要先定位到0A FF AA 1B对应的字节串(b'\x0a\xff\xaa\x1b'),再从该位置开始读取并保存后续内容(包含标识符本身)。直接用Seek()不是完整解决方案——它只能用于定位已知偏移,你得先找到标识符的位置,再结合Seek()完成读取。
推荐实现(分块查找,适配大文件)
这种方法避免一次性加载大文件到内存,同时处理标识符跨分块的情况:
# 定义目标标识符的字节形式 target_bytes = b'\x0a\xff\xaa\x1b' chunk_size = 4096 # 分块大小可根据文件体积调整 with open("原始文件路径", "rb") as src_file, open("输出文件路径", "wb") as dst_file: buffer = b'' while True: current_chunk = src_file.read(chunk_size) if not current_chunk: print("未找到目标标识符") break # 合并缓冲区与当前块,查找目标 combined_data = buffer + current_chunk target_pos = combined_data.find(target_bytes) if target_pos != -1: # 写入从标识符开始的所有内容(包含标识符) dst_file.write(combined_data[target_pos:]) # 写入剩余的文件内容 dst_file.write(src_file.read()) break else: # 保留最后3个字节(标识符长度-1),防止标识符跨块被截断 buffer = combined_data[-(len(target_bytes)-1):]
结合Seek()的实现方式
如果你坚持用Seek(),可以先遍历文件找到标识符偏移,再定位读取:
target_bytes = b'\x0a\xff\xaa\x1b' target_length = len(target_bytes) current_offset = 0 with open("原始文件路径", "rb") as src_file: while True: chunk = src_file.read(4096) if not chunk: print("未找到目标标识符") break target_pos = chunk.find(target_bytes) if target_pos != -1: # 计算标识符的总偏移量 total_offset = current_offset + target_pos # 定位到标识符起始位置 src_file.seek(total_offset) # 读取并保存后续内容 with open("输出文件路径", "wb") as dst_file: dst_file.write(src_file.read()) break # 回退偏移,避免跨块遗漏标识符 current_offset += len(chunk) - (target_length - 1) src_file.seek(current_offset)
关键说明
- 必须将HEX值转换为字节串处理,不能直接用字符串匹配,避免编码问题。
- 分块处理时保留缓冲区是为了防止标识符刚好跨两个读取块的情况。
内容的提问来源于stack exchange,提问作者molko
相关产品推荐
相关产品推荐

