从伪装为ZIP的Gzip文件提取PKL文件遇技术问题
问题描述
我有一批后缀为.zip但实际是Gzip格式的文件,需要从中提取.pkl文件。使用gunzip工具无法处理,因为它仅支持.gz后缀。我尝试了以下Python代码,但代码会将pkl、ekl及CRC数据全部写入同一个pkl文件中,无法正确拆分文件。
原代码
def extract_gzip_file_and_create_folder(event): source_path = event['source_path'] target_path = event['target_path'] gzip_files = [file for file in os.listdir(source_path) if file.endswith('.zip')] for gzip_file in gzip_files: source_file = os.path.join(source_path, gzip_file) # Create a folder for each file in the target path folder_name = os.path.splitext(gzip_file)[0] # Remove '.zip' folder_path = os.path.join(target_path, folder_name) os.makedirs(folder_path, exist_ok=True) # Create the folder if it doesn't exist try: with gzip.open(source_file, 'rb') as f_in: for file_name in ['pkl', 'ekl']: target_file = os.path.join(folder_path, folder_name + '.' + file_name) with open(target_file, 'wb') as f_out: shutil.copyfileobj(f_in, f_out) print(f"Extracted: {gzip_file} to {target_file}") except Exception as e: print(f"An error occurred while extracting {gzip_file}: {e}")
问题分析
- 循环逻辑错误:创建文件夹的
for循环结束后,source_file和folder_path仅保留最后一个文件的信息,try块在循环外,导致只会处理最后一个.zip文件,前面的文件都被忽略。 - 文件拆分逻辑缺失:
shutil.copyfileobj会一次性读取整个Gzip流的内容,第一次循环写入pkl文件时就会把所有解压后的内容写完,第二次写入ekl文件时流已经为空。另外,如果你的Gzip文件实际是tar打包后再压缩的(即tar.gz伪装成zip),直接用gzip.open读取无法识别内部的多个文件结构,必须用tarfile模块处理。
解决方案
情况1:Gzip内是tar打包的多文件(最常见场景)
如果你的文件是tar打包后再用gzip压缩的,直接用tarfile模块处理,自动识别内部文件:
import os import tarfile def extract_gzip_tar_file(event): source_path = event['source_path'] target_path = event['target_path'] # 筛选后缀为.zip但实际是gzip的文件 gzip_files = [f for f in os.listdir(source_path) if f.endswith('.zip')] for gzip_file in gzip_files: source_file = os.path.join(source_path, gzip_file) folder_name = os.path.splitext(gzip_file)[0] folder_path = os.path.join(target_path, folder_name) os.makedirs(folder_path, exist_ok=True) try: # 打开tar.gz文件(自动识别gzip压缩) with tarfile.open(source_file, 'r:gz') as tar: # 提取所有文件到目标文件夹 tar.extractall(path=folder_path) # 如果只需要提取.pkl文件,可以用下面的代码: # for member in tar.getmembers(): # if member.name.endswith('.pkl'): # tar.extract(member, path=folder_path) print(f"成功提取文件: {gzip_file} 到 {folder_path}") except Exception as e: print(f"提取 {gzip_file} 时出错: {e}")
情况2:Gzip内是自定义拼接的多文件(已知每个文件大小)
如果你的Gzip解压后是多个文件直接拼接的,且知道每个文件的大小,可以按字节数拆分:
import os import gzip import shutil def extract_spliced_gzip_file(event): source_path = event['source_path'] target_path = event['target_path'] # 假设你知道pkl和ekl文件的大小(单位:字节) file_sizes = {'pkl': 102400, 'ekl': 51200} # 替换为实际大小 gzip_files = [f for f in os.listdir(source_path) if f.endswith('.zip')] for gzip_file in gzip_files: source_file = os.path.join(source_path, gzip_file) folder_name = os.path.splitext(gzip_file)[0] folder_path = os.path.join(target_path, folder_name) os.makedirs(folder_path, exist_ok=True) try: with gzip.open(source_file, 'rb') as f_in: for ext, size in file_sizes.items(): target_file = os.path.join(folder_path, f"{folder_name}.{ext}") with open(target_file, 'wb') as f_out: # 按指定大小读取并写入 shutil.copyfileobj(f_in, f_out, length=size) print(f"成功提取: {target_file}") # 如果还有CRC数据需要处理,可以继续读取剩余内容 crc_file = os.path.join(folder_path, f"{folder_name}.crc") with open(crc_file, 'wb') as f_out: shutil.copyfileobj(f_in, f_out) print(f"成功提取CRC文件: {crc_file}") except Exception as e: print(f"提取 {gzip_file} 时出错: {e}")
内容的提问来源于stack exchange,提问作者random23
相关产品推荐
相关产品推荐

