如何用Python基于数据而非扩展名识别压缩文件真实类型
批量识别归档文件真实格式的高效方案
不用逐个尝试解压,核心思路是通过文件签名(文件头)识别真实格式——这是最可靠且高效的方式,因为每种归档格式都有固定的特征字节序列,和后缀名无关。以下是具体实现方案:
1. 直接读取文件头签名识别
常见归档格式的特征签名:
- ZIP:文件前2字节为
PK(十六进制50 4B) - RAR:文件前4字节为
Rar!(十六进制52 61 72 21,主流RAR 3.x及以上版本通用) - 7Z:文件前6字节为
7zBCAF27(十六进制37 7A BC AF 27 1C)
用Python实现的批量识别示例:
import os def detect_archive_type(file_path): signature_map = { b'PK': 'zip', b'Rar!': 'rar', b'7z\xBC\xAF\x27\x1C': '7z' } with open(file_path, 'rb') as f: header = f.read(6) # 读取足够覆盖所有目标格式的字节数 for sig, typ in signature_map.items(): if header.startswith(sig): return typ return None # 批量处理指定目录下的文件 target_dir = './archives' for filename in os.listdir(target_dir): file_path = os.path.join(target_dir, filename) if os.path.isfile(file_path): arch_type = detect_archive_type(file_path) if arch_type: print(f"{filename} 真实格式为 {arch_type}") # 自动调用对应解压工具(根据实际环境调整命令) if arch_type == 'zip': os.system(f'unzip "{file_path}" -d "{os.path.splitext(file_path)[0]}"') elif arch_type == 'rar': os.system(f'unrar x "{file_path}" "{os.path.splitext(file_path)[0]}"') elif arch_type == '7z': os.system(f'7z x "{file_path}" -o"{os.path.splitext(file_path)[0]}"')
2. 利用系统工具批量检测
如果是Linux/macOS环境,直接用file命令就能识别文件真实类型,无需自己写签名匹配逻辑:
# 单个文件检测 file --mime-type a.gho # 输出示例:a.gho: application/zip # 批量检测并标记格式 for file in ./archives/*; do if [ -f "$file" ]; then mime_type=$(file --mime-type "$file" | awk -F': ' '{print $2}') case $mime_type in application/zip) echo "$file -> ZIP格式" ;; application/x-rar) echo "$file -> RAR格式" ;; application/x-7z-compressed) echo "$file -> 7Z格式" ;; esac fi done
Windows环境可通过PowerShell结合文件头判断,或使用第三方工具(如GNUWin32的file.exe)实现类似功能。
3. 封装自动化解压脚本
把识别逻辑和解压操作整合,写一个批量处理脚本,一次性完成「识别真实格式→调用对应工具解压」的全流程,彻底替代逐个尝试的低效方式。
内容的提问来源于stack exchange,提问作者cloudinstone
相关产品推荐
相关产品推荐

