在Azure Databricks中解压ADLS Gen2容器内ZIP文件遇错求助
在Azure Databricks中用PySpark解压ADLS Gen2的ZIP文件时遇到BadZipFile或FileNotFoundError
问题背景
我能正常读取ADLS Gen2容器同一文件夹下的CSV文件,但解压ZIP文件时遇到异常。ZIP文件路径通过dbutils.fs.ls(blob_folder_url)获取,使用zipfile.ZipFile时触发两种错误:BadZipFile或FileNotFoundError。
错误详情
- BadZipFile错误:代码尝试将读取到的ZIP内容转成字节流初始化
ZipFile时,提示File is not a zip file。 - FileNotFoundError错误:直接使用ADLS的
abfss://路径初始化ZipFile时,提示[Errno 2] No such file or directory。 - 正常读取CSV的情况:通过相同的ADLS路径配置,能成功读取文件夹内的CSV文件,说明存储权限和路径配置无问题。
我的代码
import zipfile, os, io, re # Azure Blob Storage details storage_account_name = "<>" container_name = "<>" folder_path = "<>" blob_folder_url = f"abfss://{container_name}@{storage_account_name}.dfs.core.windows.net/{folder_path}" zip_file = blob_folder_url + 'batch1_weekly_catman_20241109.zip' # List files in the specified blob folder files = dbutils.fs.ls(blob_folder_url) for file in files: # Check if the file is a ZIP file if file.name.endswith('.zip'): print(f"Processing ZIP file: {file.name}") # Read the ZIP file into memory zip_file_path = file.path zip_blob_data = dbutils.fs.head(zip_file_path) # Read the ZIP file content # Unzip the file with zipfile.ZipFile(io.BytesIO(zip_blob_data.encode('utf-8')), 'r') as z: print('zipppppppper') # with zipfile.ZipFile(zip_file, 'r') as z: # print('zipppppppper')
错误信息
BadZipFile: File is not a zip fileFileNotFoundError: [Errno 2] No such file or directory
解决方法
1. 错误根源分析
- FileNotFoundError:
zipfile.ZipFile仅支持本地文件路径或内存字节流,无法直接识别ADLS的abfss://分布式路径,直接传入该路径会触发文件找不到的错误。 - BadZipFile:
dbutils.fs.head()默认仅读取文件前1024字节,且返回字符串格式,对于二进制的ZIP文件来说,读取的内容不完整,转成字节流后不是有效的ZIP格式。
2. 可行解决方案
方案一:使用本地临时文件处理
将ADLS上的ZIP文件复制到Databricks本地临时目录,再用zipfile处理,适合大文件场景:
import zipfile # 遍历目标文件夹下的文件 for file in files: if file.name.endswith('.zip'): print(f"Processing ZIP file: {file.name}") # 定义本地临时文件路径 temp_zip_path = f"/tmp/{file.name}" # 从ADLS复制文件到本地临时路径(注意前缀file:) dbutils.fs.cp(file.path, f"file:{temp_zip_path}") # 解压并处理ZIP文件 with zipfile.ZipFile(temp_zip_path, 'r') as zf: # 打印ZIP内的文件列表 print("ZIP包含文件:", zf.namelist()) # 示例:读取ZIP内的第一个文件内容 with zf.open(zf.namelist()[0]) as f: content = f.read() print("文件前100字节内容:", content[:100]) # 清理临时文件 dbutils.fs.rm(f"file:{temp_zip_path}")
方案二:读取完整二进制内容到内存处理
通过Spark的binaryFile格式读取完整的ZIP二进制内容,适合小文件场景:
import zipfile import io from pyspark.sql.functions import col # 遍历目标文件夹下的文件 for file in files: if file.name.endswith('.zip'): print(f"Processing ZIP file: {file.name}") # 读取ZIP文件的完整二进制内容 df = spark.read.format("binaryFile").load(file.path) binary_content = df.select(col("content")).first()[0] # 用字节流初始化ZipFile并处理 with zipfile.ZipFile(io.BytesIO(binary_content), 'r') as zf: print("ZIP包含文件:", zf.namelist()) # 示例:处理ZIP内的CSV文件 for filename in zf.namelist(): if filename.endswith('.csv'): with zf.open(filename) as f: csv_content = f.read().decode('utf-8') print("CSV文件前500字符内容:", csv_content[:500])
关键注意事项
- 禁止用
dbutils.fs.head()读取二进制文件,该方法仅适用于文本文件的内容预览。 - 处理大ZIP文件时优先选择本地临时文件方案,避免内存溢出。
- 确保Databricks集群对ADLS Gen2容器有读写权限(已通过CSV读取验证,此点可忽略)。
内容的提问来源于stack exchange,提问作者Mariah Akinbi
相关产品推荐
相关产品推荐

