You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用多线程结合requests批量下载并提取Gzip文件中的JSON内容

问题原因

gzip.GzipFile的fileobj入参要求传入支持read()、seek()操作的类文件对象,你直接传入字节类型的html.content不符合参数要求,导致解压读取不到内容,最终生成空文件。

修复方案

首先导入依赖的io模块:

import io

调整fetch_file函数实现,先把字节内容转为BytesIO类文件对象再传入gzip.GzipFile:

def fetch_file(url,filename):
    try:
        html = requests.get(url, stream=True, allow_redirects=True)
        # 校验请求是否成功,避免把错误响应当做gzip处理
        html.raise_for_status()
        # 转为BytesIO类文件对象后解压
        with gzip.GzipFile(fileobj=io.BytesIO(html.content)) as gz:
            with open(f'test/{filename}.json','wb') as out:
                out.write(gz.read())
        return html.status_code
    except Exception as e:
       return e

可选优化建议

  • 流式解压降内存:如果单文件体积较大,可以直接用response.raw类文件对象传入解压,不需要把整个压缩包加载到内存,性能更高:
def fetch_file(url,filename):
    try:
        html = requests.get(url, stream=True, allow_redirects=True)
        html.raise_for_status()
        with gzip.GzipFile(fileobj=html.raw) as gz:
            with open(f'test/{filename}.json','wb') as out:
                out.write(gz.read())
        return html.status_code
    except Exception as e:
       return e
  • 调整并发数:当前线程池最大工作数为10,若带宽充足且目标接口无严格频率限制,可将max_workers调整为20~50,进一步提升62000个文件的处理效率。
  • 路径兼容处理:使用os.path.join拼接文件路径,避免不同操作系统的路径分隔符不兼容问题。
  • 异常分类处理:可根据需要拆分请求异常、解压异常、文件写入异常的捕获逻辑,方便排查处理失败的任务。

内容的提问来源于stack exchange,提问作者Niels

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.03 20:18:02