使用多线程结合requests批量下载并提取Gzip文件中的JSON内容
问题原因
gzip.GzipFile的fileobj入参要求传入支持read()、seek()操作的类文件对象,你直接传入字节类型的html.content不符合参数要求,导致解压读取不到内容,最终生成空文件。
修复方案
首先导入依赖的io模块:
import io
调整fetch_file函数实现,先把字节内容转为BytesIO类文件对象再传入gzip.GzipFile:
def fetch_file(url,filename): try: html = requests.get(url, stream=True, allow_redirects=True) # 校验请求是否成功,避免把错误响应当做gzip处理 html.raise_for_status() # 转为BytesIO类文件对象后解压 with gzip.GzipFile(fileobj=io.BytesIO(html.content)) as gz: with open(f'test/{filename}.json','wb') as out: out.write(gz.read()) return html.status_code except Exception as e: return e
可选优化建议
- 流式解压降内存:如果单文件体积较大,可以直接用
response.raw类文件对象传入解压,不需要把整个压缩包加载到内存,性能更高:
def fetch_file(url,filename): try: html = requests.get(url, stream=True, allow_redirects=True) html.raise_for_status() with gzip.GzipFile(fileobj=html.raw) as gz: with open(f'test/{filename}.json','wb') as out: out.write(gz.read()) return html.status_code except Exception as e: return e
- 调整并发数:当前线程池最大工作数为10,若带宽充足且目标接口无严格频率限制,可将
max_workers调整为20~50,进一步提升62000个文件的处理效率。 - 路径兼容处理:使用
os.path.join拼接文件路径,避免不同操作系统的路径分隔符不兼容问题。 - 异常分类处理:可根据需要拆分请求异常、解压异常、文件写入异常的捕获逻辑,方便排查处理失败的任务。
内容的提问来源于stack exchange,提问作者Niels
相关产品推荐
相关产品推荐

