You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python批量下载Zip文件并提取内容的代码优化咨询

优化建议

针对你处理13000个Zip文件的IO密集型任务,以下是具体的代码优化方向及修改示例:

1. 复用Requests Session,减少连接开销

当前每个任务都创建新的requests.Session,会重复建立TCP连接池,浪费资源。用线程本地存储实现每个线程只初始化一次Session:

import threading
from requests.adapters import HTTPAdapter, Retry
import requests

# 线程本地存储,每个线程维护一个独立Session
thread_local = threading.local()

def get_session():
    if not hasattr(thread_local, "session"):
        retry_strategy = Retry(
            total=3,
            backoff_factor=1,
            status_forcelist=[429, 500, 502, 503, 504]
        )
        adapter = HTTPAdapter(max_retries=retry_strategy)
        session = requests.Session()
        session.mount('https://', adapter)
        session.mount('http://', adapter)
        thread_local.session = session
    return thread_local.session

def download_code_helper(url):
    session = get_session()
    # 后续逻辑复用该session

2. 调整线程池大小,避免过度并发

默认ThreadPoolExecutor的max_workers是os.cpu_count() * 5,针对大规模任务,建议根据目标网站的并发限制调整(比如设为30-50),避免触发限流:

with ThreadPoolExecutor(max_workers=40) as executor:
    # 提交任务逻辑

3. 修复HTTP状态码判断逻辑

当前代码对成功状态码的判断完全错误,直接用response.raise_for_status()自动处理HTTP错误:

try:
    response = session.get(url, timeout=10)  # 添加超时避免阻塞
    response.raise_for_status()  # 自动抛出4xx/5xx类型错误
    zfile = zipfile.ZipFile(io.BytesIO(response.content))
    # 后续解压逻辑
except requests.exceptions.HTTPError as e:
    return (url, f"HTTP错误: {str(e)}")
except requests.exceptions.ConnectionError:
    return (url, "达到最大重试次数,无法获取代码")

4. 用Pathlib优化路径处理

替换字符串拼接的路径方式,避免跨平台问题和路径错误:

from pathlib import Path

download_dir = Path('tools/zytools/downloads')
download_dir.mkdir(parents=True, exist_ok=True)  # 确保目录存在

# 生成文件路径
file_name = url.split('/')[-1].removesuffix('.zip')
path = download_dir / f"{file_name}.cpp"

5. 优化Zip文件读取的资源管理

读取Zip内文件时用with语句自动关闭句柄,避免资源泄漏,同时添加编码错误处理:

with zfile.open(filenames[0], 'r') as f:
    content = f.read()
result = content.decode('utf-8', errors='replace')  # 替换无效编码字符,避免崩溃

6. 增强缓存的健壮性

检查缓存文件时,若文件为空或损坏,自动删除并重下载:

if path.exists():
    with open(path, 'r') as file:
        result = file.read()
    if result.strip():  # 文件非空则返回
        return (url, result)
    else:  # 空文件则删除,触发重新下载
        path.unlink()
# 执行下载逻辑

7. 升级为异步IO(推荐Web服务器环境)

线程池在高并发下有切换开销,用asyncio+aiohttp+aiofiles实现异步下载/写入,效率更高:

import asyncio
import aiohttp
import aiofiles
from zipfile import ZipFile
import io
from pathlib import Path

async def async_download_helper(url, download_dir):
    file_name = url.split('/')[-1].removesuffix('.zip')
    path = download_dir / f"{file_name}.cpp"
    
    if path.exists():
        async with aiofiles.open(path, 'r') as f:
            result = await f.read()
        if result.strip():
            return (url, result)
        else:
            path.unlink()
    
    timeout = aiohttp.ClientTimeout(total=30)
    async with aiohttp.ClientSession(timeout=timeout) as session:
        for attempt in range(3):
            try:
                async with session.get(url) as response:
                    response.raise_for_status()
                    content = await response.read()
                    zfile = ZipFile(io.BytesIO(content))
                    with zfile.open(zfile.namelist()[0], 'r') as f:
                        cpp_content = f.read().decode('utf-8', errors='replace')
                    async with aiofiles.open(path, 'w') as f:
                        await f.write(cpp_content)
                    return (url, cpp_content)
            except (aiohttp.ClientError, Exception) as e:
                if attempt == 2:
                    return (url, f"下载失败: {str(e)}")
                await asyncio.sleep(1 * (2 ** attempt))  # 指数退避重试

async def async_download_code(logfile):
    urls = logfile.zip_location.to_list()
    download_dir = Path('tools/zytools/downloads')
    download_dir.mkdir(parents=True, exist_ok=True)
    
    tasks = [async_download_helper(url, download_dir) for url in urls]
    student_code = await asyncio.gather(*tasks)
    
    df = pd.DataFrame(student_code, columns=['zip_location', 'student_code'])
    return pd.merge(left=logfile, right=df, on=['zip_location'])

# 调用方式
# logfile = asyncio.run(async_download_code(logfile))

8. 添加进度跟踪

用tqdm库添加进度条,方便实时跟踪任务进度:

from tqdm import tqdm

with ThreadPoolExecutor(max_workers=40) as executor:
    futures = [executor.submit(download_code_helper, url) for url in urls]
    student_code = []
    for future in tqdm(as_completed(futures), total=len(futures)):
        student_code.append(future.result())

9. 避免文件名冲突

如果不同URL对应的Zip文件名相同,会覆盖缓存文件。可以用URL的哈希值作为文件名前缀:

import hashlib

url_hash = hashlib.md5(url.encode()).hexdigest()[:8]
file_name = f"{url_hash}_{url.split('/')[-1].removesuffix('.zip')}"
path = download_dir / f"{file_name}.cpp"

内容的提问来源于stack exchange,提问作者Tetraxcode

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 03:55:29