Python批量下载Zip文件并提取内容的代码优化咨询
优化建议
针对你处理13000个Zip文件的IO密集型任务,以下是具体的代码优化方向及修改示例:
1. 复用Requests Session,减少连接开销
当前每个任务都创建新的requests.Session,会重复建立TCP连接池,浪费资源。用线程本地存储实现每个线程只初始化一次Session:
import threading from requests.adapters import HTTPAdapter, Retry import requests # 线程本地存储,每个线程维护一个独立Session thread_local = threading.local() def get_session(): if not hasattr(thread_local, "session"): retry_strategy = Retry( total=3, backoff_factor=1, status_forcelist=[429, 500, 502, 503, 504] ) adapter = HTTPAdapter(max_retries=retry_strategy) session = requests.Session() session.mount('https://', adapter) session.mount('http://', adapter) thread_local.session = session return thread_local.session def download_code_helper(url): session = get_session() # 后续逻辑复用该session
2. 调整线程池大小,避免过度并发
默认ThreadPoolExecutor的max_workers是os.cpu_count() * 5,针对大规模任务,建议根据目标网站的并发限制调整(比如设为30-50),避免触发限流:
with ThreadPoolExecutor(max_workers=40) as executor: # 提交任务逻辑
3. 修复HTTP状态码判断逻辑
当前代码对成功状态码的判断完全错误,直接用response.raise_for_status()自动处理HTTP错误:
try: response = session.get(url, timeout=10) # 添加超时避免阻塞 response.raise_for_status() # 自动抛出4xx/5xx类型错误 zfile = zipfile.ZipFile(io.BytesIO(response.content)) # 后续解压逻辑 except requests.exceptions.HTTPError as e: return (url, f"HTTP错误: {str(e)}") except requests.exceptions.ConnectionError: return (url, "达到最大重试次数,无法获取代码")
4. 用Pathlib优化路径处理
替换字符串拼接的路径方式,避免跨平台问题和路径错误:
from pathlib import Path download_dir = Path('tools/zytools/downloads') download_dir.mkdir(parents=True, exist_ok=True) # 确保目录存在 # 生成文件路径 file_name = url.split('/')[-1].removesuffix('.zip') path = download_dir / f"{file_name}.cpp"
5. 优化Zip文件读取的资源管理
读取Zip内文件时用with语句自动关闭句柄,避免资源泄漏,同时添加编码错误处理:
with zfile.open(filenames[0], 'r') as f: content = f.read() result = content.decode('utf-8', errors='replace') # 替换无效编码字符,避免崩溃
6. 增强缓存的健壮性
检查缓存文件时,若文件为空或损坏,自动删除并重下载:
if path.exists(): with open(path, 'r') as file: result = file.read() if result.strip(): # 文件非空则返回 return (url, result) else: # 空文件则删除,触发重新下载 path.unlink() # 执行下载逻辑
7. 升级为异步IO(推荐Web服务器环境)
线程池在高并发下有切换开销,用asyncio+aiohttp+aiofiles实现异步下载/写入,效率更高:
import asyncio import aiohttp import aiofiles from zipfile import ZipFile import io from pathlib import Path async def async_download_helper(url, download_dir): file_name = url.split('/')[-1].removesuffix('.zip') path = download_dir / f"{file_name}.cpp" if path.exists(): async with aiofiles.open(path, 'r') as f: result = await f.read() if result.strip(): return (url, result) else: path.unlink() timeout = aiohttp.ClientTimeout(total=30) async with aiohttp.ClientSession(timeout=timeout) as session: for attempt in range(3): try: async with session.get(url) as response: response.raise_for_status() content = await response.read() zfile = ZipFile(io.BytesIO(content)) with zfile.open(zfile.namelist()[0], 'r') as f: cpp_content = f.read().decode('utf-8', errors='replace') async with aiofiles.open(path, 'w') as f: await f.write(cpp_content) return (url, cpp_content) except (aiohttp.ClientError, Exception) as e: if attempt == 2: return (url, f"下载失败: {str(e)}") await asyncio.sleep(1 * (2 ** attempt)) # 指数退避重试 async def async_download_code(logfile): urls = logfile.zip_location.to_list() download_dir = Path('tools/zytools/downloads') download_dir.mkdir(parents=True, exist_ok=True) tasks = [async_download_helper(url, download_dir) for url in urls] student_code = await asyncio.gather(*tasks) df = pd.DataFrame(student_code, columns=['zip_location', 'student_code']) return pd.merge(left=logfile, right=df, on=['zip_location']) # 调用方式 # logfile = asyncio.run(async_download_code(logfile))
8. 添加进度跟踪
用tqdm库添加进度条,方便实时跟踪任务进度:
from tqdm import tqdm with ThreadPoolExecutor(max_workers=40) as executor: futures = [executor.submit(download_code_helper, url) for url in urls] student_code = [] for future in tqdm(as_completed(futures), total=len(futures)): student_code.append(future.result())
9. 避免文件名冲突
如果不同URL对应的Zip文件名相同,会覆盖缓存文件。可以用URL的哈希值作为文件名前缀:
import hashlib url_hash = hashlib.md5(url.encode()).hexdigest()[:8] file_name = f"{url_hash}_{url.split('/')[-1].removesuffix('.zip')}" path = download_dir / f"{file_name}.cpp"
内容的提问来源于stack exchange,提问作者Tetraxcode
相关产品推荐
相关产品推荐

