Python requests批量下载局域网CSV文件的GET性能优化咨询
局域网CSV批量下载脚本性能优化咨询
背景说明
- 基于Python3编写脚本,调用
requests模块从局域网(LAN)内部署的Web服务器下载指定数量的CSV文件,本次测试共下载61个文件 - 当前串行执行模式总耗时约2分54秒,单文件平均耗时超过2.85秒
- 所有CSV文件遵循
f'name{index}'命名规则,可通过索引遍历访问 - 核心诉求:确认是否存在可行方案进一步缩短脚本总执行时长
现有实现方案
1. 基础for循环串行方案(耗时2分54秒)
for file in range(0, files_to_process): nome_CSV = f'costantName_{file}.csv' # 若文件已存在则先删除再重新创建 if os.path.exists(nome_CSV): os.remove(nome_CSV) url = f"http://{self.ip}/this_is_the_path/{nome_CSV}" try: # 该Web服务器不支持cookies r = requests.get( url=url, auth=(self.user, self.password), allow_redirects=True, headers=headers ) except Exception as e: print(e) continue # 本地保存文件 open(nome_CSV, "wb").write(r.content)
2. 多线程/多进程并发方案
分别测试Multithreading(多线程)、Multiprocessing(多进程)两种并发模式与优化后串行版本的性能差异,核心代码分为三个部分:
线程/进程调度部分
# 启用8工作线程的线程池 #with ThreadPoolExecutor(max_workers=8) as executor: # 启用8工作进程的进程池 with ProcessPoolExecutor(max_workers=8) as executor: # CSV_names是包含60个文件名的列表,由列表推导式生成 for name in CSV_names: url = "This_is_an_URL" + name futures = [executor.submit(self.get_file, url, name)] done, not_done = wait(futures, return_when=concurrent.futures.ALL_COMPLETED)
请求发送部分
def download_url(self, url: str): headers = {'User-Agent': 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_10_1)' ' AppleWebKit/537.36 (KHTML, like Gecko)' ' Chrome/39.0.2171.95 Safari/537.36'} try: data = requests.get( url=url, auth=(self.user, self.password), allow_redirects=True, headers=headers ) except Exception as e: print(e) return None else: return data.content
文件保存部分
def save_file(file_name: str, data, path): out_path = os.path.join(path, file_name) with open(out_path, 'wb') as file: file.write(data) return out_path
测试环境与性能数据
- 测试硬件配置:Intel Core I5-8265U 1.60GHz CPU、16GB RAM
- 不同并发方案耗时统计(单位:分:秒:毫秒):
| 并发数 | 多进程(Multiprocessing)耗时 | 多线程(Multithreading)耗时 |
|---|---|---|
| 8 | 02:39:122703 | 02:39:305354 |
| 16 | 02:33:569016 | 02:29:909762 |
- 优化后的串行版本稳定耗时为02:44:662881,仅存在毫秒级波动。
目前仅通过使用with关键字、优先使用推导式、提升并发数获得了秒级的性能提升,需要可进一步压缩总执行时长的优化思路。
内容的提问来源于stack exchange,提问作者bersi
相关产品推荐
相关产品推荐

