多线程下载文件/视频时文件体积异常增大的问题求助
多线程下载文件体积过大问题排查与修复
问题原因
- 服务器分片支持未验证:未检查服务器是否支持
Range分片请求(通过响应头Accept-Ranges判断),若服务器不支持,每个线程都会下载完整文件,多次写入导致文件体积变为原文件的N倍(N为线程数) - 文件写入逻辑错误:使用
ab追加模式写入,当分片请求失效时,所有线程的完整文件内容都会被追加到同一文件中,直接导致体积膨胀 - 进度条线程安全问题:多线程共享
tqdm进度条,可能出现重复更新,导致进度显示异常(虽不直接影响文件体积,但属于潜在问题)
修复后的代码
import os import requests from tqdm import tqdm import threading from requests.exceptions import RequestException def download_chunk(url, save_path, start_byte, end_byte, progress_bar, lock): headers = {'Range': f'bytes={start_byte}-{end_byte}'} try: response = requests.get(url, headers=headers, stream=True, allow_redirects=True) response.raise_for_status() # 检查请求是否成功 with open(save_path, 'rb+') as file: file.seek(start_byte) for data in response.iter_content(chunk_size=4096): file.write(data) # 使用锁保证进度条更新线程安全 with lock: progress_bar.update(len(data)) except RequestException as e: print(f"Chunk download failed: {e}") raise def download_with_progress(url, save_folder, filename, num_threads=4): os.makedirs(save_folder, exist_ok=True) save_path = os.path.join(save_folder, filename) # 先验证服务器是否支持分片下载 try: head_response = requests.head(url, allow_redirects=True) head_response.raise_for_status() except RequestException as e: print(f"Head request failed: {e}") return file_size = int(head_response.headers.get('content-length', 0)) accept_ranges = head_response.headers.get('Accept-Ranges') if accept_ranges != 'bytes' or file_size == 0: # 不支持分片或无法获取文件大小,退化为单线程下载 print("Server doesn't support range requests or file size unknown, using single thread download") progress_bar = tqdm(total=file_size, unit='B', unit_scale=True) try: response = requests.get(url, stream=True, allow_redirects=True) response.raise_for_status() with open(save_path, 'wb') as file: for data in response.iter_content(chunk_size=4096): file.write(data) progress_bar.update(len(data)) progress_bar.close() print(f"Downloaded {filename}") except RequestException as e: print(f"Download failed: {e}") return # 支持分片下载,启动多线程 block_size = file_size // num_threads progress_bar = tqdm(total=file_size, unit='B', unit_scale=True) lock = threading.Lock() # 用于进度条线程安全更新 # 预创建空文件并设置大小 with open(save_path, 'wb') as file: file.truncate(file_size) threads = [] for i in range(num_threads): start_byte = i * block_size end_byte = start_byte + block_size - 1 if i < num_threads - 1 else file_size - 1 print(f"Thread {i}: Downloading bytes {start_byte}-{end_byte}") thread = threading.Thread( target=download_chunk, args=(url, save_path, start_byte, end_byte, progress_bar, lock) ) threads.append(thread) thread.start() # 等待所有线程完成 for thread in threads: thread.join() progress_bar.close() # 验证最终文件大小是否正确 final_size = os.path.getsize(save_path) if final_size == file_size: print(f"Downloaded {filename} successfully, size matches original") else: print(f"Warning: Downloaded file size {final_size} doesn't match original {file_size}") # Example usage: download_with_progress("https://example.com/download.mp4", "GoogleColab", "hellyrev.mp4", num_threads=20)
关键修复点
- 增加分片支持验证:通过
head请求检查Accept-Ranges头,不支持则自动切换单线程 - 修改文件写入方式:使用
rb+模式并通过seek()定位到分片起始位置写入,避免追加导致的重复内容 - 线程安全的进度条:添加线程锁
Lock(),保证多线程更新进度条时不会出现混乱 - 增加错误处理:捕获请求异常,避免程序崩溃
- 文件大小验证:下载完成后对比最终文件大小与原文件大小,确保下载完整
内容的提问来源于stack exchange,提问作者Justahelper
相关产品推荐
相关产品推荐

