如何优化Paramiko批量小文件SSH上传速度?
提升Paramiko批量小文件SFTP上传速度的方案
针对你用Paramiko上传1200个小文件(总大小约40MB)耗时近17秒且阻塞应用的问题,可从以下几个方向优化:
1. 压缩打包后上传再解压
小文件的SSH/SFTP传输核心开销在连接握手、文件元数据交互上,将整个文件夹打包成单个压缩文件能彻底减少这类重复开销,是提升效率最明显的方案。
实现示例:
import zipfile import os import paramiko def zip_folder(source_folder, zip_path): with zipfile.ZipFile(zip_path, 'w', zipfile.ZIP_DEFLATED) as zipf: for root, dirs, files in os.walk(source_folder): for file in files: file_path = os.path.join(root, file) # 保留相对路径,保证解压后结构一致 arcname = os.path.relpath(file_path, source_folder) zipf.write(file_path, arcname) def FullBackupSSH(self): destination_folder = "/path/to/destination" temp_zip = "/tmp/backup_temp.zip" # 先打包本地文件夹 zip_folder(source_folder, temp_zip) try: transport = paramiko.Transport((self.ssh_hostname, int(self.ssh_port))) transport.connect(username=self.ssh_username, password=self.ssh_password) # 上传单个压缩包 sftp = paramiko.SFTPClient.from_transport(transport) remote_zip_path = os.path.join(destination_folder, "backup.zip") sftp.put(temp_zip, remote_zip_path) sftp.close() # 通过SSH命令在服务器端解压并清理压缩包 ssh = paramiko.SSHClient() ssh._transport = transport ssh.exec_command(f"cd {destination_folder} && unzip -o backup.zip && rm backup.zip") ssh.close() # 清理本地临时压缩包 os.remove(temp_zip) except Exception as error: print(error) if os.path.exists(temp_zip): os.remove(temp_zip)
2. 多线程并行上传
用单线程逐个上传小文件会浪费大量等待时间,通过线程池并行上传多个文件,能充分利用网络带宽。注意线程数不要过多(建议10-20个),避免服务器连接过载。
实现示例:
import os import paramiko from concurrent.futures import ThreadPoolExecutor class SFTP_Operations(paramiko.SFTPClient): def mkdir(self, path, mode=511, ignore_existing=False): try: super().mkdir(path, mode) except IOError: if not ignore_existing: raise def _upload_single_file(self, local_path, remote_path): """单个文件上传的原子函数,供线程调用""" try: self.put(local_path, remote_path) except Exception as e: print(f"上传失败 {local_path}: {str(e)}") def put_dir_parallel(self, source, target, max_workers=15): # 先遍历所有文件,收集本地-远程路径对,同时创建远程目录结构 file_tasks = [] for root, dirs, files in os.walk(source): rel_root = os.path.relpath(root, source) remote_dir = os.path.join(target, rel_root) self.mkdir(remote_dir, ignore_existing=True) for file in files: local_file = os.path.join(root, file) remote_file = os.path.join(remote_dir, file) file_tasks.append((local_file, remote_file)) # 线程池并行执行上传 with ThreadPoolExecutor(max_workers=max_workers) as executor: for local, remote in file_tasks: executor.submit(self._upload_single_file, local, remote) # 使用并行上传的备份函数 def FullBackupSSH(self): destination_folder = "/path/to/destination" try: transport = paramiko.Transport((self.ssh_hostname, int(self.ssh_port))) transport.connect(username=self.ssh_username, password=self.ssh_password) sftp = SFTP_Operations.from_transport(transport) sftp.mkdir(destination_folder, ignore_existing=True) # 启动并行上传 sftp.put_dir_parallel(source_folder, destination_folder) sftp.close() except Exception as error: print(error)
3. 调整传输缓冲区大小
Paramiko默认的窗口缓冲区较小,增大缓冲区能减少数据传输的交互次数,提升整体传输效率。
实现方式:
在创建Transport对象后添加窗口大小设置:
transport = paramiko.Transport((self.ssh_hostname, int(self.ssh_port))) # 设置窗口大小为10MB(单位:字节),可根据网络情况调整 transport.set_window_size(10 * 1024 * 1024) transport.connect(username=self.ssh_username, password=self.ssh_password)
4. 避免应用阻塞:异步执行上传
如果你的应用是GUI或需要持续响应的服务,把上传逻辑放到单独线程/进程中,避免主线程被阻塞。
线程示例:
import threading # 触发备份的入口函数,不阻塞主线程 def trigger_backup(self): backup_thread = threading.Thread(target=self.FullBackupSSH) backup_thread.daemon = True backup_thread.start() # 原FullBackupSSH方法保持不变,作为线程执行的目标函数
内容的提问来源于stack exchange,提问作者Johnson
相关产品推荐
相关产品推荐

