You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何优化Paramiko批量小文件SSH上传速度?

提升Paramiko批量小文件SFTP上传速度的方案

针对你用Paramiko上传1200个小文件(总大小约40MB)耗时近17秒且阻塞应用的问题,可从以下几个方向优化:

1. 压缩打包后上传再解压

小文件的SSH/SFTP传输核心开销在连接握手、文件元数据交互上,将整个文件夹打包成单个压缩文件能彻底减少这类重复开销,是提升效率最明显的方案。

实现示例:

import zipfile
import os
import paramiko

def zip_folder(source_folder, zip_path):
    with zipfile.ZipFile(zip_path, 'w', zipfile.ZIP_DEFLATED) as zipf:
        for root, dirs, files in os.walk(source_folder):
            for file in files:
                file_path = os.path.join(root, file)
                # 保留相对路径,保证解压后结构一致
                arcname = os.path.relpath(file_path, source_folder)
                zipf.write(file_path, arcname)

def FullBackupSSH(self):
    destination_folder = "/path/to/destination"
    temp_zip = "/tmp/backup_temp.zip"
    
    # 先打包本地文件夹
    zip_folder(source_folder, temp_zip)
    
    try:
        transport = paramiko.Transport((self.ssh_hostname, int(self.ssh_port)))
        transport.connect(username=self.ssh_username, password=self.ssh_password)
        
        # 上传单个压缩包
        sftp = paramiko.SFTPClient.from_transport(transport)
        remote_zip_path = os.path.join(destination_folder, "backup.zip")
        sftp.put(temp_zip, remote_zip_path)
        sftp.close()
        
        # 通过SSH命令在服务器端解压并清理压缩包
        ssh = paramiko.SSHClient()
        ssh._transport = transport
        ssh.exec_command(f"cd {destination_folder} && unzip -o backup.zip && rm backup.zip")
        ssh.close()
        
        # 清理本地临时压缩包
        os.remove(temp_zip)
    except Exception as error:
        print(error)
        if os.path.exists(temp_zip):
            os.remove(temp_zip)

2. 多线程并行上传

用单线程逐个上传小文件会浪费大量等待时间,通过线程池并行上传多个文件,能充分利用网络带宽。注意线程数不要过多(建议10-20个),避免服务器连接过载。

实现示例:

import os
import paramiko
from concurrent.futures import ThreadPoolExecutor

class SFTP_Operations(paramiko.SFTPClient):
    def mkdir(self, path, mode=511, ignore_existing=False):
        try:
            super().mkdir(path, mode)
        except IOError:
            if not ignore_existing:
                raise

    def _upload_single_file(self, local_path, remote_path):
        """单个文件上传的原子函数,供线程调用"""
        try:
            self.put(local_path, remote_path)
        except Exception as e:
            print(f"上传失败 {local_path}: {str(e)}")

    def put_dir_parallel(self, source, target, max_workers=15):
        # 先遍历所有文件,收集本地-远程路径对,同时创建远程目录结构
        file_tasks = []
        for root, dirs, files in os.walk(source):
            rel_root = os.path.relpath(root, source)
            remote_dir = os.path.join(target, rel_root)
            self.mkdir(remote_dir, ignore_existing=True)
            
            for file in files:
                local_file = os.path.join(root, file)
                remote_file = os.path.join(remote_dir, file)
                file_tasks.append((local_file, remote_file))
        
        # 线程池并行执行上传
        with ThreadPoolExecutor(max_workers=max_workers) as executor:
            for local, remote in file_tasks:
                executor.submit(self._upload_single_file, local, remote)

# 使用并行上传的备份函数
def FullBackupSSH(self):
    destination_folder = "/path/to/destination"
    try:
        transport = paramiko.Transport((self.ssh_hostname, int(self.ssh_port)))
        transport.connect(username=self.ssh_username, password=self.ssh_password)
        sftp = SFTP_Operations.from_transport(transport)
        sftp.mkdir(destination_folder, ignore_existing=True)
        # 启动并行上传
        sftp.put_dir_parallel(source_folder, destination_folder)
        sftp.close()
    except Exception as error:
        print(error)

3. 调整传输缓冲区大小

Paramiko默认的窗口缓冲区较小,增大缓冲区能减少数据传输的交互次数,提升整体传输效率。

实现方式:

在创建Transport对象后添加窗口大小设置:

transport = paramiko.Transport((self.ssh_hostname, int(self.ssh_port)))
# 设置窗口大小为10MB(单位:字节),可根据网络情况调整
transport.set_window_size(10 * 1024 * 1024)
transport.connect(username=self.ssh_username, password=self.ssh_password)

4. 避免应用阻塞:异步执行上传

如果你的应用是GUI或需要持续响应的服务,把上传逻辑放到单独线程/进程中,避免主线程被阻塞。

线程示例:

import threading

# 触发备份的入口函数,不阻塞主线程
def trigger_backup(self):
    backup_thread = threading.Thread(target=self.FullBackupSSH)
    backup_thread.daemon = True
    backup_thread.start()

# 原FullBackupSSH方法保持不变,作为线程执行的目标函数

内容的提问来源于stack exchange,提问作者Johnson

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 18:55:34