使用pysftp下载超500MB大文件遇SSHException连接断开求助
解决pysftp大文件下载中断的问题
我之前处理过类似的大文件SFTP下载中断问题,大概率是服务器端的连接超时或者传输限制导致的——毕竟单次传几百MB很容易触发这类防护机制。咱们试试下面这几个方案,应该能解决你的问题:
方案1:分块流式下载(最推荐)
pysftp默认的get方法是一次性传输整个文件,大文件很容易被服务器断开连接。我们可以用底层的paramiko SFTP客户端实现分块流式下载,每次只传输一小段数据,降低被断开的概率。
import pysftp myHostname = "dbfiles.xyz.org" myUsername = "XXXX" myPassword = "YYYY" cnopts = pysftp.CnOpts() cnopts.hostkeys = None # 分块大小设置为1MB(可根据服务器情况调整,比如改成512KB) CHUNK_SIZE = 1024 * 1024 with pysftp.Connection(host=myHostname, username=myUsername, password=myPassword, cnopts=cnopts) as sftp: print("Connection successfully established ... ") localFilePath = 'c:/..../' # 注意末尾加斜杠 remoteFolderPath = '/folder/' remoteFiles = sftp.listdir(remoteFolderPath) for filename in remoteFiles: if 'string_to_match' in filename: local_path = localFilePath + filename remote_path = remoteFolderPath + filename print(f"Starting download: {filename}") # 用底层SFTP客户端打开文件,流式读写 with sftp.sftp_client.open(remote_path, 'rb') as remote_file, open(local_path, 'wb') as local_file: while True: chunk = remote_file.read(CHUNK_SIZE) if not chunk: break local_file.write(chunk) # 可选:如果还是断,加个微小延迟避免触发速率限制 # import time # time.sleep(0.01) print(f"Download completed: {filename}")
这个方法的优势是:
- 不会把整个大文件加载到内存,内存占用极低
- 分块传输不容易触发服务器的连接超时限制
- 可以灵活调整分块大小适配服务器
方案2:添加心跳包+延长超时时间
有些服务器会断开长时间没有交互的连接,我们可以给SFTP连接添加心跳包,同时延长超时时间,避免连接被主动断开。
在你的连接代码里添加这两个设置:
with pysftp.Connection( host=myHostname, username=myUsername, password=myPassword, cnopts=cnopts, timeout=300 # 设置连接超时为5分钟(根据需要调整) ) as sftp: # 每30秒发送一次心跳包,保持连接活跃 sftp._transport.set_keepalive(30) print("Connection successfully established ... ") # 后续的下载代码(可以配合方案1的分块下载) # ...
方案3:断点续传+重试机制
如果偶尔还是会断开,我们可以实现断点续传,下次启动时从上次中断的位置继续下载,不用从头开始。
import os import pysftp CHUNK_SIZE = 1024 * 1024 myHostname = "dbfiles.xyz.org" myUsername = "XXXX" myPassword = "YYYY" cnopts = pysftp.CnOpts() cnopts.hostkeys = None with pysftp.Connection(host=myHostname, username=myUsername, password=myPassword, cnopts=cnopts) as sftp: print("Connection successfully established ... ") localFilePath = 'c:/..../' remoteFolderPath = '/folder/' remoteFiles = sftp.listdir(remoteFolderPath) for filename in remoteFiles: if 'string_to_match' in filename: local_path = localFilePath + filename remote_path = remoteFolderPath + filename # 获取远程文件总大小 remote_file_size = sftp.stat(remote_path).st_size # 获取本地已下载的大小(如果文件存在) local_file_size = os.path.getsize(local_path) if os.path.exists(local_path) else 0 if local_file_size >= remote_file_size: print(f"File {filename} is already fully downloaded.") continue print(f"Resuming download: {filename} from {local_file_size}/{remote_file_size} bytes") # 从断点位置开始读写 with sftp.sftp_client.open(remote_path, 'rb') as remote_file, open(local_path, 'ab') as local_file: remote_file.seek(local_file_size) # 移动到已下载的位置 while True: chunk = remote_file.read(CHUNK_SIZE) if not chunk: break local_file.write(chunk) print(f"Download completed: {filename}")
额外建议
如果以上方法还是不行,可以尝试:
- 调小
CHUNK_SIZE(比如改成256KB),减少单次传输的数据量 - 检查服务器端是否有明确的传输限制(比如单文件大小、连接时长),如果有对应调整策略
内容的提问来源于stack exchange,提问作者Hari_pb
相关产品推荐
相关产品推荐

