You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用pysftp下载超500MB大文件遇SSHException连接断开求助

解决pysftp大文件下载中断的问题

我之前处理过类似的大文件SFTP下载中断问题,大概率是服务器端的连接超时或者传输限制导致的——毕竟单次传几百MB很容易触发这类防护机制。咱们试试下面这几个方案,应该能解决你的问题:

方案1:分块流式下载(最推荐)

pysftp默认的get方法是一次性传输整个文件,大文件很容易被服务器断开连接。我们可以用底层的paramiko SFTP客户端实现分块流式下载,每次只传输一小段数据,降低被断开的概率。

import pysftp

myHostname = "dbfiles.xyz.org"
myUsername = "XXXX"
myPassword = "YYYY"
cnopts = pysftp.CnOpts()
cnopts.hostkeys = None

# 分块大小设置为1MB(可根据服务器情况调整,比如改成512KB)
CHUNK_SIZE = 1024 * 1024

with pysftp.Connection(host=myHostname, username=myUsername, password=myPassword, cnopts=cnopts) as sftp:
    print("Connection successfully established ... ")
    localFilePath = 'c:/..../'  # 注意末尾加斜杠
    remoteFolderPath = '/folder/'
    remoteFiles = sftp.listdir(remoteFolderPath)
    
    for filename in remoteFiles:
        if 'string_to_match' in filename:
            local_path = localFilePath + filename
            remote_path = remoteFolderPath + filename
            print(f"Starting download: {filename}")
            
            # 用底层SFTP客户端打开文件,流式读写
            with sftp.sftp_client.open(remote_path, 'rb') as remote_file, open(local_path, 'wb') as local_file:
                while True:
                    chunk = remote_file.read(CHUNK_SIZE)
                    if not chunk:
                        break
                    local_file.write(chunk)
                    # 可选:如果还是断,加个微小延迟避免触发速率限制
                    # import time
                    # time.sleep(0.01)
            
            print(f"Download completed: {filename}")

这个方法的优势是:

  • 不会把整个大文件加载到内存,内存占用极低
  • 分块传输不容易触发服务器的连接超时限制
  • 可以灵活调整分块大小适配服务器

方案2:添加心跳包+延长超时时间

有些服务器会断开长时间没有交互的连接,我们可以给SFTP连接添加心跳包,同时延长超时时间,避免连接被主动断开。

在你的连接代码里添加这两个设置:

with pysftp.Connection(
    host=myHostname, 
    username=myUsername, 
    password=myPassword,
    cnopts=cnopts,
    timeout=300  # 设置连接超时为5分钟(根据需要调整)
) as sftp:
    # 每30秒发送一次心跳包,保持连接活跃
    sftp._transport.set_keepalive(30)
    print("Connection successfully established ... ")
    
    # 后续的下载代码(可以配合方案1的分块下载)
    # ...

方案3:断点续传+重试机制

如果偶尔还是会断开,我们可以实现断点续传,下次启动时从上次中断的位置继续下载,不用从头开始。

import os
import pysftp

CHUNK_SIZE = 1024 * 1024

myHostname = "dbfiles.xyz.org"
myUsername = "XXXX"
myPassword = "YYYY"
cnopts = pysftp.CnOpts()
cnopts.hostkeys = None

with pysftp.Connection(host=myHostname, username=myUsername, password=myPassword, cnopts=cnopts) as sftp:
    print("Connection successfully established ... ")
    localFilePath = 'c:/..../'
    remoteFolderPath = '/folder/'
    remoteFiles = sftp.listdir(remoteFolderPath)
    
    for filename in remoteFiles:
        if 'string_to_match' in filename:
            local_path = localFilePath + filename
            remote_path = remoteFolderPath + filename
            
            # 获取远程文件总大小
            remote_file_size = sftp.stat(remote_path).st_size
            # 获取本地已下载的大小(如果文件存在)
            local_file_size = os.path.getsize(local_path) if os.path.exists(local_path) else 0
            
            if local_file_size >= remote_file_size:
                print(f"File {filename} is already fully downloaded.")
                continue
            
            print(f"Resuming download: {filename} from {local_file_size}/{remote_file_size} bytes")
            
            # 从断点位置开始读写
            with sftp.sftp_client.open(remote_path, 'rb') as remote_file, open(local_path, 'ab') as local_file:
                remote_file.seek(local_file_size)  # 移动到已下载的位置
                while True:
                    chunk = remote_file.read(CHUNK_SIZE)
                    if not chunk:
                        break
                    local_file.write(chunk)
            
            print(f"Download completed: {filename}")

额外建议

如果以上方法还是不行,可以尝试:

  • 调小CHUNK_SIZE(比如改成256KB),减少单次传输的数据量
  • 检查服务器端是否有明确的传输限制(比如单文件大小、连接时长),如果有对应调整策略

内容的提问来源于stack exchange,提问作者Hari_pb

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 12:17:35