You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于Paramiko实现SFTP文件分批处理(指定每批数量)?

Adding Batch Processing to Your SFTP File Transfer Code

Hey there! Let's get that batch processing logic added to your existing SFTP workflow. You've already got the core file migration and metadata parsing working, so we just need to adjust things to handle files in batches—whether that's 100 at a time or 5 per batch. Here's a step-by-step solution:

First: Refactor to Collect All Files First

Your current code processes files as it finds them. To add batching, we'll first collect all the files we need to transfer into a single list, then process that list in chunks. This makes it easy to control batch size and track progress.

Option 1: Process Files in Batches of 100

This is ideal if you want to handle larger chunks at once, reducing the overhead of frequent small batches.

Modified Code with 100-File Batch Logic

import os
from stat import S_ISDIR
import paramiko
import tika
from tika import parser

tika.initVM()
paramiko.util.log_to_file('D:/Paramiko/logs/paramiko.log')
ssh = paramiko.SSHClient()
ssh.set_missing_host_key_policy(paramiko.AutoAddPolicy())
ssh.connect('xxxx', username='ubuntu', password='', key_filename='D:/Paramiko/test.pem')
sftp = ssh.open_sftp()
depth = 2

# Get list of target directories
stdin, stdout, stderr = ssh.exec_command(f"find . -mindepth {depth} -maxdepth {depth} -type d ")
dlist = list(stdout.readlines())
dir_count = len(dlist)  # Renamed from 'len' to avoid conflict with built-in function

def getfiles_from_directory(subfolder, filename, sftp, local_p):
    remote_path = f"{subfolder}/{filename}"
    local_path = os.path.join(local_p, filename)
    sftp.get(remote_path, local_path)
    try:
        parsed = parser.from_file(local_path)
        print(f'MetaData: {str(parsed["metadata"])}')
    except KeyError as k:
        print(f"KeyError encountered: {k}")

# Step 1: Collect all files we need to process first
all_files_to_process = []
folders_scanned = []

for direc in dlist:
    direcname = direc.replace('./', '').strip()
    if not direcname.startswith('.'):
        local_path = 'D:/Paramiko/Test'
        # Fixed command to properly change directory and get path
        stdin, stdout, stderr = ssh.exec_command(f"cd {direcname} && pwd")
        subfolder = stdout.readlines()[0].strip()
        local_p = os.path.join(local_path, direcname)
        
        # Create local folder if it doesn't exist
        if not os.path.exists(local_p):
            os.mkdir(local_p)
            
        sftp.chdir(subfolder)
        children = sftp.listdir()
        
        for content in children:
            content_sub = content.strip()
            # Check if it's a file (not a folder)
            if '.' in content_sub:
                print(f'Found a file: {content_sub}')
                # Store file info for later batch processing
                all_files_to_process.append( (subfolder, content_sub, local_p) )
            else:
                print(f'Found a folder: {content_sub}')
                folders_scanned.append(subfolder)

# Step 2: Process files in batches of 100
BATCH_SIZE = 100
files_copied = []

for i in range(0, len(all_files_to_process), BATCH_SIZE):
    batch = all_files_to_process[i:i+BATCH_SIZE]
    start_num = i + 1
    end_num = min(i + BATCH_SIZE, len(all_files_to_process))
    print(f"Processing batch {i//BATCH_SIZE + 1} (files {start_num} to {end_num})")
    
    for file_info in batch:
        subfolder, filename, local_p = file_info
        getfiles_from_directory(subfolder, filename, sftp, local_p)
        files_copied.append(f"{subfolder}/{filename}")
    
    print(f"Completed batch {i//BATCH_SIZE + 1}\n")

# Clean up connections
sftp.close()
ssh.close()

print(f'Folders Scanned: {str(folders_scanned)}')
print(f'Files Copied: {str(files_copied)}')

Key Changes:

  • Collected all files first: We build a list all_files_to_process with tuples of (remote subfolder, filename, local path) for every file to transfer—no immediate processing.
  • Batch loop with slicing: Used Python's list slicing (all_files_to_process[i:i+BATCH_SIZE]) to split the list into manageable chunks.
  • Fixed minor bugs: Renamed len to dir_count to avoid overriding the built-in function, fixed the cd command (added && to run pwd after changing directories), and simplified file/folder detection logic.

Option 2: Process Files in Batches of 5

For smaller, more frequent batches (e.g., to monitor progress closely or reduce resource load), just update the BATCH_SIZE variable:

# Change this line to use 5 files per batch
BATCH_SIZE = 5

Bonus: Add a Delay Between Batches (Optional)

If you want to avoid overwhelming the server, add a pause between batches using the time module:

import time

# Inside the batch loop, after completing a batch
print(f"Completed batch {i//BATCH_SIZE + 1}\n")
time.sleep(3)  # Pause for 3 seconds before next batch

Quick Note on Your Original Code

I noticed your original code had a if subfolder not in folders_copied: check that would only process the first file in each subfolder. I removed that in the modified code to ensure all files in a subfolder are transferred—let me know if you intended to only process one file per subfolder, and I can adjust that back!

内容的提问来源于stack exchange,提问作者Lasit Pant

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 04:08:43