如何基于Paramiko实现SFTP文件分批处理(指定每批数量)?
Hey there! Let's get that batch processing logic added to your existing SFTP workflow. You've already got the core file migration and metadata parsing working, so we just need to adjust things to handle files in batches—whether that's 100 at a time or 5 per batch. Here's a step-by-step solution:
First: Refactor to Collect All Files First
Your current code processes files as it finds them. To add batching, we'll first collect all the files we need to transfer into a single list, then process that list in chunks. This makes it easy to control batch size and track progress.
Option 1: Process Files in Batches of 100
This is ideal if you want to handle larger chunks at once, reducing the overhead of frequent small batches.
Modified Code with 100-File Batch Logic
import os from stat import S_ISDIR import paramiko import tika from tika import parser tika.initVM() paramiko.util.log_to_file('D:/Paramiko/logs/paramiko.log') ssh = paramiko.SSHClient() ssh.set_missing_host_key_policy(paramiko.AutoAddPolicy()) ssh.connect('xxxx', username='ubuntu', password='', key_filename='D:/Paramiko/test.pem') sftp = ssh.open_sftp() depth = 2 # Get list of target directories stdin, stdout, stderr = ssh.exec_command(f"find . -mindepth {depth} -maxdepth {depth} -type d ") dlist = list(stdout.readlines()) dir_count = len(dlist) # Renamed from 'len' to avoid conflict with built-in function def getfiles_from_directory(subfolder, filename, sftp, local_p): remote_path = f"{subfolder}/{filename}" local_path = os.path.join(local_p, filename) sftp.get(remote_path, local_path) try: parsed = parser.from_file(local_path) print(f'MetaData: {str(parsed["metadata"])}') except KeyError as k: print(f"KeyError encountered: {k}") # Step 1: Collect all files we need to process first all_files_to_process = [] folders_scanned = [] for direc in dlist: direcname = direc.replace('./', '').strip() if not direcname.startswith('.'): local_path = 'D:/Paramiko/Test' # Fixed command to properly change directory and get path stdin, stdout, stderr = ssh.exec_command(f"cd {direcname} && pwd") subfolder = stdout.readlines()[0].strip() local_p = os.path.join(local_path, direcname) # Create local folder if it doesn't exist if not os.path.exists(local_p): os.mkdir(local_p) sftp.chdir(subfolder) children = sftp.listdir() for content in children: content_sub = content.strip() # Check if it's a file (not a folder) if '.' in content_sub: print(f'Found a file: {content_sub}') # Store file info for later batch processing all_files_to_process.append( (subfolder, content_sub, local_p) ) else: print(f'Found a folder: {content_sub}') folders_scanned.append(subfolder) # Step 2: Process files in batches of 100 BATCH_SIZE = 100 files_copied = [] for i in range(0, len(all_files_to_process), BATCH_SIZE): batch = all_files_to_process[i:i+BATCH_SIZE] start_num = i + 1 end_num = min(i + BATCH_SIZE, len(all_files_to_process)) print(f"Processing batch {i//BATCH_SIZE + 1} (files {start_num} to {end_num})") for file_info in batch: subfolder, filename, local_p = file_info getfiles_from_directory(subfolder, filename, sftp, local_p) files_copied.append(f"{subfolder}/{filename}") print(f"Completed batch {i//BATCH_SIZE + 1}\n") # Clean up connections sftp.close() ssh.close() print(f'Folders Scanned: {str(folders_scanned)}') print(f'Files Copied: {str(files_copied)}')
Key Changes:
- Collected all files first: We build a list
all_files_to_processwith tuples of (remote subfolder, filename, local path) for every file to transfer—no immediate processing. - Batch loop with slicing: Used Python's list slicing (
all_files_to_process[i:i+BATCH_SIZE]) to split the list into manageable chunks. - Fixed minor bugs: Renamed
lentodir_countto avoid overriding the built-in function, fixed thecdcommand (added&&to runpwdafter changing directories), and simplified file/folder detection logic.
Option 2: Process Files in Batches of 5
For smaller, more frequent batches (e.g., to monitor progress closely or reduce resource load), just update the BATCH_SIZE variable:
# Change this line to use 5 files per batch BATCH_SIZE = 5
Bonus: Add a Delay Between Batches (Optional)
If you want to avoid overwhelming the server, add a pause between batches using the time module:
import time # Inside the batch loop, after completing a batch print(f"Completed batch {i//BATCH_SIZE + 1}\n") time.sleep(3) # Pause for 3 seconds before next batch
Quick Note on Your Original Code
I noticed your original code had a if subfolder not in folders_copied: check that would only process the first file in each subfolder. I removed that in the modified code to ensure all files in a subfolder are transferred—let me know if you intended to only process one file per subfolder, and I can adjust that back!
内容的提问来源于stack exchange,提问作者Lasit Pant

