如何优化Python读取FTP多子文件夹文件名的代码以提升执行速度
Absolutely, your loop is tied to the slowdown—but not the loop structure itself. The real culprit is the repeated FTP login and connection setup inside every iteration. Establishing a TLS-encrypted FTP connection involves handshake, authentication, and session setup overhead, which adds up exponentially when you repeat it for hundreds or thousands of directories.
Here are two actionable optimization levels to fix this:
1. Basic Optimization: Reuse a Single FTP Connection (Highest Impact)
The simplest and most effective fix is to log in once, then reuse that connection for all directory listings. This eliminates connection overhead entirely.
import ftplib ftp_subdir_list = ['example_folder/2021/01/01', 'example_folder/2021/01/02', ..., 'example_folder/2021/08/09'] ftp_file_list = [] # Login ONCE before processing directories ftp = ftplib.FTP_TLS(host) ftp.login(username, password) try: for dir_path in ftp_subdir_list: try: ftp.cwd(dir_path) # Fetch file list for the current directory file_list = ftp.nlst() ftp_file_list.append(file_list) # Optional: Navigate back to parent if using relative paths, skip if paths are absolute # ftp.cwd('..') except ftplib.error_perm as e: print(f"Failed to access directory {dir_path}: {e}") ftp_file_list.append([]) # Handle error gracefully finally: # Always close the connection when done ftp.quit() print(ftp_file_list)
Key Details:
- Connection reuse cuts out TLS handshake and login steps for every directory—this will give you an immediate, massive speed boost.
- Added
try/exceptblocks to handle missing directories or permission issues, preventing the entire script from crashing. - If your directory paths are absolute, you don’t need
ftp.cwd('..')—ftp.cwd(dir_path)will jump directly to the target each time.
2. Advanced Optimization: Parallelize Directory Listings (For Very Large Datasets)
If you have thousands of directories and still need more speed, use parallel processing (threads) to fetch multiple listings at once. Note that FTP connections aren’t thread-safe, so each thread needs its own dedicated connection.
Here’s a thread-based example using concurrent.futures:
import ftplib from concurrent.futures import ThreadPoolExecutor ftp_subdir_list = ['example_folder/2021/01/01', 'example_folder/2021/01/02', ..., 'example_folder/2021/08/09'] def get_files_for_dir(dir_path): """Helper to fetch files for a single directory with its own FTP connection""" ftp = ftplib.FTP_TLS(host) ftp.login(username, password) try: ftp.cwd(dir_path) return ftp.nlst() except ftplib.error_perm as e: print(f"Error accessing {dir_path}: {e}") return [] finally: ftp.quit() # Parallelize with a thread pool with ThreadPoolExecutor(max_workers=10) as executor: ftp_file_list = list(executor.map(get_files_for_dir, ftp_subdir_list)) print(ftp_file_list)
Key Details:
- Adjust
max_workersbased on your FTP server’s concurrent connection limits (start with 5-10—too many threads might get you blocked). - This adds minor thread management overhead, but it’s worth it for large-scale directory processing.
Why Your Original Code Was Slow
Each loop iteration forced your script to:
- Establish a new TLS connection (handshake + encryption setup)
- Re-authenticate with the server
- Navigate to the target directory
- Fetch the file list
- Implicitly close the connection
Multiply those steps by hundreds of directories, and you’re wasting minutes on unnecessary connection setup alone.
内容的提问来源于stack exchange,提问作者wawawa

