WebDav多线程下载出现文件描述符(FDS)超限问题,请求定位泄漏点
Let's break down what's happening here and how to fix it. Your core issue is that recursive retry logic is causing delayed resource cleanup, combined with potential gaps in how the easywebdav2 session is being closed. Here's the step-by-step analysis and solution:
Why This Is Happening
Recursive Retry Creates Unreleased Resources
Your_downloadfunction uses recursion to handle retries. When a connection error occurs, you callclean_webdav(webdav)then immediately recurse intoself._download. The problem? The current function'sfinallyblock won't execute until the recursive call completes. Even though you're deleting thewebdavreference, Python's garbage collector might not immediately reclaim the object (and its associated file descriptors/sockets) while the recursive stack is active. With 25 retries, this means you could have up to 25 lingeringwebdavinstances holding open sockets at once—on top of your 24 threads, this easily pushes you over the file descriptor limit.Incomplete Session Cleanup
Callingwebdav.session.close()closes the requests session, but it might not fully clean up all underlying connections in the adapter's connection pool, especially if the download failed mid-request. The connection pool could be holding onto sockets that aren't properly closed.
Fixes to Implement
1. Replace Recursion with a Loop for Retries
Recursion is unnecessary here and causes resource buildup. Switch to a while loop to ensure each retry fully cleans up resources before starting the next attempt:
def _download(self, local, remote): """Thread function using loop-based retries for cleaner resource management.""" retry_count = 0 max_retry = MAX_RETRY # Use your existing MAX_RETRY constant while retry_count <= max_retry: webdav_client = None try: webdav_client = WebDavClient(**kw) webdav_client.download(remote, local) if retry_count > 0: logger.info(f'Retry number {retry_count} was successful') break # Exit loop on success except (ConnectionError, requests.exceptions.ConnectionError) as err: retry_count += 1 if retry_count > max_retry: logger.exception(f'Failed after {max_retry} retries: {err}') return logger.warning(f'Retrying download (attempt {retry_count}) due to connection error: {err}') except Exception as err: logger.error(f'Unhandled exception during download: {err}') break finally: if webdav_client is not None: self.clean_webdav(webdav_client) raise SystemExit
2. Enhance Session Cleanup Logic
Update your clean_webdav function to ensure all connections in the requests adapter pool are closed, not just the session itself:
def clean_webdav(self, webdav): """Ensure full cleanup of WebDAV client and underlying connections.""" try: # Close all adapter connection pools first for adapter in webdav.session.adapters.values(): adapter.close() # Then close the session webdav.session.close() except Exception as err: logger.error(f'Error closing WebDAV session: {err}') finally: # Force reference deletion to help garbage collection del webdav # Optional: Trigger manual GC for immediate cleanup (use sparingly) import gc gc.collect()
3. Verify Connection Pool Behavior
If you're still seeing issues, you can disable HTTP keep-alive in the WebDavClient to prevent lingering connections. Modify how you create the client to pass a custom session with keep-alive disabled:
import requests from requests.adapters import HTTPAdapter def create_webdav_client(**kw): session = requests.Session() # Disable keep-alive by setting pool_connections and pool_maxsize to 0 adapter = HTTPAdapter(pool_connections=0, pool_maxsize=0) session.mount('http://', adapter) session.mount('https://', adapter) # Pass the custom session to WebDavClient kw['session'] = session return WebDavClient(**kw)
Then use this function in your _download loop instead of directly calling WebDavClient(**kw).
How to Validate the Fix
- Use
lsof -p <your_process_id>while the program is running during retries to check open file descriptors. You should see sockets being closed immediately after each failed retry. - Monitor the number of open files with
watch -n 1 "lsof -p <pid> | wc -l"to confirm it doesn't spike beyond your thread count + a small buffer.
内容的提问来源于stack exchange,提问作者Jan Janáček

