切换至httpx后爬虫引擎出现未捕获的httpcore/anyio内部异常请求解
BrokenResourceError in httpx/httpcore After Switching from aiohttp Hey there! Let's unpack why you're seeing those persistent BrokenResourceError entries in your logs after moving from aiohttp to httpx—even though your crawler keeps running just fine.
What's Causing This Exception?
Let's break down the traceback you shared to get to the root:
- The core issue starts with
ConnectionResetError: [Errno 104] Connection reset by peer—this is a low-level TCP error where the remote end (either your proxy server or the target website) abruptly closed the connection. This can happen for all the reasons you noted: flaky proxies, timeouts, server-side connection limits, or random network blips. - httpx relies on
httpcore(its underlying HTTP core library) andanyio(the async runtime it uses), which wrap this low-level TCP error into ananyio.BrokenResourceError. This is their way of abstracting messy network failure details into a consistent, higher-level exception type.
Why Did This Start Happening After Switching to httpx?
The key difference lies in how aiohttp vs. httpx/httpcore handle uncaught exceptions in async futures:
- aiohttp likely swallows or internally handles some of these low-level connection failures without propagating them as unretrieved future exceptions. Its design might absorb these errors when your retry logic kicks in, so they never show up in your logs.
- httpx/httpcore, on the other hand, doesn't silently suppress these exceptions. When your retry mechanism spins up a new request after a failed one, the original failed request's future still holds the exception—and since your code doesn't explicitly retrieve or handle that exception, Python's asyncio logs it as "Future exception was never retrieved".
Why Isn't This Breaking Your Crawler?
Your retry logic is doing its job perfectly! When a proxy fails (triggering the connection reset and subsequent BrokenResourceError), your engine automatically retries with another proxy. The exception is from the failed request attempt, but since you're not relying on that failed future's result, your crawler keeps chugging along. The only side effect is the noisy log entries.
Quick Fixes to Quiet the Logs
If you want to stop these exceptions from cluttering your logs, you have a couple practical options:
- Explicitly handle exceptions per request: Wrap your httpx async request calls in a
try/exceptblock that catcheshttpx.TransportError(whichBrokenResourceErrorfalls under) to ensure the future's exception is properly retrieved and handled. - Adjust logging filters: Add a custom filter to your logger to ignore the "Future exception was never retrieved" messages specifically related to
BrokenResourceError. - Tweak httpx settings: Experiment with httpx's timeout and connection pool configurations to reduce the frequency of these failures (though with flaky proxies, some will still be unavoidable).
内容的提问来源于stack exchange,提问作者aikipooh

