使用Python requests批量爬取多URL邮箱时的超时失效问题咨询
使用Python requests批量爬取多URL邮箱时的超时失效问题咨询
Hey there! Let's dig into why your timeout isn't working as expected and how to fix that annoying hang issue when scraping lots of URLs.
1. 为什么设置了timeout还是会卡住?
There are a few common culprits that might be slipping past requests' timeout settings:
- DNS解析不受requests timeout限制:requests的
timeout=(connect, read)参数只负责TCP连接建立后的等待时间,以及读取响应时的间隔超时。但如果某个域名的DNS服务器无响应,系统自带的DNS查询机制可能会挂很久(比如默认几十秒甚至更久),这时候requests根本没机会触发自己的超时。 - TCP连接阶段的系统重试:在某些网络环境下,当发送TCP SYN包后服务器不回复SYN-ACK,系统的TCP栈可能会自动重试多次,这个过程的时间可能会超过你设置的connect timeout,导致请求卡住。
- 边缘情况的慢响应:如果服务器一直在断断续续地发送极小的数据包,且两次数据包的间隔不超过你的read timeout(10秒),requests会一直等待,不会触发超时——因为read timeout是指两次数据包之间的空闲时间,不是总读取时间。
2. 解决办法:让脚本再也不无限挂起
Here are some practical fixes you can try, ordered by ease of implementation:
(1)全局设置socket超时,覆盖DNS和TCP阶段
requests基于socket运行,所以可以先给所有socket操作设置一个全局超时,这样DNS解析、TCP连接等阶段都会被限制时间:
import socket import requests # 全局设置15秒超时,覆盖所有socket操作 socket.setdefaulttimeout(15) urls = [ "http://example.com", "http://anotherexample.com", # ... more URLs ] HEADERS={"user-agent":"Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36"} for url in urls: try: response = requests.get(url, headers=HEADERS, timeout=(10, 10)) if response.status_code == 200: print(f"Emails from {url}: ", response.text) except requests.exceptions.Timeout: print(f"Timeout occurred for {url}") except requests.exceptions.RequestException as e: print(f"Error occurred for {url}: {e}")
注意:这个设置会影响所有使用socket的Python代码,如果你的脚本还有其他网络操作,需要谨慎调整时长。
(2)改用总超时,而非分阶段超时
把timeout=(10,10)改成单个数值,比如timeout=15,这代表请求的总耗时上限(从发起请求到接收完响应的总时间),可以避免某些分阶段超时覆盖不到的边缘情况:
response = requests.get(url, headers=HEADERS, timeout=15)
(3)用线程池批量处理,隔离卡住的请求
批量爬取时用线程池是个好主意,即使某个请求卡住,其他请求也能继续执行,而且可以给每个任务单独设置超时:
import requests from concurrent.futures import ThreadPoolExecutor, as_completed urls = [ "http://example.com", "http://anotherexample.com", # ... more URLs ] HEADERS={"user-agent":"Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36"} def scrape_single_url(url): try: response = requests.get(url, headers=HEADERS, timeout=(10, 10)) if response.status_code == 200: return f"Emails from {url}: ", response.text except requests.exceptions.Timeout: return f"Timeout occurred for {url}" except requests.exceptions.RequestException as e: return f"Error occurred for {url}: {e}" # 最多同时运行10个线程,根据你的网络情况调整 with ThreadPoolExecutor(max_workers=10) as executor: # 提交所有任务 future_to_url = {executor.submit(scrape_single_url, url): url for url in urls} for future in as_completed(future_to_url): url = future_to_url[future] try: # 给每个任务单独设置总超时15秒 result = future.result(timeout=15) print(result) except TimeoutError: print(f"Task for {url} took too long and was terminated") future.cancel()
(4)改用异步请求库(进阶方案)
如果线程池还不够,试试异步的aiohttp——它的事件循环机制不会被单个慢请求卡住,超时处理也更可靠:
import aiohttp import asyncio urls = [ "http://example.com", "http://anotherexample.com", # ... more URLs ] HEADERS={"user-agent":"Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36"} async def scrape_url(session, url): try: async with session.get(url, headers=HEADERS, timeout=15) as response: if response.status == 200: text = await response.text() return f"Emails from {url}: ", text except asyncio.TimeoutError: return f"Timeout occurred for {url}" except Exception as e: return f"Error occurred for {url}: {e}" async def main(): async with aiohttp.ClientSession() as session: tasks = [scrape_url(session, url) for url in urls] results = await asyncio.gather(*tasks) for result in results: print(result) asyncio.run(main())
备注:内容来源于stack exchange,提问作者Mominur Rahman
相关产品推荐
相关产品推荐

