循环内使用requests.get的Python脚本挂起问题排查及优化方案咨询
循环内使用requests.get的Python脚本挂起问题排查及优化方案咨询
兄弟,我太懂批量请求API时脚本突然挂住的糟心了!之前做数据同步和监控脚本时天天碰这种情况,给你拆解下问题原因,再把正确的用法和最佳实践说清楚:
一、脚本挂起的常见诱因
- 未设置超时(或超时配置不合理):你最开始的代码完全没加超时参数,而requests默认是没有超时限制的!也就是说如果API服务器没响应、网络断了,脚本会一直死等,这是挂起的头号原因。
- 服务器端或网络波动:比如API服务器过载限流、DNS解析卡住、中间网络设备丢包,都会导致请求卡在半路上。
- 连接池耗尽:requests默认用连接池复用TCP连接,但如果请求量太大,连接池里的空闲连接被用光,新请求就得等着,看起来就像挂起。
二、超时参数的正确打开方式
你后来加了timeout=10方向是对的,但这个参数还有更严谨的用法:
- 只传单个数字(比如
timeout=10):代表连接超时+读取超时的总时长,也就是从发起请求到拿到完整响应的总时间不能超过10秒。 - 分开设置更稳妥:
timeout=(3,10),第一个数字是「连接超时」(3秒内没和服务器建立好TCP连接就报错),第二个是「读取超时」(10秒内没拿到服务器返回的完整数据就报错)。
给你改好的基础版代码,一定要加异常捕获,不然一个请求失败整个脚本就崩了:
import requests urls = ['http://example.com/1', 'http://example.com/2'] for url in urls: try: # 3秒连接超时,10秒读取超时 response = requests.get(url, timeout=(3, 10)) response.raise_for_status() # 主动抛出4xx/5xx的HTTP错误 print(response.text) except requests.exceptions.RequestException as e: # 捕获所有请求相关的异常,保证循环继续 print(f"请求 {url} 失败: {str(e)}")
三、批量请求的进阶最佳实践
如果你的URL数量多,或者对稳定性要求高,这些技巧能帮你彻底解决问题:
- 细分异常捕获:可以针对性捕获
ConnectTimeout(连接超时)、ReadTimeout(读取超时)、HTTPError(HTTP错误),方便精准定位问题:except requests.exceptions.ConnectTimeout: print(f"连接 {url} 超时,可能是服务器挂了或网络断了") except requests.exceptions.ReadTimeout: print(f"读取 {url} 响应超时,服务器太慢了") except requests.exceptions.HTTPError as e: print(f"请求 {url} 返回HTTP错误: {e.response.status_code}") except requests.exceptions.RequestException as e: print(f"请求 {url} 遇到未知错误: {str(e)}") - 自动重试机制:网络偶尔波动很正常,用
tenacity库做指数退避重试(避免频繁重试打崩服务器),或者用requests自带的连接池重试:from tenacity import retry, stop_after_attempt, wait_exponential, retry_if_exception_type import requests from requests.exceptions import RequestException # 最多重试3次,每次重试等待时间指数增长(2秒、4秒、8秒) @retry( stop=stop_after_attempt(3), wait=wait_exponential(multiplier=1, min=2, max=10), retry=retry_if_exception_type(RequestException) ) def safe_request(url): response = requests.get(url, timeout=(3, 10)) response.raise_for_status() return response # 调用示例 for url in urls: try: resp = safe_request(url) print(resp.text) except RequestException as e: print(f"请求 {url} 最终失败: {str(e)}") - 并发请求控制:串行跑太慢的话,用
concurrent.futures搞线程并发,但要控制并发数(别一下发几十上百个请求被服务器拉黑):import requests from concurrent.futures import ThreadPoolExecutor, as_completed def fetch_url(url): try: resp = requests.get(url, timeout=(3, 10)) resp.raise_for_status() return (url, resp.text, "成功") except RequestException as e: return (url, str(e), "失败") urls = ['http://example.com/1', 'http://example.com/2'] # 控制并发数为5,根据服务器承受能力调整 with ThreadPoolExecutor(max_workers=5) as executor: futures = [executor.submit(fetch_url, url) for url in urls] for future in as_completed(futures): url, result, status = future.result() print(f"URL: {url} | 状态: {status} | 结果: {result}") - 用Session复用连接:requests的Session对象会自动复用TCP连接,减少握手开销,还能统一设置超时、重试、headers,比每次用get高效多了:
from requests.adapters import HTTPAdapter from urllib3.util.retry import Retry import requests # 初始化Session,配置重试和连接池 session = requests.Session() retry_strategy = Retry( total=3, # 总重试次数 backoff_factor=1, # 重试等待时间系数(1秒、2秒、4秒...) status_forcelist=[429, 500, 502, 503, 504] # 遇到这些状态码自动重试 ) # 连接池最多10个连接,同时最多10个请求 adapter = HTTPAdapter(max_retries=retry_strategy, pool_connections=10, pool_maxsize=10) session.mount("http://", adapter) session.mount("https://", adapter) # 用Session发请求 for url in urls: try: resp = session.get(url, timeout=(3, 10)) resp.raise_for_status() print(resp.text) except RequestException as e: print(f"请求 {url} 失败: {str(e)}")
核心总结一下:超时设置是必须的,异常捕获是基础,批量请求一定要加重试和并发控制,用Session能大幅提升性能和稳定性。
内容来源于stack exchange
相关产品推荐
相关产品推荐

