使用Python requests下载特定URL图片时遇超时错误求助
解决net-a-porter图片下载超时问题
针对你遇到的特定域名图片下载超时问题,以下是几个可行的排查和解决方向:
1. 分离连接超时与读取超时
requests.get的timeout参数若传单个值,会同时作用于连接超时和读取超时。部分网站在建立连接后,数据传输阶段可能较慢,建议分别设置两个超时值,给读取阶段预留更充足的时间:
response = requests.get( image_url, timeout=(10, 300), # 连接超时10秒,读取超时300秒 stream=True, headers=headers )
2. 补充完整的浏览器请求头
net-a-porter这类电商网站反爬机制较严格,仅User-Agent和Accept不足以模拟真实会话,建议补充更多常见请求头:
headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36', 'Accept': 'image/avif,image/webp,image/apng,image/svg+xml,image/*,*/*;q=0.8', 'Accept-Language': 'zh-CN,zh;q=0.9', 'Accept-Encoding': 'gzip, deflate, br', 'Referer': 'https://www.net-a-porter.com/', 'Connection': 'keep-alive' }
注意将User-Agent更新为较新版本的浏览器标识,避免被识别为老旧爬虫。
3. 使用代理IP绕过限制
部分网站会对特定地区或IP段限流,尝试使用代理IP访问:
proxies = { 'http': 'http://your-proxy-ip:port', 'https': 'http://your-proxy-ip:port' } response = requests.get(image_url, timeout=(10, 300), stream=True, headers=headers, proxies=proxies)
4. 改用httpx替代requests
httpx原生支持HTTP/2,部分网站对HTTP/2的兼容性更好,可尝试替换下载库:
首先安装依赖:
pip install httpx[http2]
修改下载逻辑:
import httpx import tempfile import os def upload_image_to_s3_from_url(self, image_url, filename, download_timeout=300): try: headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36', 'Accept': 'image/avif,image/webp,image/apng,image/svg+xml,image/*,*/*;q=0.8', 'Accept-Language': 'zh-CN,zh;q=0.9', 'Referer': 'https://www.net-a-porter.com/' } with httpx.Client(http2=True, timeout=download_timeout) as client: response = client.get(image_url, headers=headers, stream=True) response.raise_for_status() with tempfile.NamedTemporaryFile(delete=False) as tmp_file: for chunk in response.iter_bytes(chunk_size=8192): tmp_file.write(chunk) file_url = self.upload_image_to_s3(tmp_file.name, filename) os.unlink(tmp_file.name) return file_url except httpx.HTTPError as e: raise Exception(f"Failed to download or upload image. Error: {e}")
5. 排查网络环境
- 在服务器终端用
curl测试下载,确认网络层面是否能正常访问目标域名:
curl -I "https://www.net-a-porter.com/variants/images/17266703523615883/in/w920_a3-4_q60.jpg" -A "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
如果curl也超时,说明是网络策略问题(如防火墙拦截、地区限制),需联系运维调整。
内容的提问来源于stack exchange,提问作者Usman Rafiq
相关产品推荐
相关产品推荐

