AWS EC2执行Python脚本下载图片遇ReadTimeout,本地运行正常求助
解决EC2上从net-a-porter下载图片的ReadTimeout问题
针对你在EC2实例运行Python脚本时出现的requests.exceptions.ReadTimeout错误,结合你已尝试的方法,提供以下解决思路:
1. 排查EC2网络IP被目标站点拦截的可能
net-a-porter这类电商站点通常有反爬机制,AWS EC2的IP段可能被标记为爬虫IP,导致请求被限流或屏蔽:
- 尝试为EC2实例更换弹性IP,或者通过AWS NAT网关出站,更换请求的IP来源后再测试。
- 确认EC2安全组和网络ACL已开放出站443端口(虽然本地正常,但仍需排除网络策略问题)。
2. 拆分requests的超时参数
当前timeout=download_timeout是设置的总超时时间,建议拆分连接超时和读取超时,避免连接阶段占用过多时间:
# 替换原requests.get的timeout参数,10秒连接超时,120秒读取超时 response = requests.get(image_url, timeout=(10, 120), stream=True, headers=headers)
3. 添加请求重试机制
网络波动或临时限流可能导致单次请求超时,通过requests.Session和重试适配器实现自动重试:
import requests from requests.adapters import HTTPAdapter from urllib3.util.retry import Retry def upload_image_to_s3_from_url(self, image_url, filename, download_timeout=120): try: headers = { "User-Agent": "Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36", 'Accept': 'image/avif,image/webp,image/apng,image/*,*/*;q=0.8', 'Accept-Language': 'en-US,en;q=0.9', 'Referer': 'https://www.net-a-porter.com/', 'Cache-Control': 'no-cache' } # 配置会话和重试策略 session = requests.Session() retry_strategy = Retry( total=3, # 总重试次数 backoff_factor=1, # 重试间隔(1秒、2秒、4秒...) status_forcelist=[429, 500, 502, 503, 504], # 需要重试的状态码 allowed_methods=["GET"] # 仅对GET请求重试 ) adapter = HTTPAdapter(max_retries=retry_strategy) session.mount("https://", adapter) session.mount("http://", adapter) # 使用会话发起请求 response = session.get(image_url, timeout=(10, download_timeout), stream=True, headers=headers) response.raise_for_status() # 后续临时文件处理、S3上传逻辑保持不变 content_type = response.headers.get('Content-Type', 'image/jpeg') with tempfile.NamedTemporaryFile(delete=False) as tmp_file: for chunk in response.iter_content(chunk_size=8192): tmp_file.write(chunk) file_url = self.upload_image_to_s3(tmp_file.name, filename, content_type) os.unlink(tmp_file.name) return file_url except requests.RequestException as e: raise Exception(f"Failed to download or upload image. Error: {e}")
4. 完善请求头模拟真实浏览器
除了User-Agent,补充更多浏览器常见请求头,降低被识别为爬虫的概率:
headers = { "User-Agent": "Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36", 'Accept': 'image/avif,image/webp,image/apng,image/*,*/*;q=0.8', 'Accept-Language': 'en-US,en;q=0.9', 'Referer': 'https://www.net-a-porter.com/', 'Cache-Control': 'no-cache', 'Connection': 'keep-alive' }
5. 检查目标站点的访问规则
查看net-a-porter的robots.txt,确认是否允许抓取图片资源;同时确保你的抓取行为符合站点的使用条款,避免触发更严格的拦截。
内容的提问来源于stack exchange,提问作者Usman Rafiq
相关产品推荐
相关产品推荐

