You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

AWS EC2执行Python脚本下载图片遇ReadTimeout,本地运行正常求助

解决EC2上从net-a-porter下载图片的ReadTimeout问题

针对你在EC2实例运行Python脚本时出现的requests.exceptions.ReadTimeout错误,结合你已尝试的方法,提供以下解决思路:

1. 排查EC2网络IP被目标站点拦截的可能

net-a-porter这类电商站点通常有反爬机制,AWS EC2的IP段可能被标记为爬虫IP,导致请求被限流或屏蔽:

  • 尝试为EC2实例更换弹性IP,或者通过AWS NAT网关出站,更换请求的IP来源后再测试。
  • 确认EC2安全组和网络ACL已开放出站443端口(虽然本地正常,但仍需排除网络策略问题)。

2. 拆分requests的超时参数

当前timeout=download_timeout是设置的总超时时间,建议拆分连接超时和读取超时,避免连接阶段占用过多时间:

# 替换原requests.get的timeout参数,10秒连接超时,120秒读取超时
response = requests.get(image_url, timeout=(10, 120), stream=True, headers=headers)

3. 添加请求重试机制

网络波动或临时限流可能导致单次请求超时,通过requests.Session和重试适配器实现自动重试:

import requests
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry

def upload_image_to_s3_from_url(self, image_url, filename, download_timeout=120):
    try:
        headers = {
            "User-Agent": "Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36",
            'Accept': 'image/avif,image/webp,image/apng,image/*,*/*;q=0.8',
            'Accept-Language': 'en-US,en;q=0.9',
            'Referer': 'https://www.net-a-porter.com/',
            'Cache-Control': 'no-cache'
        }
        # 配置会话和重试策略
        session = requests.Session()
        retry_strategy = Retry(
            total=3,  # 总重试次数
            backoff_factor=1,  # 重试间隔(1秒、2秒、4秒...)
            status_forcelist=[429, 500, 502, 503, 504],  # 需要重试的状态码
            allowed_methods=["GET"]  # 仅对GET请求重试
        )
        adapter = HTTPAdapter(max_retries=retry_strategy)
        session.mount("https://", adapter)
        session.mount("http://", adapter)

        # 使用会话发起请求
        response = session.get(image_url, timeout=(10, download_timeout), stream=True, headers=headers)
        response.raise_for_status()
        
        # 后续临时文件处理、S3上传逻辑保持不变
        content_type = response.headers.get('Content-Type', 'image/jpeg')
        with tempfile.NamedTemporaryFile(delete=False) as tmp_file:
            for chunk in response.iter_content(chunk_size=8192):
                tmp_file.write(chunk)
            file_url = self.upload_image_to_s3(tmp_file.name, filename, content_type)
        os.unlink(tmp_file.name)
        return file_url
    except requests.RequestException as e:
        raise Exception(f"Failed to download or upload image. Error: {e}")

4. 完善请求头模拟真实浏览器

除了User-Agent,补充更多浏览器常见请求头,降低被识别为爬虫的概率:

headers = {
    "User-Agent": "Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36",
    'Accept': 'image/avif,image/webp,image/apng,image/*,*/*;q=0.8',
    'Accept-Language': 'en-US,en;q=0.9',
    'Referer': 'https://www.net-a-porter.com/',
    'Cache-Control': 'no-cache',
    'Connection': 'keep-alive'
}

5. 检查目标站点的访问规则

查看net-a-porter的robots.txt,确认是否允许抓取图片资源;同时确保你的抓取行为符合站点的使用条款,避免触发更严格的拦截。


内容的提问来源于stack exchange,提问作者Usman Rafiq

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.27 01:35:29