You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python下载丝芙兰网站图片?请求遭服务器拦截求解决方案

解决Sephora图片下载被服务器拦截的方案

Sephora的反爬机制不只是校验User-Agent,还会检查请求来源(Referer)、会话Cookie,甚至验证访问流程的合法性,单纯设置User-Agent大概率会被拦截。以下是两种可行的解决方法:

方法1:完善请求头+会话保持(Requests库)

用requests.Session()维持会话,先访问Sephora主页获取合法Cookie,再带着完整请求头请求图片,模拟真实用户的访问路径:

import requests

# 初始化会话对象,自动管理Cookie
session = requests.Session()

# 模拟真实Chrome浏览器的请求头
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36',
    'Referer': 'https://www.sephora.com/',  # 必须携带Sephora域名作为来源
    'Accept': 'image/webp,image/apng,image/svg+xml,image/*,*/*;q=0.8',
    'Accept-Language': 'zh-CN,zh;q=0.9',
    'Connection': 'keep-alive'
}

image_url = "https://www.sephora.com/productimages/sku/s2586261-main-zoom.jpg"
save_path = 'test.jpg'

try:
    # 先访问主页,获取服务器认可的会话Cookie
    session.get('https://www.sephora.com/', headers=headers)
    # 流式请求图片,避免内存占用过高
    resp = session.get(image_url, headers=headers, stream=True)
    resp.raise_for_status()  # 抛出HTTP请求异常
    
    # 写入文件
    with open(save_path, 'wb') as f:
        for chunk in resp.iter_content(chunk_size=8192):
            f.write(chunk)
    print(f"图片已保存到:{save_path}")
except Exception as e:
    print(f"下载失败:{str(e)}")

方法2:Selenium模拟真实浏览器下载

如果Requests方法仍失败,用Selenium启动真实浏览器加载图片,完全模拟用户操作,绕过绝大多数反爬:

from selenium import webdriver
from selenium.webdriver.chrome.options import Options
import base64
import time

# Chrome浏览器配置
chrome_options = Options()
chrome_options.add_argument('--headless=new')  # 无头模式,后台运行
chrome_options.add_argument('--disable-gpu')
chrome_options.add_argument('user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36')

driver = webdriver.Chrome(options=chrome_options)
image_url = "https://www.sephora.com/productimages/sku/s2586261-main-zoom.jpg"
save_path = 'test.jpg'

try:
    driver.get(image_url)
    time.sleep(2)  # 等待图片加载完成
    
    # 通过JS获取图片的base64编码(避免直接下载被拦截)
    base64_data = driver.execute_script("""
        const img = document.querySelector('img');
        const canvas = document.createElement('canvas');
        canvas.width = img.naturalWidth;
        canvas.height = img.naturalHeight;
        canvas.getContext('2d').drawImage(img, 0, 0);
        return canvas.toDataURL('image/jpeg').split(',')[1];
    """)
    
    # 解码并保存
    with open(save_path, 'wb') as f:
        f.write(base64.b64decode(base64_data))
    print(f"图片已保存到:{save_path}")
except Exception as e:
    print(f"下载失败:{str(e)}")
finally:
    driver.quit()  # 关闭浏览器

注意事项

  • 不要短时间内大量请求,建议每次请求后加1-3秒的随机延迟,避免IP被封禁。
  • 定期更新User-Agent,使用当前主流浏览器的标识。
  • 使用Selenium时,要保证ChromeDriver版本和本地Chrome浏览器版本匹配。

内容的提问来源于stack exchange,提问作者chalukya reddy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 15:12:50