如何用Python下载丝芙兰网站图片?请求遭服务器拦截求解决方案
解决Sephora图片下载被服务器拦截的方案
Sephora的反爬机制不只是校验User-Agent,还会检查请求来源(Referer)、会话Cookie,甚至验证访问流程的合法性,单纯设置User-Agent大概率会被拦截。以下是两种可行的解决方法:
方法1:完善请求头+会话保持(Requests库)
用requests.Session()维持会话,先访问Sephora主页获取合法Cookie,再带着完整请求头请求图片,模拟真实用户的访问路径:
import requests # 初始化会话对象,自动管理Cookie session = requests.Session() # 模拟真实Chrome浏览器的请求头 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36', 'Referer': 'https://www.sephora.com/', # 必须携带Sephora域名作为来源 'Accept': 'image/webp,image/apng,image/svg+xml,image/*,*/*;q=0.8', 'Accept-Language': 'zh-CN,zh;q=0.9', 'Connection': 'keep-alive' } image_url = "https://www.sephora.com/productimages/sku/s2586261-main-zoom.jpg" save_path = 'test.jpg' try: # 先访问主页,获取服务器认可的会话Cookie session.get('https://www.sephora.com/', headers=headers) # 流式请求图片,避免内存占用过高 resp = session.get(image_url, headers=headers, stream=True) resp.raise_for_status() # 抛出HTTP请求异常 # 写入文件 with open(save_path, 'wb') as f: for chunk in resp.iter_content(chunk_size=8192): f.write(chunk) print(f"图片已保存到:{save_path}") except Exception as e: print(f"下载失败:{str(e)}")
方法2:Selenium模拟真实浏览器下载
如果Requests方法仍失败,用Selenium启动真实浏览器加载图片,完全模拟用户操作,绕过绝大多数反爬:
from selenium import webdriver from selenium.webdriver.chrome.options import Options import base64 import time # Chrome浏览器配置 chrome_options = Options() chrome_options.add_argument('--headless=new') # 无头模式,后台运行 chrome_options.add_argument('--disable-gpu') chrome_options.add_argument('user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36') driver = webdriver.Chrome(options=chrome_options) image_url = "https://www.sephora.com/productimages/sku/s2586261-main-zoom.jpg" save_path = 'test.jpg' try: driver.get(image_url) time.sleep(2) # 等待图片加载完成 # 通过JS获取图片的base64编码(避免直接下载被拦截) base64_data = driver.execute_script(""" const img = document.querySelector('img'); const canvas = document.createElement('canvas'); canvas.width = img.naturalWidth; canvas.height = img.naturalHeight; canvas.getContext('2d').drawImage(img, 0, 0); return canvas.toDataURL('image/jpeg').split(',')[1]; """) # 解码并保存 with open(save_path, 'wb') as f: f.write(base64.b64decode(base64_data)) print(f"图片已保存到:{save_path}") except Exception as e: print(f"下载失败:{str(e)}") finally: driver.quit() # 关闭浏览器
注意事项
- 不要短时间内大量请求,建议每次请求后加1-3秒的随机延迟,避免IP被封禁。
- 定期更新User-Agent,使用当前主流浏览器的标识。
- 使用Selenium时,要保证ChromeDriver版本和本地Chrome浏览器版本匹配。
内容的提问来源于stack exchange,提问作者chalukya reddy
相关产品推荐
相关产品推荐

