Selenium获取图片src为Base64,仅能抓取少量图片URL问题
Selenium无头模式抓取图片仅获取少量真实URL的问题解决
问题原因
- 无头浏览器检测:多数网站通过JavaScript检测
navigator.webdriver、window.chrome等属性,识别出无头模式后,返回Base64占位图而非真实图片链接。 - 图片懒加载机制:页面初始仅加载可视区域的图片,其余图片用Base64占位,需滚动页面触发JS加载真实URL;无头模式默认未执行滚动操作,因此未触发替换。
- 无头模式配置差异:默认窗口尺寸过小、User-Agent带有"HeadlessChrome"标识,导致页面渲染逻辑与正常浏览器不一致。
解决方法
1. 伪装无头浏览器参数,消除检测特征
修改ChromeOptions配置,让无头浏览器更接近真实环境:
from selenium import webdriver from selenium.webdriver.chrome.options import Options options = Options() options.add_argument('--headless=new') # 使用新版无头模式,兼容性更好 options.add_argument('--disable-blink-features=AutomationControlled') # 禁用自动化检测标识 options.add_argument('--user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36') # 替换为真实浏览器UA options.add_argument('--window-size=1920,1080') # 设置正常窗口尺寸 driver = webdriver.Chrome(options=options)
2. 模拟页面滚动,触发懒加载
通过JavaScript滚动页面,触发图片加载逻辑:
# 多次滚动到页面底部,确保所有懒加载图片被触发 for _ in range(3): driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") driver.implicitly_wait(2) # 等待图片加载完成
3. 等待图片src属性完成替换
针对img元素,等待其src不再是Base64格式后再提取:
from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC # 等待所有img的src不包含Base64前缀(适配多数场景) WebDriverWait(driver, 10).until( lambda d: all(not img.get_attribute('src').startswith('data:image/') for img in d.find_elements(By.TAG_NAME, 'img')) ) # 提取所有带https前缀的图片URL img_urls = [ img.get_attribute('src') for img in driver.find_elements(By.TAG_NAME, 'img') if img.get_attribute('src').startswith('https') ]
内容的提问来源于stack exchange,提问作者keyboardNoob
相关产品推荐
相关产品推荐

