You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Selenium获取图片src为Base64,仅能抓取少量图片URL问题

Selenium无头模式抓取图片仅获取少量真实URL的问题解决

问题原因

  • 无头浏览器检测:多数网站通过JavaScript检测navigator.webdriver、window.chrome等属性,识别出无头模式后,返回Base64占位图而非真实图片链接。
  • 图片懒加载机制:页面初始仅加载可视区域的图片,其余图片用Base64占位,需滚动页面触发JS加载真实URL;无头模式默认未执行滚动操作,因此未触发替换。
  • 无头模式配置差异:默认窗口尺寸过小、User-Agent带有"HeadlessChrome"标识,导致页面渲染逻辑与正常浏览器不一致。

解决方法

1. 伪装无头浏览器参数,消除检测特征

修改ChromeOptions配置,让无头浏览器更接近真实环境:

from selenium import webdriver
from selenium.webdriver.chrome.options import Options

options = Options()
options.add_argument('--headless=new')  # 使用新版无头模式,兼容性更好
options.add_argument('--disable-blink-features=AutomationControlled')  # 禁用自动化检测标识
options.add_argument('--user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36')  # 替换为真实浏览器UA
options.add_argument('--window-size=1920,1080')  # 设置正常窗口尺寸

driver = webdriver.Chrome(options=options)

2. 模拟页面滚动,触发懒加载

通过JavaScript滚动页面,触发图片加载逻辑:

# 多次滚动到页面底部,确保所有懒加载图片被触发
for _ in range(3):
    driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
    driver.implicitly_wait(2)  # 等待图片加载完成

3. 等待图片src属性完成替换

针对img元素,等待其src不再是Base64格式后再提取:

from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

# 等待所有img的src不包含Base64前缀(适配多数场景)
WebDriverWait(driver, 10).until(
    lambda d: all(not img.get_attribute('src').startswith('data:image/') for img in d.find_elements(By.TAG_NAME, 'img'))
)

# 提取所有带https前缀的图片URL
img_urls = [
    img.get_attribute('src') 
    for img in driver.find_elements(By.TAG_NAME, 'img') 
    if img.get_attribute('src').startswith('https')
]

内容的提问来源于stack exchange,提问作者keyboardNoob

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.11 02:03:17