You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

网页爬取:如何获取懒加载图片的真实src属性?

解决ESPNcricinfo懒加载图片爬取问题

问题分析

你的代码核心问题是:先用Selenium加载了渲染后的页面,但随后又通过urllib.request.urlopen(url)重新获取了原始未渲染的HTML,导致拿到的是懒加载的占位图地址。此外,该网站的真实图片地址并不在src属性中,而是存储在data-src属性里。

修改方案

  1. 使用Selenium获取已经渲染完成的页面源码,替代urllib的原始请求
  2. 提取img标签的data-src属性作为真实图片地址
  3. 优化等待逻辑(用显式等待替代固定sleep,提升稳定性)

修改后的代码

import urllib.request
from bs4 import BeautifulSoup
from selenium import webdriver
from selenium.webdriver.chrome.service import Service
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import time

url = "https://www.espncricinfo.com/series/indian-premier-league-2022-1298423/squads"
s = Service("M:\WebScraping\chromedriver.exe")

driver = webdriver.Chrome(service=s)
driver.maximize_window()
driver.get(url)

# 显式等待页面元素加载完成,替代固定sleep
try:
    WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.CLASS_NAME, "ds-mb-4"))
    )
    # 滚动页面触发懒加载(确保所有图片地址被加载)
    driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
    time.sleep(2)  # 给滚动加载留缓冲时间
except Exception as e:
    print(f"等待页面加载失败: {e}")

# 获取Selenium渲染后的页面源码
page_source = driver.page_source
doc = BeautifulSoup(page_source, "html.parser")

teams = doc.find(class_="ds-p-0").find(class_="ds-mb-4")

for team in teams:
    img_tag = team.find("img")
    if img_tag and "data-src" in img_tag.attrs:
        real_img_src = img_tag["data-src"]
        print(real_img_src)
        file_name = img_tag["alt"].replace("/", "-")  # 处理文件名中的非法字符
        try:
            img_file = open(file_name + ".png", "wb")
            img_file.write(urllib.request.urlopen(real_img_src).read())
            img_file.close()
            print(f"成功保存图片: {file_name}.png")
        except Exception as e:
            print(f"保存图片失败 {file_name}: {e}")

driver.quit()

关键修改点说明

  • 用driver.page_source获取渲染后源码:确保拿到的是经过JS加载后的页面内容,包含真实图片地址
  • 提取data-src属性:该网站的懒加载图片将真实地址存在data-src中,页面渲染时JS会把这个值替换到src里
  • 显式等待:相比固定time.sleep,显式等待会在元素出现后立即继续执行,更高效稳定
  • 文件名处理:替换文件名中的/等非法字符,避免保存失败

内容的提问来源于stack exchange,提问作者Rayyan Alam

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 13:25:24