网页爬取:如何获取懒加载图片的真实src属性?
解决ESPNcricinfo懒加载图片爬取问题
问题分析
你的代码核心问题是:先用Selenium加载了渲染后的页面,但随后又通过urllib.request.urlopen(url)重新获取了原始未渲染的HTML,导致拿到的是懒加载的占位图地址。此外,该网站的真实图片地址并不在src属性中,而是存储在data-src属性里。
修改方案
- 使用Selenium获取已经渲染完成的页面源码,替代urllib的原始请求
- 提取img标签的
data-src属性作为真实图片地址 - 优化等待逻辑(用显式等待替代固定sleep,提升稳定性)
修改后的代码
import urllib.request from bs4 import BeautifulSoup from selenium import webdriver from selenium.webdriver.chrome.service import Service from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import time url = "https://www.espncricinfo.com/series/indian-premier-league-2022-1298423/squads" s = Service("M:\WebScraping\chromedriver.exe") driver = webdriver.Chrome(service=s) driver.maximize_window() driver.get(url) # 显式等待页面元素加载完成,替代固定sleep try: WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CLASS_NAME, "ds-mb-4")) ) # 滚动页面触发懒加载(确保所有图片地址被加载) driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") time.sleep(2) # 给滚动加载留缓冲时间 except Exception as e: print(f"等待页面加载失败: {e}") # 获取Selenium渲染后的页面源码 page_source = driver.page_source doc = BeautifulSoup(page_source, "html.parser") teams = doc.find(class_="ds-p-0").find(class_="ds-mb-4") for team in teams: img_tag = team.find("img") if img_tag and "data-src" in img_tag.attrs: real_img_src = img_tag["data-src"] print(real_img_src) file_name = img_tag["alt"].replace("/", "-") # 处理文件名中的非法字符 try: img_file = open(file_name + ".png", "wb") img_file.write(urllib.request.urlopen(real_img_src).read()) img_file.close() print(f"成功保存图片: {file_name}.png") except Exception as e: print(f"保存图片失败 {file_name}: {e}") driver.quit()
关键修改点说明
- 用
driver.page_source获取渲染后源码:确保拿到的是经过JS加载后的页面内容,包含真实图片地址 - 提取
data-src属性:该网站的懒加载图片将真实地址存在data-src中,页面渲染时JS会把这个值替换到src里 - 显式等待:相比固定
time.sleep,显式等待会在元素出现后立即继续执行,更高效稳定 - 文件名处理:替换文件名中的
/等非法字符,避免保存失败
内容的提问来源于stack exchange,提问作者Rayyan Alam
相关产品推荐
相关产品推荐

