Python爬虫抓取动态网页时如何自动执行JS加载内容无需手动滚动
问题成因
你访问的CoinMarketCap站点采用了**懒加载(虚拟滚动)**的前端渲染策略:只有当内容区块进入浏览器可视区域时,才会触发接口请求、渲染对应的DOM节点。
你当前的代码仅等待了固定时间,没有触发内容加载的条件:
- 非无头模式下,默认打开的浏览器窗口可视区域有限,仅顶部未超出窗口的币种数据会被加载,未展示的区域对应HTML为空
- 无头模式下默认窗口尺寸更小,触发加载的内容更少,因此最终抓取结果和直接请求静态页一致
- 手动滚动时内容会依次进入可视区,触发加载逻辑,因此可以拿到完整数据
解决方案
可以通过Selenium执行JS脚本模拟滚动、或者强制修改窗口尺寸的方式,自动触发所有懒加载内容的渲染,不需要手动操作。
以下是修改后的可直接运行的代码:
import time from bs4 import BeautifulSoup from selenium import webdriver from selenium.webdriver.chrome.service import Service if __name__ == "__main__": options = webdriver.ChromeOptions() options.add_argument('--ignore-certificate-errors') options.add_argument('--incognito') # 无头模式下需要指定窗口尺寸,否则默认极小窗口 options.add_argument('--window-size=1920,3000') # 开启无头模式也可正常使用 # options.add_argument('--headless=new') # 高版本Selenium需要用Service传入chromedriver路径,低版本可保留原有写法 service = Service("C:\Program Files (x86)\chromedriver.exe") driver = webdriver.Chrome(service=service, options=options) driver.get('https://coinmarketcap.com/') time.sleep(2) # 模拟滚动到底部的逻辑,直到页面高度不再变化说明加载完成 last_height = driver.execute_script("return document.body.scrollHeight") while True: # 滚动到页面底部 driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") time.sleep(1) # 计算新的页面高度,和上一次高度对比 new_height = driver.execute_script("return document.body.scrollHeight") if new_height == last_height: break last_height = new_height html = driver.page_source soup = BeautifulSoup(html, "html.parser") tbody = soup.tbody trs = tbody.contents for tr in trs: print(tr) driver.close()
如果你只需要抓取默认展示的前100条币种数据,也可以直接把窗口高度设置为10000,不需要循环滚动,打开页面等待2-3秒即可直接拿到全部渲染后的DOM。
内容的提问来源于stack exchange,提问作者ybrenning
相关产品推荐
相关产品推荐

