You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup、Selenium爬取CoinMarketCap仅返回10条数据求助

问题根因
  • 页面源码抓取时机错误:你在执行滚动操作前就已经把初始页面源码存到了变量里,后续滚动加载出的内容不会更新到该变量中
  • 懒加载机制未适配:CoinMarketCap的列表采用懒加载渲染,只有元素进入可视区域才会加载数据,单次滚动到底且未预留渲染时间的操作,无法触发全量数据加载
  • 隐式等待配置不生效:implicitly_wait仅作用于Selenium查找元素的场景,直接获取页面源码的操作不会触发该等待逻辑
修正代码
from selenium import webdriver
from webdriver_manager.chrome import ChromeDriverManager
from bs4 import BeautifulSoup
import time

options = webdriver.ChromeOptions()
options.add_experimental_option("excludeSwitches", ["enable-logging"])
driver = webdriver.Chrome(ChromeDriverManager().install(), options=options)
driver.maximize_window()
driver.get('https://coinmarketcap.com/')
# 等待页面基础内容加载完成
time.sleep(2)

# 逐次滚动触发懒加载,覆盖100条数据的加载需求
for i in range(10):
    driver.execute_script("window.scrollBy(0, 700)")
    time.sleep(0.8)

# 所有内容加载完成后再获取完整源码
page_source = driver.page_source
driver.quit()

soup = BeautifulSoup(page_source, 'html.parser')
status_today = soup.find_all('div', {'class': 'sc-16r8icm-0 escjiH'})

for x in status_today:
    if x.a and 'href' in x.a.attrs:
        print('x.a[href]=', x.a['href'])

运行上述代码即可输出完整100条加密货币的对应链接。如果后续网站更新修改了容器类名,重新在浏览器F12控制台定位对应元素替换类名参数即可。

内容的提问来源于stack exchange,提问作者Danny

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.30 03:45:02