使用BeautifulSoup(BS4)爬取网站时无法获取全部img标签
问题原因
你遇到的情况是因为CoinMarketCap采用了动态内容加载机制:初始请求返回的HTML只包含页面顶部的部分数据,后续的代币条目(包括对应的<img>标签)是在用户滚动页面时通过AJAX请求异步加载的。requests库只能获取服务器初始返回的静态HTML,无法捕获JavaScript动态渲染的内容,所以只能拿到前10条左右的数据。
解决方案
方案1:使用CoinMarketCap官方API(推荐)
直接调用官方API获取数据,比爬取页面更稳定合规,还能直接拿到所需的图片URL。
首先需要申请CoinMarketCap API密钥(免费版足够满足基础需求),然后编写代码:
import requests API_KEY = "你的API密钥" url = "https://pro-api.coinmarketcap.com/v1/cryptocurrency/listings/latest" params = { "start": "1", "limit": "500", # 按需设置返回的代币数量 "category": "real-world-assets" } headers = { "Accepts": "application/json", "X-CMC_PRO_API_KEY": API_KEY } response = requests.get(url, params=params, headers=headers) data = response.json() # 提取每个代币的logo URL for crypto in data["data"]: print(f"{crypto['name']} ({crypto['symbol']}): {crypto['logo']}")
方案2:使用Selenium模拟浏览器渲染
如果不想使用API,可以用Selenium模拟浏览器打开页面,等待所有内容加载完成后再解析HTML:
先安装依赖:
pip install selenium
然后编写代码(需提前下载对应浏览器的驱动,比如ChromeDriver):
from selenium import webdriver from selenium.webdriver.chrome.options import Options from bs4 import BeautifulSoup import time # 配置Chrome无头模式(不显示浏览器窗口) chrome_options = Options() chrome_options.add_argument("--headless=new") chrome_options.add_argument("--disable-gpu") driver = webdriver.Chrome(options=chrome_options) driver.get("https://coinmarketcap.com/view/real-world-assets/") # 滚动页面到底部,触发所有内容加载 last_height = driver.execute_script("return document.body.scrollHeight") while True: driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") time.sleep(2) # 等待内容加载 new_height = driver.execute_script("return document.body.scrollHeight") if new_height == last_height: break last_height = new_height # 解析完整页面的HTML soup = BeautifulSoup(driver.page_source, "lxml") driver.quit() # 提取所有<img>标签的src属性 table = soup.find("table", {"class": "sc-482c3d57-3 iTyfmj cmc-table"}) rows = table.find_all("tr") for row in rows: for img in row.find_all("img"): print(img.get("src"))
内容的提问来源于stack exchange,提问作者BorangeOrange1337
相关产品推荐
相关产品推荐

