You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python抓取CoinGecko页面中处于加载状态的开发者GitHub链接

解决CoinGecko开发者链接抓取失败的问题

Hey,我帮你分析下问题所在,再给你两个可行的解决方案:

为什么会拿到"loading"内容?

CoinGecko的开发者板块是动态加载的——你用requests.get()拿到的只是页面的初始静态HTML,这部分里开发者区域只有个占位的"loading"提示,真实的GitHub链接是等页面加载完成后,通过JavaScript异步请求渲染出来的。BeautifulSoup只能解析静态HTML,自然找不到你要的内容。

方案1:用CoinGecko官方API(强烈推荐)

直接调用官方API比爬网页靠谱多了,不仅不会遇到动态加载的问题,还能避免反爬限制,数据格式也更规整。

试试这段代码:

import requests

# 先搜索币种获取唯一ID
search_res = requests.get("https://api.coingecko.com/api/v3/search", params={"query": "solpatrol-bail"})
coin_id = search_res.json()["coins"][0]["id"]

# 拉取开发者数据
dev_res = requests.get(f"https://api.coingecko.com/api/v3/coins/{coin_id}/developer_data")
dev_data = dev_res.json()

# 提取GitHub链接
github_link = dev_data.get("github", {}).get("url", "未找到GitHub链接")
print(github_link)

方案2:用Selenium模拟浏览器渲染

如果你一定要从网页抓取,那就得模拟真实浏览器的渲染过程,等动态内容加载完再提取。

示例代码:

from selenium import webdriver
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By
from bs4 import BeautifulSoup

# 初始化Chrome浏览器(记得装对应版本的chromedriver)
driver = webdriver.Chrome()
driver.get("https://www.coingecko.com/en/coins/solpatrol-bail#developer")

# 等待开发者区域加载完成(最多等10秒)
wait = WebDriverWait(driver, 10)
wait.until(EC.presence_of_element_located((By.ID, "developer")))

# 获取渲染后的页面源码
page_html = driver.page_source
driver.quit()

# 解析提取GitHub链接
soup = BeautifulSoup(page_html, "html.parser")
github_links = soup.select("#developer a[href*='github.com']")
for link in github_links:
    print(link["href"])

小提醒

  • 使用API时注意免费版的速率限制:每分钟最多50次请求,别刷太猛。
  • Selenium可以加--headless=new参数开启无头模式,不用弹出浏览器窗口,更适合服务器运行。

内容的提问来源于stack exchange,提问作者Arcos

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.28 22:37:35