You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python urllib+bs4抓取CoinMarketCap数据报AttributeError排查

问题现象

编写Python程序爬取CoinMarketCap官网的加密货币名称、价格数据时,运行持续抛出如下错误:

AttributeError: 'NoneType' object has no attribute 'text'

已手动确认目标页面内,class为priceHeading的h1标签、class为priceValue的div标签分别对应币种名称、价格的实际展示内容,但现有代码无法正常获取两个值,初始代码如下:

# import libraries
import urllib.request as ur
from bs4 import BeautifulSoup

quote_page = 'https://coinmarketcap.com/currencies/bitcoin/'

page = ur.urlopen(quote_page)

soup = BeautifulSoup(page, 'html.parser')

name_box = soup.find('h1', attrs={'class': 'priceHeading'})

name = name_box.text.strip() 
print (name)

price_box = soup.find('div', attrs={'class':'priceValue'})
price = price_box.text
print (price)
报错原因
  • 原生urllib默认请求头的User-Agent特征明显,会被网站反爬规则识别为爬虫,返回的页面内容缺失目标DOM节点
  • CoinMarketCap核心数据部分依赖前端JS动态渲染,直接请求拿到的初始静态HTML里根本没有对应class的标签,soup.find()匹配不到元素就会返回None,后续调用.text属性自然会抛出属性错误
解决方法

方案1:添加浏览器请求头做静态爬取(轻量无额外依赖)

给请求加上正常浏览器的User-Agent头,模拟普通用户访问,目前测试可以直接拿到包含目标数据的完整静态页,修复后代码:

import urllib.request as ur
from bs4 import BeautifulSoup

quote_page = 'https://coinmarketcap.com/currencies/bitcoin/'
# 模拟Chrome浏览器请求头
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
}
req = ur.Request(quote_page, headers=headers)
page = ur.urlopen(req)

soup = BeautifulSoup(page, 'html.parser')

name_box = soup.find('h1', attrs={'class': 'priceHeading'})
name = name_box.text.strip() 
print(name)

price_box = soup.find('div', attrs={'class':'priceValue'})
price = price_box.text
print(price)

提示:如果后续网站升级反爬策略,静态请求拿不到数据就换方案2

方案2:用Selenium加载完整渲染页面(适配所有动态加载场景)

如果静态请求始终拿不到完整内容,直接用浏览器自动化工具加载JS渲染完成后的页面源码即可,先安装依赖:

pip install selenium webdriver-manager

对应代码:

from selenium import webdriver
from selenium.webdriver.chrome.service import Service
from webdriver_manager.chrome import ChromeDriverManager
from bs4 import BeautifulSoup

quote_page = 'https://coinmarketcap.com/currencies/bitcoin/'
# 启动Chrome加载页面
driver = webdriver.Chrome(service=Service(ChromeDriverManager().install()))
driver.get(quote_page)
# 等待页面元素加载
driver.implicitly_wait(5)

soup = BeautifulSoup(driver.page_source, 'html.parser')
name_box = soup.find('h1', attrs={'class': 'priceHeading'})
name = name_box.text.strip() 
print(name)

price_box = soup.find('div', attrs={'class':'priceValue'})
price = price_box.text
print(price)

# 关闭浏览器
driver.quit()

内容的提问来源于stack exchange,提问作者Meer Modi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.31 06:06:25