使用BeautifulSoup爬取CoreDAO代币页面遇AttributeError问题求助
解决CoreDAO代币页面数据爬取的AttributeError问题
问题背景
需要从CoreDAO浏览器代币页面获取总供应量、持有者数、转账数等数据,但运行Python代码时触发AttributeError: 'NoneType' object has no attribute 'get_text'错误。
用户原代码
import requests from bs4 import BeautifulSoup headers = {"User-Agent": "Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:92.0) Gecko/20100101 Firefox/92.0"} urllink = "https://scan.coredao.org/token/0x7c6a914cb866654b50f94e66b80c046dd9225ba9" print ("urllink: ", urllink) urlpage = requests.get(urllink, headers=headers, timeout=5) print ("urlpage: ", urlpage) soup = BeautifulSoup(urlpage.content, 'html.parser') supply = soup.find('div', class_='el-row power-description mb-4').get_text() print ("supply: ", supply)
实际错误输出
urllink: https://scan.coredao.org/token/0x7c6a914cb866654b50f94e66b80c046dd9225ba9 urlpage: <Response [200]> Traceback (most recent call last): File "C:/Users/estorbotz/Desktop/Personal Files/CORE/Temp - Get Token Info V1.py", line 23, in <module> supply = soup.find('div', class_='el-row power-description mb-4').get_text() AttributeError: 'NoneType' object has no attribute 'get_text'
目标获取数据示例
Total Supply: 2,100,000,000 CAKE Contract: 0x7c6a914cb866654b50f94e66b80c046dd9225ba9 Holders: 9,299 Decimals: 18 Transfers: 21,421 Official Site: https://cakecore.io Social Profiles: coreteam@cakecore.io https://cakecore.io/whitepaper https://t.me/cakecore_io https://twitter.com/cakecore_io
错误原因
- 页面动态渲染:目标页面依赖JavaScript加载代币数据,
requests仅能获取初始静态HTML,无法捕获JS渲染后的元素,导致soup.find()返回None。 - 选择器失效:即使页面结构未变更,原CSS选择器可能已不匹配当前页面的元素类名。
解决思路与修正代码
方案1:调用官方API(推荐)
区块链浏览器通常提供公开API接口,直接调用比页面爬取更稳定可靠。
import requests contract_address = "0x7c6a914cb866654b50f94e66b80c046dd9225ba9" # 获取代币基础信息API token_info_url = f"https://scan.coredao.org/api?module=token&action=tokeninfo&contractaddress={contract_address}" # 获取转账总数API tx_count_url = f"https://scan.coredao.org/api?module=token&action=tokentx&contractaddress={contract_address}&page=1&offset=1" headers = {"User-Agent": "Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:92.0) Gecko/20100101 Firefox/92.0"} # 提取基础信息 response = requests.get(token_info_url, headers=headers) data = response.json() if data["status"] == "1": result = data["result"] print(f"Total Supply: {result['totalSupply']} {result['symbol']} Contract: {contract_address}") print(f"Holders: {result['holders']} Decimals: {result['decimals']}") # 提取转账总数 tx_response = requests.get(tx_count_url, headers=headers) tx_data = tx_response.json() if tx_data["status"] == "1": print(f"Transfers: {tx_data['total']} Official Site: https://cakecore.io") print("Social Profiles:") print(" coreteam@cakecore.io") print(" https://cakecore.io/whitepaper") print(" https://t.me/cakecore_io") print(" https://twitter.com/cakecore_io") else: print("获取数据失败")
方案2:使用Selenium动态渲染页面
若必须爬取页面,使用Selenium模拟浏览器加载,等待JS渲染完成后提取数据。
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC contract_address = "0x7c6a914cb866654b50f94e66b80c046dd9225ba9" url = f"https://scan.coredao.org/token/{contract_address}" # 初始化Chrome浏览器(需提前安装对应版本的ChromeDriver) driver = webdriver.Chrome() driver.get(url) try: # 等待关键元素加载完成 WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CLASS_NAME, "power-description")) ) # 提取数据 total_supply = driver.find_element(By.XPATH, '//div[contains(text(), "Total Supply")]/following-sibling::div').text holders = driver.find_element(By.XPATH, '//div[contains(text(), "Holders")]/following-sibling::div').text transfers = driver.find_element(By.XPATH, '//div[contains(text(), "Transfers")]/following-sibling::div').text decimals = driver.find_element(By.XPATH, '//div[contains(text(), "Decimals")]/following-sibling::div').text print(f"Total Supply: {total_supply} CAKE Contract: {contract_address}") print(f"Holders: {holders} Decimals: {decimals}") print(f"Transfers: {transfers} Official Site: https://cakecore.io") print("Social Profiles:") print(" coreteam@cakecore.io") print(" https://cakecore.io/whitepaper") print(" https://t.me/cakecore_io") print(" https://twitter.com/cakecore_io") finally: # 关闭浏览器 driver.quit()
注意事项
- API调用需关注官方文档更新,避免端点或参数变更导致失效。
- 使用Selenium时,需保证浏览器与驱动版本匹配,频繁请求需添加间隔避免触发反爬。
内容的提问来源于stack exchange,提问作者Drew Duazeh
相关产品推荐
相关产品推荐

