You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何编写正确循环爬取CoinGecko历史数据(解决类名重复问题)

CoinGecko比特币历史数据爬取列值重复问题

需要爬取2021年1月1日至9月30日的比特币历史数据,目标是提取Date、Market Cap、Volume、Open、Close各列的独立列表。但遇到的问题是,Market Cap、Volume、Open、Close对应的<td>标签类名均为text-center,示例HTML结构如下:

<td class="text-center">
$161,716,193,676
</td>
<td class="text-center">
$16,571,161,476
</td>
<td class="text-center">
$1,340.02
</td>
<td class="text-center">
N/A
</td>

尝试的代码如下:

import requests

url_get = requests.get('https://www.coingecko.com/en/coins/ethereum/historical_data#panel')

from bs4 import BeautifulSoup 

soup = BeautifulSoup(url_get.content,"html.parser")

table = soup.find('div', attrs={'class':'card-block'})

row = table.find_all('th', attrs={'class':'font-semibold text-center'})
row_length = len(row)
temp = []

for i in range(0, row_length):
    Date = table.find_all('th', attrs={'class':'font-semibold text-center'})[i].text

    Market_Cap = table.find_all('td', attrs={'class':'text-center'})[i].text
    Market_Cap = Market_Cap.strip()

    Volume = table.find_all('td', attrs={'class':'text-center'})[i].text
    Volume = Market_Cap.strip()

    Open = table.find_all('td', attrs={'class':'text-center'})[i].text
    Open = Open.strip()

    Close = table.find_all('td', attrs={'class':'text-center'})[i].text
    Close = Close.strip()

    temp.append((Date,Market_Cap,Volume,Open,Close))

print(temp)

输出结果不符合预期,每行的Market Cap、Volume、Open、Close值完全相同:

[('2022-09-29',
'$161,716,193,676',
'$161,716,193,676',
'$161,716,193,676',
'$161,716,193,676'),
('2022-09-28',
'$16,571,161,476',
'$16,571,161,476',
'$16,571,161,476',
'$16,571,161,476'),
...
]

修正方案

问题出在循环逻辑上:每一行对应1个Date(th标签)和4个td标签(Market Cap、Volume、Open、Close),所有td标签按行连续排列,第i行的4个td对应的索引应为i*4、i*4+1、i*4+2、i*4+3。同时修正请求URL为比特币的目标地址。

修正后的代码:

import requests
from bs4 import BeautifulSoup 

# 修正为比特币目标数据地址
url_get = requests.get('https://www.coingecko.com/en/coins/bitcoin/historical_data/usd?start_date=2021-01-01&end_date=2021-09-30')
soup = BeautifulSoup(url_get.content,"html.parser")

table = soup.find('div', attrs={'class':'card-block'})
# 获取所有日期行(th标签)
date_rows = table.find_all('th', attrs={'class':'font-semibold text-center'})
row_length = len(date_rows)
# 获取所有数据列(td标签)
data_cells = table.find_all('td', attrs={'class':'text-center'})

temp = []

for i in range(row_length):
    # 获取当前行的日期
    date = date_rows[i].text.strip()
    # 计算当前行对应的4个td的起始索引
    start_idx = i * 4
    # 提取对应列的数据
    market_cap = data_cells[start_idx].text.strip()
    volume = data_cells[start_idx+1].text.strip()
    open_val = data_cells[start_idx+2].text.strip()
    close_val = data_cells[start_idx+3].text.strip()
    
    temp.append((date, market_cap, volume, open_val, close_val))

# 查看结果(输出前5行验证)
for item in temp[:5]:
    print(item)

说明

  1. 提前一次性获取所有date_rows和data_cells,避免循环中重复调用find_all,提升效率。
  2. 通过start_idx = i *4定位每行对应的4个td标签,确保提取的是当前行的Market Cap、Volume、Open、Close值。
  3. 修正请求URL为目标比特币数据地址,避免爬取以太坊的数据。

内容的提问来源于stack exchange,提问作者DSI Andre

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 18:15:39