You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python爬取CoinGecko时如何绕过HTTP 403错误?

问题描述

尝试用Python爬取CoinGecko比特币市场板块时持续遭遇HTTP 403错误,已使用requests搭配自定义请求头模拟浏览器,但问题未解决。

所用代码

import requests
import pandas as pd

# Base URL for Bitcoin markets on CoinGecko
base_url = "https://www.coingecko.com/en/coins/bitcoin"

# Function to fetch a single page
def fetch_page(url, page):
    headers = {
        "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36",
        "X-Requested-With": "XMLHttpRequest"
    }
    response = requests.get(f"{url}?page={page}", headers=headers)
    if response.status_code != 200:
        print(f"Failed to fetch page {page}: Status code {response.status_code}")
        return None
    return response.text

# Function to extract market data from a page
def extract_markets(html):
    dfs = pd.read_html(html)
    return dfs[0] if dfs else pd.DataFrame()

# Main function to scrape all pages
def scrape_all_pages(base_url, max_pages=10):
    all_markets = []
    for page in range(1, max_pages + 1):
        print(f"Scraping page {page}...")
        html = fetch_page(base_url, page)
        if html is None:
            break
        df = extract_markets(html)
        if df.empty:
            break
        all_markets.append(df)

    return pd.concat(all_markets, ignore_index=True) if all_markets else pd.DataFrame()

# Scrape data and store in a DataFrame
max_pages = 10  # Adjust this to scrape more pages if needed
df = scrape_all_pages(base_url, max_pages)

# Display the DataFrame
print(df)

错误信息

Scraping page 1...
Failed to fetch page 1: Status code 403
Empty DataFrame
Columns: []
Index: []

已尝试StackOverflow上的建议方案,但未解决问题,求可行解决方法或更有效的爬取方式。


可行解决方法

方法一:使用CoinGecko官方API(推荐)

CoinGecko提供免费官方API,无需爬取网页即可获取结构化数据,完全规避反爬限制。

示例代码:

import requests
import pandas as pd

# 官方API端点,获取比特币市场数据
api_url = "https://api.coingecko.com/api/v3/coins/bitcoin/tickers"

# 发送请求
response = requests.get(api_url)
if response.status_code == 200:
    data = response.json()
    # 提取核心市场字段转为DataFrame
    markets_df = pd.DataFrame(data['tickers'])
    print(markets_df[['base', 'target', 'market', 'last', 'volume']])
else:
    print(f"API请求失败,状态码:{response.status_code}")
  • 优势:稳定合规,无需处理反爬逻辑,数据结构清晰易解析
  • 注意:免费版API有请求频率限制,每分钟最多50次请求,避免高频调用

方法二:优化网页爬取的反爬规避策略

若坚持爬取网页,可尝试以下调整:

  • 更新请求头:使用更贴近现代浏览器的User-Agent,补充更多真实请求头字段
  • 用Session维持连接:模拟浏览器会话,避免每次请求重建连接
  • 添加请求延迟:降低爬取频率,避免触发频率限制
  • 预获取Cookies:先请求首页获取站点Cookies,再请求目标页面

示例优化代码:

import requests
import pandas as pd
import time

base_url = "https://www.coingecko.com/en/coins/bitcoin"

def fetch_page(url, page):
    # 初始化会话
    session = requests.Session()
    # 先请求首页获取Cookies
    session.get("https://www.coingecko.com", headers={
        "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/128.0.0.0 Safari/537.36"
    })
    # 补充真实请求头字段
    headers = {
        "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/128.0.0.0 Safari/537.36",
        "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8",
        "Accept-Language": "en-US,en;q=0.5",
        "Connection": "keep-alive",
        "Upgrade-Insecure-Requests": "1"
    }
    # 添加2秒延迟
    time.sleep(2)
    response = session.get(f"{url}?page={page}", headers=headers)
    if response.status_code != 200:
        print(f"Failed to fetch page {page}: Status code {response.status_code}")
        return None
    return response.text

# 后续extract_markets和scrape_all_pages函数保持不变

方法三:使用无头浏览器

若网站有JavaScript渲染或更严格反爬机制,可使用Playwright/Selenium等工具模拟真实浏览器行为:

示例(使用Playwright):

from playwright.sync_api import sync_playwright
import pandas as pd

def fetch_page_with_playwright(url, page_num):
    with sync_playwright() as p:
        browser = p.chromium.launch(headless=True)
        page = browser.new_page()
        # 等待网络空闲后获取页面内容
        page.goto(f"{url}?page={page_num}", wait_until="networkidle")
        html = page.content()
        browser.close()
        return html

# 调用方式替换原fetch_page即可
  • 注意:无头浏览器资源消耗较大,需控制爬取频率,避免被识别为爬虫

内容的提问来源于stack exchange,提问作者HamidBee

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.21 18:14:59