使用Python爬取CoinGecko时如何绕过HTTP 403错误?
问题描述
尝试用Python爬取CoinGecko比特币市场板块时持续遭遇HTTP 403错误,已使用requests搭配自定义请求头模拟浏览器,但问题未解决。
所用代码
import requests import pandas as pd # Base URL for Bitcoin markets on CoinGecko base_url = "https://www.coingecko.com/en/coins/bitcoin" # Function to fetch a single page def fetch_page(url, page): headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36", "X-Requested-With": "XMLHttpRequest" } response = requests.get(f"{url}?page={page}", headers=headers) if response.status_code != 200: print(f"Failed to fetch page {page}: Status code {response.status_code}") return None return response.text # Function to extract market data from a page def extract_markets(html): dfs = pd.read_html(html) return dfs[0] if dfs else pd.DataFrame() # Main function to scrape all pages def scrape_all_pages(base_url, max_pages=10): all_markets = [] for page in range(1, max_pages + 1): print(f"Scraping page {page}...") html = fetch_page(base_url, page) if html is None: break df = extract_markets(html) if df.empty: break all_markets.append(df) return pd.concat(all_markets, ignore_index=True) if all_markets else pd.DataFrame() # Scrape data and store in a DataFrame max_pages = 10 # Adjust this to scrape more pages if needed df = scrape_all_pages(base_url, max_pages) # Display the DataFrame print(df)
错误信息
Scraping page 1... Failed to fetch page 1: Status code 403 Empty DataFrame Columns: [] Index: []
已尝试StackOverflow上的建议方案,但未解决问题,求可行解决方法或更有效的爬取方式。
可行解决方法
方法一:使用CoinGecko官方API(推荐)
CoinGecko提供免费官方API,无需爬取网页即可获取结构化数据,完全规避反爬限制。
示例代码:
import requests import pandas as pd # 官方API端点,获取比特币市场数据 api_url = "https://api.coingecko.com/api/v3/coins/bitcoin/tickers" # 发送请求 response = requests.get(api_url) if response.status_code == 200: data = response.json() # 提取核心市场字段转为DataFrame markets_df = pd.DataFrame(data['tickers']) print(markets_df[['base', 'target', 'market', 'last', 'volume']]) else: print(f"API请求失败,状态码:{response.status_code}")
- 优势:稳定合规,无需处理反爬逻辑,数据结构清晰易解析
- 注意:免费版API有请求频率限制,每分钟最多50次请求,避免高频调用
方法二:优化网页爬取的反爬规避策略
若坚持爬取网页,可尝试以下调整:
- 更新请求头:使用更贴近现代浏览器的
User-Agent,补充更多真实请求头字段 - 用Session维持连接:模拟浏览器会话,避免每次请求重建连接
- 添加请求延迟:降低爬取频率,避免触发频率限制
- 预获取Cookies:先请求首页获取站点Cookies,再请求目标页面
示例优化代码:
import requests import pandas as pd import time base_url = "https://www.coingecko.com/en/coins/bitcoin" def fetch_page(url, page): # 初始化会话 session = requests.Session() # 先请求首页获取Cookies session.get("https://www.coingecko.com", headers={ "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/128.0.0.0 Safari/537.36" }) # 补充真实请求头字段 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/128.0.0.0 Safari/537.36", "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8", "Accept-Language": "en-US,en;q=0.5", "Connection": "keep-alive", "Upgrade-Insecure-Requests": "1" } # 添加2秒延迟 time.sleep(2) response = session.get(f"{url}?page={page}", headers=headers) if response.status_code != 200: print(f"Failed to fetch page {page}: Status code {response.status_code}") return None return response.text # 后续extract_markets和scrape_all_pages函数保持不变
方法三:使用无头浏览器
若网站有JavaScript渲染或更严格反爬机制,可使用Playwright/Selenium等工具模拟真实浏览器行为:
示例(使用Playwright):
from playwright.sync_api import sync_playwright import pandas as pd def fetch_page_with_playwright(url, page_num): with sync_playwright() as p: browser = p.chromium.launch(headless=True) page = browser.new_page() # 等待网络空闲后获取页面内容 page.goto(f"{url}?page={page_num}", wait_until="networkidle") html = page.content() browser.close() return html # 调用方式替换原fetch_page即可
- 注意:无头浏览器资源消耗较大,需控制爬取频率,避免被识别为爬虫
内容的提问来源于stack exchange,提问作者HamidBee
相关产品推荐
相关产品推荐

