如何使用Python Playwright(或替代方案)高效抓取网站数据至CSV文件并规避访问错误
Hey there! Let's tackle your scraping problems step by step—we'll speed up your Playwright code, beat those access blocks, and even look at a way better alternative for CoinGecko specifically (since that's the site you're targeting).
一、优化现有Playwright代码(解决速度慢问题)
Your current setup is slow mainly because of headless=False (running a visible browser) and slow_mo=2000 (adding 2-second delays on every action). Let's fix that, plus add smarter waiting and resource blocking:
关键优化点:
- 启用无头模式(
headless=True):这是提升速度的最大杠杆,无头浏览器比可视化的快得多 - 移除
slow_mo参数:除非你需要调试,否则完全没必要 - 替换固定
asyncio.sleep(5)为等待目标元素加载:硬编码等待时间要么浪费时间要么不够,等表格加载完成再继续 - 禁用不必要的资源加载:比如图片、CSS,减少页面加载时间
- 模拟真实浏览器的User-Agent:避免被网站识别为爬虫
优化后的Playwright代码:
import asyncio import random from playwright.async_api import async_playwright import pandas as pd from io import StringIO URL = "https://www.coingecko.com/en/coins/1/markets/spot" async def fetch_page(page, url): print(f"Fetching: {url}") try: # 等待页面加载完成,超时10秒 await page.goto(url, wait_until="networkidle", timeout=10000) # 等待表格元素出现,确保数据加载完成 await page.wait_for_selector('table', timeout=8000) # 随机短延迟,避免请求过于规律 await asyncio.sleep(random.uniform(1, 3)) return await page.content() except Exception as e: print(f"Failed to fetch {url}: {str(e)}") return None async def scrape_all_pages(url, max_pages=10): async with async_playwright() as p: # 启用无头模式,移除slow_mo browser = await p.chromium.launch(headless=True) # 设置浏览器上下文:模拟真实UA,禁用不必要资源 context = await browser.new_context( viewport={"width": 1280, "height": 900}, user_agent="Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36", block_resources=["image", "stylesheet"] # 不加载图片和CSS,提速 ) page = await context.new_page() markets = [] for page_num in range(1, max_pages + 1): html = await fetch_page(page, f"{url}?page={page_num}") if not html: continue dfs = pd.read_html(StringIO(html)) if dfs: markets.extend(dfs) await page.close() await context.close() await browser.close() return pd.concat(markets, ignore_index=True) if markets else pd.DataFrame() def run_async(coro): try: loop = asyncio.get_running_loop() except RuntimeError: loop = None if loop and loop.is_running(): return asyncio.create_task(coro) else: return asyncio.run(coro) async def main(): max_pages = 10 df = await scrape_all_pages(URL, max_pages) df = df.dropna(how='all') print(df) # 保存到CSV df.to_csv('coin_gecko_markets.csv', index=False) print("Data saved to coin_gecko_markets.csv") run_async(main())
二、规避访问错误(403/404)
If you still hit blocks even with optimized Playwright or when using requests, try these fixes:
- 使用真实User-Agent: 如上代码所示,设置和主流浏览器一致的UA,避免被反爬规则识别
- 添加请求延迟: 用随机延迟(
random.uniform(1,3))代替固定等待,模拟人类浏览节奏 - 重试机制: 对失败的请求进行重试(比如用
tenacity库),处理临时的网络波动或封禁 - 代理IP: 如果你的IP被网站封禁,可以使用代理服务,在Playwright的
browser.new_context()中添加proxy参数
三、最优替代方案:使用CoinGecko官方API
Wait a minute—you're scraping CoinGecko, which has a free, official API that gives you all market data directly, no web scraping needed! This is way faster, more reliable, and won't get you blocked.
示例代码(用CoinGecko API):
First, install the official Python client:
pip install pycoingecko
Then scrape market data and save to CSV:
from pycoingecko import CoinGeckoAPI import pandas as pd cg = CoinGeckoAPI() def scrape_markets(coin_id='bitcoin', max_pages=10): markets = [] for page_num in range(1, max_pages + 1): # 调用API获取交易市场列表(对应你之前爬的页面数据) data = cg.get_coin_tickers(id=coin_id, page=page_num) if 'tickers' in data: markets.extend(data['tickers']) # 转换为DataFrame并保存 df = pd.DataFrame(markets) return df if __name__ == '__main__': df = scrape_markets(coin_id='bitcoin', max_pages=10) df.to_csv('coin_gecko_markets_api.csv', index=False) print("Data saved to coin_gecko_markets_api.csv")
注意:CoinGecko API有免费版请求限制(每分钟10-30次),但完全足够个人使用,且数据格式规范,无需解析HTML。
四、用requests+BeautifulSoup的替代方案
If you prefer not to use Playwright, fix the 403 errors by adding proper headers:
import requests import time import pandas as pd from io import StringIO import random URL = "https://www.coingecko.com/en/coins/1/markets/spot" HEADERS = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36', 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8' } def scrape_with_requests(max_pages=10): session = requests.Session() session.headers.update(HEADERS) markets = [] for page_num in range(1, max_pages + 1): url = f"{URL}?page={page_num}" print(f"Fetching: {url}") try: response = session.get(url) response.raise_for_status() # 抛出HTTP错误 dfs = pd.read_html(StringIO(response.text)) markets.extend(dfs) # 随机延迟 time.sleep(random.uniform(1, 2)) except Exception as e: print(f"Failed to fetch {url}: {str(e)}") continue return pd.concat(markets, ignore_index=True) if markets else pd.DataFrame() # 运行并保存 df = scrape_with_requests(max_pages=10) df.to_csv('markets_requests.csv', index=False)
备注:内容来源于stack exchange,提问作者HamidBee

