You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python Playwright(或替代方案)高效抓取网站数据至CSV文件并规避访问错误

如何使用Python Playwright(或替代方案)高效抓取网站数据至CSV文件并规避访问错误

Hey there! Let's tackle your scraping problems step by step—we'll speed up your Playwright code, beat those access blocks, and even look at a way better alternative for CoinGecko specifically (since that's the site you're targeting).

一、优化现有Playwright代码(解决速度慢问题)

Your current setup is slow mainly because of headless=False (running a visible browser) and slow_mo=2000 (adding 2-second delays on every action). Let's fix that, plus add smarter waiting and resource blocking:

关键优化点:

  • 启用无头模式(headless=True):这是提升速度的最大杠杆,无头浏览器比可视化的快得多
  • 移除slow_mo参数:除非你需要调试,否则完全没必要
  • 替换固定asyncio.sleep(5)为等待目标元素加载:硬编码等待时间要么浪费时间要么不够,等表格加载完成再继续
  • 禁用不必要的资源加载:比如图片、CSS,减少页面加载时间
  • 模拟真实浏览器的User-Agent:避免被网站识别为爬虫

优化后的Playwright代码:

import asyncio
import random
from playwright.async_api import async_playwright
import pandas as pd
from io import StringIO

URL = "https://www.coingecko.com/en/coins/1/markets/spot"

async def fetch_page(page, url):
    print(f"Fetching: {url}")
    try:
        # 等待页面加载完成,超时10秒
        await page.goto(url, wait_until="networkidle", timeout=10000)
        # 等待表格元素出现,确保数据加载完成
        await page.wait_for_selector('table', timeout=8000)
        # 随机短延迟,避免请求过于规律
        await asyncio.sleep(random.uniform(1, 3))
        return await page.content()
    except Exception as e:
        print(f"Failed to fetch {url}: {str(e)}")
        return None

async def scrape_all_pages(url, max_pages=10):
    async with async_playwright() as p:
        # 启用无头模式,移除slow_mo
        browser = await p.chromium.launch(headless=True)
        # 设置浏览器上下文:模拟真实UA,禁用不必要资源
        context = await browser.new_context(
            viewport={"width": 1280, "height": 900},
            user_agent="Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36",
            block_resources=["image", "stylesheet"]  # 不加载图片和CSS,提速
        )
        page = await context.new_page()

        markets = []
        for page_num in range(1, max_pages + 1):
            html = await fetch_page(page, f"{url}?page={page_num}")
            if not html:
                continue
            dfs = pd.read_html(StringIO(html))
            if dfs:
                markets.extend(dfs)

        await page.close()
        await context.close()
        await browser.close()

    return pd.concat(markets, ignore_index=True) if markets else pd.DataFrame()

def run_async(coro):
    try:
        loop = asyncio.get_running_loop()
    except RuntimeError:
        loop = None

    if loop and loop.is_running():
        return asyncio.create_task(coro)
    else:
        return asyncio.run(coro)

async def main():
    max_pages = 10
    df = await scrape_all_pages(URL, max_pages)
    df = df.dropna(how='all')
    print(df)
    # 保存到CSV
    df.to_csv('coin_gecko_markets.csv', index=False)
    print("Data saved to coin_gecko_markets.csv")

run_async(main())

二、规避访问错误(403/404)

If you still hit blocks even with optimized Playwright or when using requests, try these fixes:

  • 使用真实User-Agent: 如上代码所示,设置和主流浏览器一致的UA,避免被反爬规则识别
  • 添加请求延迟: 用随机延迟(random.uniform(1,3))代替固定等待,模拟人类浏览节奏
  • 重试机制: 对失败的请求进行重试(比如用tenacity库),处理临时的网络波动或封禁
  • 代理IP: 如果你的IP被网站封禁,可以使用代理服务,在Playwright的browser.new_context()中添加proxy参数

三、最优替代方案:使用CoinGecko官方API

Wait a minute—you're scraping CoinGecko, which has a free, official API that gives you all market data directly, no web scraping needed! This is way faster, more reliable, and won't get you blocked.

示例代码(用CoinGecko API):

First, install the official Python client:

pip install pycoingecko

Then scrape market data and save to CSV:

from pycoingecko import CoinGeckoAPI
import pandas as pd

cg = CoinGeckoAPI()

def scrape_markets(coin_id='bitcoin', max_pages=10):
    markets = []
    for page_num in range(1, max_pages + 1):
        # 调用API获取交易市场列表(对应你之前爬的页面数据)
        data = cg.get_coin_tickers(id=coin_id, page=page_num)
        if 'tickers' in data:
            markets.extend(data['tickers'])
    
    # 转换为DataFrame并保存
    df = pd.DataFrame(markets)
    return df

if __name__ == '__main__':
    df = scrape_markets(coin_id='bitcoin', max_pages=10)
    df.to_csv('coin_gecko_markets_api.csv', index=False)
    print("Data saved to coin_gecko_markets_api.csv")

注意:CoinGecko API有免费版请求限制(每分钟10-30次),但完全足够个人使用,且数据格式规范,无需解析HTML。

四、用requests+BeautifulSoup的替代方案

If you prefer not to use Playwright, fix the 403 errors by adding proper headers:

import requests
import time
import pandas as pd
from io import StringIO
import random

URL = "https://www.coingecko.com/en/coins/1/markets/spot"
HEADERS = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36',
    'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8'
}

def scrape_with_requests(max_pages=10):
    session = requests.Session()
    session.headers.update(HEADERS)
    markets = []
    
    for page_num in range(1, max_pages + 1):
        url = f"{URL}?page={page_num}"
        print(f"Fetching: {url}")
        try:
            response = session.get(url)
            response.raise_for_status()  # 抛出HTTP错误
            dfs = pd.read_html(StringIO(response.text))
            markets.extend(dfs)
            # 随机延迟
            time.sleep(random.uniform(1, 2))
        except Exception as e:
            print(f"Failed to fetch {url}: {str(e)}")
            continue
    
    return pd.concat(markets, ignore_index=True) if markets else pd.DataFrame()

# 运行并保存
df = scrape_with_requests(max_pages=10)
df.to_csv('markets_requests.csv', index=False)

备注:内容来源于stack exchange,提问作者HamidBee

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.15 03:20:22