You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

爬取https://www.etf.com/channels时遭遇503错误的解决方法

解决ETF.com 503反爬限制的可行方案

从你遇到的503错误和浏览器的安全检查页面来看,该网站大概率使用了Cloudflare类的JS反爬机制,普通HTTP请求无法绕过JS验证。以下是几种可行的解决思路及代码调整方案:

1. 使用无头浏览器模拟真实访问

这类反爬需要执行JS验证,无头浏览器能完全模拟真人浏览器行为,绕过JS挑战。

Scrapy + Playwright 实现

先安装依赖:pip install scrapy-playwright,然后修改爬虫代码:

from scrapy_playwright.page import PageCoroutine

class BrandETFs(scrapy.Spider):
    name = "etfs"
    start_urls = ['https://www.etf.com/channels']

    custom_settings = {
        'DOWNLOAD_DELAY': 2,  # 增加延迟,模拟真人访问
        "CONCURRENT_REQUESTS": 1,  # 降低并发避免触发限制
        'PLAYWRIGHT_LAUNCH_OPTIONS': {
            'headless': True,
            'args': ['--no-sandbox', '--disable-dev-shm-usage'],
        },
        'DOWNLOAD_HANDLERS': {
            "http": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler",
            "https": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler",
        },
        'PLAYWRIGHT_ABORT_REQUEST': lambda request: request.resource_type in ['image', 'font', 'stylesheet'],  # 减少无关请求
    }

    def start_requests(self):
        yield scrapy.Request(
            url=self.start_urls[0],
            meta={
                'playwright': True,
                'playwright_page_coroutines': [
                    PageCoroutine('wait_for_selector', 'div.discovery-slat'),  # 等待目标元素加载完成
                ],
            },
        )

    def parse(self, response):
        test = response.css('div.discovery-slat').getall()
        yield {
            "test": test
        }

Requests + Playwright 实现

from playwright.sync_api import sync_playwright

url = 'https://www.etf.com/channels'

with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    page = browser.new_page(user_agent="Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36")
    page.goto(url, wait_until='networkidle')  # 等待网络稳定
    page.wait_for_selector('div.discovery-slat')  # 等待目标元素
    html_content = page.content()
    # 后续可解析html_content提取数据
    print(html_content)
    browser.close()

2. 修复现有代码的基础问题

你的现有代码存在两个明显错误,先修正这些再尝试:

  • Scrapy未传递自定义请求头:原代码中定义了headers但start_requests里没传入,导致自定义UA等参数无效,修改如下:
def start_requests(self):
    url = self.start_urls[0]
    yield scrapy.Request(url=url, headers=self.headers)
  • Requests使用错误请求方法:该页面是GET请求,原代码用了POST,改成GET:
r = requests.get(url, headers=headers)

3. 使用Cloudflare专用爬虫库

cloudscraper库专门处理Cloudflare的JS挑战,无需手动模拟浏览器:
安装依赖:pip install cloudscraper,代码示例:

import cloudscraper

scraper = cloudscraper.create_scraper()
url = 'https://www.etf.com/channels'
response = scraper.get(url)
# 检查响应状态
if response.status_code == 200:
    print(response.text)
else:
    print(f"请求失败,状态码:{response.status_code}")

4. 配合代理IP与请求频率控制

如果你的IP被暂时封禁,可添加代理IP并进一步降低请求频率:

  • Scrapy中在settings.py配置代理中间件:
DOWNLOADER_MIDDLEWARES = {
    'scrapy.downloadermiddlewares.httpproxy.HttpProxyMiddleware': 110,
}
  • 在请求中添加代理:
yield scrapy.Request(
    url=url,
    headers=self.headers,
    meta={'proxy': 'http://your-proxy-ip:port'}
)

内容的提问来源于stack exchange,提问作者YungOne

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 12:15:42