You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Scrapy中绕过Yahoo财经新闻的数据授权同意墙?

解决Scrapy爬取Yahoo财经新闻时的同意墙问题

我来帮你搞定这个头疼的同意墙问题——这类Cookie授权验证在欧美站点里很常见,咱们一步步来解决:

方法1:手动携带已同意的Cookie(快速临时方案)

首先在浏览器里完成一次同意操作:打开Yahoo财经新闻页面,点击同意按钮后,按F12打开开发者工具,切换到Application标签,找到finance.yahoo.com下的Cookie(比如B、GUCS这两个核心值)。

然后修改你的爬虫,把这些Cookie直接带入请求:

class YfinNewsSpider(scrapy.Spider):
    name = 'yfin_news_spider'
    custom_settings = {
        'DOWNLOAD_DELAY': 1.5,  # 加大延迟,避免反爬
        'COOKIES_ENABLED': True
    }
    
    def __init__(self, month, year, **kwargs):
        self.start_urls = ['https://finance.yahoo.com/sitemap/2020_03_all']
        self.allowed_domains = ['finance.yahoo.com']
        # 从浏览器复制的Cookie值
        self.consent_cookies = {
            'B': '你的B值',
            'GUCS': '你的GUCS值'
        }
        super().__init__(**kwargs)
    
    def start_requests(self):
        # 给初始请求带上Cookie和浏览器UA
        headers = {
            'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
        }
        for url in self.start_urls:
            yield scrapy.Request(
                url,
                headers=headers,
                cookies=self.consent_cookies,
                callback=self.parse
            )
    
    # 剩下的parse、parse_news方法保留原有逻辑,注意把相对URL补全
    def parse(self, response):
        all_news_urls = response.xpath('//ul/li[@class="List(n) Py(3px) Lh(1.2)"]')
        for news in all_news_urls:
            news_url = news.xpath('.//a[@class="Td(n) Td(u):h C($c-fuji-grey-k)"]/@href').extract_first()
            # 补全相对路径
            if news_url.startswith('/'):
                news_url = f'https://finance.yahoo.com{news_url}'
            yield scrapy.Request(news_url, callback=self.parse_news, dont_filter=True)
    
    def parse_news(self, response):
        news_url = str(response.url)
        title = response.xpath('//title/text()').extract_first()
        paragraphs = response.xpath('//div[@class="caas-body"]/p/text()').extract()
        date_time = response.xpath('//div[@class="caas-attr-time-style"]/time/@datetime').extract_first()
        yield {'title': title, 'url': news_url, 'body_text': paragraphs, 'timestamp': date_time}

⚠️ 注意:Cookie会过期,过段时间需要重新从浏览器复制更新。

方法2:自动处理同意墙请求(纯Scrapy方案)

让爬虫自动识别同意墙页面并提交同意请求,无需手动复制Cookie:

def parse(self, response):
    # 先判断是否是同意墙页面
    if 'collectConsent' in response.url:
        # 定位同意表单的提交地址和参数
        consent_form = response.xpath('//form[@id="consent-form"]')
        if consent_form:
            consent_url = consent_form.xpath('@action').get()
            # 构造同意请求(参数根据实际表单调整)
            yield scrapy.FormRequest(
                url=consent_url,
                formdata={'agree': 'agree'},
                headers={'Referer': response.url},
                callback=self.parse
            )
        return
    
    # 原有解析逻辑
    all_news_urls = response.xpath('//ul/li[@class="List(n) Py(3px) Lh(1.2)"]')
    for news in all_news_urls:
        news_url = news.xpath('.//a[@class="Td(n) Td(u):h C($c-fuji-grey-k)"]/@href').extract_first()
        if news_url.startswith('/'):
            news_url = f'https://finance.yahoo.com{news_url}'
        yield scrapy.Request(news_url, callback=self.parse_news, dont_filter=True)

这个方案适合静态同意墙,如果遇到JS渲染的动态同意按钮,可能需要更复杂的处理。

方法3:Scrapy + Playwright(最稳定的长期方案)

如果上面两种方法都失效,推荐用Playwright模拟真实浏览器行为,彻底绕过JS渲染和同意墙:

  1. 先安装依赖:
pip install scrapy-playwright
  1. 在项目的settings.py里配置Playwright:
DOWNLOAD_HANDLERS = {
    "http": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler",
    "https": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler",
}
PLAYWRIGHT_LAUNCH_OPTIONS = {
    "headless": True,  # 设为False可以看到浏览器操作过程
    "args": ["--no-sandbox"],
}
  1. 修改爬虫代码:
class YfinNewsSpider(scrapy.Spider):
    name = 'yfin_news_spider'
    custom_settings = {
        'DOWNLOAD_DELAY': 2,
        'COOKIES_ENABLED': True
    }
    
    def __init__(self, month, year, **kwargs):
        self.start_urls = ['https://finance.yahoo.com/sitemap/2020_03_all']
        self.allowed_domains = ['finance.yahoo.com']
        super().__init__(**kwargs)
    
    def parse(self, response):
        all_news_urls = response.xpath('//ul/li[@class="List(n) Py(3px) Lh(1.2)"]')
        for news in all_news_urls:
            news_url = news.xpath('.//a[@class="Td(n) Td(u):h C($c-fuji-grey-k)"]/@href').extract_first()
            if news_url.startswith('/'):
                news_url = f'https://finance.yahoo.com{news_url}'
            # 给新闻请求加上Playwright参数
            yield scrapy.Request(
                news_url,
                callback=self.parse_news,
                meta={"playwright": True, "playwright_include_page": True},
                dont_filter=True
            )
    
    async def parse_news(self, response):
        page = response.meta["playwright_page"]
        # 检查并点击同意按钮
        consent_button = await page.query_selector('button:has-text("Accept")')
        if consent_button:
            await consent_button.click()
            await page.wait_for_load_state("networkidle")
        
        # 用Playwright获取页面内容
        title = await page.title()
        paragraphs = await page.query_selector_all('div.caas-body p')
        body_text = [await p.text_content() for p in paragraphs]
        date_time_elem = await page.query_selector('div.caas-attr-time-style time')
        timestamp = await date_time_elem.get_attribute('datetime') if date_time_elem else None
        
        await page.close()
        
        yield {
            'title': title,
            'url': response.url,
            'body_text': body_text,
            'timestamp': timestamp
        }

这种方法完全模拟真实用户操作,几乎能绕过所有反爬和验证机制,是长期爬取的最优解。

内容的提问来源于stack exchange,提问作者Nuttapong Mekvipad

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 07:48:35