如何在Scrapy中绕过Yahoo财经新闻的数据授权同意墙?
解决Scrapy爬取Yahoo财经新闻时的同意墙问题
我来帮你搞定这个头疼的同意墙问题——这类Cookie授权验证在欧美站点里很常见,咱们一步步来解决:
方法1:手动携带已同意的Cookie(快速临时方案)
首先在浏览器里完成一次同意操作:打开Yahoo财经新闻页面,点击同意按钮后,按F12打开开发者工具,切换到Application标签,找到finance.yahoo.com下的Cookie(比如B、GUCS这两个核心值)。
然后修改你的爬虫,把这些Cookie直接带入请求:
class YfinNewsSpider(scrapy.Spider): name = 'yfin_news_spider' custom_settings = { 'DOWNLOAD_DELAY': 1.5, # 加大延迟,避免反爬 'COOKIES_ENABLED': True } def __init__(self, month, year, **kwargs): self.start_urls = ['https://finance.yahoo.com/sitemap/2020_03_all'] self.allowed_domains = ['finance.yahoo.com'] # 从浏览器复制的Cookie值 self.consent_cookies = { 'B': '你的B值', 'GUCS': '你的GUCS值' } super().__init__(**kwargs) def start_requests(self): # 给初始请求带上Cookie和浏览器UA headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' } for url in self.start_urls: yield scrapy.Request( url, headers=headers, cookies=self.consent_cookies, callback=self.parse ) # 剩下的parse、parse_news方法保留原有逻辑,注意把相对URL补全 def parse(self, response): all_news_urls = response.xpath('//ul/li[@class="List(n) Py(3px) Lh(1.2)"]') for news in all_news_urls: news_url = news.xpath('.//a[@class="Td(n) Td(u):h C($c-fuji-grey-k)"]/@href').extract_first() # 补全相对路径 if news_url.startswith('/'): news_url = f'https://finance.yahoo.com{news_url}' yield scrapy.Request(news_url, callback=self.parse_news, dont_filter=True) def parse_news(self, response): news_url = str(response.url) title = response.xpath('//title/text()').extract_first() paragraphs = response.xpath('//div[@class="caas-body"]/p/text()').extract() date_time = response.xpath('//div[@class="caas-attr-time-style"]/time/@datetime').extract_first() yield {'title': title, 'url': news_url, 'body_text': paragraphs, 'timestamp': date_time}
⚠️ 注意:Cookie会过期,过段时间需要重新从浏览器复制更新。
方法2:自动处理同意墙请求(纯Scrapy方案)
让爬虫自动识别同意墙页面并提交同意请求,无需手动复制Cookie:
def parse(self, response): # 先判断是否是同意墙页面 if 'collectConsent' in response.url: # 定位同意表单的提交地址和参数 consent_form = response.xpath('//form[@id="consent-form"]') if consent_form: consent_url = consent_form.xpath('@action').get() # 构造同意请求(参数根据实际表单调整) yield scrapy.FormRequest( url=consent_url, formdata={'agree': 'agree'}, headers={'Referer': response.url}, callback=self.parse ) return # 原有解析逻辑 all_news_urls = response.xpath('//ul/li[@class="List(n) Py(3px) Lh(1.2)"]') for news in all_news_urls: news_url = news.xpath('.//a[@class="Td(n) Td(u):h C($c-fuji-grey-k)"]/@href').extract_first() if news_url.startswith('/'): news_url = f'https://finance.yahoo.com{news_url}' yield scrapy.Request(news_url, callback=self.parse_news, dont_filter=True)
这个方案适合静态同意墙,如果遇到JS渲染的动态同意按钮,可能需要更复杂的处理。
方法3:Scrapy + Playwright(最稳定的长期方案)
如果上面两种方法都失效,推荐用Playwright模拟真实浏览器行为,彻底绕过JS渲染和同意墙:
- 先安装依赖:
pip install scrapy-playwright
- 在项目的
settings.py里配置Playwright:
DOWNLOAD_HANDLERS = { "http": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler", "https": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler", } PLAYWRIGHT_LAUNCH_OPTIONS = { "headless": True, # 设为False可以看到浏览器操作过程 "args": ["--no-sandbox"], }
- 修改爬虫代码:
class YfinNewsSpider(scrapy.Spider): name = 'yfin_news_spider' custom_settings = { 'DOWNLOAD_DELAY': 2, 'COOKIES_ENABLED': True } def __init__(self, month, year, **kwargs): self.start_urls = ['https://finance.yahoo.com/sitemap/2020_03_all'] self.allowed_domains = ['finance.yahoo.com'] super().__init__(**kwargs) def parse(self, response): all_news_urls = response.xpath('//ul/li[@class="List(n) Py(3px) Lh(1.2)"]') for news in all_news_urls: news_url = news.xpath('.//a[@class="Td(n) Td(u):h C($c-fuji-grey-k)"]/@href').extract_first() if news_url.startswith('/'): news_url = f'https://finance.yahoo.com{news_url}' # 给新闻请求加上Playwright参数 yield scrapy.Request( news_url, callback=self.parse_news, meta={"playwright": True, "playwright_include_page": True}, dont_filter=True ) async def parse_news(self, response): page = response.meta["playwright_page"] # 检查并点击同意按钮 consent_button = await page.query_selector('button:has-text("Accept")') if consent_button: await consent_button.click() await page.wait_for_load_state("networkidle") # 用Playwright获取页面内容 title = await page.title() paragraphs = await page.query_selector_all('div.caas-body p') body_text = [await p.text_content() for p in paragraphs] date_time_elem = await page.query_selector('div.caas-attr-time-style time') timestamp = await date_time_elem.get_attribute('datetime') if date_time_elem else None await page.close() yield { 'title': title, 'url': response.url, 'body_text': body_text, 'timestamp': timestamp }
这种方法完全模拟真实用户操作,几乎能绕过所有反爬和验证机制,是长期爬取的最优解。
内容的提问来源于stack exchange,提问作者Nuttapong Mekvipad
相关产品推荐
相关产品推荐

