You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Scrapy爬取URL时遭遇403错误,请求技术解决方案

解决Scrapy请求403错误并提取目标URL

问题概述

请求页面 https://www.askgamblers.com/online-casinos/countries/ca 时返回403错误:

Ignoring response <403 https://www.askgamblers.com/online-casinos/countries/ca>: HTTP status code is not handled or not allowed

原Scrapy代码如下:

import scrapy
from scrapy.http import Request
from bs4 import BeautifulSoup
from selenium import webdriver
import time
from scrapy_selenium import SeleniumRequest

class TestSpider(scrapy.Spider):
    name = 'test'
    start_urls = ['https://www.askgamblers.com/online-casinos/countries/ca']

    def parse(self, response):
            books = response.xpath("//div[@class='card__desc']//a[starts-with(@href, '/online')]").extract()
            for book in books:
                    url = response.urljoin(book)
                    print(url)

解决方法

1. 替换默认请求为SeleniumRequest

网站反爬机制识别了Scrapy默认请求的特征,返回403。你已导入SeleniumRequest但未使用,改用它模拟真实浏览器请求,绕过检测。

2. 添加真实请求头伪装

在请求中加入主流浏览器的User-Agent,降低被识别为爬虫的概率。

3. 修正XPath提取逻辑

原代码用extract()获取的是完整<a>标签的HTML字符串,应直接提取href属性值,避免处理冗余内容。

修改后的完整代码

import scrapy
from scrapy_selenium import SeleniumRequest

class TestSpider(scrapy.Spider):
    name = 'test'

    def start_requests(self):
        yield SeleniumRequest(
            url='https://www.askgamblers.com/online-casinos/countries/ca',
            wait_time=3,  # 等待页面完全加载
            headers={
                'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
            },
            callback=self.parse
        )

    def parse(self, response):
        # 直接提取a标签的href属性值列表
        casino_hrefs = response.xpath("//div[@class='card__desc']//a[starts-with(@href, '/online')]/@href").getall()
        for href in casino_hrefs:
            full_url = response.urljoin(href)
            # 可选择将结果存入Item或直接打印
            yield {'casino_url': full_url}
            # print(full_url)

必要配置(settings.py)

确保在项目的settings.py中配置Scrapy-Selenium中间件和驱动信息:

DOWNLOADER_MIDDLEWARES = {
    'scrapy_selenium.SeleniumMiddleware': 800
}

# 根据使用的浏览器驱动配置
SELENIUM_DRIVER_NAME = 'chrome'
SELENIUM_DRIVER_EXECUTABLE_PATH = '/path/to/chromedriver'  # 替换为你的ChromeDriver实际路径
SELENIUM_DRIVER_ARGUMENTS = ['--headless=new']  # 可选,无头模式运行浏览器

内容的提问来源于stack exchange,提问作者developer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 09:50:43