You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Scrapy Crawl Spider中正确覆盖默认User-Agent?

问题解决:Scrapy Crawl Spider设置User-Agent报错处理

问题描述

使用Scrapy Crawl Spider模板时,尝试通过start_requests定义初始请求的User-Agent,并在Rule中指定process_request为set_user_agent方法统一后续请求的UA,却触发错误:TypeError: BestMovieSpider.set_user_agent() takes 2 positional arguments but 3 were given

错误原因

Scrapy中,Rule的process_request回调方法会自动传入三个参数:self(spider实例)、request(待处理的请求对象)、response(生成该请求的响应对象)。你的set_user_agent方法只定义了self和request两个参数,参数数量不匹配导致报错。

修复方法

方法1:修正set_user_agent方法的参数

给set_user_agent添加response参数(即使不需要使用它):

user_agent = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/110.0.0.0 Safari/537.36'

def start_requests(self):
    yield scrapy.Request(
        url="https://www.imdb.com/search/title/?genres=drama&groups=top_250&sort=user_rating",
        headers={'User-Agent': self.user_agent}
    )

rules = (
    Rule(
        LinkExtractor(restrict_xpaths='//h3[@class="lister-item-header"]/a'),
        callback="parse_item",
        follow=True,
        process_request='set_user_agent'
    ),
)

# 新增response参数
def set_user_agent(self, request, response):
    request.headers['User-Agent'] = self.user_agent
    return request

def parse_item(self, response):
    yield {
        'title': response.xpath('//div[@class="sc-b5e8e7ce-1 kNhUtn"]/h1[@class="sc-b73cd867-0 gLtJub"]/text()').get()
    }

方法2:全局统一设置User-Agent(更推荐)

不需要单独写process_request方法,直接通过全局配置统一UA,所有请求会自动使用该值:

方式A:在spider类中添加custom_settings

class BestMovieSpider(CrawlSpider):
    name = 'best_movie'
    allowed_domains = ['imdb.com']
    
    # 全局配置UA
    custom_settings = {
        'USER_AGENT': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/110.0.0.0 Safari/537.36'
    }

    start_urls = ["https://www.imdb.com/search/title/?genres=drama&groups=top_250&sort=user_rating"]

    rules = (
        Rule(
            LinkExtractor(restrict_xpaths='//h3[@class="lister-item-header"]/a'),
            callback="parse_item",
            follow=True
        ),
    )

    def parse_item(self, response):
        yield {
            'title': response.xpath('//div[@class="sc-b5e8e7ce-1 kNhUtn"]/h1[@class="sc-b73cd867-0 gLtJub"]/text()').get()
        }

方式B:在项目settings.py中设置

找到USER_AGENT配置项,替换为你的UA:

USER_AGENT = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/110.0.0.0 Safari/537.36'

内容的提问来源于stack exchange,提问作者Asib Ahmed

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 04:42:52