You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Scrapy CrawlSpider内部为所有请求配置X-Forwarded-For头

实现方案

完全可以在爬虫代码内部实现该需求,无需依赖全局settings配置,修改后的完整代码如下:

from scrapy.spiders import CrawlSpider, Rule
from scrapy.linkextractors import LinkExtractor
from scrapy.http.request import Request

class MySpider(CrawlSpider):
    name = 'spidy'
    allowed_domains = ['website.com', 'www.website.com']
    start_urls = ['http://www.website.com/']
    rules = (
        # 给Rule添加process_request参数,指定处理请求的方法
        Rule(LinkExtractor(allow=('/uk/', )), callback='parse_item', follow=True, process_request='add_x_forwarded_for'),
    )

    # 自定义给请求加X-Forwarded-For头的方法
    def add_x_forwarded_for(self, request, response=None):
        # 替换为你需要的IP值,也可以在这里写逻辑随机生成IP实现轮换
        request.headers['X-Forwarded-For'] = '192.168.1.1'
        return request

    # 重写start_requests方法,给初始请求也加上头
    def start_requests(self):
        for url in self.start_urls:
            yield self.add_x_forwarded_for(Request(url))

    def parse_item(self, response):
        print(response.url)

代码说明

  • 自定义的add_x_forwarded_for方法会处理所有Rule匹配生成的请求(包括follow参数触发的跟进爬取请求),统一添加X-Forwarded-For头,你可以在这个方法内添加随机IP逻辑,实现每次请求携带不同的伪源IP。
  • 重写start_requests时直接调用统一的加头方法,避免重复配置,后续修改头参数只需要调整add_x_forwarded_for一处即可。
  • 如果需要针对不同请求设置不同的头,也可以在上述方法中加判断逻辑按需配置。

内容的提问来源于stack exchange,提问作者Pavel Nikolaev

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.30 09:27:07