You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Scrapy的CrawlSpider时如何获取当前正在爬取的URL?

CrawlSpider获取当前访问页面URL的实现方案
  • 无需引入Selenium、Scrapy-Selenium等额外工具,Scrapy原生Response对象已内置URL相关属性,在CrawlSpider的规则回调函数中可直接调用使用。

常用取值场景

1. 获取最终加载页面的URL

直接调用response.url即可,该属性返回经过所有重定向、跳转后最终加载页面的完整URL,是绝大多数场景下的首选取值方式。

2. 重定向场景获取原始请求URL

如果目标页面存在3XX重定向,需要获取发起请求时的原始URL,可调用response.request.url;如果需要完整的重定向链路所有URL,可从meta中提取:response.meta.get('redirect_urls', [])。

代码示例

import scrapy
from scrapy.spiders import CrawlSpider, Rule
from scrapy.linkextractors import LinkExtractor

class DemoSpider(CrawlSpider):
    name = 'demo_spider'
    allowed_domains = ['example.com']
    start_urls = ['https://example.com/list']

    # 原有规则不用修改
    rules = (
        Rule(LinkExtractor(allow=r'/detail/\d+'), callback='parse_detail', follow=True),
    )

    def parse_detail(self, response):
        # 新增URL提取逻辑即可
        final_url = response.url
        origin_url = response.request.url
        redirect_history = response.meta.get('redirect_urls', [])

        # 原有数据提取逻辑不变,和URL一起返回
        yield {
            'product_name': response.xpath('//h2[@class="product-title"]/text()').get(),
            'price': response.xpath('//span[@class="price"]/text()').get(),
            'current_page_url': final_url,
            'origin_request_url': origin_url,
            'redirect_url_list': redirect_history
        }

注意事项

  • 上述属性在所有绑定到Rule的回调函数中都可以直接使用,无需额外配置,完全适配CrawlSpider的自动链接提取、页面跟进逻辑。
  • 不需要对现有CrawlSpider的代码结构做大幅修改,仅需在数据提取的回调逻辑中加入对应属性的取值即可。

内容的提问来源于stack exchange,提问作者Jonathan Simpson

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.05 13:27:04