You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Scrapy中追踪URL的访问路径及重定向历史?

Scrapy中追踪URL访问路径的方法

Scrapy完全支持追踪URL的完整访问路径,不管是HTTP重定向过程,还是爬虫通过页面链接跳转的路径,具体实现方式如下:

1. 追踪HTTP重定向的完整路径

每个Response对象对应的请求(response.request)里,meta['redirect_urls']属性会记录所有重定向步骤的URL。这个列表包含的是每次被重定向的地址,结合初始请求URL和最终响应URL,就能得到完整的重定向链条。

示例代码:

def parse(self, response):
    # 获取重定向过程中的URL列表
    redirect_steps = response.request.meta.get('redirect_urls', [])
    # 拼接完整路径:初始请求URL → 所有重定向URL → 最终响应URL
    full_redirect_path = [response.request.url] + redirect_steps + [response.url]
    print("完整重定向路径:", full_redirect_path)

2. 追踪爬虫页面跳转的访问路径(非重定向)

如果需要记录爬虫从起始页面,通过点击页面链接一步步到达目标页面的路径,需要手动在请求的meta字段中传递路径信息,每跳转一次就更新路径:

示例代码:

def parse_start(self, response):
    # 起始页面,初始化访问路径
    next_url = "https://test.com/resource_1/"
    yield scrapy.Request(
        next_url,
        callback=self.parse_resource1,
        meta={'visit_path': [response.url]}
    )

def parse_resource1(self, response):
    # 更新访问路径,添加当前页面URL
    current_path = response.meta['visit_path'] + [response.url]
    next_url = "https://test.com/resource_1/resource_2/"
    yield scrapy.Request(
        next_url,
        callback=self.parse_resource2,
        meta={'visit_path': current_path}
    )

def parse_resource2(self, response):
    # 得到完整的页面跳转路径
    full_visit_path = response.meta['visit_path'] + [response.url]
    print("页面跳转路径:", full_visit_path)

内容的提问来源于stack exchange,提问作者Ilias Koritsas

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 20:45:14