You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy项目如何将触发请求的源URL作为输出字段添加到结果中

实现方案

你要在返回结果里新增源URL字段非常简单,直接在parse方法的返回字典中新增对应字段,调用response.url即可拿到当前请求的原始URL。另外你的原有代码存在几处会导致运行失败的问题,同步给你修复了:

  • 原有方法名start_request拼写错误,Scrapy官方约定的请求入口方法名是start_requests(末尾带s)
  • allowed_domains只需填写域名本体,不能带https://协议头,否则会触发请求过滤
  • Request对象一次只能接收单个URL,不能直接传入start_urls列表,需要遍历列表逐个生成请求对象

修改后的完整代码如下:

import scrapy
from scrapy import Request

class GmapsclosedlocationsSpider(scrapy.Spider):
    name = 'gmapsclosedlocations'
    # 修复allowed_domains格式
    allowed_domains = ['google.com']
    
    with open('urls.csv') as file:
        start_urls = [line.strip() for line in file]

    # 修复方法名拼写错误
    def start_requests(self):
        # 遍历所有url逐个生成请求
        for url in self.start_urls:
            yield Request(url=url, callback=self.parse)

    def parse(self, response):
        yield {
            'name' : response.css('.qrShPb').extract(),
            'closed' : response.css('p.wlxxf::text').extract(),
            'address' : response.css('.LrzXr::text').extract(),
            'phone' : response.xpath('//*[contains(text(), "+46")]').extract_first(),
            'website' : response.css('a.ab_button::attr(href)').extract(),
            'firsttitle' : response.css('.DKV0Md::text').extract_first(),
            # 新增源URL字段
            'source_url': response.url
        }

内容的提问来源于stack exchange,提问作者djmystica

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.28 10:36:01