You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Scrapy中用Link Extractor限制XPath?遇参数多值错误求指导

问题排查与修复

错误根源

你碰到的「multiple values for argument」错误,核心是调用Request时参数传递完全错误:

  • scrapy.Request的第二个参数是callback,但你把rules放在了这个位置,同时又显式传了callback='base_url',相当于给callback传了两个值,直接触发参数冲突。
  • 另外你混淆了普通Spider和CrawlSpider的用法:Rule和LinkExtractor是CrawlSpider专属的类属性规则,不是在parse方法里临时创建后传给Request的。

修复方案

根据你的爬取需求,提供两种可行的修复方式:

方案1:继续使用普通Spider(适合简单爬取逻辑)

直接删掉无用的rules相关代码,修正Request的调用格式:

import scrapy
from scrapy.http import Request

class BarSpider(scrapy.Spider):
    name = 'bar'
    start_urls=["https://www.veteranownedbusiness.com/?mode=geo#BrowseByState"]

    def parse(self, response):
        # 提取分类页面链接
        books = response.xpath('//table[@class="categories"]//tr//td//a[@class="category"]//@href').extract()
        for book in books:
            url = response.urljoin(book)
            # 正确调用Request,仅传递必要参数
            yield Request(url, callback='base_url')

    def base_url(self,response):
        # 提取详情页链接并输出
        links = response.xpath('//table[@class="listings"]//a//@href').extract()
        for link in links:
            b_link = response.urljoin(link)
            yield{
                'url':b_link,
            }

方案2:改用CrawlSpider(适合复杂链接规则爬取)

如果想利用LinkExtractor自动匹配链接,需要继承CrawlSpider,并把爬取规则定义为类属性:

import scrapy
from scrapy.linkextractors import LinkExtractor
from scrapy.spiders import CrawlSpider, Rule

class BarSpider(CrawlSpider):
    name = 'bar'
    start_urls=["https://www.veteranownedbusiness.com/?mode=geo#BrowseByState"]

    # 定义爬取规则:先跟进分类页,再提取详情页链接
    rules = (
        # 提取分类链接,跟进页面并交给base_url处理
        Rule(LinkExtractor(restrict_xpaths='//table[@class="categories"]//tr//td//a[@class="category"]'), callback='base_url', follow=True),
    )

    def base_url(self,response):
        links = response.xpath('//table[@class="listings"]//a//@href').extract()
        for link in links:
            b_link = response.urljoin(link)
            yield{
                'url':b_link,
            }

额外注意点

  • LinkExtractor的restrict_xpaths参数需要传入节点的XPath,而非属性的XPath(你之前写的//@href是提取属性值,应该去掉,直接写指向<a>标签的路径)。
  • 如果你需要用Selenium渲染动态页面,后续可以把普通Request替换为SeleniumRequest,但当前错误和Selenium无关,先搞定参数问题再整合即可。

内容的提问来源于stack exchange,提问作者Amen Aziz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.16 23:45:19