You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy中urljoin生成链接不完整/重复问题排查求助

Scrapy爬虫生成不完整/重复XML链接问题

问题场景

在Scrapy Shell中执行XPath可以正确获取预期的XML相对链接:

>>> response.xpath('//div[@class="toolbar"]/a[contains(@href, ".xml")]/@href').extract()[1]
'/dhq/vol/16/3/000642.xml'

但运行爬虫输出CSV时,却生成大量不完整链接或重复根链接。

爬虫代码

import scrapy
from scrapy.spiders import CrawlSpider, Rule
from scrapy.linkextractors import LinkExtractor

class DhqSpider(CrawlSpider):
    name = 'dhq'
    allowed_domains = ['digitalhumanities.org']
    start_urls = ['http://www.digitalhumanities.org/dhq/vol/16/3/index.html']

    rules = (
            Rule(LinkExtractor(allow = 'index.html')), 
            Rule(LinkExtractor(allow = 'vol'), callback='parse_xml'),        
        )
    
    def parse_xml(self, response):
        xmllinks = response.xpath('//div[@class="toolbar"]/a[contains(@href, ".xml")]/@href').extract()[1]
        for link in xmllinks:
                yield{
                    'file_urls': [response.urljoin(link)]
                }

问题原因

  1. 循环逻辑错误:extract()[1]会取出单个字符串类型的链接(比如'/dhq/vol/16/3/000642.xml'),之后对这个字符串进行for link in xmllinks循环时,会将字符串拆分为单个字符逐个遍历。response.urljoin(link)传入单个字符后,会生成根域名拼接单个字符的无效链接(如http://www.digitalhumanities.org/d),这就是不完整链接的来源。
  2. Rule规则匹配范围过宽:第二个Rule使用allow='vol'会匹配所有包含vol字符的链接,可能导致无关页面被传入parse_xml方法,这类页面中找不到符合条件的XML链接,或处理后生成重复的根链接。

解决方案

1. 修复循环逻辑

  • 如果需要抓取所有符合条件的XML链接:去掉[1],直接用extract()获取链接列表,再循环处理每个链接:
def parse_xml(self, response):
    # 获取所有XML链接
    xmllinks = response.xpath('//div[@class="toolbar"]/a[contains(@href, ".xml")]/@href').extract()
    for link in xmllinks:
        yield {
            'file_urls': [response.urljoin(link)]
        }
  • 如果确实只需要第二个XML链接:无需循环,直接处理单个链接即可:
def parse_xml(self, response):
    # 仅获取第二个XML链接
    xmllink = response.xpath('//div[@class="toolbar"]/a[contains(@href, ".xml")]/@href').extract()[1]
    yield {
        'file_urls': [response.urljoin(xmllink)]
    }

2. 优化Rule匹配规则

缩小Rule的匹配范围,避免无关页面被处理:

rules = (
    # 精准匹配index.html页面,跟进其内部链接
    Rule(LinkExtractor(allow=r'index\.html$'), follow=True),
    # 精准匹配卷号页面(如/vol/16/3/这类格式),避免匹配含vol的无关链接
    Rule(LinkExtractor(allow=r'/dhq/vol/\d+/\d+/'), callback='parse_xml'),
)

内容的提问来源于stack exchange,提问作者David Beales

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 01:55:37