Scrapy中urljoin生成链接不完整/重复问题排查求助
Scrapy爬虫生成不完整/重复XML链接问题
问题场景
在Scrapy Shell中执行XPath可以正确获取预期的XML相对链接:
>>> response.xpath('//div[@class="toolbar"]/a[contains(@href, ".xml")]/@href').extract()[1] '/dhq/vol/16/3/000642.xml'
但运行爬虫输出CSV时,却生成大量不完整链接或重复根链接。
爬虫代码
import scrapy from scrapy.spiders import CrawlSpider, Rule from scrapy.linkextractors import LinkExtractor class DhqSpider(CrawlSpider): name = 'dhq' allowed_domains = ['digitalhumanities.org'] start_urls = ['http://www.digitalhumanities.org/dhq/vol/16/3/index.html'] rules = ( Rule(LinkExtractor(allow = 'index.html')), Rule(LinkExtractor(allow = 'vol'), callback='parse_xml'), ) def parse_xml(self, response): xmllinks = response.xpath('//div[@class="toolbar"]/a[contains(@href, ".xml")]/@href').extract()[1] for link in xmllinks: yield{ 'file_urls': [response.urljoin(link)] }
问题原因
- 循环逻辑错误:
extract()[1]会取出单个字符串类型的链接(比如'/dhq/vol/16/3/000642.xml'),之后对这个字符串进行for link in xmllinks循环时,会将字符串拆分为单个字符逐个遍历。response.urljoin(link)传入单个字符后,会生成根域名拼接单个字符的无效链接(如http://www.digitalhumanities.org/d),这就是不完整链接的来源。 - Rule规则匹配范围过宽:第二个Rule使用
allow='vol'会匹配所有包含vol字符的链接,可能导致无关页面被传入parse_xml方法,这类页面中找不到符合条件的XML链接,或处理后生成重复的根链接。
解决方案
1. 修复循环逻辑
- 如果需要抓取所有符合条件的XML链接:去掉
[1],直接用extract()获取链接列表,再循环处理每个链接:
def parse_xml(self, response): # 获取所有XML链接 xmllinks = response.xpath('//div[@class="toolbar"]/a[contains(@href, ".xml")]/@href').extract() for link in xmllinks: yield { 'file_urls': [response.urljoin(link)] }
- 如果确实只需要第二个XML链接:无需循环,直接处理单个链接即可:
def parse_xml(self, response): # 仅获取第二个XML链接 xmllink = response.xpath('//div[@class="toolbar"]/a[contains(@href, ".xml")]/@href').extract()[1] yield { 'file_urls': [response.urljoin(xmllink)] }
2. 优化Rule匹配规则
缩小Rule的匹配范围,避免无关页面被处理:
rules = ( # 精准匹配index.html页面,跟进其内部链接 Rule(LinkExtractor(allow=r'index\.html$'), follow=True), # 精准匹配卷号页面(如/vol/16/3/这类格式),避免匹配含vol的无关链接 Rule(LinkExtractor(allow=r'/dhq/vol/\d+/\d+/'), callback='parse_xml'), )
内容的提问来源于stack exchange,提问作者David Beales
相关产品推荐
相关产品推荐

