Scrapy分页链接追踪失效问题求助
问题分析与解决方案
核心问题拆解
- 分页链接提取错误:
原代码用li.pagination__item a匹配所有分页链接,会命中上一页、页码、下一页等多个元素,取第一个匹配项的href必然出错。该网站的下一页按钮专属类是pagination__item--next,需要精准定位。 - 分页逻辑位置错误:
将下一页请求逻辑放在房源遍历循环内部,导致每处理一个房源就触发一次下一页请求,既造成重复请求,还干扰了第一页房源的正常抓取流程,最终导致第一页内容未被完整爬取。 - 房源链接拼接错误:
原base_url包含/apartments/amsterdam,而房源的href本身就是/apartments/amsterdam/xxx格式,拼接后会生成重复路径的无效链接。 - 手动sleep阻塞异步流程:
Scrapy是异步框架,手动添加sleep(1)会破坏并发逻辑,降低爬取效率,应该用框架内置的延迟配置替代。
修正后的代码
import scrapy class ParariusScraper(scrapy.Spider): name = 'pararius' start_urls = ['https://www.pararius.com/apartments/amsterdam/'] # 用框架内置配置控制请求间隔,替代手动sleep custom_settings = { 'DOWNLOAD_DELAY': 1, } def parse(self, response): base_url = 'https://www.pararius.com' # 遍历当前页所有房源 for section in response.css('section.listing-search-item'): yield { 'Title': section.css('h2.listing-search-item__title > a::text').get().strip(), 'Location': section.css('div.listing-search-item__sub-title::text').get().strip(), 'Price': section.css('div.listing-search-item__price::text').get().strip(), 'Size': section.css('li.illustrated-features__item::text').get().strip(), 'Link': f"{base_url}{section.css('h2.listing-search-item__title a').attrib['href']}" } # 提取下一页链接,放在房源遍历循环外部,仅执行一次 next_page = response.css('li.pagination__item--next a::attr(href)').get() if next_page: # response.follow自动处理相对路径,无需手动拼接 yield response.follow(next_page, self.parse)
关键修正说明
- 精准定位下一页:通过
li.pagination__item--next a::attr(href)直接提取下一页的相对路径,避免匹配错误链接。 - 分页逻辑移至循环外:确保当前页所有房源处理完成后,再发起下一页请求,避免流程干扰。
- 修正房源链接拼接:将
base_url简化为域名,避免生成重复路径的无效链接。 - 替换手动sleep:使用Scrapy的
DOWNLOAD_DELAY配置控制请求间隔,保留框架异步特性。
内容的提问来源于stack exchange,提问作者Kamal Moha
相关产品推荐
相关产品推荐

