You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy增量爬取多页仅返回首页问题求助

Scrapy爬虫仅爬取首页数据的修复建议

我看了你的代码,主要有两个关键问题导致爬虫只返回首页数据,下面一步步给你解决:

1. allowed_domains 配置错误

你的allowed_domains设置成了完整的URL,但Scrapy要求这里只填域名,不能带https://和后续路径。错误的配置会导致后续翻页请求被框架自动过滤掉。

修复方式:

allowed_domains = ['daft.ie']

2. 类变量管理页码的线程安全问题

你直接修改DaftieSpiderSpider.page_number这个类属性,但Scrapy是多线程异步运行的,多个请求同时修改这个值会导致页码混乱,甚至出现跳过页面、重复请求的情况。正确的做法是用response.meta来传递当前页码,保证每个请求的页码独立可控。

修改后的完整代码

import scrapy

class DaftieSpiderSpider(scrapy.Spider):
    name = 'daftie_spider'
    allowed_domains = ['daft.ie']
    start_urls = ['https://www.daft.ie/dublin-city/property-for-sale/dublin-4/']

    def parse(self, response):
        # 处理当前页面的房源数据
        listings = response.xpath('//div[@class="PropertyCardContainer__container"]')
        for listing in listings:
            price = listing.xpath('.//a/strong[@class="PropertyInformationCommonStyles__costAmountCopy"]/text()').extract_first()
            address = listing.xpath('.//*[@class="PropertyInformationCommonStyles__addressCopy--link"]/text()').extract_first()
            bedrooms = listing.xpath('.//*[@class="QuickPropertyDetails__iconCopy"]/text()').extract_first()
            bathrooms = listing.xpath('.//*[@class="QuickPropertyDetails__iconCopy--WithBorder"]/text()').extract_first()
            prop_type = listing.xpath('.//*[@class="QuickPropertyDetails__propertyType"]/text()').extract_first()
            agent = listing.xpath('.//div[@class="BrandedHeader__agentLogoContainer"]/img/@alt').extract_first()
            
            yield {
                'price': price,
                'address': address,
                'bedrooms': bedrooms,
                'bathrooms': bathrooms,
                'prop_type': prop_type,
                'agent': agent
            }

        # 获取当前偏移量:从meta中取,默认0对应首页
        current_offset = response.meta.get('offset', 0)
        next_offset = current_offset + 20

        # 控制爬取到第10页(offset=180)
        if next_offset <= 180:
            next_page_url = f'https://www.daft.ie/dublin-city/property-for-sale/dublin-4/?offset={next_offset}'
            # 传递下一页偏移量到meta中
            yield scrapy.Request(
                url=next_page_url,
                callback=self.parse,
                meta={'offset': next_offset}
            )

额外优化建议

  • 用f-string拼接URL比字符串相加更清晰易读,我已经在代码里做了替换。
  • 建议在settings.py中添加浏览器UA模拟,避免被网站反爬拦截:
    USER_AGENT = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
    

这样修改后,爬虫就能正确遍历从首页到第10页的所有房源数据了。

内容的提问来源于stack exchange,提问作者Robert Chestnutt

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 13:22:37