Scrapy增量爬取多页仅返回首页问题求助
Scrapy爬虫仅爬取首页数据的修复建议
我看了你的代码,主要有两个关键问题导致爬虫只返回首页数据,下面一步步给你解决:
1. allowed_domains 配置错误
你的allowed_domains设置成了完整的URL,但Scrapy要求这里只填域名,不能带https://和后续路径。错误的配置会导致后续翻页请求被框架自动过滤掉。
修复方式:
allowed_domains = ['daft.ie']
2. 类变量管理页码的线程安全问题
你直接修改DaftieSpiderSpider.page_number这个类属性,但Scrapy是多线程异步运行的,多个请求同时修改这个值会导致页码混乱,甚至出现跳过页面、重复请求的情况。正确的做法是用response.meta来传递当前页码,保证每个请求的页码独立可控。
修改后的完整代码
import scrapy class DaftieSpiderSpider(scrapy.Spider): name = 'daftie_spider' allowed_domains = ['daft.ie'] start_urls = ['https://www.daft.ie/dublin-city/property-for-sale/dublin-4/'] def parse(self, response): # 处理当前页面的房源数据 listings = response.xpath('//div[@class="PropertyCardContainer__container"]') for listing in listings: price = listing.xpath('.//a/strong[@class="PropertyInformationCommonStyles__costAmountCopy"]/text()').extract_first() address = listing.xpath('.//*[@class="PropertyInformationCommonStyles__addressCopy--link"]/text()').extract_first() bedrooms = listing.xpath('.//*[@class="QuickPropertyDetails__iconCopy"]/text()').extract_first() bathrooms = listing.xpath('.//*[@class="QuickPropertyDetails__iconCopy--WithBorder"]/text()').extract_first() prop_type = listing.xpath('.//*[@class="QuickPropertyDetails__propertyType"]/text()').extract_first() agent = listing.xpath('.//div[@class="BrandedHeader__agentLogoContainer"]/img/@alt').extract_first() yield { 'price': price, 'address': address, 'bedrooms': bedrooms, 'bathrooms': bathrooms, 'prop_type': prop_type, 'agent': agent } # 获取当前偏移量:从meta中取,默认0对应首页 current_offset = response.meta.get('offset', 0) next_offset = current_offset + 20 # 控制爬取到第10页(offset=180) if next_offset <= 180: next_page_url = f'https://www.daft.ie/dublin-city/property-for-sale/dublin-4/?offset={next_offset}' # 传递下一页偏移量到meta中 yield scrapy.Request( url=next_page_url, callback=self.parse, meta={'offset': next_offset} )
额外优化建议
- 用
f-string拼接URL比字符串相加更清晰易读,我已经在代码里做了替换。 - 建议在
settings.py中添加浏览器UA模拟,避免被网站反爬拦截:USER_AGENT = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
这样修改后,爬虫就能正确遍历从首页到第10页的所有房源数据了。
内容的提问来源于stack exchange,提问作者Robert Chestnutt
相关产品推荐
相关产品推荐

