Scrapy Link Extractor仅抓取部分房源链接问题排查与修复
问题:Funda爬虫遗漏房源的原因分析与修复方案
我使用Scrapy 2.6.2抓取Funda网站阿姆斯特丹已售房源页面(https://www.funda.nl/koop/amsterdam/verkocht/sorteer-afmelddatum-af/)的第1至617页所有房源链接,但结果仅获取到部分房源,例如第209页的"Eef Kamerbeekstraat 504 + PP"就被遗漏了。以下是我的爬虫脚本:
import scrapy from scrapy.linkextractors import LinkExtractor from scrapy.spiders import CrawlSpider, Rule class FundaSpider(CrawlSpider): name = 'funda_verkocht' allowed_domains = ['funda.nl'] start_urls = [] user_agent = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/76.0.3809.100 Safari/537.36' def start_requests(self): yield scrapy.Request(url='https://www.funda.nl/koop/amsterdam/verkocht/sorteer-afmelddatum-af/', headers={ 'User-Agent': self.user_agent }) rules = ( Rule(LinkExtractor(restrict_xpaths="//a[@data-object-url-tracking='resultlist']"), callback='parse_item', follow=True), Rule(LinkExtractor(restrict_xpaths="//a[@rel='next']"), follow = True) ) def set_user_agent(self, request): request.headers['User-Agent'] = self.user_agent return request def parse_item(self, response): yield{ 'address': response.xpath("normalize-space(//span[@class='object-header__title']/text())").get(), 'postal_code': response.xpath("normalize-space(//span[@class='object-header__subtitle fd-color-dark-3']/text())").get(), 'offered_since': response.xpath("normalize-space(//dt[.='Aangeboden sinds']/following-sibling::dd[1]/span[1]/text())").get(), 'asking-price': response.xpath("normalize-space(//div/strong[@class='object-header__price--historic']/text())").get(), 'surface': response.xpath("normalize-space(//dt[.='Wonen']/following-sibling::dd[1]/span/text())").get(), 'energy_label': response.xpath("normalize-space(//dt[.='Energielabel']/following-sibling::dd[1]/span[1]/text())").get(), 'housing_type': response.xpath("normalize-space(//dt[.='Soort appartement']/following-sibling::dd[1]/span[1]/text())").get(), 'build_year': response.xpath("normalize-space(//dt[.='Bouwjaar']/following-sibling::dd[1]/span[1]/text())").get(), 'number_rooms': response.xpath("normalize-space(//dt[.='Aantal kamers']/following-sibling::dd[1]/span[1]/text())").get(), 'number_bathrooms': response.xpath("normalize-space(//dt[.='Aantal badkamers']/following-sibling::dd[1]/span[1]/text())").get(), 'bathroom_facilities': response.xpath("normalize-space(//dt[.='Badkamervoorzieningen']/following-sibling::dd[1]/span[1]/text())").get(), 'total_floors': response.xpath("normalize-space(//dt[.='Aantal woonlagen']/following-sibling::dd[1]/span[1]/text())").get(), 'isolation': response.xpath("normalize-space(//dt[.='Isolatie']/following-sibling::dd[1]/span[1]/text())").get(), 'heating': response.xpath("normalize-space(//dt[.='Verwarming']/following-sibling::dd[1]/span[1]/text())").get(), 'warm_water': response.xpath("normalize-space(//dt[.='Warm water']/following-sibling::dd[1]/span[1]/text())").get(), 'cv_ketel': response.xpath("normalize-space(//dt[.='Cv-ketel']/following-sibling::dd[1]/span[1]/text())").get(), 'land_ownership': response.xpath("normalize-space(//dt[.='Eigendomssituatie']/following-sibling::dd[1]/span[1]/text())").get(), 'erfpacht': response.xpath("normalize-space(//dt[.='Lasten']/following-sibling::dd[1]/span[1]/text())").get(), 'location': response.xpath("normalize-space(//dt[.='Ligging']/following-sibling::dd[1]/span[1]/text())").get(), 'balcony_terrace': response.xpath("normalize-space(//dt[.='Balkon/dakterras']/following-sibling::dd[1]/span[1]/text())").get(), 'parking': response.xpath("normalize-space(//dt[.='Soort parkeergelegenheid']/following-sibling::dd[1]/span[1]/text())").get(), 'garden': response.xpath("normalize-space(//dt[.='Tuin']/following-sibling::dd[1]/span[1]/text())").get(), 'inhabitants_neighborhood': response.xpath("normalize-space(//div[.='Inwoners']/following-sibling::div[1]//text())").get(), 'families_with_kids_perc': response.xpath("normalize-space(//div[.='Gezin met kinderen']/following-sibling::div[1]//text())").get(), 'neighborhood_price_sqm': response.xpath("normalize-space(//div[.='Gem. vraagprijs / m²']/following-sibling::div[1]//text())").get() }
原因分析
- User-Agent未全局生效:你定义了
set_user_agent方法但未启用,仅start_requests里的初始请求携带了UA,Rule生成的翻页、房源详情请求会使用Scrapy默认UA,可能被网站识别为爬虫,返回不完整内容或限制访问。 - 房源链接选择器容错性差:仅依赖
data-object-url-tracking='resultlist'单一属性定位链接,部分房源可能该属性缺失或存在变体(如动态加载的房源),导致LinkExtractor漏抓。 - 翻页逻辑依赖不稳定:依赖
next按钮的rel属性翻页,若某页按钮加载异常或属性变更,会直接中断后续页面抓取。 - 反爬机制触发:高频请求可能触发网站临时封禁,导致部分页面返回空内容或不完整的房源列表。
修复方案
1. 全局配置合法User-Agent
在Spider中通过custom_settings全局配置UA,确保所有请求都携带合法标识,替代未启用的set_user_agent方法:
custom_settings = { 'USER_AGENT': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36', 'DOWNLOAD_DELAY': 2, # 添加下载延迟,降低反爬风险 'CONCURRENT_REQUESTS': 4 # 限制并发请求数 }
2. 优化房源链接选择器
放宽定位条件,结合房源卡片容器+链接特征双重匹配,提高容错率:
Rule(LinkExtractor(restrict_xpaths="//div[contains(@class, 'search-result__content')]/a[contains(@href, '/koop/') and contains(@href, '/verkocht/')]"), callback='parse_item', follow=True),
3. 替换翻页逻辑为直接构造URL
Funda的分页URL格式固定为https://www.funda.nl/koop/amsterdam/verkocht/sorteer-afmelddatum-af/p<页码>/,直接遍历页码生成请求,避免依赖不稳定的翻页按钮:
def start_requests(self): base_url = "https://www.funda.nl/koop/amsterdam/verkocht/sorteer-afmelddatum-af/p{}/" for page_num in range(1, 618): yield scrapy.Request(url=base_url.format(page_num))
修改后的完整爬虫脚本
import scrapy from scrapy.linkextractors import LinkExtractor from scrapy.spiders import CrawlSpider, Rule class FundaSpider(CrawlSpider): name = 'funda_verkocht' allowed_domains = ['funda.nl'] user_agent = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' custom_settings = { 'USER_AGENT': user_agent, 'DOWNLOAD_DELAY': 2, 'CONCURRENT_REQUESTS': 4 } def start_requests(self): base_url = "https://www.funda.nl/koop/amsterdam/verkocht/sorteer-afmelddatum-af/p{}/" for page_num in range(1, 618): yield scrapy.Request(url=base_url.format(page_num)) rules = ( Rule(LinkExtractor(restrict_xpaths="//div[contains(@class, 'search-result__content')]/a[contains(@href, '/koop/') and contains(@href, '/verkocht/')]"), callback='parse_item', follow=True), ) def parse_item(self, response): yield{ 'address': response.xpath("normalize-space(//span[@class='object-header__title']/text())").get(), 'postal_code': response.xpath("normalize-space(//span[@class='object-header__subtitle fd-color-dark-3']/text())").get(), 'offered_since': response.xpath("normalize-space(//dt[.='Aangeboden sinds']/following-sibling::dd[1]/span[1]/text())").get(), 'asking-price': response.xpath("normalize-space(//div/strong[@class='object-header__price--historic']/text())").get(), 'surface': response.xpath("normalize-space(//dt[.='Wonen']/following-sibling::dd[1]/span/text())").get(), 'energy_label': response.xpath("normalize-space(//dt[.='Energielabel']/following-sibling::dd[1]/span[1]/text())").get(), 'housing_type': response.xpath("normalize-space(//dt[.='Soort appartement']/following-sibling::dd[1]/span[1]/text())").get(), 'build_year': response.xpath("normalize-space(//dt[.='Bouwjaar']/following-sibling::dd[1]/span[1]/text())").get(), 'number_rooms': response.xpath("normalize-space(//dt[.='Aantal kamers']/following-sibling::dd[1]/span[1]/text())").get(), 'number_bathrooms': response.xpath("normalize-space(//dt[.='Aantal badkamers']/following-sibling::dd[1]/span[1]/text())").get(), 'bathroom_facilities': response.xpath("normalize-space(//dt[.='Badkamervoorzieningen']/following-sibling::dd[1]/span[1]/text())").get(), 'total_floors': response.xpath("normalize-space(//dt[.='Aantal woonlagen']/following-sibling::dd[1]/span[1]/text())").get(), 'isolation': response.xpath("normalize-space(//dt[.='Isolatie']/following-sibling::dd[1]/span[1]/text())").get(), 'heating': response.xpath("normalize-space(//dt[.='Verwarming']/following-sibling::dd[1]/span[1]/text())").get(), 'warm_water': response.xpath("normalize-space(//dt[.='Warm water']/following-sibling::dd[1]/span[1]/text())").get(), 'cv_ketel': response.xpath("normalize-space(//dt[.='Cv-ketel']/following-sibling::dd[1]/span[1]/text())").get(), 'land_ownership': response.xpath("normalize-space(//dt[.='Eigendomssituatie']/following-sibling::dd[1]/span[1]/text())").get(), 'erfpacht': response.xpath("normalize-space(//dt[.='Lasten']/following-sibling::dd[1]/span[1]/text())").get(), 'location': response.xpath("normalize-space(//dt[.='Ligging']/following-sibling::dd[1]/span[1]/text())").get(), 'balcony_terrace': response.xpath("normalize-space(//dt[.='Balkon/dakterras']/following-sibling::dd[1]/span[1]/text())").get(), 'parking': response.xpath("normalize-space(//dt[.='Soort parkeergelegenheid']/following-sibling::dd[1]/span[1]/text())").get(), 'garden': response.xpath("normalize-space(//dt[.='Tuin']/following-sibling::dd[1]/span[1]/text())").get(), 'inhabitants_neighborhood': response.xpath("normalize-space(//div[.='Inwoners']/following-sibling::div[1]//text())").get(), 'families_with_kids_perc': response.xpath("normalize-space(//div[.='Gezin met kinderen']/following-sibling::div[1]//text())").get(), 'neighborhood_price_sqm': response.xpath("normalize-space(//div[.='Gem. vraagprijs / m²']/following-sibling::div[1]//text())").get() }
内容的提问来源于stack exchange,提问作者Lisa Herzog
相关产品推荐
相关产品推荐

