You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy Link Extractor仅抓取部分房源链接问题排查与修复

问题:Funda爬虫遗漏房源的原因分析与修复方案

我使用Scrapy 2.6.2抓取Funda网站阿姆斯特丹已售房源页面(https://www.funda.nl/koop/amsterdam/verkocht/sorteer-afmelddatum-af/)的第1至617页所有房源链接,但结果仅获取到部分房源,例如第209页的"Eef Kamerbeekstraat 504 + PP"就被遗漏了。以下是我的爬虫脚本:

import scrapy
from scrapy.linkextractors import LinkExtractor
from scrapy.spiders import CrawlSpider, Rule


class FundaSpider(CrawlSpider):
    name = 'funda_verkocht'
    allowed_domains = ['funda.nl']
    start_urls = []
    user_agent = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/76.0.3809.100 Safari/537.36'

    def start_requests(self):
        yield scrapy.Request(url='https://www.funda.nl/koop/amsterdam/verkocht/sorteer-afmelddatum-af/', headers={
            'User-Agent': self.user_agent
        })


    rules = (
        Rule(LinkExtractor(restrict_xpaths="//a[@data-object-url-tracking='resultlist']"), callback='parse_item', follow=True),
        Rule(LinkExtractor(restrict_xpaths="//a[@rel='next']"), follow = True)
    )

    def set_user_agent(self, request):
        request.headers['User-Agent'] = self.user_agent
        return request

    def parse_item(self, response):
        yield{
            'address': response.xpath("normalize-space(//span[@class='object-header__title']/text())").get(),
            'postal_code': response.xpath("normalize-space(//span[@class='object-header__subtitle fd-color-dark-3']/text())").get(),
            'offered_since': response.xpath("normalize-space(//dt[.='Aangeboden sinds']/following-sibling::dd[1]/span[1]/text())").get(),
            'asking-price': response.xpath("normalize-space(//div/strong[@class='object-header__price--historic']/text())").get(),
            'surface': response.xpath("normalize-space(//dt[.='Wonen']/following-sibling::dd[1]/span/text())").get(),
            'energy_label': response.xpath("normalize-space(//dt[.='Energielabel']/following-sibling::dd[1]/span[1]/text())").get(),
            'housing_type': response.xpath("normalize-space(//dt[.='Soort appartement']/following-sibling::dd[1]/span[1]/text())").get(),
            'build_year': response.xpath("normalize-space(//dt[.='Bouwjaar']/following-sibling::dd[1]/span[1]/text())").get(),
            'number_rooms': response.xpath("normalize-space(//dt[.='Aantal kamers']/following-sibling::dd[1]/span[1]/text())").get(),
            'number_bathrooms': response.xpath("normalize-space(//dt[.='Aantal badkamers']/following-sibling::dd[1]/span[1]/text())").get(),
            'bathroom_facilities': response.xpath("normalize-space(//dt[.='Badkamervoorzieningen']/following-sibling::dd[1]/span[1]/text())").get(),
            'total_floors': response.xpath("normalize-space(//dt[.='Aantal woonlagen']/following-sibling::dd[1]/span[1]/text())").get(),
            'isolation': response.xpath("normalize-space(//dt[.='Isolatie']/following-sibling::dd[1]/span[1]/text())").get(),
            'heating': response.xpath("normalize-space(//dt[.='Verwarming']/following-sibling::dd[1]/span[1]/text())").get(),
            'warm_water': response.xpath("normalize-space(//dt[.='Warm water']/following-sibling::dd[1]/span[1]/text())").get(),
            'cv_ketel': response.xpath("normalize-space(//dt[.='Cv-ketel']/following-sibling::dd[1]/span[1]/text())").get(),
            'land_ownership': response.xpath("normalize-space(//dt[.='Eigendomssituatie']/following-sibling::dd[1]/span[1]/text())").get(),
            'erfpacht': response.xpath("normalize-space(//dt[.='Lasten']/following-sibling::dd[1]/span[1]/text())").get(),
            'location': response.xpath("normalize-space(//dt[.='Ligging']/following-sibling::dd[1]/span[1]/text())").get(),
            'balcony_terrace': response.xpath("normalize-space(//dt[.='Balkon/dakterras']/following-sibling::dd[1]/span[1]/text())").get(),
            'parking': response.xpath("normalize-space(//dt[.='Soort parkeergelegenheid']/following-sibling::dd[1]/span[1]/text())").get(),
            'garden': response.xpath("normalize-space(//dt[.='Tuin']/following-sibling::dd[1]/span[1]/text())").get(),
            'inhabitants_neighborhood': response.xpath("normalize-space(//div[.='Inwoners']/following-sibling::div[1]//text())").get(),
            'families_with_kids_perc': response.xpath("normalize-space(//div[.='Gezin met kinderen']/following-sibling::div[1]//text())").get(),
            'neighborhood_price_sqm': response.xpath("normalize-space(//div[.='Gem. vraagprijs / m²']/following-sibling::div[1]//text())").get()

        }

原因分析

  • User-Agent未全局生效:你定义了set_user_agent方法但未启用,仅start_requests里的初始请求携带了UA,Rule生成的翻页、房源详情请求会使用Scrapy默认UA,可能被网站识别为爬虫,返回不完整内容或限制访问。
  • 房源链接选择器容错性差:仅依赖data-object-url-tracking='resultlist'单一属性定位链接,部分房源可能该属性缺失或存在变体(如动态加载的房源),导致LinkExtractor漏抓。
  • 翻页逻辑依赖不稳定:依赖next按钮的rel属性翻页,若某页按钮加载异常或属性变更,会直接中断后续页面抓取。
  • 反爬机制触发:高频请求可能触发网站临时封禁,导致部分页面返回空内容或不完整的房源列表。

修复方案

1. 全局配置合法User-Agent

在Spider中通过custom_settings全局配置UA,确保所有请求都携带合法标识,替代未启用的set_user_agent方法:

custom_settings = {
    'USER_AGENT': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36',
    'DOWNLOAD_DELAY': 2,  # 添加下载延迟,降低反爬风险
    'CONCURRENT_REQUESTS': 4  # 限制并发请求数
}

2. 优化房源链接选择器

放宽定位条件,结合房源卡片容器+链接特征双重匹配,提高容错率:

Rule(LinkExtractor(restrict_xpaths="//div[contains(@class, 'search-result__content')]/a[contains(@href, '/koop/') and contains(@href, '/verkocht/')]"), callback='parse_item', follow=True),

3. 替换翻页逻辑为直接构造URL

Funda的分页URL格式固定为https://www.funda.nl/koop/amsterdam/verkocht/sorteer-afmelddatum-af/p<页码>/,直接遍历页码生成请求,避免依赖不稳定的翻页按钮:

def start_requests(self):
    base_url = "https://www.funda.nl/koop/amsterdam/verkocht/sorteer-afmelddatum-af/p{}/"
    for page_num in range(1, 618):
        yield scrapy.Request(url=base_url.format(page_num))

修改后的完整爬虫脚本

import scrapy
from scrapy.linkextractors import LinkExtractor
from scrapy.spiders import CrawlSpider, Rule


class FundaSpider(CrawlSpider):
    name = 'funda_verkocht'
    allowed_domains = ['funda.nl']
    user_agent = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
    
    custom_settings = {
        'USER_AGENT': user_agent,
        'DOWNLOAD_DELAY': 2,
        'CONCURRENT_REQUESTS': 4
    }

    def start_requests(self):
        base_url = "https://www.funda.nl/koop/amsterdam/verkocht/sorteer-afmelddatum-af/p{}/"
        for page_num in range(1, 618):
            yield scrapy.Request(url=base_url.format(page_num))

    rules = (
        Rule(LinkExtractor(restrict_xpaths="//div[contains(@class, 'search-result__content')]/a[contains(@href, '/koop/') and contains(@href, '/verkocht/')]"), callback='parse_item', follow=True),
    )

    def parse_item(self, response):
        yield{
            'address': response.xpath("normalize-space(//span[@class='object-header__title']/text())").get(),
            'postal_code': response.xpath("normalize-space(//span[@class='object-header__subtitle fd-color-dark-3']/text())").get(),
            'offered_since': response.xpath("normalize-space(//dt[.='Aangeboden sinds']/following-sibling::dd[1]/span[1]/text())").get(),
            'asking-price': response.xpath("normalize-space(//div/strong[@class='object-header__price--historic']/text())").get(),
            'surface': response.xpath("normalize-space(//dt[.='Wonen']/following-sibling::dd[1]/span/text())").get(),
            'energy_label': response.xpath("normalize-space(//dt[.='Energielabel']/following-sibling::dd[1]/span[1]/text())").get(),
            'housing_type': response.xpath("normalize-space(//dt[.='Soort appartement']/following-sibling::dd[1]/span[1]/text())").get(),
            'build_year': response.xpath("normalize-space(//dt[.='Bouwjaar']/following-sibling::dd[1]/span[1]/text())").get(),
            'number_rooms': response.xpath("normalize-space(//dt[.='Aantal kamers']/following-sibling::dd[1]/span[1]/text())").get(),
            'number_bathrooms': response.xpath("normalize-space(//dt[.='Aantal badkamers']/following-sibling::dd[1]/span[1]/text())").get(),
            'bathroom_facilities': response.xpath("normalize-space(//dt[.='Badkamervoorzieningen']/following-sibling::dd[1]/span[1]/text())").get(),
            'total_floors': response.xpath("normalize-space(//dt[.='Aantal woonlagen']/following-sibling::dd[1]/span[1]/text())").get(),
            'isolation': response.xpath("normalize-space(//dt[.='Isolatie']/following-sibling::dd[1]/span[1]/text())").get(),
            'heating': response.xpath("normalize-space(//dt[.='Verwarming']/following-sibling::dd[1]/span[1]/text())").get(),
            'warm_water': response.xpath("normalize-space(//dt[.='Warm water']/following-sibling::dd[1]/span[1]/text())").get(),
            'cv_ketel': response.xpath("normalize-space(//dt[.='Cv-ketel']/following-sibling::dd[1]/span[1]/text())").get(),
            'land_ownership': response.xpath("normalize-space(//dt[.='Eigendomssituatie']/following-sibling::dd[1]/span[1]/text())").get(),
            'erfpacht': response.xpath("normalize-space(//dt[.='Lasten']/following-sibling::dd[1]/span[1]/text())").get(),
            'location': response.xpath("normalize-space(//dt[.='Ligging']/following-sibling::dd[1]/span[1]/text())").get(),
            'balcony_terrace': response.xpath("normalize-space(//dt[.='Balkon/dakterras']/following-sibling::dd[1]/span[1]/text())").get(),
            'parking': response.xpath("normalize-space(//dt[.='Soort parkeergelegenheid']/following-sibling::dd[1]/span[1]/text())").get(),
            'garden': response.xpath("normalize-space(//dt[.='Tuin']/following-sibling::dd[1]/span[1]/text())").get(),
            'inhabitants_neighborhood': response.xpath("normalize-space(//div[.='Inwoners']/following-sibling::div[1]//text())").get(),
            'families_with_kids_perc': response.xpath("normalize-space(//div[.='Gezin met kinderen']/following-sibling::div[1]//text())").get(),
            'neighborhood_price_sqm': response.xpath("normalize-space(//div[.='Gem. vraagprijs / m²']/following-sibling::div[1]//text())").get()
        }

内容的提问来源于stack exchange,提问作者Lisa Herzog

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 03:25:29