You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy爬取realtor.com无法翻页:CSS选择器获取下一页href失败求助

解决Realtor.com Scrapy爬虫翻页失败的核心问题

核心失败原因

  • 动态类名导致选择器失效
    你代码里依赖的jsx-1709448077、styles__StyledPaginator-rui__sc-1vqyfdo-0这类类名是React框架动态生成的,网站每次更新部署后这类类名会随机变化,直接用它们写CSS选择器,大概率找不到目标分页元素,这是翻页失败的最核心原因。

  • 未处理选择器为空的异常
    你直接通过.attrib['href']获取链接,一旦选择器没匹配到元素,会直接抛出KeyError导致爬虫中断,而非跳过或容错处理。

  • User-Agent设置错误
    你在Spider类里定义的user_agent属性不会被Scrapy自动识别使用,默认情况下爬虫会用Scrapy自带的UA,容易被网站反爬机制识别,返回的页面可能缺失分页组件。

  • 可能存在JS动态渲染问题
    Realtor.com部分内容依赖JavaScript动态加载,Scrapy默认下载器不执行JS,会导致response中没有完整的分页HTML元素。

修复步骤与优化代码

1. 改用稳定的选择器定位下一页

放弃依赖动态类名,改用页面中稳定的属性(如aria-label)或语义化文本定位:

# 用aria-label定位下一页链接
next_page = response.css('a[aria-label="Go to next page"]::attr(href)').get()
# 或用文本定位备选:
# next_page = response.xpath('//a[contains(text(), "Next")]/@href').get()

2. 正确配置User-Agent

通过custom_settings在Spider类中配置UA,确保请求使用浏览器级别的UA:

class RealtorScrape(scrapy.Spider):
    # ... 其他属性 ...
    custom_settings = {
        'USER_AGENT': "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/106.0.0.0 Safari/537.36"
    }

3. 容错处理选择器为空的情况

使用.get()方法安全获取链接,避免直接访问attrib触发异常:

next_page = response.css('a[aria-label="Go to next page"]::attr(href)').get()
if next_page:
    yield response.follow(next_page, callback=self.parse)

4. 应对JS动态渲染(可选)

如果确认页面依赖JS加载,可使用scrapy-splash或playwright扩展渲染页面,获取完整HTML内容。

完整优化代码示例

import scrapy

class RealtorScrape(scrapy.Spider):
    name = 'realtor'
    allowed_domains = ['realtor.com']
    start_urls = ['https://www.realtor.com/realestateandhomes-search/Minneapolis_MN/']
    custom_settings = {
        'USER_AGENT': "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/106.0.0.0 Safari/537.36"
    }

    def parse(self, response):      
        # 改用稳定的data-testid属性定位房源卡片
        for house in response.css('li[data-testid="result-card"]'):
            status = house.css('div[data-testid="card-status"]::text').get()
            if status == 'For Sale':
                yield {
                    'Status': status,
                    'Price': house.css('div[data-testid="card-price"]::text').get(),
                    'Beds': ' '.join(house.css('li[data-testid="property-meta-beds"] span::text').getall()),
                    'Baths': ' '.join(house.css('li[data-testid="property-meta-baths"] span::text').getall()),
                    'Square_feet': ' '.join(house.css('li[data-testid="property-meta-sqft"] span::text').getall()),
                    'Accre_lot': ' '.join(house.css('li[data-testid="property-meta-lot-size"] span::text').getall()),
                    'Location': house.css('div[data-testid="card-address"]::text').get()
                }

        # 稳定的下一页选择器
        next_page = response.css('a[aria-label="Go to next page"]::attr(href)').get()
        if next_page:
            yield response.follow(next_page, callback=self.parse)

内容的提问来源于stack exchange,提问作者Kamal Moha

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 22:16:30