Scrapy爬取realtor.com无法翻页:CSS选择器获取下一页href失败求助
解决Realtor.com Scrapy爬虫翻页失败的核心问题
核心失败原因
动态类名导致选择器失效
你代码里依赖的jsx-1709448077、styles__StyledPaginator-rui__sc-1vqyfdo-0这类类名是React框架动态生成的,网站每次更新部署后这类类名会随机变化,直接用它们写CSS选择器,大概率找不到目标分页元素,这是翻页失败的最核心原因。未处理选择器为空的异常
你直接通过.attrib['href']获取链接,一旦选择器没匹配到元素,会直接抛出KeyError导致爬虫中断,而非跳过或容错处理。User-Agent设置错误
你在Spider类里定义的user_agent属性不会被Scrapy自动识别使用,默认情况下爬虫会用Scrapy自带的UA,容易被网站反爬机制识别,返回的页面可能缺失分页组件。可能存在JS动态渲染问题
Realtor.com部分内容依赖JavaScript动态加载,Scrapy默认下载器不执行JS,会导致response中没有完整的分页HTML元素。
修复步骤与优化代码
1. 改用稳定的选择器定位下一页
放弃依赖动态类名,改用页面中稳定的属性(如aria-label)或语义化文本定位:
# 用aria-label定位下一页链接 next_page = response.css('a[aria-label="Go to next page"]::attr(href)').get() # 或用文本定位备选: # next_page = response.xpath('//a[contains(text(), "Next")]/@href').get()
2. 正确配置User-Agent
通过custom_settings在Spider类中配置UA,确保请求使用浏览器级别的UA:
class RealtorScrape(scrapy.Spider): # ... 其他属性 ... custom_settings = { 'USER_AGENT': "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/106.0.0.0 Safari/537.36" }
3. 容错处理选择器为空的情况
使用.get()方法安全获取链接,避免直接访问attrib触发异常:
next_page = response.css('a[aria-label="Go to next page"]::attr(href)').get() if next_page: yield response.follow(next_page, callback=self.parse)
4. 应对JS动态渲染(可选)
如果确认页面依赖JS加载,可使用scrapy-splash或playwright扩展渲染页面,获取完整HTML内容。
完整优化代码示例
import scrapy class RealtorScrape(scrapy.Spider): name = 'realtor' allowed_domains = ['realtor.com'] start_urls = ['https://www.realtor.com/realestateandhomes-search/Minneapolis_MN/'] custom_settings = { 'USER_AGENT': "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/106.0.0.0 Safari/537.36" } def parse(self, response): # 改用稳定的data-testid属性定位房源卡片 for house in response.css('li[data-testid="result-card"]'): status = house.css('div[data-testid="card-status"]::text').get() if status == 'For Sale': yield { 'Status': status, 'Price': house.css('div[data-testid="card-price"]::text').get(), 'Beds': ' '.join(house.css('li[data-testid="property-meta-beds"] span::text').getall()), 'Baths': ' '.join(house.css('li[data-testid="property-meta-baths"] span::text').getall()), 'Square_feet': ' '.join(house.css('li[data-testid="property-meta-sqft"] span::text').getall()), 'Accre_lot': ' '.join(house.css('li[data-testid="property-meta-lot-size"] span::text').getall()), 'Location': house.css('div[data-testid="card-address"]::text').get() } # 稳定的下一页选择器 next_page = response.css('a[aria-label="Go to next page"]::attr(href)').get() if next_page: yield response.follow(next_page, callback=self.parse)
内容的提问来源于stack exchange,提问作者Kamal Moha
相关产品推荐
相关产品推荐

