You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Selenium爬取Centris房源:价格与MLS编号提取问题

Centris.ca房源价格与MLS编号爬取优化方案

需求说明

需要爬取页面https://www.centris.ca/en/properties~for-sale~brossard?view=Thumbnail的两项核心信息:

  • 房源价格
  • MLS编号

原代码通过拆分文本提取价格,操作繁琐且易受页面结构变动影响;同时无法定位MLS编号元素,多次尝试均失败。

优化后的完整代码

from selenium import webdriver
from webdriver_manager.chrome import ChromeDriverManager
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import time

url = 'https://www.centris.ca/en/properties~for-sale~brossard?view=Thumbnail'

def scrap_pages(driver, wait):
    # 等待房源列表加载完成
    wait.until(EC.presence_of_all_elements_located((By.CLASS_NAME, 'description')))
    listings = driver.find_elements(By.CLASS_NAME, 'description')

    # 过滤空条目
    listings = [listing for listing in listings if listing.text.strip()]

    for listing in listings:
        # 提取价格:直接定位description下的price子元素
        try:
            price = listing.find_element(By.CLASS_NAME, 'price').text
        except:
            price = "N/A"

        # 提取MLS编号:定位当前房源下的MlsNumberNoStealth元素
        try:
            mls_container = listing.find_element(By.CLASS_NAME, 'MlsNumberNoStealth')
            mls = mls_container.find_element(By.TAG_NAME, 'p').text.strip()
        except:
            mls = "N/A"

        # 可选:保留原代码中的其他字段提取(按需调整)
        text_lines = listing.text.split('\n')
        prop_type = text_lines[1] if len(text_lines)>=2 else "N/A"
        addr = text_lines[2] if len(text_lines)>=3 else "N/A"

        listing_item = {
            'price': price,
            'MLS': mls,
            'Address': addr,
            'property Type': prop_type
        }
        centris_list.append(listing_item)
        print(listing_item)

if __name__ == '__main__':
    chrome_options = Options()
    chrome_options.add_experimental_option("detach", True)
    # chrome_options.add_argument("headless")
    # 无头模式下建议添加窗口尺寸参数:chrome_options.add_argument("--window-size=1920,1080")

    driver = webdriver.Chrome(ChromeDriverManager().install(), options=chrome_options)
    wait = WebDriverWait(driver, 10)
    centris_list = []

    driver.get(url)

    # 获取总页数:等待分页元素加载完成
    total_pages = wait.until(EC.visibility_of_element_located((By.CLASS_NAME, 'pager-current'))).text.split('/')[1].strip()

    for i in range(int(total_pages)):
        scrap_pages(driver, wait)
        # 最后一页无需点击下一页
        if i < int(total_pages)-1:
            next_btn = wait.until(EC.element_to_be_clickable((By.CSS_SELECTOR, 'li.next> a')))
            next_btn.click()
            time.sleep(1)

    driver.quit()

关键优化点说明

1. 房源价格提取优化

  • 原问题:之前的定位失败是因为未使用相对定位,且未等待元素完全加载。
  • 解决方案:
    • 在每个description元素范围内,通过By.CLASS_NAME, 'price'直接定位价格子元素,无需拆分文本,逻辑更简洁稳定。
    • 增加异常捕获,避免单个房源价格元素缺失导致程序中断。

2. MLS编号获取

  • 原问题:错误使用By.ID定位(页面中每个房源的MLS元素ID唯一,无法批量获取),且未基于当前房源做相对定位。
  • 解决方案:
    • 基于当前description元素,通过By.CLASS_NAME, 'MlsNumberNoStealth'定位到MLS编号的容器。
    • 再获取容器内的<p>标签文本,即为MLS编号,确保每个房源对应正确的编号。

额外优化建议

  • 替换固定time.sleep为WebDriverWait显式等待,提升爬取稳定性和效率。
  • 若启用无头模式,需添加窗口尺寸参数保证元素布局正常,避免元素定位失败。

内容的提问来源于stack exchange,提问作者D.Zou

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 00:45:39