使用Selenium爬取Centris房源:价格与MLS编号提取问题
Centris.ca房源价格与MLS编号爬取优化方案
需求说明
需要爬取页面https://www.centris.ca/en/properties~for-sale~brossard?view=Thumbnail的两项核心信息:
- 房源价格
- MLS编号
原代码通过拆分文本提取价格,操作繁琐且易受页面结构变动影响;同时无法定位MLS编号元素,多次尝试均失败。
优化后的完整代码
from selenium import webdriver from webdriver_manager.chrome import ChromeDriverManager from selenium.webdriver.chrome.options import Options from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import time url = 'https://www.centris.ca/en/properties~for-sale~brossard?view=Thumbnail' def scrap_pages(driver, wait): # 等待房源列表加载完成 wait.until(EC.presence_of_all_elements_located((By.CLASS_NAME, 'description'))) listings = driver.find_elements(By.CLASS_NAME, 'description') # 过滤空条目 listings = [listing for listing in listings if listing.text.strip()] for listing in listings: # 提取价格:直接定位description下的price子元素 try: price = listing.find_element(By.CLASS_NAME, 'price').text except: price = "N/A" # 提取MLS编号:定位当前房源下的MlsNumberNoStealth元素 try: mls_container = listing.find_element(By.CLASS_NAME, 'MlsNumberNoStealth') mls = mls_container.find_element(By.TAG_NAME, 'p').text.strip() except: mls = "N/A" # 可选:保留原代码中的其他字段提取(按需调整) text_lines = listing.text.split('\n') prop_type = text_lines[1] if len(text_lines)>=2 else "N/A" addr = text_lines[2] if len(text_lines)>=3 else "N/A" listing_item = { 'price': price, 'MLS': mls, 'Address': addr, 'property Type': prop_type } centris_list.append(listing_item) print(listing_item) if __name__ == '__main__': chrome_options = Options() chrome_options.add_experimental_option("detach", True) # chrome_options.add_argument("headless") # 无头模式下建议添加窗口尺寸参数:chrome_options.add_argument("--window-size=1920,1080") driver = webdriver.Chrome(ChromeDriverManager().install(), options=chrome_options) wait = WebDriverWait(driver, 10) centris_list = [] driver.get(url) # 获取总页数:等待分页元素加载完成 total_pages = wait.until(EC.visibility_of_element_located((By.CLASS_NAME, 'pager-current'))).text.split('/')[1].strip() for i in range(int(total_pages)): scrap_pages(driver, wait) # 最后一页无需点击下一页 if i < int(total_pages)-1: next_btn = wait.until(EC.element_to_be_clickable((By.CSS_SELECTOR, 'li.next> a'))) next_btn.click() time.sleep(1) driver.quit()
关键优化点说明
1. 房源价格提取优化
- 原问题:之前的定位失败是因为未使用相对定位,且未等待元素完全加载。
- 解决方案:
- 在每个
description元素范围内,通过By.CLASS_NAME, 'price'直接定位价格子元素,无需拆分文本,逻辑更简洁稳定。 - 增加异常捕获,避免单个房源价格元素缺失导致程序中断。
- 在每个
2. MLS编号获取
- 原问题:错误使用
By.ID定位(页面中每个房源的MLS元素ID唯一,无法批量获取),且未基于当前房源做相对定位。 - 解决方案:
- 基于当前
description元素,通过By.CLASS_NAME, 'MlsNumberNoStealth'定位到MLS编号的容器。 - 再获取容器内的
<p>标签文本,即为MLS编号,确保每个房源对应正确的编号。
- 基于当前
额外优化建议
- 替换固定
time.sleep为WebDriverWait显式等待,提升爬取稳定性和效率。 - 若启用无头模式,需添加窗口尺寸参数保证元素布局正常,避免元素定位失败。
内容的提问来源于stack exchange,提问作者D.Zou
相关产品推荐
相关产品推荐

