You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Selenium爬取房产网站存Excel遇问题:重复写入+无法获最新房源

房产网站爬取问题修复

问题原因

  1. 重复写入第一页数据:

    • 错误混用requests和Selenium实例:Selenium已打开并筛选最新房源页面,但后续用requests.get请求了未筛选的默认页面,解析的是默认房源数据
    • 循环逻辑错误:for page in range(1,3)中每次都遍历同一组results(第一页数据),导致重复写入
    • 翻页代码放错位置:在详情页解析时尝试找列表页的翻页按钮,完全无效
  2. 无法获取最新房源:

    • 核心错误是用requests请求了未经过筛选的页面,而非Selenium操作后已筛选为“最新房源”的页面源码

修正后的代码

import pandas as pd
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.common.by import By
from selenium.common.exceptions import NoSuchElementException
import time

# 初始化浏览器实例,复用同一个窗口
options = Options()
options.add_argument("--disable-notifications")
chrome = webdriver.Chrome('./chromedriver', options=options)
chrome.get("https://www.28hse.com/en/buy/")
chrome.maximize_window()

# 切换到最新房源筛选
time.sleep(2)
sort_btn = chrome.find_element(By.XPATH, '//*[@id="mainMenuDiv"]/form[2]/div/div[5]/div[2]/div')
sort_btn.click()
time.sleep(1)
latest_option = chrome.find_element(By.XPATH, '//*[@id="mainMenuDiv"]/form[2]/div/div[5]/div[2]/div/div[2]/*[@data-value="latest"]')
latest_option.click()
time.sleep(3)  # 等待筛选结果加载完成

df = pd.DataFrame(columns=["title", "district", "url", "phone", "address"])

while True:
    # 解析当前页面的房源列表
    results = chrome.find_elements(By.CLASS_NAME, "property_item")
    for item in results:
        # 提取标题
        title = item.find_element(By.CLASS_NAME, "property_title").text.strip()
        # 提取区域
        district = item.find_element(By.CLASS_NAME, "district_area").find_element(By.TAG_NAME, "a").text.strip()
        # 提取详情页URL
        detail_url = item.find_element(By.CLASS_NAME, "detail_page").get_attribute("href")
        
        # 打开详情页(新标签页)
        chrome.execute_script(f"window.open('{detail_url}');")
        chrome.switch_to.window(chrome.window_handles[-1])
        time.sleep(2)
        
        # 提取电话(点击显示完整号码)
        try:
            chrome.find_element(By.CLASS_NAME, "shortPhone").click()
            time.sleep(1)
            phone = chrome.find_element(By.CLASS_NAME, "fullPhone").text.strip()
        except NoSuchElementException:
            phone = "无公开电话"
        
        # 提取地址
        try:
            address = chrome.find_element(By.CLASS_NAME, "table_right.last").text.strip().replace("\n", "")
        except NoSuchElementException:
            address = "无公开地址"
        
        # 添加到DataFrame
        df.loc[len(df)] = [title, district, detail_url, phone, address]
        
        # 关闭当前详情页,切回列表页
        chrome.close()
        chrome.switch_to.window(chrome.window_handles[0])
        time.sleep(1)
    
    # 尝试翻页
    try:
        next_page_btn = chrome.find_element(By.XPATH, '//*[@id="block_search_results"]/div[4]/div/a[contains(@class,"next")]')
        # 判断按钮是否可点击(有些网站会用disabled类)
        if "disabled" in next_page_btn.get_attribute("class"):
            break
        next_page_btn.click()
        time.sleep(3)  # 等待下一页加载
    except NoSuchElementException:
        # 没有下一页,退出循环
        break

# 最后统一写入CSV,避免重复写入
df.to_csv("28hse_latest.csv", encoding="utf-8-sig", index=False)
chrome.quit()

关键修复点

  • 统一使用Selenium实例:全程用同一个浏览器窗口,不再混用requests,确保拿到的是筛选后的最新房源数据
  • 修正翻页逻辑:用while True循环判断是否有下一页,翻页操作放在当前页所有房源解析完成后
  • 优化详情页处理:用新标签页打开详情页,复用浏览器实例,避免重复创建驱动,提升效率
  • 修复变量错误:补充了之前未定义的title变量,添加异常处理避免页面元素缺失导致崩溃
  • 统一数据写入时机:最后一次性写入CSV,避免循环中重复写入导致数据重复

内容的提问来源于stack exchange,提问作者gou joe

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.22 09:29:55