Selenium爬取房产网站存Excel遇问题:重复写入+无法获最新房源
房产网站爬取问题修复
问题原因
重复写入第一页数据:
- 错误混用
requests和Selenium实例:Selenium已打开并筛选最新房源页面,但后续用requests.get请求了未筛选的默认页面,解析的是默认房源数据 - 循环逻辑错误:
for page in range(1,3)中每次都遍历同一组results(第一页数据),导致重复写入 - 翻页代码放错位置:在详情页解析时尝试找列表页的翻页按钮,完全无效
- 错误混用
无法获取最新房源:
- 核心错误是用
requests请求了未经过筛选的页面,而非Selenium操作后已筛选为“最新房源”的页面源码
- 核心错误是用
修正后的代码
import pandas as pd from selenium import webdriver from selenium.webdriver.chrome.options import Options from selenium.webdriver.common.by import By from selenium.common.exceptions import NoSuchElementException import time # 初始化浏览器实例,复用同一个窗口 options = Options() options.add_argument("--disable-notifications") chrome = webdriver.Chrome('./chromedriver', options=options) chrome.get("https://www.28hse.com/en/buy/") chrome.maximize_window() # 切换到最新房源筛选 time.sleep(2) sort_btn = chrome.find_element(By.XPATH, '//*[@id="mainMenuDiv"]/form[2]/div/div[5]/div[2]/div') sort_btn.click() time.sleep(1) latest_option = chrome.find_element(By.XPATH, '//*[@id="mainMenuDiv"]/form[2]/div/div[5]/div[2]/div/div[2]/*[@data-value="latest"]') latest_option.click() time.sleep(3) # 等待筛选结果加载完成 df = pd.DataFrame(columns=["title", "district", "url", "phone", "address"]) while True: # 解析当前页面的房源列表 results = chrome.find_elements(By.CLASS_NAME, "property_item") for item in results: # 提取标题 title = item.find_element(By.CLASS_NAME, "property_title").text.strip() # 提取区域 district = item.find_element(By.CLASS_NAME, "district_area").find_element(By.TAG_NAME, "a").text.strip() # 提取详情页URL detail_url = item.find_element(By.CLASS_NAME, "detail_page").get_attribute("href") # 打开详情页(新标签页) chrome.execute_script(f"window.open('{detail_url}');") chrome.switch_to.window(chrome.window_handles[-1]) time.sleep(2) # 提取电话(点击显示完整号码) try: chrome.find_element(By.CLASS_NAME, "shortPhone").click() time.sleep(1) phone = chrome.find_element(By.CLASS_NAME, "fullPhone").text.strip() except NoSuchElementException: phone = "无公开电话" # 提取地址 try: address = chrome.find_element(By.CLASS_NAME, "table_right.last").text.strip().replace("\n", "") except NoSuchElementException: address = "无公开地址" # 添加到DataFrame df.loc[len(df)] = [title, district, detail_url, phone, address] # 关闭当前详情页,切回列表页 chrome.close() chrome.switch_to.window(chrome.window_handles[0]) time.sleep(1) # 尝试翻页 try: next_page_btn = chrome.find_element(By.XPATH, '//*[@id="block_search_results"]/div[4]/div/a[contains(@class,"next")]') # 判断按钮是否可点击(有些网站会用disabled类) if "disabled" in next_page_btn.get_attribute("class"): break next_page_btn.click() time.sleep(3) # 等待下一页加载 except NoSuchElementException: # 没有下一页,退出循环 break # 最后统一写入CSV,避免重复写入 df.to_csv("28hse_latest.csv", encoding="utf-8-sig", index=False) chrome.quit()
关键修复点
- 统一使用Selenium实例:全程用同一个浏览器窗口,不再混用
requests,确保拿到的是筛选后的最新房源数据 - 修正翻页逻辑:用
while True循环判断是否有下一页,翻页操作放在当前页所有房源解析完成后 - 优化详情页处理:用新标签页打开详情页,复用浏览器实例,避免重复创建驱动,提升效率
- 修复变量错误:补充了之前未定义的
title变量,添加异常处理避免页面元素缺失导致崩溃 - 统一数据写入时机:最后一次性写入CSV,避免循环中重复写入导致数据重复
内容的提问来源于stack exchange,提问作者gou joe
相关产品推荐
相关产品推荐

