Python Selenium始终返回首个广告结果,如何获取全部广告详情?
问题分析与解决方案
核心问题
你的代码中,循环内使用的XPath是绝对路径(//div[@class='propertydetails']),这会从整个HTML文档的根节点开始查找匹配元素,所以每次都返回页面中第一个符合条件的元素,而非当前广告项(one)下的子元素。
修复步骤
1. 使用相对路径定位子元素
将循环内的XPath改为相对路径,在开头添加.,表示从当前one元素的上下文开始查找:
title = one.find_element(By.XPATH, ".//div[@class='propertydetails']")
2. 正确获取广告内的meta数据
如果要获取<meta itemprop="position">或<meta itemprop="price">,同样需要使用相对路径在当前广告元素下查找:
# 获取position position_meta = one.find_element(By.XPATH, ".//meta[@itemprop='position']") position = position_meta.get_attribute('content') # 获取price price_meta = one.find_element(By.XPATH, ".//meta[@itemprop='price']") price = price_meta.get_attribute('content')
3. 优化懒加载处理
页面广告是懒加载的,仅滚动到1000px可能无法加载全部广告,建议循环滚动到底部直到所有广告加载完成:
# 滚动加载所有广告 last_height = driver.execute_script("return document.body.scrollHeight") while True: driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") time.sleep(3) new_height = driver.execute_script("return document.body.scrollHeight") if new_height == last_height: break last_height = new_height
完整修复后的代码
from selenium.common.exceptions import TimeoutException from selenium.webdriver import ActionChains, Keys from selenium.webdriver.common.by import By from selenium.webdriver.support.wait import WebDriverWait from selenium.webdriver.support.ui import Select from selenium.webdriver.support import expected_conditions as EC import time from seleniumbase import Driver from selenium.webdriver.remote.webelement import WebElement driver = Driver(uc=True) wait = WebDriverWait(driver, 60) URL = "https://www.nepremicnine.net/oglasi-prodaja/gorenjska/kranj/kranj/stanovanje/letnik-od-1980-do-1989/" driver.open(URL) # 拒绝cookie wait.until(EC.element_to_be_clickable((By.ID,"CybotCookiebotDialogBodyButtonDecline"))).click() # 获取广告数量 cnt_oglasov = wait.until(EC.visibility_of_element_located((By.XPATH, "//div[@class='oglasi_cnt']"))) print(cnt_oglasov.text) st_oglasov_str = "" if "Št. ustreznih oglasov: " in cnt_oglasov.text: st_oglasov_str = cnt_oglasov.text.replace("Št. ustreznih oglasov: ", "") print(f"st. najdenih oglasov: {st_oglasov_str}") # 滚动加载所有广告 last_height = driver.execute_script("return document.body.scrollHeight") while True: driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") time.sleep(3) new_height = driver.execute_script("return document.body.scrollHeight") if new_height == last_height: break last_height = new_height # 获取所有广告项 vsebina = wait.until(EC.visibility_of_element_located((By.XPATH, "//div[@class='seznam']"))) posamezni = vsebina.find_elements(By.CLASS_NAME, 'property-details') # 遍历每个广告 for idx, one in enumerate(posamezni, 1): # 获取广告链接 title_elem = one.find_element(By.XPATH, ".//div[@class='propertydetails']") ad_url = title_elem.get_attribute('data-href') print(f"广告 {idx} 链接: {ad_url}") # 获取position try: position_meta = one.find_element(By.XPATH, ".//meta[@itemprop='position']") position = position_meta.get_attribute('content') print(f"广告 {idx} 位置: {position}") except: print(f"广告 {idx} 未找到position数据") # 获取price try: price_meta = one.find_element(By.XPATH, ".//meta[@itemprop='price']") price = price_meta.get_attribute('content') print(f"广告 {idx} 价格: {price}") except: print(f"广告 {idx} 未找到price数据") print("-"*50) time.sleep(1) driver.quit()
关键说明
- 相对路径的作用:
.开头的XPath会限制查找范围在当前元素的子节点内,确保每个循环都获取当前广告项的对应数据。 - 懒加载处理:通过循环滚动到底部并判断页面高度是否变化,确保所有广告都被加载出来,避免遗漏。
- 异常处理:添加try-except块防止部分广告缺少某些meta标签导致程序崩溃。
内容的提问来源于stack exchange,提问作者Rok Golob
相关产品推荐
相关产品推荐

