You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Selenium分页爬取Zoopla数据仅获取第一页内容的问题排查

Selenium分页爬取Zoopla数据仅获取第一页内容的问题排查

嘿,我仔细看了你的代码,发现几个关键问题导致脚本只爬取第一页就停止了,咱们一步步来梳理和修正:

1. 浏览器实例冲突(最核心的问题)

你先初始化了undetected_chromedriver的实例,但之后又用with Chrome() as driver:启动了普通的Chrome浏览器——这相当于同时开了两个浏览器,而实际爬取用的是普通Chrome。Zoopla有反爬机制,普通Chrome很容易被检测到,导致下一页按钮无法正常加载、点击无响应,甚至页面直接被拦截。

修正方案:统一使用undetected_chromedriver的上下文管理器,替换掉普通Chrome的调用:

# 删掉这句无用的初始化:driver = uc.Chrome()
with uc.Chrome() as driver:
    driver.get(url)
    wait = WebDriverWait(driver, 10)
    # 后续逻辑保持不变(其他问题再调整)

2. 下一页按钮定位不准确

你当前用的CSS选择器#main-content div._12n2exy2 nav div._14xj7k72可能没有定位到真正可点击的按钮——这类动态生成的类名很容易变化,而且这个元素可能只是按钮的容器,不是可点击的按钮本身。

修正方案:改用更稳定的定位方式,比如通过按钮的data-testid属性或者文本内容定位:

# 方式1:通过data-testid定位(Zoopla的分页按钮通常有这个标识)
next_button = wait.until(EC.element_to_be_clickable((By.CSS_SELECTOR, 'button[data-testid="pagination-next"]')))
# 方式2:通过"Next"文本定位
next_button = wait.until(EC.element_to_be_clickable((By.XPATH, '//button[contains(text(), "Next")]')))

这里一定要用element_to_be_clickable等待按钮可见且可点击,而不是直接找元素,避免按钮还没加载完成就点击失败。

3. 页面加载等待逻辑不够可靠

你用wait.until(EC.staleness_of(houses[0]))等待旧页面元素失效,这个逻辑有时候会出现误判——比如页面部分刷新但旧元素还没被移除,导致脚本提前开始提取数据,或者等待超时。

修正方案:改为等待新页面的结果列表加载完成,确保页面完全切换后再继续:

next_button.click()
# 等待新的结果列表加载出来
wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, "div[data-testid=result-item]")))

修正后的完整代码

from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By
import pandas as pd
import undetected_chromedriver as uc

url = "https://www.zoopla.co.uk/house-prices/england/?new_homes=include&q=england+&orig_q=united+kingdom&view_type=list&pn=1"

# Handle elements that may not be currently visible
def etext(e):
    """Extracts text from an element, handling visibility issues."""
    if e:
        if t := e.text.strip():
            return t
        if (p := e.get_property("textContent")) and isinstance(p, str):
            return p.strip()
    return ""

# Initialize result list to store data
result = []

with uc.Chrome() as driver:
    driver.get(url)
    wait = WebDriverWait(driver, 15)  # 稍微延长等待时间,应对页面加载慢的情况
    
    while True:
        # Wait for the main content to load
        sel = (By.CSS_SELECTOR, "div[data-testid=result-item]")
        houses = wait.until(EC.presence_of_all_elements_located(sel))
        
        # Extract and store data from the current page
        for house in houses:
            try:
                item = {
                    "address": etext(house.find_element(By.CSS_SELECTOR, "h2")),
                    "DateLast_sold": etext(house.find_element(By.CSS_SELECTOR, "._1hzil3o9._1hzil3o8._194zg6t7"))
                }
                result.append(item)
            except Exception as e:
                print(f"Error extracting address or date: {e}")
        
        # Check for "Next" button and move to the next page
        try:
            # 用更稳定的方式定位下一页按钮,等待可点击
            next_button = wait.until(EC.element_to_be_clickable((By.CSS_SELECTOR, 'button[data-testid="pagination-next"]')))
            next_button.click()
            # 等待新页面的结果列表加载完成
            wait.until(EC.presence_of_all_elements_located(sel))
        except Exception as e:
            print("No more pages to scrape or error:", e)
            break  # Stop if no more pages

# Convert results to a DataFrame and display
df = pd.DataFrame(result)
print(df)

额外小提示

  • 可以在每次点击下一页后加个短暂的休眠(比如time.sleep(1)),避免请求过于频繁触发反爬;
  • 定期保存爬取到的数据到文件,避免中途出错丢失数据;
  • 如果还是遇到反爬,可以尝试给undetected_chromedriver添加一些配置,比如设置浏览器窗口大小、启用无痕模式等。

备注:内容来源于stack exchange,提问作者Chioma Okoroafor

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.14 08:29:28