You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Selenium+BeautifulSoup爬虫遇Cookie弹窗:无法获取目标元素

问题分析与解决

你的代码里有几个关键问题导致Cookie弹窗没被正确关闭,最终rooms返回空值:

  1. 无效冗余代码:soup = BeautifulSoup(r_more_housing.text, 'html.parser')这行完全没用,且r_more_housing变量未定义,直接删掉即可。
  2. 未等待元素加载:页面刚打开就直接查找关闭按钮,可能元素还没渲染完成,导致点击操作根本没生效。
  3. 错误的刷新操作:点击关闭弹窗后调用driver.refresh(),等于重新加载页面,Cookie弹窗会再次出现,之前的点击操作完全白做了。
  4. 解析器缺失:最后解析page_source时未指定解析器,可能导致页面解析异常。

修正后的代码

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from bs4 import BeautifulSoup
import time

PATH = "你的ChromeDriver路径"
link = "https://sturents.com/s/newcastle/newcastle?ne=54.9972%2C-1.5544&sw=54.9546%2C-1.6508"

driver = webdriver.Chrome(PATH)
driver.get(link)

# 显式等待Cookie弹窗关闭按钮可点击,最多等待10秒
try:
    close_btn = WebDriverWait(driver, 10).until(
        EC.element_to_be_clickable((By.CLASS_NAME, "new--icon-cross"))
    )
    close_btn.click()
    # 短暂等待页面状态稳定
    time.sleep(2)
except Exception as e:
    print("未找到Cookie弹窗或点击失败:", e)

# 获取弹窗关闭后的页面源码并解析
soup = BeautifulSoup(driver.page_source, 'html.parser')
rooms = soup.find_all('a', class_="new--listing-item js-listing-item")

print(f"共找到{len(rooms)}个房源")
driver.quit()

关键修改说明

  • 显式等待替代直接查找:用WebDriverWait确保按钮加载完成且可点击,比固定time.sleep更适配网络延迟情况,避免操作失效。
  • 移除刷新操作:刷新会重置页面状态,Cookie弹窗会重新弹出,完全没必要执行这一步。
  • 明确指定解析器:解析page_source时加上html.parser,避免BeautifulSoup自动选择解析器可能出现的兼容问题。
  • 增加异常捕获:避免因Cookie弹窗未出现(比如已保存过Cookie)导致脚本直接崩溃。

内容的提问来源于stack exchange,提问作者explorevaluechain

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 00:20:28