使用Selenium爬取Benzinga分析师页面时无法跳过初始广告启动页
解决Benzinga分析师页面爬取时的广告弹窗跳过问题
问题概述
用Selenium爬取Benzinga分析师股票评级页面时,遇到以下问题:
- 依赖
sleep函数能处理前两个广告,但第三个订阅弹窗需要鼠标悬停或点击才会加载 - 部分按钮ID带唯一编码(如
prosper-ButtonElement--0pVDSU9IuxoaeJZCJ7Qz),跨设备运行时点击失败 - 最终触发
NoSuchElementException,找不到订阅弹窗的关闭按钮
当前代码
import time import copy from lxml import html import numpy as np import pandas as pd import requests from bs4 import BeautifulSoup from datetime import datetime from selenium import webdriver def get_benzinga_data(): ffox_options = webdriver.FirefoxOptions() #ffox_options.set_headless() tol_amount = 5 tols_curr = 0 bp = False ff = webdriver.Firefox(options = ffox_options) ff.get('https://benzinga.com/analyst-stock-ratings') time.sleep(40) # wait for the large circle advertisment ff.find_element("xpath",'//*[@id="prosper-ButtonElement--0pVDSU9IuxoaeJZCJ7Qz"]').click() time.sleep(30) # wait for the 2nd advertisment ff.find_element("xpath",'//*[@id="__next"]/div[2]/div/div[3]/div/div/div[1]').click() time.sleep(20) # click to try to get mouse over/click on text to force next popup ff.find_element("xpath",'/html/body/div[1]/div[2]/div/div[1]/div[2]/div[1]/div/p').click() time.sleep(20) # fails here as selenium can't find the next popup since it may have not been generated ff.find_element("xpath",'//*[@id="__next"]/div[2]/div/div[1]/div[2]/div[1]/div/div/div[2]/div/div[2]/div/div[2]/div[4]/button').click() get_benzinga_data()
错误信息
NoSuchElementException: Unable to locate element: //*[@id="__next"]/div[2]/div/div[1]/div[2]/div[1]/div/div/div[2]/div/div[2]/div/div[2]/div[4]/button
解决方案
1. 替换固定ID定位为相对定位
放弃依赖带唯一编码的ID,改用按钮文本、属性或层级关系定位,避免跨设备失效:
# 示例:通过关闭按钮的通用属性定位 close_btn = WebDriverWait(ff, 10).until( EC.element_to_be_clickable((By.XPATH, '//button[contains(@aria-label, "Close") or contains(text(), "Close")]')) ) close_btn.click()
2. 用显式等待替代time.sleep()
固定等待时长易受加载速度影响,改用Selenium显式等待,直到元素可交互再操作:
from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC # 封装等待并点击的通用函数 def wait_and_click(driver, locator, timeout=15): try: element = WebDriverWait(driver, timeout).until( EC.element_to_be_clickable(locator) ) element.click() return True except: return False
3. 处理需要触发的订阅弹窗
针对需鼠标悬停或点击才加载的弹窗,先触发元素再等待弹窗出现:
from selenium.webdriver.common.action_chains import ActionChains # 定位触发弹窗的正文元素 trigger_element = WebDriverWait(ff, 10).until( EC.presence_of_element_located((By.XPATH, '/html/body/div[1]/div[2]/div/div[1]/div[2]/div[1]/div/p')) ) # 模拟鼠标悬停触发弹窗 ActionChains(ff).move_to_element(trigger_element).perform() # 等待并关闭订阅弹窗 wait_and_click(ff, (By.XPATH, '//button[contains(text(), "No Thanks") or contains(text(), "Skip")]'))
4. 优化后的完整代码
import time from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.action_chains import ActionChains def wait_and_click(driver, locator, timeout=15): try: element = WebDriverWait(driver, timeout).until( EC.element_to_be_clickable(locator) ) element.click() return True except Exception as e: print(f"点击失败: {e}") return False def get_benzinga_data(): ffox_options = webdriver.FirefoxOptions() # 可选:添加反检测参数 # ffox_options.add_argument("--disable-blink-features=AutomationControlled") # ffox_options.add_argument("--headless") ff = webdriver.Firefox(options=ffox_options) ff.get('https://benzinga.com/analyst-stock-ratings') # 处理第一个广告弹窗 wait_and_click(ff, (By.XPATH, '//button[contains(@aria-label, "Close") or contains(text(), "Close")]')) # 处理第二个广告弹窗 wait_and_click(ff, (By.XPATH, '//div[contains(@class, "close") or contains(text(), "Close")]')) # 触发订阅弹窗 trigger_element = WebDriverWait(ff, 10).until( EC.presence_of_element_located((By.XPATH, '/html/body/div[1]/div[2]/div/div[1]/div[2]/div[1]/div/p')) ) ActionChains(ff).move_to_element(trigger_element).perform() # 关闭订阅弹窗 wait_and_click(ff, (By.XPATH, '//button[contains(text(), "No Thanks") or contains(text(), "Skip") or contains(text(), "Decline")]')) # 后续爬取逻辑... print("弹窗处理完成,开始爬取数据") get_benzinga_data()
额外建议
- 添加
--disable-blink-features=AutomationControlled参数,降低网站对Selenium的检测概率 - 若弹窗仍不稳定,可遍历页面中所有带关闭标识的元素,逐一尝试点击
- 定期检查页面元素结构,避免网站更新导致定位失效
内容的提问来源于stack exchange,提问作者sumabeach
相关产品推荐
相关产品推荐

