Selenium实现Web scraping返回0条记录且弹窗元素定位报错
问题修复方案
1. 抓取结果为0条数据的问题
问题原因
你使用Selenium完成登录后,直接调用requests库发起目标页面请求,requests和Selenium的浏览器会话完全独立,不会携带登录态Cookie,请求返回的是未登录状态的无权限页面,自然解析不到有效数据。
另外你登录时切换到了登录iframe,后续如果不切回默认页面上下文,所有主页面的元素定位都会失败。
修复方案
放弃requests请求,直接用已登录的Selenium浏览器实例访问目标页面,获取渲染完成的页面源码后再交给BeautifulSoup解析。
2. 弹窗无法定位的问题
问题原因
- CSS选择器语法错误:你写的
button[id='popup-actions fn-close-popup btn autotests__popup-btn-close']是把class属性的多类名错放到了id选择器中,id属性是唯一单值,不可能包含空格分隔的多段内容 - 未从登录iframe切回默认上下文,无法定位主页面的弹窗元素
- 未添加等待逻辑,弹窗还未渲染完成就执行了定位操作
修复方案
- 登录完成后先调用
driver.switch_to.default_content()切回主页面上下文 - 调整正确的元素选择器,比如用类名匹配:
button.autotests__popup-btn-close - 增加显式等待,等弹窗可点击后再执行点击
修正后完整代码
import pandas as pd from bs4 import BeautifulSoup from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.common.by import By from selenium.webdriver.support import expected_conditions as EC from selenium import webdriver # 账号配置 username = "xxxxxxxxxxxxxx@gmail.com" password = "xxxxxxxxxxxxxxxxx" # 初始化Chrome驱动 driver = webdriver.Chrome(executable_path=r'C:\Users\snyder\seaborn-data\chromedriver.exe') driver.maximize_window() driver.implicitly_wait(5) driver.get("https://www.finq.com/en/login") wait = WebDriverWait(driver, 10) # 切换到登录iframe完成登录 wait.until(EC.frame_to_be_available_and_switch_to_it((By.CSS_SELECTOR, "iframe[id='login']"))) driver.find_element(By.ID, "login_username").send_keys(username) driver.find_element(By.ID, "login_password").send_keys(password) driver.find_element(By.CSS_SELECTOR, "button[id='submit_login']").click() # 等待页面加载完成 WebDriverWait(driver, 10).until( lambda x: x.execute_script("return document.readyState === 'complete'") ) # 校验登录结果 error_message = "Incorrect username or password." errors = driver.find_elements(By.CLASS_NAME, "flash-error") if any(error_message in e.text for e in errors): print("[!] 登录失败") else: print("[+] 登录成功") # 切回主页面上下文 driver.switch_to.default_content() # 处理弹窗:等待弹窗可点击后关闭 try: close_btn = wait.until(EC.element_to_be_clickable((By.CSS_SELECTOR, "button.autotests__popup-btn-close"))) close_btn.click() print("[+] 弹窗已关闭") except Exception as e: print("[*] 未检测到需要关闭的弹窗") # 用Selenium访问目标数据页面 target_url = "https://live-cosmos.finq.com/trading-platform/#trading/Shares/Global/USA/All/FACEBOOK" driver.get(target_url) # 等待表格数据加载完成 wait.until(EC.presence_of_element_located((By.TAG_NAME, "tr"))) # 获取渲染后的页面源码解析 soup = BeautifulSoup(driver.page_source, 'html5lib') df = pd.DataFrame(columns=["Instrument", "Sell", "Buy", "Change"]) for row in soup.find_all('tr'): col = row.find_all("td") if len(col) >=4: # 过滤空行/表头行 Instrument = col[0].text.strip() Sell = col[1].text.strip() Buy = col[2].text.strip() Change = col[3].text.strip() df = df.append({"Instrument":Instrument,"Sell":Sell,"Buy":Buy,"Change":Change}, ignore_index=True) print(df) # 退出驱动 driver.quit()
内容的提问来源于stack exchange,提问作者Snyder Fox
相关产品推荐
相关产品推荐

