如何使用Selenium获取VPS证券网VN30衍生品页Bid、Ask列价格数据
解决方法
根因说明
你拿到空tbody标签是因为该站点的行情数据是异步加载的:driver.get仅等待页面基础HTML结构加载完成就会继续执行后续代码,此时行情接口还未返回数据,表格内部的行、列内容还没被渲染到DOM中,所以只能拿到空的tbody容器。
可直接运行的修改后代码
from selenium import webdriver from selenium.webdriver.chrome.service import Service from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC # 配置Chrome参数 options = webdriver.ChromeOptions() options.headless = True # 避免无头模式下的反爬检测 options.add_argument("--disable-blink-features=AutomationControlled") options.add_experimental_option("excludeSwitches", ["enable-automation"]) options.add_experimental_option("useAutomationExtension", False) # 适配Selenium 4+的driver初始化方式,旧版本Selenium可保留原有写法 path = 'C:/Users/quank/PycharmProjects/pythonProject2/chromedriver.exe' s = Service(executable_path=path) driver = webdriver.Chrome(service=s, options=options) url = 'https://banggia.vps.com.vn/chung-khoan/derivative-VN30' driver.get(url=url) # 显式等待最多10秒,直到tbody下出现至少一行数据再继续执行 wait = WebDriverWait(driver, 10) wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, 'tbody.price-table-content tr'))) # 提取所有行数据 rows = driver.find_elements(By.CSS_SELECTOR, 'tbody.price-table-content tr') for row in rows: # Bid列对应第7个td,Ask列对应第9个td,可根据实际页面结构调整索引 tds = row.find_elements(By.TAG_NAME, 'td') if len(tds) >= 9: bid_price = tds[6].text.strip() ask_price = tds[8].text.strip() print(f"买价(Bid): {bid_price}, 卖价(Ask): {ask_price}") driver.quit()
关键修改说明
- 新增显式等待逻辑:强制等待表格行数据渲染完成后再提取内容,比固定
time.sleep更高效,不会受网络波动影响 - 新增反爬参数:避免站点检测到无头浏览器拦截请求,导致始终拿不到数据
- 直接定位Bid、Ask列对应的单元格提取内容,适配该站点的表格结构
- 适配Selenium 4+的官方推荐初始化写法,避免旧API弃用带来的报错
内容的提问来源于stack exchange,提问作者Quan Khuc
相关产品推荐
相关产品推荐

