Python动态爬取问题:用BeautifulSoup/Selenium无法获取股票开盘价
无法通过BeautifulSoup/Selenium获取股票开盘价数值
问题描述
尝试用BeautifulSoup获取TradingView页面中US30股票的开盘价(目标数值33931.0),执行代码后仅返回HTML标签而非实际数值;使用Selenium时也无法正确抓取数据。原代码如下:
# 目标元素: <div class="tv-fundamental-block__value js-symbol-open">33931.0</div> import requests from bs4 import BeautifulSoup url = requests.get('https://www.tradingview.com/symbols/PEPPERSTONE-US30/') response = url.content soup = BeautifulSoup(response, 'html.parser') open = soup.find('div', {'class': 'js-symbol-open'}) print(open)
解决方案
1. BeautifulSoup 修复(仅适用于静态渲染内容)
你的代码只获取到了HTML标签对象,需要提取标签内的文本内容。修改代码如下:
import requests from bs4 import BeautifulSoup url = requests.get('https://www.tradingview.com/symbols/PEPPERSTONE-US30/') response = url.content soup = BeautifulSoup(response, 'html.parser') open_element = soup.find('div', {'class': 'js-symbol-open'}) # 提取标签内文本并去除多余空白 if open_element: open_price = open_element.text.strip() print(open_price) else: print("未找到目标元素")
注意:如果运行后仍返回"未找到目标元素",说明该数据是通过JavaScript动态渲染的,requests无法获取到动态加载的内容,此时必须使用Selenium。
2. Selenium 正确实现(处理动态渲染页面)
使用Selenium时需要等待页面元素加载完成,避免因元素未渲染就抓取导致失败。代码示例:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC # 初始化浏览器驱动(需对应你的浏览器版本,比如ChromeDriver) driver = webdriver.Chrome() driver.get('https://www.tradingview.com/symbols/PEPPERSTONE-US30/') try: # 等待目标元素加载完成,最长等待10秒 open_element = WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CLASS_NAME, 'js-symbol-open')) ) open_price = open_element.text.strip() print(open_price) finally: # 关闭浏览器 driver.quit()
该方法会等待页面动态加载完成后再抓取元素,确保能获取到实际数值。
内容的提问来源于stack exchange,提问作者Teddy
相关产品推荐
相关产品推荐

