如何用Selenium和Python正确提取网页td元素及股票数据
解决TradingView土耳其股票页面数据提取问题
核心问题分析
目标页面的前3列(股票代码、名称、价格)被包裹在同一个<td>元素内,需要通过子元素拆分提取;行业和市值则是独立的<td>,可直接定位。
静态页面提取方案(使用Requests+BeautifulSoup)
如果页面内容是静态渲染的,可直接通过HTTP请求抓取并解析:
import requests from bs4 import BeautifulSoup url = "https://www.tradingview.com/markets/stocks-turkey/market-movers-all-stocks/" headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" } # 请求页面 response = requests.get(url, headers=headers) soup = BeautifulSoup(response.text, "html.parser") # 提取所有数据行 stock_data = [] rows = soup.select("tbody tr") for row in rows: # 定位包含前3列的主td main_col = row.select_one("td:first-of-type") if not main_col: continue # 拆分提取股票代码、名称、价格 ticker = main_col.select_one(".tv-screener__symbol span:first-of-type").text.strip() name = main_col.select_one(".tv-screener__symbol span:last-of-type").text.strip() price = main_col.select_one(".tv-screener__price").text.strip() # 提取行业和市值(注意列索引可能随页面结构调整,需验证) sector = row.select_one("td:nth-child(4)").text.strip() market_cap = row.select_one("td:nth-child(5)").text.strip() stock_data.append({ "ticker": ticker, "name": name, "price": price, "sector": sector, "market_cap": market_cap }) # 打印前5条数据验证 for item in stock_data[:5]: print(item)
动态页面提取方案(使用Selenium)
如果页面是JS动态加载的,Requests无法获取完整数据,需用Selenium模拟浏览器:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from bs4 import BeautifulSoup import time url = "https://www.tradingview.com/markets/stocks-turkey/market-movers-all-stocks/" driver = webdriver.Chrome() try: driver.get(url) # 等待数据行加载完成 WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CSS_SELECTOR, "tbody tr")) ) # 滚动加载更多数据(按需执行) driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") time.sleep(2) # 解析页面源码 soup = BeautifulSoup(driver.page_source, "html.parser") rows = soup.select("tbody tr") stock_data = [] for row in rows: main_col = row.select_one("td:first-of-type") if not main_col: continue ticker = main_col.select_one(".tv-screener__symbol span:first-of-type").text.strip() name = main_col.select_one(".tv-screener__symbol span:last-of-type").text.strip() price = main_col.select_one(".tv-screener__price").text.strip() sector = row.select_one("td:nth-child(4)").text.strip() market_cap = row.select_one("td:nth-child(5)").text.strip() stock_data.append({ "ticker": ticker, "name": name, "price": price, "sector": sector, "market_cap": market_cap }) finally: driver.quit() # 输出结果 for item in stock_data[:5]: print(item)
关键注意事项
- 元素选择器验证:页面元素类名可能随网站更新变化,需用浏览器开发者工具(F12)查看最新结构,调整
select/select_one中的选择器。 - 空值处理:部分行可能存在缺失数据,需添加
if判断避免AttributeError。 - 反爬机制:频繁请求可能触发反爬,需添加请求间隔、更换User-Agent。
内容的提问来源于stack exchange,提问作者Learner
相关产品推荐
相关产品推荐

