You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Selenium和Python正确提取网页td元素及股票数据

解决TradingView土耳其股票页面数据提取问题

核心问题分析

目标页面的前3列(股票代码、名称、价格)被包裹在同一个<td>元素内,需要通过子元素拆分提取;行业和市值则是独立的<td>,可直接定位。

静态页面提取方案(使用Requests+BeautifulSoup)

如果页面内容是静态渲染的,可直接通过HTTP请求抓取并解析:

import requests
from bs4 import BeautifulSoup

url = "https://www.tradingview.com/markets/stocks-turkey/market-movers-all-stocks/"
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"
}

# 请求页面
response = requests.get(url, headers=headers)
soup = BeautifulSoup(response.text, "html.parser")

# 提取所有数据行
stock_data = []
rows = soup.select("tbody tr")

for row in rows:
    # 定位包含前3列的主td
    main_col = row.select_one("td:first-of-type")
    if not main_col:
        continue
    
    # 拆分提取股票代码、名称、价格
    ticker = main_col.select_one(".tv-screener__symbol span:first-of-type").text.strip()
    name = main_col.select_one(".tv-screener__symbol span:last-of-type").text.strip()
    price = main_col.select_one(".tv-screener__price").text.strip()
    
    # 提取行业和市值(注意列索引可能随页面结构调整,需验证)
    sector = row.select_one("td:nth-child(4)").text.strip()
    market_cap = row.select_one("td:nth-child(5)").text.strip()
    
    stock_data.append({
        "ticker": ticker,
        "name": name,
        "price": price,
        "sector": sector,
        "market_cap": market_cap
    })

# 打印前5条数据验证
for item in stock_data[:5]:
    print(item)

动态页面提取方案(使用Selenium)

如果页面是JS动态加载的,Requests无法获取完整数据,需用Selenium模拟浏览器:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from bs4 import BeautifulSoup
import time

url = "https://www.tradingview.com/markets/stocks-turkey/market-movers-all-stocks/"
driver = webdriver.Chrome()

try:
    driver.get(url)
    # 等待数据行加载完成
    WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.CSS_SELECTOR, "tbody tr"))
    )
    # 滚动加载更多数据(按需执行)
    driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
    time.sleep(2)
    
    # 解析页面源码
    soup = BeautifulSoup(driver.page_source, "html.parser")
    rows = soup.select("tbody tr")
    
    stock_data = []
    for row in rows:
        main_col = row.select_one("td:first-of-type")
        if not main_col:
            continue
        
        ticker = main_col.select_one(".tv-screener__symbol span:first-of-type").text.strip()
        name = main_col.select_one(".tv-screener__symbol span:last-of-type").text.strip()
        price = main_col.select_one(".tv-screener__price").text.strip()
        sector = row.select_one("td:nth-child(4)").text.strip()
        market_cap = row.select_one("td:nth-child(5)").text.strip()
        
        stock_data.append({
            "ticker": ticker,
            "name": name,
            "price": price,
            "sector": sector,
            "market_cap": market_cap
        })

finally:
    driver.quit()

# 输出结果
for item in stock_data[:5]:
    print(item)

关键注意事项

  • 元素选择器验证:页面元素类名可能随网站更新变化,需用浏览器开发者工具(F12)查看最新结构,调整select/select_one中的选择器。
  • 空值处理:部分行可能存在缺失数据,需添加if判断避免AttributeError。
  • 反爬机制:频繁请求可能触发反爬,需添加请求间隔、更换User-Agent。

内容的提问来源于stack exchange,提问作者Learner

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.03 07:55:30