使用pandas爬取雅虎财经World Indices列表报错No tables found如何解决
报错原因
雅虎财经的World Indices页面数据为前端JavaScript动态渲染,直接通过HTTP请求获取的原始HTML中不存在完整表格结构,因此pd.read_html()直接读取URL会提示找不到表格。你现有代码中调用driver的逻辑也存在问题:未提前初始化Selenium驱动、未请求目标页面、未等待页面元素加载完成,自然无法定位到对应的表格元素。
方案选择建议
方案1:优先使用pandas-datareader实现需求(最简便)
不需要爬取页面,直接适配雅虎财经数据接口即可拿到全球指数数据,稳定性更高,无需处理页面渲染逻辑:
首先安装依赖库:
pip install pandas-datareader yfinance
参考代码:
import pandas_datareader.data as web import yfinance as yf # 修复pandas-datareader对雅虎最新接口的适配问题 yf.pdr_override() # 全球常用指数代码列表,可按需补充 indices_tickers = ["^GSPC", "^DJI", "^IXIC", "^FTSE", "^N225", "^HSI", "000001.SS", "^STOXX50E"] df = web.get_quote_table(indices_tickers) print(df.head())
方案2:如果需要严格匹配页面展示的表格内容,使用Selenium爬取
需要等待页面动态加载完成后再提取表格,参考代码如下:
import pandas as pd from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC # 初始化Chrome驱动,需提前安装对应版本的ChromeDriver driver = webdriver.Chrome() driver.get("https://finance.yahoo.com/world-indices/") # 等待表格加载完成,最长等待10秒 try: table = WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.TAG_NAME, "table")) ) # 直接提取表格转换为dataframe,无需单独用BeautifulSoup解析 df = pd.read_html(table.get_attribute('outerHTML'))[0] print(df.head()) finally: driver.quit()
补充说明:如果仅需要指数数据,建议优先选择方案1,爬取动态页面的方案受前端结构调整影响较大,维护成本更高。
内容的提问来源于stack exchange,提问作者John
相关产品推荐
相关产品推荐

