You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup抓取Yahoo全球指数表时名称为空的问题

解决Yahoo Finance全球指数爬取的名称缺失与全称获取问题

问题背景

尝试从Yahoo Finance全球指数页面爬取指数代码(ticker)和全称,但导出的CSV中部分名称为空,且已输出的名称仅为页面显示的缩写而非全称。

原爬取代码

from bs4 import BeautifulSoup
import pandas as pd

# URL of the Yahoo Finance world indices page
url = 'https://finance.yahoo.com/world-indices/'

# Send HTTP request to the URL
response = requests.get(url)
response.raise_for_status()  # Raise an exception for HTTP errors

# Parse the HTML content of the page using BeautifulSoup
soup = BeautifulSoup(response.content, 'html.parser')

# Find the table containing the indices data
table = soup.find('table', {'class': 'W(100%)'})

# Initialize lists to store ticker symbols and names
tickers = []
names = []

# Extract data from the table rows
for row in table.find_all('tr')[1:]:  # Skip the header row
    cells = row.find_all('td')
    ticker = cells[0].text.strip()
    name = cells[1].text.strip()
    tickers.append(ticker)
    names.append(name)

# Create a dataframe using pandas
data = {'Ticker': tickers, 'Name': names}
df = pd.DataFrame(data)

# Save the dataframe to a CSV file
df.to_csv('yahoo_finance_world_indices.csv', index=False)

print('Data saved to yahoo_finance_world_indices.csv')

原导出CSV内容

^GSPC,S&P 500
^DJI,Dow 30
^IXIC,Nasdaq
^NYA,
^XAX,
^BUK100P,
^RUT,Russell 2000
^VIX,
^FTSE,FTSE 100
^GDAXI,
^FCHI,
^STOXX50E,
^N100,
^BFX,
IMOEX.ME,
^N225,Nikkei 225
^HSI,
000001.SS,
399001.SZ,
^STI,
^AXJO,
^AORD,
^BSESN,
^JKSE,
^KLSE,
^NZ50,
^KS11,
^TWII,
^GSPTSE,
^BVSP,
^MXX,
^IPSA,
^MERV,
^TA125.TA,
^CASE30,
^JN0U.JO,

问题原因

  1. 名称存储位置错误:Yahoo Finance页面中,指数的全称实际存储在第二列td内部的span标签的title属性中,而非td的直接文本内容。直接取cells[1].text只能获取页面显示的缩写,部分行的缩写为空,导致CSV中对应名称缺失。
  2. 缺少依赖库导入:原代码中使用了requests库但未导入,运行时会报错。

修改后的代码

from bs4 import BeautifulSoup
import pandas as pd
import requests  # 补充导入requests库

# URL of the Yahoo Finance world indices page
url = 'https://finance.yahoo.com/world-indices/'

# Send HTTP request to the URL
response = requests.get(url)
response.raise_for_status()  # Raise an exception for HTTP errors

# Parse the HTML content of the page using BeautifulSoup
soup = BeautifulSoup(response.content, 'html.parser')

# Find the table containing the indices data
table = soup.find('table', {'class': 'W(100%)'})

# Initialize lists to store ticker symbols and names
tickers = []
names = []

# Extract data from the table rows
for row in table.find_all('tr')[1:]:  # Skip the header row
    cells = row.find_all('td')
    if len(cells) >= 2:  # 确保存在至少两个单元格
        ticker = cells[0].text.strip()
        # 获取第二列span标签的title属性,即全称;如果没有则用文本内容兜底
        name_span = cells[1].find('span')
        name = name_span.get('title', '').strip() if name_span else cells[1].text.strip()
        tickers.append(ticker)
        names.append(name)

# Create a dataframe using pandas
data = {'Ticker': tickers, 'Name': names}
df = pd.DataFrame(data)

# Save the dataframe to a CSV file
df.to_csv('yahoo_finance_world_indices.csv', index=False)

print('Data saved to yahoo_finance_world_indices.csv')

修改说明

  • 补充导入requests库,解决运行报错问题。
  • 调整名称提取逻辑:从第二列的span标签的title属性中获取指数全称,若span不存在或title为空,则使用td的文本内容作为兜底,避免名称缺失。
  • 添加单元格数量判断,防止因页面结构变化导致的索引越界错误。

内容的提问来源于stack exchange,提问作者D.Kang

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.26 06:24:55