使用BeautifulSoup抓取Yahoo全球指数表时名称为空的问题
解决Yahoo Finance全球指数爬取的名称缺失与全称获取问题
问题背景
尝试从Yahoo Finance全球指数页面爬取指数代码(ticker)和全称,但导出的CSV中部分名称为空,且已输出的名称仅为页面显示的缩写而非全称。
原爬取代码
from bs4 import BeautifulSoup import pandas as pd # URL of the Yahoo Finance world indices page url = 'https://finance.yahoo.com/world-indices/' # Send HTTP request to the URL response = requests.get(url) response.raise_for_status() # Raise an exception for HTTP errors # Parse the HTML content of the page using BeautifulSoup soup = BeautifulSoup(response.content, 'html.parser') # Find the table containing the indices data table = soup.find('table', {'class': 'W(100%)'}) # Initialize lists to store ticker symbols and names tickers = [] names = [] # Extract data from the table rows for row in table.find_all('tr')[1:]: # Skip the header row cells = row.find_all('td') ticker = cells[0].text.strip() name = cells[1].text.strip() tickers.append(ticker) names.append(name) # Create a dataframe using pandas data = {'Ticker': tickers, 'Name': names} df = pd.DataFrame(data) # Save the dataframe to a CSV file df.to_csv('yahoo_finance_world_indices.csv', index=False) print('Data saved to yahoo_finance_world_indices.csv')
原导出CSV内容
^GSPC,S&P 500 ^DJI,Dow 30 ^IXIC,Nasdaq ^NYA, ^XAX, ^BUK100P, ^RUT,Russell 2000 ^VIX, ^FTSE,FTSE 100 ^GDAXI, ^FCHI, ^STOXX50E, ^N100, ^BFX, IMOEX.ME, ^N225,Nikkei 225 ^HSI, 000001.SS, 399001.SZ, ^STI, ^AXJO, ^AORD, ^BSESN, ^JKSE, ^KLSE, ^NZ50, ^KS11, ^TWII, ^GSPTSE, ^BVSP, ^MXX, ^IPSA, ^MERV, ^TA125.TA, ^CASE30, ^JN0U.JO,
问题原因
- 名称存储位置错误:Yahoo Finance页面中,指数的全称实际存储在第二列
td内部的span标签的title属性中,而非td的直接文本内容。直接取cells[1].text只能获取页面显示的缩写,部分行的缩写为空,导致CSV中对应名称缺失。 - 缺少依赖库导入:原代码中使用了
requests库但未导入,运行时会报错。
修改后的代码
from bs4 import BeautifulSoup import pandas as pd import requests # 补充导入requests库 # URL of the Yahoo Finance world indices page url = 'https://finance.yahoo.com/world-indices/' # Send HTTP request to the URL response = requests.get(url) response.raise_for_status() # Raise an exception for HTTP errors # Parse the HTML content of the page using BeautifulSoup soup = BeautifulSoup(response.content, 'html.parser') # Find the table containing the indices data table = soup.find('table', {'class': 'W(100%)'}) # Initialize lists to store ticker symbols and names tickers = [] names = [] # Extract data from the table rows for row in table.find_all('tr')[1:]: # Skip the header row cells = row.find_all('td') if len(cells) >= 2: # 确保存在至少两个单元格 ticker = cells[0].text.strip() # 获取第二列span标签的title属性,即全称;如果没有则用文本内容兜底 name_span = cells[1].find('span') name = name_span.get('title', '').strip() if name_span else cells[1].text.strip() tickers.append(ticker) names.append(name) # Create a dataframe using pandas data = {'Ticker': tickers, 'Name': names} df = pd.DataFrame(data) # Save the dataframe to a CSV file df.to_csv('yahoo_finance_world_indices.csv', index=False) print('Data saved to yahoo_finance_world_indices.csv')
修改说明
- 补充导入
requests库,解决运行报错问题。 - 调整名称提取逻辑:从第二列的
span标签的title属性中获取指数全称,若span不存在或title为空,则使用td的文本内容作为兜底,避免名称缺失。 - 添加单元格数量判断,防止因页面结构变化导致的索引越界错误。
内容的提问来源于stack exchange,提问作者D.Kang
相关产品推荐
相关产品推荐

