请求协助:从美股涨幅页提取前5最高涨幅存入数组并打印
嘿,我来帮你搞定这个抓取WSJ涨幅前5股票数据的问题!你现在的代码会输出所有行列,核心问题是没有对涨幅数据进行排序和截取前5。下面是完整的解决方案,用Python的requests和BeautifulSoup实现:
完整实现代码
import requests from bs4 import BeautifulSoup # 目标页面URL target_url = "http://www.wsj.com/mdc/public/page/2_3021-gainnyse-gainer.html" try: # 模拟浏览器请求,避免被反爬拦截 browser_headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } response = requests.get(target_url, headers=browser_headers) response.raise_for_status() # 检查请求是否成功 # 解析HTML页面 soup = BeautifulSoup(response.text, "html.parser") # 定位涨幅表格(WSJ页面的表格有固定的class标识) gain_table = soup.find("table", class_="mdcTable") if not gain_table: print("抱歉,未找到目标涨幅表格") exit() # 提取表头和行数据 table_headers = [th.get_text(strip=True) for th in gain_table.find_all("th")] stock_data = [] # 遍历表格行(跳过表头行) for row in gain_table.find_all("tr")[1:]: columns = [td.get_text(strip=True) for td in row.find_all("td")] if not columns: continue # 跳过空行 # 处理涨幅数据:把百分比字符串转为浮点数,方便排序 try: change_pct_idx = table_headers.index("% Chg") change_pct = float(columns[change_pct_idx].replace("%", "")) except (ValueError, IndexError): continue # 跳过数据格式异常的行 # 整理单条股票数据 stock_info = { "股票名称": columns[table_headers.index("Name")], "股票代码": columns[table_headers.index("Symbol")], "涨幅(%)": change_pct, "最新价格": columns[table_headers.index("Last")], "成交量": columns[table_headers.index("Volume")] } stock_data.append(stock_info) # 按涨幅从高到低排序 sorted_stocks = sorted(stock_data, key=lambda x: x["涨幅(%)"], reverse=True) # 截取前5条存入数组 top5_gainers = sorted_stocks[:5] # 打印结果 print("📈 涨幅最高的前5条股票数据:") for rank, stock in enumerate(top5_gainers, 1): print(f"{rank}. {stock['股票名称']} ({stock['股票代码']}) | 涨幅: {stock['涨幅(%)']}% | 最新价: {stock['最新价格']} | 成交量: {stock['成交量']}") except requests.exceptions.RequestException as e: print(f"请求页面时出错:{str(e)}")
关键步骤说明
- 模拟浏览器请求:添加
User-Agent请求头,避免被WSJ的反爬机制直接拦截 - 数据清洗与转换:把涨幅的百分比字符串转为浮点数,这样才能进行数值排序
- 排序与筛选:用
sorted()函数按涨幅降序排列,再通过切片[:5]获取前5条数据存入top5_gainers数组 - 异常处理:处理请求失败、表格找不到、数据格式异常等情况,让代码更健壮
你可以直接运行这段代码,它会自动抓取页面数据、筛选出涨幅最高的前5条并打印出来~
内容的提问来源于stack exchange,提问作者Sam Maloney
相关产品推荐
相关产品推荐

